From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from mail02.haj.ipfire.org (localhost [IPv6:::1]) by mail02.haj.ipfire.org (Postfix) with ESMTP id 4hWGwK64z9z36Wb for ; Thu, 27 Aug 2026 22:51:29 +0000 (UTC) Received: from mail01.ipfire.org (mail01.haj.ipfire.org [172.28.1.202]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange x25519 server-signature ECDSA (secp384r1 raw public key) server-digest SHA384 client-signature RSA-PSS (4096 bits) client-digest SHA256) (Client CN "mail01.haj.ipfire.org", Issuer "YR2" (not verified)) by mail02.haj.ipfire.org (Postfix) with ESMTPS id 4hWGwG2ff3z2xLJ for ; Thu, 27 Aug 2026 22:51:26 +0000 (UTC) Received: from layka.disroot.org (layka.disroot.org [178.21.23.139]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange x25519 server-signature RSA-PSS (4096 bit raw public key) server-digest SHA256) (Client did not present a certificate) by mail01.ipfire.org (Postfix) with ESMTPS id 4hWGw53dcdz1TJ for ; Thu, 27 Aug 2026 22:51:17 +0000 (UTC) Authentication-Results: mail01.ipfire.org; dkim=pass header.d=disroot.org header.s=mail header.b=T3fb0LuX; spf=pass (mail01.ipfire.org: domain of robin.roevens@disroot.org designates 178.21.23.139 as permitted sender) smtp.mailfrom=robin.roevens@disroot.org; dmarc=pass (policy=reject) header.from=disroot.org ARC-Seal: i=1; a=rsa-sha256; d=lists.ipfire.org; s=202003rsa; cv=none; t=1787871077; b=KeuW8l4mmJxFlzmuXlM+qCMCB7FDfN696KigR6K6LbEAYJ2Jr7UGzDAX8Il51u7yeMqIKK 7EDZp04QPuCtAOBaOrRNwOCYBVYIDea2Dk2nBq4pVJnSchTbEJj6OUXGKdo9x8oCeWwQv6 Z0TSleREBJz+9CFM1heNDY023Z7zi9AiDViJvCb4FFRCjZtGClD68i7Vty0sdzoYC8XlGF RQJ6TufDwaS5tybqiP+HTC4r+9Fipnx/GX/nzjepoIoe/sVWuqI2GubjiwTHoO0yGAZtdV JSWhFtp5sfLpRQicLh4saF8KbOmInQZGAHCH/pBwdublO/m0py75FoCmoMEWBg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=lists.ipfire.org; s=202003rsa; t=1787871077; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=4Nym7NNYqC8Yo8a84pfrn8UXWYaenAR6dc48LqIzMWs=; b=JwvGIUHRnIDLZl/xxEj4f/j7rrXhFgGqtizJoVJJOKPmcQ+p2runRxZi9E86vd1lXzYclX qufjZnwW1wg02OlM/XHo2sNGP+S40dT7WKXIRrDhES/xOf/r+kZemvYGyz+xlmbzyFlrlr z1muWquBlWGW9tvc3FGH9cXDic/81U4GDxOuhFXqwdTTwdv3OqIw4lT25q29yJj7rvkWD1 DmB9H/NRsyHnZUQrpcsN1lUSTbZNtj7XM26CFtKl57AVbupBBXx5I5TgVlT6hsCiVwXmhN a32wbN4lV5JyNubQMiz3Sr1hVHoa2WfyKmE5yqevhe4FpoDLvj38t9liFcHmpw== ARC-Authentication-Results: i=1; mail01.ipfire.org; dkim=pass header.d=disroot.org header.s=mail header.b=T3fb0LuX; spf=pass (mail01.ipfire.org: domain of robin.roevens@disroot.org designates 178.21.23.139 as permitted sender) smtp.mailfrom=robin.roevens@disroot.org; dmarc=pass (policy=reject) header.from=disroot.org Received: from mail01.layka.lan (localhost [127.0.0.1]) by disroot.org (Postfix) with ESMTP id 0312088282; Fri, 28 Aug 2026 00:51:17 +0200 (CEST) X-Virus-Scanned: SPAM Filter at disroot.org Received: from layka.disroot.org ([127.0.0.1]) by localhost (disroot.org [127.0.0.1]) (amavis, port 10024) with ESMTP id zr2ozUxjb6Dm; Fri, 28 Aug 2026 00:51:15 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=disroot.org; s=mail; t=1787871075; bh=MgD4TV5RJ4F7BMHbXXFqANgzWp9BYD7tsGwpc7eU5C0=; h=Subject:From:To:Cc:Date:In-Reply-To:References; b=T3fb0LuX2Ssln36ARDSYSI1ZhuozH+EVbAuDJ/HpbsYvryQ06Wz6dCn6iP786Y3Pd qIZ78Gi2EpVVxkNkp1/7K5e9Mdq8HMKh/Cl9iUF2Fcig+3gJFnhPZA6w1aw58jbbRo xWJu0Sf0M/k8Zr5uNPjipvTs1Cky66WrR1SK7aBmM3X1BNQuzPWYJLn/wg3aTAETT4 DCJzg2YNpsO6bdEfwo+SApyPonWoSf0c8Hexns9rvy1hmm/onkAx0oQ1MyxBki7j+F 3Hm4UA4k3hYNJPvXeEUv7oSbPm/c6zjmBEx/oNHTEzuvuXL3vbNMpteC5pWsfDgmSC dzgprzYBL5Cng== Received: from chojin.roevenslambrechts.be (chojin.roevenslambrechts.be [192.168.0.50]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits)) (no client certificate requested) (Authenticated sender) by hachiman (MailScanner Milter) with SMTP id D2FFA5A8649; Fri, 28 Aug 2026 00:51:12 +0200 (CEST) Message-ID: <5c15819a5f758f2700cf725402c4d563699bd5bf.camel@disroot.org> Subject: Re: [PATCH 0/5] Add Zabbix functionality to suricata-reporter From: Robin Roevens To: Michael Tremer Cc: development@lists.ipfire.org Date: Fri, 28 Aug 2026 00:51:11 +0200 In-Reply-To: References: <20260730195148.3278295-1-robin.roevens@disroot.org> Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable User-Agent: Evolution 3.60.2 Precedence: list List-Id: List-Subscribe: , List-Unsubscribe: , List-Post: List-Help: Sender: Mail-Followup-To: MIME-Version: 1.0 X-RoevensLambrechts-MailScanner-ID: D2FFA5A8649.A4267 X-RoevensLambrechts-MailScanner: Found to be clean X-RoevensLambrechts-MailScanner-From: robin.roevens@disroot.org X-RoevensLambrechts-MailScanner-Watermark: 1788475873.59212@0LPXST4TZv6+PJzYDRumFg X-Rspamd-Action: no action X-Spamd-Result: default: False [-7.62 / 11.00]; NEURAL_HAM(-3.00)[-1.000]; BAYES_HAM(-3.00)[100.00%]; SPF_REPUTATION_HAM(-0.62)[-0.61867959927556]; DMARC_POLICY_ALLOW(-0.50)[disroot.org,reject]; R_SPF_ALLOW(-0.20)[+a]; R_DKIM_ALLOW(-0.20)[disroot.org:s=mail]; MIME_GOOD(-0.10)[text/plain]; ARC_SIGNED(0.00)[lists.ipfire.org:s=202003rsa:i=1]; MIME_TRACE(0.00)[0:+]; TO_DN_SOME(0.00)[]; ARC_NA(0.00)[]; ASN(0.00)[asn:50673, ipnet:178.21.23.0/24, country:NL]; MX_INFLIGHT(0.00)[disroot.org]; DKIM_REPUTATION(0.00)[0]; IP_REPUTATION_HAM(0.00)[asn: 50673(0.00), country: NL(-0.01), ip: 178.21.23.139(0.00)]; RCPT_COUNT_TWO(0.00)[2]; FROM_EQ_ENVFROM(0.00)[]; FROM_HAS_DN(0.00)[]; MID_RHS_MATCH_FROM(0.00)[]; TO_MATCH_ENVRCPT_SOME(0.00)[]; RCVD_COUNT_THREE(0.00)[3]; RCVD_TLS_LAST(0.00)[]; DKIM_TRACE(0.00)[disroot.org:+] X-Rspamd-Server: mail01.haj.ipfire.org X-Rspamd-Queue-Id: 4hWGw53dcdz1TJ Hi Michael Vacation period here is officially over.. So I have no more excuses and I'm ready to dive into this again :-) Michael Tremer schreef op vr 31-07-2026 om 11:24 [+0100]: > Hello Robin, >=20 > Thank you very much for sending these patches. >=20 > Before we dig into the code, I have a couple of questions about the > design... Ok, I will try to answer them first, as discussing this may result in significant design changes :-) >=20 > > On 30 Jul 2026, at 20:15, Robin Roevens > > wrote: > >=20 > > Hi all, > >=20 > > As discussed here earlier, I've worked on implementing sending > > Suricata alerts straight to Zabbix from within suricata-reporter > > instead > > of trying to parse the suricata logging separately using the Zabbix > > agent. > >=20 > > For this I use the zabbix-utils python library, which I submited > > here > > also as a separate pak (but meanwhile already requires an update, > > which > > I will post soon). This set of patches makes suricata-reporter able > > to > > directly communicate to a Zabbix server without having the > > zabbix_agentd > > pak installed, sending suricata alerts in real-time. >=20 > Yes, this is a good choice and I like that suricate-reporter will try > to load support for Zabbix and if the module is not available, it > simply disables support for Zabbix. That allows us to have a smaller > configuration file if things like this are auto-detected. >=20 > > As Zabbix supports sending items in bulk, I have opted to create an > > async background task that will send all events from last 1 second > > in > > bulk so that even in the case that there are hundreds of incoming > > alerts, Zabbix server is only contacted once per second. >=20 > Okay, this makes sense. But I believe that there is already a small > race in the implementation: >=20 > If the client side (in this case suricata-reporter) does not finish > the call of flush_pending_to_zabbix() within that second, it will be > called again which will result in the same rows being selected again, > transmitted again, and assuming that there are just thousands of > alarms it will take over a second again, the function will be called > again, and so on. So the application will stall very quickly. >=20 > Although we should not see thousands of alerts per second under > normal conditions, there could be other reasons why this is taking > some time. For example, the Zabbix host could be in a different > location and round-trips around half the planet are taking some time; > it could be busy writing other things to its database or the database > has just decided to do a little cleanup job. One second isn=E2=80=99t a l= ot > of time then and we will have to make the system a little bit more > resilient against this. Is this really the case? For what I understood of the Python async methods is that by using await in: async def _periodic_zabbix_sender_flush(self): ... if await self.flush_pending_to_zabbix(): ... await asyncio.sleep(1) The task will wait for the flush/sending to complete before waiting 1 sec so the next itteration should not be able to begin until flush_pending_to_zabbix effectively returns. Hence if the sending would take 20s the flow would be: flush starts=C2=A0 -> send pending events, taking 20s -> wait 20s until flush_pending_to_zabbix finishes -> wait an additional 1s -> select pending rows again=20=20 So there should be no overlaps with previous calls and the same rows are not concurrently selected and transmitted by this task. >=20 > > When for some reason sending to Zabbix server fails, it will be > > retried > > 3 times and then the background task will be suspended until a new > > suricata event comes in. That will wake the task again and retry to > > send all > > pending events. In environments with many events, that may actually > > not > > have that much of an effect. But in the average environment, this > > will > > give the Zabbix Server some breathing space as it failing to > > receive our > > events, may indicate a Zabbix server overload. >=20 > Good thinking here. >=20 > > For this I have to keep track which events are sent and which are > > pending. So I added a column in the database that keeps track of > > that. >=20 > So, this is a very crucial thing we probably need to discuss :) >=20 > What is the rationale behind this? Obviously there are some easy > answers: >=20 > 1) We don=E2=80=99t want to loose any history if the network or Zabbix is > down >=20 > 2) We can even restart the reporter without losing any alerts I would add that it can even be killed, crash, powerfail, kernelfail or any other disaster may happen. When it restarts (and still has its database) the events won't be lost :-) >=20 > But then I am already running out of ideas why this could be a good > idea. The cons that I can see are: >=20 > * A lot of additional I/O on the database. Although we would be > updating rows very briefly after they have been written to the > database, it will create a copy of the row and change the append-only > architecture of the database. It will have a lot more cleaning up to > do to evict all updated rows. True >=20 > * You will only ever go back by about 1h by default. Could we just > not keep things in RAM for that long? I would rather not only keep it in memory. On systems with only one alert every x time, that won't be a problem, but on systems with many alerts per second, we risk losing many alerts by any failure.=20 Maybe postponing DB writes a few alert-batch sends is possible, but with memory-only the risk of lost alerts is too high for me. I considered only updating the DB on shutdown, but an unexpected shutdown/crash would then potentially cause large replay bursts depending on how long reporter has been running, which could be days, months, years (hopefully not, as they should upgrade their IPFire regularly ;-)) >=20 > I am not saying that I hate the idea, but I am not sure whether it is > worth paying the price. The good side is that if people are not using > Zabbix, there is no overhead except the space for the extra column. > But if we would add another monitoring solution, we would potentially > have to add another field, and another, and another? That is indeed one of the goals: no extra overhead when Zabbix is not used. But I do think if people bother to set up a proper monitoring and/or logging system, they generally would like the data flow as robust as possible. Such systems can also be configured to react on incoming data, possible starting whole workflows, making it even more important that there is no data missing. I may have a solution for adding more alert consumers a bit further in this mail.=20 >=20 > So a possible other solution that I can come up with would be: > Creating a separate table with all pending events that have to be > transmitted. And every once in a while we truncate it should it > become too long. We could even keep a list of IDs in memory only if > we want to go down that route. >=20 Considering your valid remarks, I have been rethinking possible other methods, trying to keep the alerts-table append-only and SQLite work minimal and came up with 3 alternatives to the current method: * A separate table for keeping the queue of event id's to be sent -> Pending queries should be cheaper and the table would remain bounded if delivery keeps up. However, it would require deletes for every event that is sent, still causing extra SQLite writes and cleanup work, especially with high event counts. I don't see how to implement once-in-a-while truncating on this kind of table if we are not updating every event once it is sent to actually mark it as sent. So I don't think this would be less work for SQLite? Also a crash between inserting an alert in the alert table and adding the alert to the queue table would cause missing deliveries or it should be done in one transaction, possible causing longer lock times. So I don't think that is a good solution. * A separate table for keeping successfully sent event id's -> This makes the separate table also append-only and would allow for once-in-a-while truncating the table as you suggest. It would require a query on the alerts table using NOT EXISTS which will probably be a little more costly for SQLite than current method or the pending table method, although an indexed primary key should make this inexpensive enough? Worst case after a crash between sending a batch of events succesfully and recording it in this table, would be that those events would be sent again, causing duplicate entries in Zabbix. But in my opinion it is better to have an alert reported twice than missing it. * Using a high-watermark in a separate "state" table -> This would record only the last successfully sent alert ID and update it when a new alert or batch of alerts is successfully sent. So even with a high alert rate, only a single update is required after send a batch of alerts succesfully. So in worst case there is still only a single update once per second. The query on the alerts table would also be cheap only using a ">" equation on the primary key. It could look something like: zabbix_state ------------ last_sent_id INTEGER NOT NULL it could even be used for possible additional future consumers or maybe even for tracking succesfully sent alerts by syslog or email? consumer_state ------------ consumer TEXT PRIMARY KEY -> in this case consumer =3D "zabbix" last_sent_id INTEGER NOT NULL Caveats are that I must make sure alerts are always sent in order of their ID and when an alert failed to be sent, it would block newer alerts to be sent. And as I send alerts in bulk, which in turn is chopped into chunks by zabbix_utils itself, it is possible that out of 3 chunks the middle chunk failed and in that case there are events with higher ID's sent and lower ID's that failed, rendering this method unusable.=C2=A0So I will then need to split the batches into chunks <=3D zabbix_utils chunk-size and send them separately myself. Possibly generating more DB writes within a second. (However current default chunk-size is 250, so it takes > 250 alerts within a second to cause an extra DB update) But then I still risk that if, for some reason a single alert consequently fails to be sent (I don't think this should ever happen, but you never know..Maybe a bug in Zabbix failing to parse the event due to some unexpected character or something like that?), would still block any subsequent alert to be sent.=C2=A0 And by batch-sending, there is unfortunately no way of knowing which alert(s) in the batch failed, and which where successfully accepted. Zabbix server only returns how much have failed and how much have succeeded. Currently I retry sending the whole chunk for an hour (by default), but I don't block newer events. In this method a failed batch would be resent indefinitely so some sane threshold should also be implemented here. Possibly something like retries=3D3.. however this would risk dropping alerts too soon when Zabbix or the network is effectively having troubles itself.. Not sure yet how to handle this, maybe retries=3D3 if there are any succesfully sent alerts in the batch and up to 1 hour if all events in the batch fail. Or something like that.. Or I could try implementing to recursively split a failed batch until such a possible single failing alert is separated, and then drop that alert, but I think that feels quite overkill as this situation should actually never happen. This last method introduces some challenges, but I think it has a very high potential of being extremely cheap, both on database operations and size as only ever one single row is maintained, and makes it also cheap to add even more alert consumers, therefore may be worth investigating deeper? > > I have also added an alert_max_age config parameter that allows the > > user > > to set how long suricata-reporter should retry to send events to > > Zabbix. > > Events older than that set age, will no longer be sent to Zabbix. > > This also give the user the implicit option to send older events > > when > > only just enabling the zabbix sending functionality, since the DB > > column > > exists and no event was ever sent to Zabbix, all events will be > > 'pending". At first run with zabbix functionality enabled, all > > events up > > to alert_max_age that are in the database will be sent to zabbix > > immediatly. >=20 > I like the mechanism, but whenever I am building something like this, > I am never sure what would be a reasonable window. Me neither, I think this highly depends on the user's infrastructure and/or needs. If network outages or zabbix server outages are expected to be longer than 1 hour in some environments, then 1 hour is probably not a good threshold.. Therefore I would definitely make it user configurable. But a default of one hour, feels sane to me. >=20 > Locally, suricate-reporter is keeping the events for pretty much > forever. So we could even go back three days or something. Or we > could give up really quickly. After maybe a minute. I never know what > is right, but for the implementation, the length of the window plays > a role - see above. >=20 > With email and syslog we do more of a =E2=80=9Cfire and forget=E2=80=9D a= pproach. If > we send the syslog message and syslog wasn=E2=80=99t ready to receive it,= we > wouldn=E2=80=99t know and we would not try again... If you have no way of knowing, there are no other options, I think. But in the case of Zabbix, we do know. And knowing how much fuss the SoC team at my work makes when some security related logs are missing, I think many really do like it to be as reliable as possible when exporting the alerts to a monitoring or other collecting system.=20 >=20 > > All events sent to Zabbix contain the timestamp of retrieval by > > suricata-reporter, so Zabbix will register and order them as > > received on that > > timestamp independently of the actual time Zabbix itself received > > the > > event. > >=20 > > This is my first adventure in Python async programming, so I hope I > > did > > not make any flagrant mistakes. But the code has been running here > > for > > weeks now without any problem. I have not actually tested large > > bursts > > of events, as I could not simulate that.. But I did make Zabbix > > server > > slow, unavailable and finally replaced it with netcat (to accept > > the connection, but > > not react on it) and I had the connection with the server off for a > > few > > hours to then re-establish the connection to see hundereds of > > pending events=20 > > being registered in only a few milliseconds. > > I did not notice any problems with suricata-reporter in any of > > these > > cases. >=20 > This is good testing. Usually, if I need to create a lot of events, I > enable the =E2=80=9CPING=E2=80=9D rule in =E2=80=9Cicmp_info=E2=80=9D and= just send a lot of ping > packets to the firewall. You could try a flood ping with =E2=80=9Cping -f= =E2=80=9D. >=20 > I will send some more comments about the code in the other emails. For now I will await your (or maybe other list members'?) reaction to above db approach proposals before diving into your code comments. But I will make sure to review and properly implement/consider or answer them when I start changing the code based on the outcome of current discussion. Regards Robin >=20 > Best, > -Michael >=20 > >=20 > > Regards > >=20 > > Robin > >=20 > > --=20 > > Dit bericht is gescanned op virussen en andere gevaarlijke > > inhoud door MailScanner en lijkt schoon te zijn. > >=20 > >=20 >=20 --=20 Dit bericht is gescanned op virussen en andere gevaarlijke inhoud door MailScanner en lijkt schoon te zijn.