Tracking and debugging millions of IoT Packets at scale

August 18, 2026

Some of the hardest bugs we've had to chase weren't in our application logic at all — they were in the data itself. When a plant floor is generating hundreds of thousands of sensor packets per shift, "something looks off with the counts" is not a bug report you can fix by reading code. You need to see the data flow in, second by second, and spot exactly where it breaks down.

That need is what led us to build a small but critical internal tool: a dev dashboard that visualizes sensor heartbeats and packet arrivals in real time, and lets us scrub back through any shift to reconstruct exactly what happened.

The problem: debugging by guesswork

Our sensors report constantly — counters, status changes, up/down events, blowups — and at scale that adds up to an enormous volume of packets per line, per shift. When something goes wrong — a missing count, a suspicious flat line, a sensor that seems to have gone silent — the only way to investigate used to be digging through raw logs or database rows, one query at a time, trying to reconstruct a timeline in your head.

That approach doesn't scale. It's slow, it's error-prone, and it makes it nearly impossible to spot patterns like a burst of duplicate packets, a gap where a sensor stopped reporting, or a spike that doesn't line up with anything on the production floor.

We needed a way to see the shape of the data, not just the rows.

The fix: a point graph for packets and heartbeats

We built a dev-only dashboard that plots sensor activity as a point/line graph over time — heartbeats, packet counts, and status transitions all layered on the same timeline. Instead of scrolling through logs, we can now look at a graph and immediately spot:

  • Count spikes — sudden jumps that don't match expected production behavior, often the first sign of duplicate or malformed packets.
  • Gaps and flat lines — a sensor that stopped sending heartbeats, visible instantly as a break in the line rather than a silence in the logs.
  • Correlated events — downtime, pauses, and line up/down events plotted alongside the raw counts, so we can see cause and effect in one view instead of cross-referencing multiple tables by hand.

Filters let us scope down to a specific plant, line, machine, and sensor, and toggle event types (downtime, pause, line up/down, blowups) on and off so we can isolate exactly the signal we're chasing.

Solving the timezone problem

One detail that seems small but caused real debugging pain: our database stores everything in UTC, but nobody on the floor — or on our team — thinks in UTC. A shift that "started at 8am" means 8am local time, not 8am UTC. When you're trying to line up a customer's report of "something happened around 2pm" with raw UTC timestamps in the database, every investigation starts with manual time-zone arithmetic — a great way to introduce off-by-a-few-hours mistakes into an already tricky bug.

The dashboard handles this conversion automatically. You pick a date range in local time, and the tool converts it to the correct UTC window under the hood for querying — while still showing the UTC range alongside it, so nothing is hidden and we can always sanity-check the conversion. It's a small thing, but it removes an entire category of "wait, was that AM or PM UTC?" confusion from every debugging session.

Accounting for every packet

The real goal behind this tool isn't just pretty graphs — it's accountability. With volumes this high, "trust the pipeline" isn't good enough; we need to be able to prove that every packet a sensor sent was received, processed, and reflected correctly downstream. This dashboard gives us that receipt. If a customer or an internal alert flags something suspicious, we can pull up the exact window, see the raw packet stream, and confirm — packet by packet if we need to — whether the system behaved correctly.

Why this matters beyond debugging

This started as a tracing tool to solve gnarly, hard-to-reproduce issues, but it's turned into something we lean on constantly:

  • Faster incident response. What used to take an engineer an hour of log spelunking now takes a couple of minutes on the dashboard.
  • Confidence during the backend migration. As we roll out our new OEE backend in phases, this tool is one of our key ways of verifying that the new system is processing packets correctly against real production data before we cut over write traffic.
  • A shared source of truth. Instead of everyone reconstructing timelines from raw queries differently, the whole team looks at the same graph and the same conversation gets a lot shorter.

What's next

We're planning to extend the dashboard with anomaly highlighting — automatically flagging spikes, gaps, or out-of-pattern bursts rather than requiring an engineer to eyeball the graph — and hooking it into our alerting so a suspicious pattern surfaces before a customer ever notices it.

Tools like this rarely make it into a product demo, but they're often the difference between a five-minute fix and a day lost to guesswork. Given how much time it's already saved us, it's easily one of the highest-leverage things we've built this quarter.