In the span of a single week, our engineering team quietly rebuilt the foundation that powers Procheck. What started as a set of targeted fixes to our OEE (Overall Equipment Effectiveness) pipeline has grown into a full backend migration, a microservices rearchitecture, a frontend overhaul, and a broader push toward performance, reliability, and test coverage across the stack. Here's a look at what we've been building, why, and how we're rolling it out.
The problem: an aging backend under growing load
Procheck's original backend has served us well, but as usage has grown, so have the cracks. Our legacy system was showing its age in a few specific ways: data races during concurrent writes, slow read-and-compute operations (especially around OEE calculations), and a codebase that had accumulated technical debt with little documentation and almost no automated test coverage.
Rather than patch these issues piecemeal, we made the call to build a new backend — code-named procheck-oee — designed from the ground up to fix these problems properly.
Starting with data, not guesswork
Before rearchitecting anything, we wanted to know exactly how the system was actually being used — not how we assumed it was being used. We pulled six years of historical usage data to map out our real access patterns: which endpoints get hit hardest, which queries are run most often, what data is read frequently versus rarely, and where the true bottlenecks in the system have historically formed.
That analysis became the backbone of our rearchitecture decisions. Instead of guessing at where to invest engineering effort, we let six years of production behavior tell us where the pain actually lives — and it pointed us toward breaking the monolith apart.
Moving to microservices
One of the biggest structural changes underway is a shift from our existing monolithic backend to a microservices architecture. Splitting the system into smaller, independently deployable services lets us scale the components that need it (like our OEE computation pipeline) without over-provisioning the parts that don't, isolate failures so one slow or broken service doesn't take down the whole platform, and ship changes to individual services faster and with less risk than redeploying a monolith.
This rearchitecture is happening hand-in-hand with the backend migration described below, using the access-pattern data as our guide for where service boundaries should live.
A phased, low-risk rollout strategy
Migrating a production system that customers depend on every day is not something we take lightly. Instead of a single risky cutover, we designed a phased rollout across three parallel deployments:
- procheck-prod — our current production backend, running the legacy codebase.
- procheck-oee-staging — our active development branch, containing the latest code fixes, manually tested and verified.
- procheck-oee-production — a production-grade deployment of the new backend. It's connected to our real production database in read-only mode today, giving us a safe way to validate the new system against live data before it takes over write responsibilities.
This structure lets us test the new backend against real traffic and real data without any risk to our live system. It's also proven to be a major win for iteration speed — having a dedicated staging environment that mirrors production means we can test rearchitecture changes, service splits, and performance tweaks quickly, without touching anything customers depend on.
Phase one is already live: staging.procheck.io, our staging frontend, is now running against procheck-oee-staging. Because the new backend isn't writing to the production database yet, customers using staging won't notice a difference in data consistency — but they are seeing tangible improvements in read performance, particularly around OEE calculations and other read-plus-compute workloads.
Phase two, coming very soon, is the bigger step: cutting over data ingestion and stats-generation triggers from the legacy backend and routing oee.procheck.io to procheck-oee-production in full — write operations included. We're also planning to ship the enhanced UI alongside this cutover, so the improvements are felt end-to-end, not just under the hood.
To keep environments clean and predictable going forward, we're also standing up a new staging.oee.procheck.io domain, dedicated to running the enhanced UI against procheck-oee-staging — giving us a permanent, isolated space to validate future changes before they reach production.
We know that even well-tested migrations can produce a few unexpected hiccups — a handful of support tickets during the transition wouldn't surprise us, and we'd rather be upfront about that than pretend a migration of this scale is risk-free. Our team is ready to respond quickly if that happens.
Performance work: making the database do less, faster
In parallel with the migration, we've been digging into raw performance. A few changes we're making:
- Scaling up our read replica. Our read replica — previously used mainly for read traffic and disaster recovery — is getting a size bump (a modest ~$100/month increase) so we can offload our most expensive queries to it. This takes pressure off our primary database's IOPS and query load, which benefits every operation that touches it.
- Better indexing. Guided by our access-pattern analysis, we've identified and are addressing indexing gaps that were causing unnecessarily expensive queries.
- A connection pooler and lightweight caching layer, deployed where it makes the most sense, to cut down on redundant round-trips and speed up response times.
- Consolidating to a single region. Simplifying our infrastructure footprint reduces cross-region latency and operational complexity.
Together, these changes are aimed at squeezing meaningfully more performance out of our existing infrastructure before we consider scaling up hardware further.
A frontend that matches the backend
The backend isn't the only thing getting rebuilt. Our staging frontend has been substantially modernized — upgraded from an older Umi setup to the latest version of Umi paired with esbuild, along with a completely refreshed visual design, dynamic configuration services, and other under-the-hood improvements. The result is a faster, more maintainable frontend that's ready to grow with the new backend.
Investing in test coverage
One of the less visible but most important parts of this effort has been building out proper test suites. The codebase we inherited had little to no documentation and effectively no automated tests. As part of this migration, we're changing that — writing comprehensive test coverage so that future changes can be made with confidence, not guesswork.
What this means going forward
None of this is about flashy new features — it's about building a backend and frontend that are fast, reliable, and maintainable for the long run. By grounding our rearchitecture in six years of real usage data, rolling this out in phases, validating against real production data before flipping the switch, and investing in the boring-but-critical work of testing and documentation, we're aiming to make this transition as smooth as possible while setting Procheck up for the next stage of growth.
We'll share more updates as the production cutover progresses.
Regards,
Kabeer Jaffri