The first mistake I see is teams talking about real time when they actually mean same day. That sounds like a wording problem until you are in a command center and a bed manager is looking at a stale census tile from an hour ago. At that point the issue is not semantics. It is an operational lie.
I do not design these systems around a single magical data lake. I design them around the clinical events that matter: ADT, orders, results, medications, notes, claims-adjacent operational data, and the reference data that lets those facts make sense. In health systems, the warehouse is only as good as the integration layer feeding it. If HL7v2 ADT messages are late, if lab results arrive out of order, or if the same encounter comes through Epic and a downstream interface engine with different identifiers, the warehouse becomes a very expensive guessing machine.
At AST, I have seen this play out the same way more than once. A team will ask for a unified patient timeline, then discover that three departments define discharge differently. One source says the patient left the unit. Another says the encounter closed. A third says the account hit a billing-ready state. If you do not model those as distinct business events, your warehouse will flatten them into nonsense. We learned that the ugly way on an integration-heavy build where near real-time feeds looked clean in test and fell apart the first time an interface engine retried a batch after a network blip. The data was not missing. It was duplicated, and our downstream consumer had no way to know which copy was the truth.
My rule is simple: design for provenance first, querying second. Every record needs a lineage path that a data engineer can explain and a clinical analyst can audit without opening six tickets. That means I care about all of this up front:
- Source authority Which system owns the field: Epic, Cerner/Oracle Health, athenahealth, a departmental app, or the interface engine.
- Event time vs ingest time When did it happen clinically, and when did we receive it.
- Identity resolution How do we connect the patient, encounter, order, and result across systems without creating false merges.
- Late-arriving data rules What happens when a lab result shows up after the discharge summary.
- Change detection How do we handle corrections, cancellations, amended notes, and replaced results.
If you get those wrong, the warehouse will still load. That is the trap. Everything will look operational until a cardiologist asks why the troponin trend changed overnight, or a quality team finds that a denominator shifted after a patient merge. The system will not scream. It will quietly rewrite history.
The best architecture I know starts with a narrow question: what decisions need to happen within minutes, not hours? That answer defines your latency budget. A patient safety alert might need a near-immediate medication administration feed. A population health panel usually tolerates longer freshness. A bed status board sits somewhere in the middle. If you try to architect everything as ultra-low-latency, you pay for complexity you do not need. If you treat all data as batch, you break the workflows that actually need timeliness.
That is why I separate use cases into operational, clinical, and analytical tiers. Operational consumers need freshness and traceability. Clinical users need confidence that the warehouse is not contradicting the chart. Analysts need stable definitions more than speed. When those three are mixed in one model, every release becomes a compromise and every bug becomes a policy meeting.
| Layer | What it stores | Why it exists | Common failure mode |
|---|---|---|---|
| Landing | Raw HL7v2, FHIR R4, X12, vendor API payloads | Replay, audit, rollback | Teams transform too early and lose evidence |
| Canonical | Normalized clinical events with lineage | Single logic layer for patient, encounter, order, result | Over-modeling slows ingestion and breaks delta handling |
| Serving | Curated marts, metrics, dashboards, extracts | Fast queries for users and downstream apps | Business logic leaks into every dashboard |
The table looks boring. That is the point. Real-time warehouse design is mostly about refusing to let complexity leak upward. If the serving layer has to know how twelve source systems describe a canceled order, you have already lost.
One of the few things I disagree with strongly is the idea that a lakehouse by itself solves clinical analytics. It does not. It gives you storage and compute. It does not give you governance, identity resolution, clinical semantics, or a reliable replay strategy. I have watched teams buy the platform and then spend months inventing the hard parts from scratch. The platform was never the missing piece. The data contract was.
At AST, when we build this kind of platform, we treat the integration path and warehouse path as one delivery problem. That is the same way we think about our EHR and interoperability work in general: if the edge feed is messy, no downstream model can save it. I have seen clean FHIR resources from one source and malformed HL7v2 segments from another arrive on the same day, and the warehouse had to reconcile both without losing the audit trail. That is normal. The design has to assume normal is chaos.
If you want this to work, I recommend a sequence that keeps each decision testable. I use this playbook with health systems that want real-time behavior without breaking the enterprise model:
- Map the decisions, not the data Start with the operational questions that need freshness. Define who uses the output, what delay they can tolerate, and what happens if the data is wrong.
- Declare source authority Assign a source of truth for each domain field. Do not assume one system owns everything just because it has the largest contract.
- Build the landing layer for replay Preserve the raw payloads, timestamps, message IDs, and transformation metadata so you can reprocess after a vendor correction or interface outage.
- Normalize into a clinical event model Convert raw feeds into a canonical structure that can represent admissions, transfers, discharges, results, orders, administrations, and amendments without flattening important nuance.
- Add survivorship and merge rules Decide which record wins when there are duplicates, corrections, cancellations, or patient merges. Write those rules down where engineers and analysts can see them.
- Test with ugly data Feed the pipeline late messages, duplicates, out-of-order updates, and source resets. If your warehouse only works with pristine data, it does not work.
- Instrument freshness and drift Monitor ingest lag, failed transformations, row-count anomalies, and metric drift. The warehouse should tell you when it starts lying.
The last step is the one most teams skip. They build observability for the platform and ignore observability for the meaning. I want alerts when a medication feed drops, yes. I also want alerts when the definition of an encounter changes because a source system release quietly altered a field map. That is not a technical edge case. That is where trust goes to die.
We have had projects where the surprise was not scale. It was semantics. A source that looked stable for years started emitting slightly different status codes after an upgrade. The warehouse loaded them just fine. The dashboards did not. Everyone blamed the BI layer. The actual bug was upstream meaning drift, and if we had not kept the raw event layer, we would have spent weeks trying to reconstruct the lost context.
If you are evaluating this work, ask the vendor or the internal team the questions that matter:
- Can they replay a day of data after a source correction without manual intervention?
- Can they explain which layer owns patient identity and why?
- Do they treat HL7v2, FHIR R4, and API payloads as equal citizens in lineage, or do they pretend one format is enough?
- Can they separate operational freshness from analytical stability?
- Will a clinician be able to trust the result after a source system merge or amendment?
Those questions are not theoretical. They are the difference between a warehouse that helps care teams and one that makes them double-check everything in Excel. If that sounds harsh, it should. I have shipped enough of these systems to know the cost of ambiguity shows up later, in angry users and broken confidence.
That is also why I like using a service-oriented platform approach instead of a pile of one-off extracts. At AST, we build the integration and analytics path as an operational system, not as a one-time migration project. If the warehouse has to live for years, it needs maintenance disciplines: contract testing, source versioning, release coordination, and owners who know what happens when a feed goes sideways. That is the unglamorous part. It is also the part that keeps the warehouse useful after the demo ends.
If you are building this now, do not start by picking the prettiest cloud service or the fastest dashboard tool. Start by deciding what the warehouse must never get wrong. Then design the pipeline so every record can be traced, corrected, and reprocessed without hand-waving. That is the only way real-time clinical analytics stays real instead of theatrical.
Build a warehouse your care teams can trust
If your clinical data platform is still fighting late messages, duplicate identities, or dashboards that drift from the chart, I can help you design the right architecture from the pipeline up. We build these systems with the same discipline we bring to integration, cloud, and analytics delivery.





Comments
Comments are warming up. Live, no-sign-in discussion will appear here shortly.
Have a question now? Email info@allstartech.net.