I do not trust healthcare data lake projects that start with a platform diagram and end with a folder full of CSVs. I have seen that movie too many times. Someone buys storage, stands up a warehouse or lakehouse, points a few connectors at Epic, Oracle Health, athenahealth, a claims feed, and a device vendor, then assumes the work is basically done. The first month looks productive. The third month is where the pain shows up: duplicate patients, mismatched visit counts, broken timestamps, and no one able to explain where a number came from.
Healthcare data is not hard because it is big. It is hard because it arrives with different meanings, different clocks, different identifiers, and different degrees of trust. An HL7v2 ADT feed is not a lab result feed. X12 claims are not the same thing as clinical encounters. FHIR R4 resources are excellent for structure, but they do not magically solve source drift or identity problems. If you build the lake like every record is already harmonized, you will spend the next year cleaning up the consequences.
When I say multi-source, I mean the real mess: EMR extracts, FHIR APIs, HL7v2 feeds, claims and remits, eligibility responses, scheduling data, device streams, and external reference data like payers, providers, and locations. The architecture only works when you accept that each source has a different contract. Some sources are event streams. Some are snapshots. Some are noisy by design. Your job is not to flatten them immediately. Your job is to keep them usable without pretending they are the same thing.
At AST, the fastest way we have found to kill a data lake project is to treat ingestion and transformation as one step. They are not. In one rollout, we lost time because a source looked stable in dev and changed identifiers in production in a way that broke our join logic. Nothing was wrong with the source system. Our assumption was wrong. That is the undocumented failure mode in healthcare analytics: the pipeline works until the first real edge case, and then you discover you modeled the happy path, not the business.
That replay requirement changes the whole design. I want three layers at minimum:
- Landing/raw: immutable copies of source payloads, with source system, ingest time, batch or event identifier, and checksum metadata.
- Standardized/normalized: data that is mapped into a common model, but still traceable back to the source record and field.
- Curated/serving: analytics-ready datasets built for reporting, quality measures, operational dashboards, machine learning, and research use cases with stricter business rules.
The raw layer is where teams get disciplined or get hurt. If you overwrite files, compress away source metadata, or strip out time zone and message control fields because they seem unimportant, you will regret it. I have seen teams regret it when they try to explain why a claim appeared before the encounter, or why a lab landed one hour earlier than the order, or why a patient appears twice because two source systems disagree on name formatting. Once those details are gone, you are arguing from memory instead of evidence.
For identity, I do not believe in pretending one universal identifier exists across healthcare. It does not. You need a master identity strategy that can handle MRN, account number, enterprise patient ID, payer member ID, NPI, location IDs, and device IDs without forcing all of them into one brittle key. That usually means a survivorship model, deterministic matching where possible, probabilistic matching where necessary, and a clear policy on what is allowed to merge automatically versus what requires review. If the matching rules are vague, your analytics layer will absorb the ambiguity and spread it everywhere.
Here is the part people underestimate: source governance matters more than storage technology. I have seen well-funded teams argue about object storage versus warehouse compute while ignoring the actual failure points — schema drift, late data, missing reference tables, and feed ownership. Those are the things that break healthcare pipelines in production. The cloud bill is not the hard part. The semantics are.
This is why I like a source registry as a first-class asset. Every feed should have an owner, a contract, a refresh cadence, expected record counts or volume patterns, known fields, downstream consumers, and a rollback plan. If you cannot answer who owns a lab feed when it stops arriving, then your architecture is a hope, not a system. The same goes for versioning. A FHIR R4 endpoint that changes coding behavior without notice can be just as disruptive as a broken flat file export.
We have learned at AST that it helps to design the lake for both technical and clinical consumers. Engineers want schemas, partitions, and lineage. Analysts want clean measures and repeatable definitions. Clinical operations teams want numbers they can trust without calling three people to decode them. Those audiences do not care about the same thing, but the lake has to serve all of them if it is going to survive. That is why our delivery pods keep the ingestion contract, data model, and consumption layer in the same conversation from day one. It is slower at the start and dramatically faster later.
For multi-source healthcare data, I usually separate the design choices into three questions:
- What is the source type? Event stream, API, batch extract, message feed, or external reference data.
- What is the trust level? Canonical clinical record, operational helper data, or reference/lookup data.
- What is the consumption pattern? Near-real-time operations, nightly reporting, longitudinal analytics, or model training.
Those three answers determine the rest. If a feed is high-volume clinical event data with operational consumers, you treat latency and reconciliation differently than a monthly claims file for finance. If a feed is only used for analytics, you can tolerate one more transformation layer. If it drives care operations, you need stronger freshness checks and explicit exception handling.
| Layer | What it stores | Main rule | Common failure if ignored |
|---|---|---|---|
| Raw landing | Original source payloads and metadata | Never overwrite, never silently normalize | Cannot replay history or audit changes |
| Normalized | Mapped records in a common structure | Keep source lineage from field to field | Traceability breaks during corrections |
| Curated | Business-ready datasets and metrics | Only publish agreed definitions | Conflicting dashboards and measure drift |
That table is the architecture in miniature. Every time someone tries to collapse those layers into one because it feels simpler, I push back. Simpler for whom? Usually not for the people who have to defend the data later. I would rather maintain a little extra structure than spend months explaining why one feed’s transformation silently rewrote a clinical history.
If you are building this on cloud infrastructure, the operational details matter just as much as the model. Partitioning strategies should follow access patterns, not just source names. Encryption, role-based access, and audit logging should be built into the pipeline, not hung off the side. And if you have PHI in the lake, then access control has to be granular enough to support real least-privilege boundaries. Good intentions do not satisfy auditors. Logs do.
For teams that are moving from point-to-point reporting into a real architecture, I recommend a sequencing approach that avoids the usual trap of overbuilding the platform before anyone can use it.
- Inventory the sources List every inbound system, format, owner, refresh cadence, and downstream use. If you only map the obvious feeds, the hidden ones will surprise you later.
- Define record-level lineage Decide what metadata is mandatory for every ingest: source, timestamp, batch ID, file hash, and correlation keys. Without that, you cannot troubleshoot or audit.
- Keep raw immutable Store source payloads exactly as received. If the source sends bad data, preserve the bad data and annotate it. Do not destroy evidence.
- Build the canonical model selectively Normalize only the entities you actually need first — patient, encounter, order, result, claim, payer, provider, location. Do not boil the ocean.
- Version the business logic Treat every metric definition, merge rule, and mapping rule like code. If it changes, it needs review and history.
- Publish curated datasets Expose only the datasets that are stable enough for broad use. Everything else stays internal until the logic settles.
- Monitor freshness and drift Track delays, row counts, missing fields, and distribution shifts. The first symptom of a broken feed is usually a subtle shape change, not a total outage.
That playbook is boring on purpose. Boring architecture survives contact with healthcare reality. Flashy architecture does not. I would rather build a conservative lake that is replayable, governable, and inspectable than a clever one that only works when the source system behaves perfectly.
If the use case includes clinical AI, quality reporting, or claims analytics, the lake becomes even more valuable when it preserves the exact source evidence needed to explain a result. That is true whether you are feeding a BI dashboard or a co-pilot like Medexa that needs clinically grounded, payer-aware data behind the scenes. Models do not fix bad data discipline. They make bad discipline look expensive faster.
I have no patience for healthcare data lakes that only look clean in a demo. Real systems have corrections, duplicates, late feeds, and source quirks. If your architecture does not make those visible, it is hiding the truth instead of organizing it. That is the difference between analytics that inform care and analytics that just decorate a dashboard.
At AST, we build these platforms as integrated engineering work, not as a pile of disconnected deliverables. That means the cloud layer, ingestion pipelines, data model, and governance rules get designed together. That is the only way I know to keep a multi-source healthcare lake from collapsing under its own cleverness.
Build the lake around provenance, not optimism
If you want a healthcare data lake that can handle EHR extracts, HL7v2 feeds, FHIR APIs, claims, and device data without losing traceability, we can help you design the layers the right way. AST builds clinical data platforms that hold up in production, not just in architecture decks.





Comments
Comments are warming up. Live, no-sign-in discussion will appear here shortly.
Have a question now? Email info@allstartech.net.