Kastoria Health System Overview
Kastoria Health is the control plane for medical imaging. It receives DICOM imaging studies, stores them in versioned, content-addressed storage, prepares them for search and retrieval, and makes them available to viewers, EHRs and downstream systems through open standards: DICOMweb, SMART on FHIR, and event notifications delivered to a queue.
This document describes, at a high level, how Kastoria Health is built and what it provides as of release v0.9.1. It covers:
- Kastoria Core, the data plane for Kastoria Health;
- the architecture of Kastoria Health and how a study moves through it; and
- the features Kastoria Health provides, together with how the system is operated, observed and secured.
It does not describe APIs, configuration or deployment steps. Those are covered by the REST API, integration and deployment documentation.
Audience
This document is intended for CIOs, CTOs, security reviewers, enterprise architects, and senior developers evaluating, deploying, or integrating Kastoria Health.
Kastoria Core
Kastoria Core is the data plane for Kastoria Health. Every study, every piece of derived data, and every configuration record Kastoria Health keeps is held in Kastoria Core.
Most systems store data by location: a file at a path, or a row in a table, whose contents can be changed in place. Kastoria Core stores data by content. Each piece of data is identified by a cryptographic fingerprint of its own bytes, and larger structures, such as an imaging study, are built by linking those pieces together into a Merkle tree. The fingerprint of the top of the tree identifies everything beneath it.
Why Merkle tree storage
- Immutability. Stored content is never edited in place. A change produces new content with a new identifier, and the earlier content remains. Every earlier version of a record is kept and can be read, so the system retains a complete history rather than only the latest state.
- Cryptographic integrity. Content is identified by its SHA-256 hash. Any change to the bytes, no matter how small, produces a different identifier, and each version of a record is linked to the version before it by that version's hash. The structure is therefore tamper-evident: content cannot be altered while keeping its identity, and history cannot be rewritten without breaking the chain of hashes.
- Deduplication. Identical content always has the same identifier, so it is stored once. When a study is re-sent, or a new version of a study shares most of its images with the previous version, the unchanged content is not stored again.
- Safe concurrency. Because content never changes, many readers and writers can work at once without locking. The only thing that changes is which version a name points to, and that change is made with a conditional write: two writers can never silently overwrite each other.
- Safe retries. Storing the same content twice produces the same result as storing it once. An interrupted operation can simply be repeated.
Primary components
Kastoria Core is made up of four primary components.
Block Storage
A block is a unit of stored content, addressed by its content identifier (CID): a self-describing identifier derived from the SHA-256 hash of the block's bytes. Blocks hold either structured data or raw binary data, such as the pixel data of a single image frame, and blocks refer to one another by CID to form larger structures.
Block Storage is where all content lives. Because a block's identifier is derived from its content, Block Storage is naturally deduplicated, and blocks can be spread evenly across multiple storage accounts to scale capacity and throughput. In an Azure deployment, blocks are held in Azure Storage.
Named Roots
Content-addressed data needs stable names. A patient's imaging study has to be found by its Study Instance UID, not by a hash that changes every time an image is added.
A Named Root is a durable, named pointer to the top of a Merkle tree. Its name never changes; the content it points to does. Each time a Named Root is updated, Kastoria Core records a new version that links to the version before it, so a Named Root carries its full, hash-linked version history.
Named Roots provide:
- Stable identity for records whose content evolves, such as a study that receives additional series over time.
- Complete version history, from which any earlier state of the record can be read.
- Conflict-free updates, since an update is applied only if the record is still in the state the writer last read. A writer that loses a race sees the winner's content and decides again, rather than overwriting it.
- Recoverable removal. Removing a Named Root is recorded as a new version, not a deletion, so a removal can be undone and its history is never lost.
Changelog
The Changelog is an append-only, time-ordered record of every change made to Named Roots: which record changed, when, and the content it moved from and to.
The Changelog is what lets the rest of the system react to change. Rather than being called directly when a study is stored, downstream components follow the Changelog and act on what they find there. This keeps each step of the system independent: a component that is paused, slow or restarted resumes from where it left off and misses nothing. The Changelog also provides a durable history of when each record changed, which supports troubleshooting and audit.
Ranged Indexes
Named Roots find a record by its name. Many questions are not asked by name: every study for a patient, or every study performed between two dates.
A Ranged Index is a named, persistent index that maps search keys to content, kept in key order. It answers exact lookups, range queries (for example, from one date to another) and prefix queries (for example, every name beginning with a given family name), and returns its results in order, a page at a time, without reading the records themselves.
Ranged Indexes are what make search fast at scale. Kastoria Health uses them to answer DICOMweb queries, such as searches by patient, accession number or study date, and to find the study an EHR asks to launch. Unlike Named Roots, a Ranged Index is not versioned: it always reflects the current content, and Kastoria Health keeps it up to date as studies are processed.
Kastoria Health Architecture
Kastoria Health is a set of containerized services deployed in the customer's own environment. All data remains in that environment.
flowchart LR
modalities["PACS / modalities /<br/>migration tools"]
ehr["EHR"]
operators["Operators"]
consumer["Downstream<br/>systems"]
subgraph kh["Kastoria Health (customer environment)"]
subgraph services["Services"]
gateway["Gateway"]
dicomweb["DICOMweb Service"]
viewer["Web Viewer"]
smart["SMART Launch Service"]
console["Admin Console"]
processor["Study Processor<br/>(scheduled)"]
events["Event Notification<br/>(scheduled)"]
end
core[("Kastoria Core<br/>Block Storage · Named Roots ·<br/>Changelog · Ranged Indexes")]
queue[["Notification<br/>Queue"]]
collector["Telemetry<br/>Collector"]
end
grafana["Grafana"]
modalities -- "STOW-RS" --> gateway
ehr -- "SMART launch" --> gateway
operators --> gateway
gateway --> dicomweb
gateway --> smart
gateway --> viewer
gateway --> console
viewer -- "QIDO-RS / WADO-RS" --> dicomweb
dicomweb --> core
smart --> core
console --> core
processor -- "follows Changelog" --> core
events -- "follows Changelog" --> core
events --> queue
queue --> consumer
services -. "traces & metrics" .-> collector
collector -.-> grafana
Components
| Component | Role |
|---|---|
| Gateway | The single entry point. Routes requests to each service and, in a mutual-TLS configuration, admits only callers presenting a trusted client certificate, granting each certificate access to specific services. |
| DICOMweb Service | Receives studies (STOW-RS), answers queries (QIDO-RS) and serves images and metadata (WADO-RS). |
| Study Processor | A scheduled job that prepares newly stored studies for query and retrieval. |
| Event Notification | A scheduled job that publishes a Notification Event to a queue each time a study is processed. |
| SMART Launch Service | Brokers SMART on FHIR launches from an EHR into the viewer. |
| Web Viewer | A zero-footprint, browser-based diagnostic viewer built on the open-source OHIF Viewer. |
| Admin Console | A web application for operating the system. |
| Kastoria Core | Versioned, content-addressed storage for all data. |
| Telemetry Collector | Receives traces and metrics from every service and forwards them to Grafana. |
The services and scheduled jobs are delivered as signed container images; see Docker Images.
The services that answer requests (DICOMweb, SMART Launch, the viewer and the gateway) keep no state of their own between requests; all shared state is in Kastoria Core. They scale horizontally by adding replicas.
The Study Processor and Event Notification run as scheduled jobs. Each run picks up where the last one finished, and only one run of each happens at a time, so overlapping schedules are harmless.
How a study moves through the system
A study passes through the ingestion pipeline, a sequence of steps that each scale independently:
- Store. A client sends DICOM instances to the DICOMweb Service. The instances are committed to Kastoria Core as a new version of the study, and the change is recorded in the Changelog.
- Process. The Study Processor finds the change in the Changelog and prepares the study for use: it builds the study's search indexes and summary, records the study's patient as a FHIR Patient resource, and assigns the study its FHIR ImagingStudy id.
- Notify. Event Notification finds the processed study and publishes a Notification Event, carrying the study's Ingestion Manifest, to the customer's queue.
A study is ingested once it has been processed. From then on it can be found with QIDO-RS, retrieved with WADO-RS, opened in the viewer, and launched from an EHR.
Features
DICOMweb
Kastoria Health implements the DICOMweb standard (DICOM PS3.18) over HTTPS. The full detail of what is supported is given in the DICOM Conformance Statement and the DICOMweb API documentation.
Study ingestion (STOW-RS)
- High-throughput ingestion of one or many DICOM instances per request, across one or many studies.
- Incremental studies. Instances and series sent for an existing study are added to it. Each store creates a new version of the study, and every earlier version is kept.
- Deduplicated storage. Each image frame is stored as its own content-addressed block, so content that has already been stored, such as a re-sent instance, is not stored again.
- Lossless compression on ingest. Images from modalities such as CT, MR, CR, DX and MG are stored in High-Throughput JPEG 2000 (HTJ2K) lossless format, which reduces storage and lets viewers display images progressively. Compression can be turned off for an individual request.
- Per-instance results. The response reports which instances were stored and which failed, and why. Rejected instances are also recorded in an ingestion error log, which operators can search in the Admin Console by study, request correlation id, or time.
- Concurrent senders. Many clients can send instances for the same study at the same time without overwriting one another's work.
See STOW-RS.
Query (QIDO-RS)
- Study search by patient name, Patient ID, accession number, study date and time (including date ranges), study description, modalities in the study, and Study Instance UID.
- Series search within a study, by modality, Series Instance UID and series number.
- Case-insensitive matching for clinical text such as patient names and accession numbers.
- Paged results, newest study first.
See QIDO-RS.
Image retrieval (WADO-RS)
- Retrieval of whole studies, series or instances.
- Metadata at study, series and instance level, in DICOM JSON.
- Individual frames, for efficient streaming to viewers.
- Rendered images and thumbnails in consumer formats such as JPEG and PNG, with windowing and sizing.
- On-the-fly transcoding. A client can request instances or frames in the transfer syntax it supports, including uncompressed, RLE, JPEG, JPEG-LS, JPEG 2000, HTJ2K and JPEG XL.
See WADO-RS.
Study processing
Stored studies are prepared for use by the Study Processor, which runs on a schedule and follows the Changelog to find every study that has changed.
- Search indexes and study summaries are built for each study, so queries are answered from indexes rather than by scanning studies.
- FHIR identity. Each study with a Patient ID has a FHIR R4 Patient resource recorded for its patient, and each study is assigned a FHIR ImagingStudy id. The ImagingStudy id is assigned when the study is first processed and does not change; it is the id an EHR uses to launch the study.
- Always the latest version. When a study changes several times before it is processed, it is processed once, at its latest version.
- Idempotent. Processing a study that has not changed produces no new data and no new events.
- Automatic recovery from transient failures. A study that fails because of a temporary condition, such as a storage timeout, is retried automatically.
- Permanent failures are surfaced, not retried endlessly. A study whose content cannot be processed is recorded as a processing failure, visible in the Admin Console. It is processed again when a new version of the study arrives or when an operator requests it.
- Reprocessing on request. Operators can ask for a single study, or every study ingested within a time window, to be processed again on the next run.
- Run history. Every run that does work is recorded, with its outcome and the number of studies it processed, failed or skipped.
Event notification and the Ingestion Manifest
Kastoria Health tells downstream systems about the studies it has processed. Each time a study is processed, Event Notification publishes a Notification Event to a queue in the customer's own environment.
- The Ingestion Manifest. Each event carries the study's Ingestion Manifest: the DICOM identifiers and patient and study attributes a third party needs to map the study between systems, together with the FHIR Patient and ImagingStudy ids Kastoria Health holds for it.
- Complete, self-contained events. Every event carries the whole manifest as it stands after the change, so a consumer never needs to call Kastoria Health back.
- Created and updated events. A study's first processing publishes a
createdevent; each later change publishes anupdatedevent. - At-least-once delivery. Events are published durably and never expire on the queue. A consumer applies a simple ordering rule that makes duplicates and out-of-order delivery harmless.
- Secretless access. Kastoria Health publishes, and consumers read, using Microsoft Entra ID identities. No shared keys or SAS tokens are issued.
- Re-publishing. If a consumer loses events, an operator can re-publish every event from a chosen point in time from the Admin Console.
The event format, delivery guarantees and consumer guidance are specified in Ingestion Manifest.
SMART on FHIR launch
Kastoria Health lets clinicians open imaging studies directly from their EHR using the SMART App Launch standard (EHR launch, v2.2).
- In-context launch. The EHR launches Kastoria Health with the patient and the imaging study in context. Kastoria Health authorizes with the EHR, confirms that the study belongs to the patient in context, and opens the study in the viewer.
- Registered EHRs only. Each EHR is registered in the Admin Console with its own client configuration. Launches from an unregistered or deactivated EHR are refused, and an EHR can be deactivated at any time without redeploying.
- Standards-based security. Every authorization uses PKCE, and the EHR's identity token is verified against the EHR's signing keys whenever the EHR publishes them.
- Stateless and scalable. The state of an in-progress launch is carried in a short-lived encrypted, tamper-proof token, keyed per EHR. Nothing is held on the server between steps, so the service scales horizontally.
- Study-scoped viewer access. A launch can issue the viewer a short-lived signed token that authorizes it to query and retrieve only the launched study. Query and retrieval requests without a valid token, or for a study outside the token's scope, are refused. See JWT Authorization.
- Clear error pages. A launch that cannot complete ends on an error page that names what went wrong, without exposing patient data.
See SMART on FHIR — smart-launch-api.
Admin Console
The Admin Console is a web application for operating Kastoria Health.
- SMART launch configuration. Register, edit, activate, deactivate and remove the EHRs permitted to launch Kastoria Health.
- Content Explorer. Search stored studies, view a study's series and its FHIR Patient and ImagingStudy ids, open a study in the viewer, and request that a study be reprocessed.
- Study Ingestion. Search the ingestion error log for instances that could not be stored, and inspect each error.
- Study Processing. Monitor the Study Processor at a glance: whether it is running, the studies waiting to be processed, processing failures and reprocess requests. Review run history, reprocess or dismiss failures, and request reprocessing for a single study or a time window.
- Events. Monitor event publishing and pending events, and re-publish events from a point in time.
- Storage Explorer. Browse the Kastoria Core Changelog by record type, time window and kind of change, and inspect the stored records it refers to.
Access to the Admin Console is controlled at the network layer, through the gateway.
Telemetry and monitoring
Every Kastoria Health service is instrumented with OpenTelemetry, the open industry standard for observability.
- Traces follow each request through the services and into storage, showing where time is spent.
- Metrics record request rates, errors and latency for every route, together with service-specific measures such as instances stored, search results returned, processing runs and failures, and events published.
- Logs are written by every service and collected by Azure Monitor (Log Analytics).
- No PHI in DICOMweb or storage telemetry. Patient identifiers and study UIDs are not used as metric labels or trace attributes, and URL paths and query strings are scrubbed of patient data before they are recorded.
Telemetry is sent to an OpenTelemetry collector deployed inside the customer's environment. The collector can forward traces, metrics and logs to Grafana Cloud, or telemetry can be kept entirely within the environment. When Grafana is used, Azure Monitor is also connected as a Grafana data source. See the Azure Telemetry module.
Kastoria Health provides ready-made Grafana dashboards:
| Dashboard | Shows |
|---|---|
| Services Overview | Request rate, error rate and latency across all services, with the slowest routes. |
| DICOMweb Performance | Request rate, errors and latency broken down by STOW-RS, QIDO-RS and WADO-RS, with latency distributions, abandoned searches, and links to traces. |
| Admin Console Performance | Request rate, errors and latency for the Admin Console by area. |
| Performance Test | Results of Kastoria Health's load and regression test suite: latency, errors, throughput and concurrent users by scenario. |
Security
- Customer-controlled deployment. Kastoria Health is deployed in the customer's own environment. All imaging data, derived data and telemetry stays in the customer's environment.
- Network isolation. Services run on a private virtual network.
- Mutual TLS at the gateway. The mutual-TLS gateway admits only callers presenting a trusted client certificate, and grants each certificate access only to the services assigned to it. Callers with no assigned access are denied. See kastoria-proxy-mtls.
- Encryption. Data is encrypted at rest by Azure Storage, and all traffic to storage uses TLS 1.2 or later.
- No stored credentials. Services reach storage, queues and the key vault with Azure managed identities, so the storage configuration carries no secret material (see Configuration). Secrets referenced from Azure Key Vault never appear in deployment state.
- Tamper-evident storage. All data is content-addressed and versioned in Kastoria Core, as described above.
- Signed, scanned software. Every container image and deployment module Merkalis delivers is cryptographically signed and can be verified before deployment. Images run as non-privileged users, are scanned for vulnerabilities during the build, cannot be released with a known critical or high severity vulnerability, and carry SLSA provenance and a software bill of materials (SBOM). See Distribution Registry for verifying signatures.