Department: Information Technology | Campus: The Indus Hospital, Korangi Campus, Karachi
Eligibility / Qualification Required:
-
Qualification:
Bachelorʼs degree in Computer Science
-
Experience:
1–4 years in data engineering, backend data pipelines, or closely related software/data roles
Job Description:
- Job summary.
- We are building an on-premises analytical data platform to bring together clinical and operational data from multiple hospital sites into a governed lakehouse (raw, curated, and serving layers). This role focuses on pipelines, data reliability, and platform operations — turning captured data into trusted, linked, analysis-ready datasets.
- You will work as part of a data team under established architecture and quality standards. A platform steering committee and senior consultant provide design direction, phase planning, and review; you will implement, document, and operate the ingestion and transformation layer with growing independence over time.
- Key responsibilities
- Data ingestion & platform pipelines
- Design, build, and maintain batch and change-capture ingestion from operational databases across multiple sites. Implement orchestrated workflows for scheduled extraction, landing, transformation, and promotion between lakehouse layers.
- Ensure pipelines are idempotent, observable, and recoverable (retries, checkpoints, replay where appropriate).
- Work with database administrators on read-only access, change-log readiness, and safe extract windows.
- Lakehouse layers (raw → curated)
- Manage immutable raw landing and curated (silver) datasets following medallion-style layering.
- Implement validation, cleansing, typing, and conformed models using transformation-as code practices.
- Support slowly changing history for key demographic and reference entities where attributes change over time.
- Apply configuration-driven pipeline definitions (declarative specs reviewed in version control) rather than one-off scripts per table.
- Cross-site identity & data linking Implement logic to unify records across sites (e.g. patients, providers, facilities) using defined matching rules and steward review for ambiguous cases.
- Maintain bridge and reference structures that map source identifiers to enterprise identifiers with full audit trail.
- Data quality, reconciliation & operations
- Run and automate reconciliation between source systems and platform copies (counts, keys, samples, freshness).
- Respond to pipeline failures and data drift using runbooks and escalation paths.
- Contribute to schema change handling (additive vs breaking changes) with documentation and alerts.
- Monitor service levels (lag, success rate, data freshness) via operational dashboards.
- Documentation & collaboration
- Maintain technical runbooks, pipeline documentation, and change records on the same day as changes.
- Partner with the Analytics Engineer on catalog entries, data dictionary fields, and lineage metadata.
- Support clinical and operational data stewards with technical fixes; business decisions on duplicates and definitions stay with stewards.
How to Apply:
- Apply online by clicking the Apply button below