About the job
Company: Pinnacle Systems, Inc. Location: 100% Remote — Pakistan or India Working Hours: US Eastern Time zone coverage required Type: Full-time Reports to: Director, AI Experience: 6+ years
About the Role
Pinnacle Systems builds and operates modern data platforms for clients in healthcare and adjacent regulated industries. We are hiring a Senior Data Engineer to design, build, and own production-grade data pipelines on Databricks and Microsoft Azure.
This is a hands-on engineering role, not a coordination role. You will be responsible for the correctness, performance, and cost profile of the pipelines you ship. You will work directly with client stakeholders, data scientists, and application teams to turn raw operational data into reliable, governed, analytics-ready assets.
What You'll Do
Design and build scalable batch and streaming pipelines on Databricks using PySpark, Spark SQL, and Delta Lake.
Implement and maintain a medallion (bronze/silver/gold) lakehouse architecture with clear contracts between layers.
Build ingestion from heterogeneous sources — relational databases (SQL Server, Postgres), REST APIs, event streams, flat files, and SaaS systems — using Azure Data Factory, Auto Loader, and Event Hubs.
Model data for analytics and downstream consumption (dimensional models, slowly changing dimensions, semantic layers for Power BI).
Own data quality: validation rules, reconciliation checks, anomaly detection, and alerting on pipeline failures and silent data drift.
Implement governance and access control through Unity Catalog — lineage, data classification, row/column-level security, and audit trails.
Optimize Spark workloads for performance and cost: partitioning strategy, Z-ordering and liquid clustering, file compaction, cluster sizing, Photon, and job vs. all-purpose compute decisions.
Build and maintain CI/CD for data assets using Azure DevOps or GitHub Actions, with infrastructure defined in Terraform or Bicep and deployments managed through Databricks Asset Bundles.
Write documentation and runbooks that let another engineer support your pipelines without a handoff meeting.
Mentor mid-level engineers through code review, design review, and pairing.
Required Qualifications
6+ years in data engineering, with at least 3 years building production workloads on Databricks.
Deep, practical Spark expertise — you can read a Spark UI, diagnose a skewed join, and explain why a job is spilling to disk.
Strong Python and advanced SQL. You write tested, modular code, not notebooks that only run in one order.
Production experience across the Azure data stack: Azure Data Lake Storage Gen2, Azure Data Factory or Synapse Pipelines, Azure SQL, Key Vault, and Azure DevOps.
Hands-on experience with Delta Lake — ACID semantics, time travel, MERGE patterns, schema evolution, and change data feed.
Demonstrated ownership of orchestration and scheduling (Databricks Workflows, Airflow, or ADF triggers) including retry, backfill, and idempotency design.
Working knowledge of Git-based development, code review, and automated testing for data pipelines (pytest, Great Expectations, or equivalent).
Ability to communicate technical tradeoffs clearly to non-technical stakeholders and to push back on requirements that will not hold up in production.
Preferred Qualifications
Databricks Certified Data Engineer Professional, or Azure Data Engineer Associate (DP-203).
Experience with Unity Catalog migration and multi-workspace governance.
Streaming experience with Structured Streaming, Kafka, or Azure Event Hubs.
Experience in healthcare data environments — HL7, FHIR, claims, EHR extracts — and familiarity with HIPAA and PHI handling requirements.
Exposure to dbt, Power BI semantic models, or MLflow and feature-store patterns.
Experience working in a consulting or client-facing delivery model across multiple concurrent engagements.
Tech Stack
Databricks (Spark, Delta Lake, Unity Catalog, Workflows) · Azure (ADLS Gen2, ADF, Synapse, Event Hubs, Key Vault, Azure SQL) · Python · SQL · Terraform / Bicep · Azure DevOps · Power BI
What Success Looks Like
First 30 days — You understand the client's data landscape end to end, have shipped your first pipeline change to production, and can name the three biggest reliability risks in the platform.
First 90 days — You own a defined domain of the lakehouse, have measurably improved either cost or reliability in that domain, and are the person other engineers ask before making a schema change.
First year — You have shaped the platform's architecture direction, raised the engineering standard around you through mentorship and review, and made yourself the reason a client renews.
Compensation & Benefits
Pakistan: PKR 700,000 per month India: INR 242,000 per month
Both figures represent the same offer, converted at prevailing rates. This is a fully remote, full-time position; you may work from anywhere in Pakistan or India, but your working day must cover US Eastern Time business hours for collaboration with client teams and the AI leadership group.
Benefits include paid time off, sick leave, and public holidays per Pinnacle's HR policy, plus certification reimbursement for Databricks and Azure credentials.
How to Apply
Send your resume to hr@pinnaclesyst.com, along with a short note describing a data pipeline you built and what you would do differently if you built it today.
How to Interview
You can complete your first-round interview immediately — no scheduling required. Candidates who complete the AI self-evaluation are reviewed faster and stand a significantly stronger chance of moving forward.
Start here: https://cogniter.ai/apply/008eefac-b46d-4731-9b8a-162b1d706210
The evaluation takes approximately 60 minutes and can be completed at any time that suits you.
Pinnacle Systems, Inc. is an equal opportunity employer. We evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic.
Monthly based
Karachi Division,Sindh,Pakistan
Karachi Division,Sindh,Pakistan