Design, develop, and maintain complex ETL pipelines for structured and unstructured data
Build and optimize scalable data processing workflows using Python and Big Data technologies
Work extensively with Python data libraries (Pandas, NumPy, Polars, PySpark, etc.)
Develop and optimize SQL queries, indexes, and database performance for large datasets
Integrate data from multiple sources including APIs, databases, files, and streaming systems
Ensure data quality, integrity, and reliability across all pipelines
Collaborate with data analysts, data scientists, and backend teams to support data needs
Monitor, troubleshoot, and optimize data workflows in Linux-based environments
Implement logging, monitoring, and error-handling mechanisms
Maintain documentation for data pipelines, processes, and system architecture
Hands-on experience with Python data libraries (Pandas, NumPy, PySpark, Polars, etc.)
Solid experience in designing and managing complex ETL workflows
Strong understanding of Big Data concepts and tools (Spark, Hadoop, Kafka, etc.)
Advanced knowledge of SQL, including query optimization and performance tuning
Experience working in Linux/Unix environments
Familiarity with data engineering tools and platforms (Airflow, dbt, NiFi, etc.)
Apply: career@curvedigitalsolutions.com.
Monthly based
Karachi Division,Pakistan,Pakistan
Karachi Division,Pakistan,Pakistan