Data Engineer – Databricks
Company not disclosed
- Pay
- ₹20–25 lakh a year
- Experience
- 6–10 years
- Location
- Bengaluru, Pune, Hyderabad · On-site
- Job type
- Full-time
- Databricks
- PySpark
- Apache Airflow
- Python
- SQL
- data warehousing
- cloud-based data platforms
- performance tuning
About the role
Build and maintain scalable data pipelines and data processing solutions. Develop and optimize data pipelines using Databricks, Python, and SQL. Design and manage workflow orchestration using Apache Airflow. Build scalable ETL/ELT processes for data ingestion and transformation. Monitor, troubleshoot, and improve data pipeline performance. Implement data quality checks and validation frameworks. Collaborate with cross-functional teams to deliver data-driven solutions.
What you’ll do
- Develop and optimize data pipelines using Databricks, Python, and SQL.
- Design and manage workflow orchestration using Apache Airflow.
- Build scalable ETL/ELT processes for data ingestion and transformation.
- Monitor, troubleshoot, and improve data pipeline performance.
- Implement data quality checks and validation frameworks.
- Collaborate with cross-functional teams to deliver data-driven solutions.
What we’re looking for
Must have
- Databricks
- PySpark
- Apache Airflow
- Python
- SQL
- data warehousing
Nice to have
- AWS
- Azure
- GCP
- Delta Lake
- CI/CD
- Agile methodologies
What makes this role challenging
- Optimizing data pipeline performance at scale on Databricks
- Designing robust workflow orchestration with Apache Airflow
- Implementing data quality checks and validation frameworks
- Balancing multiple cross-functional data delivery demands
What success looks like
- First 30 days
- Ramped on the existing data pipelines and Databricks environment, contributing to small optimizations and fixes.
- By 90 days
- Owning a set of data pipelines end-to-end, implementing data quality checks, and collaborating with cross-functional teams on data solutions.
- First year
- Leading the design and optimization of scalable ETL/ELT processes, driving performance improvements, and setting best practices for data engineering on Databricks.