We are looking for a Python / PySpark Data Engineer to design, build, and optimize large-scale data pipelines and analytics solutions on cloud infrastructure. This role is pivotal in enabling data-driven decisions across our portfolio of SaaS businesses. The ideal candidate will partner closely with product, analytics, and engineering teams to deliver reliable, high-performance data platforms.
Responsibilities
- Design and develop scalable data pipelines using Python and PySpark on AWS (EMR, Glue, or Databricks).
- Build and maintain ETL/ELT workflows to process large volumes of structured and semi-structured data across multiple data sources.
- Collaborate with product managers and data analysts to translate business requirements into robust data models and analytical pipelines.
- Design efficient data schemas and manage data persistence using SQL/NoSQL stores including Redshift, Athena, S3, and DynamoDB.
- Implement data quality frameworks, monitoring, and alerting to ensure pipeline reliability and data integrity.
- Lead CI/CD integration for data pipelines using Bitbucket Pipelines or equivalent DevOps tooling for continuous delivery and deployment.
- Promote engineering best practices including pair programming, code reviews, unit testing, and documentation.
- Operate within Scrum teams, contributing to agile ceremonies and driving data engineering domain best practices.
- Champion an AI-first development approach by integrating AI tools and LLM-powered capabilities into data engineering workflows and pipeline automation.
- Leverage Claude Code and Anthropic’s Claude API to accelerate development velocity, automate repetitive data tasks, perform AI-assisted code reviews, and generate high-quality transformation logic and test code.
- Mentor the team on responsible and effective use of AI pair-programming tools, prompt engineering best practices, and AI-augmented data pipeline development.
Desired Qualifications
- B.E./B.Tech in Computer Science, Data Engineering, or a related field; Master’s degree preferred.
- Minimum 4–6 years of experience in data engineering with strong hands-on expertise in Python and PySpark.
- Proven experience building and optimizing large-scale data pipelines on AWS (EMR, Glue, Databricks, or equivalent).
- Strong proficiency in SQL and experience with columnar/analytical databases such as Redshift, Athena, or Snowflake.
- Experience with streaming data architectures using Apache Kafka, Spark Streaming, or AWS Kinesis.
- Familiarity with data orchestration tools such as Apache Airflow.
- Working knowledge of data lake and lakehouse architectures (Delta Lake, Apache Iceberg, or AWS Lake Formation).
- Strong understanding of DevOps principles and experience with CI/CD tools, specifically Bitbucket Pipelines.
- Demonstrated hands-on experience with an AI-first development approach, including using AI coding assistants, LLM APIs, or generative AI tools to accelerate engineering delivery and product innovation.
- Proficiency with Claude Code (Anthropic’s agentic CLI coding tool) or equivalent AI-powered development environments (GitHub Copilot, Cursor, etc.) for daily development tasks, code generation, refactoring, and debugging.
- Familiarity with prompt engineering, RAG (Retrieval-Augmented Generation), and integrating LLM APIs (Anthropic Claude, OpenAI, etc.) into data pipelines or analytical workflows is a strong plus.
CTC: Decent Hike on Current CTC
Job Category: Engineering
Job Type: Full Time
Job Location: Bangalore