PySpark Data Engineer

We are looking for a Python / PySpark Data Engineer to design, build, and optimize large-scale data pipelines and analytics solutions on cloud infrastructure. This role is pivotal in enabling data-driven decisions across our portfolio of SaaS businesses. The ideal candidate will partner closely with product, analytics, and engineering teams to deliver reliable, high-performance data platforms.

Responsibilities
  • Design and develop scalable data pipelines using Python and PySpark on AWS (EMR, Glue, or Databricks).
  • Build and maintain ETL/ELT workflows to process large volumes of structured and semi-structured data across multiple data sources.
  • Collaborate with product managers and data analysts to translate business requirements into robust data models and analytical pipelines.
  • Design efficient data schemas and manage data persistence using SQL/NoSQL stores including Redshift, Athena, S3, and DynamoDB.
  • Implement data quality frameworks, monitoring, and alerting to ensure pipeline reliability and data integrity.
  • Lead CI/CD integration for data pipelines using Bitbucket Pipelines or equivalent DevOps tooling for continuous delivery and deployment.
  • Promote engineering best practices including pair programming, code reviews, unit testing, and documentation.
  • Operate within Scrum teams, contributing to agile ceremonies and driving data engineering domain best practices.
  • Champion an AI-first development approach by integrating AI tools and LLM-powered capabilities into data engineering workflows and pipeline automation.
  • Leverage Claude Code and Anthropic’s Claude API to accelerate development velocity, automate repetitive data tasks, perform AI-assisted code reviews, and generate high-quality transformation logic and test code.
  • Mentor the team on responsible and effective use of AI pair-programming tools, prompt engineering best practices, and AI-augmented data pipeline development.
Desired Qualifications
  • B.E./B.Tech in Computer Science, Data Engineering, or a related field; Master’s degree preferred.
  • Minimum 4–6 years of experience in data engineering with strong hands-on expertise in Python and PySpark.
  • Proven experience building and optimizing large-scale data pipelines on AWS (EMR, Glue, Databricks, or equivalent).
  • Strong proficiency in SQL and experience with columnar/analytical databases such as Redshift, Athena, or Snowflake.
  • Experience with streaming data architectures using Apache Kafka, Spark Streaming, or AWS Kinesis.
  • Familiarity with data orchestration tools such as Apache Airflow.
  • Working knowledge of data lake and lakehouse architectures (Delta Lake, Apache Iceberg, or AWS Lake Formation).
  • Strong understanding of DevOps principles and experience with CI/CD tools, specifically Bitbucket Pipelines.
  • Demonstrated hands-on experience with an AI-first development approach, including using AI coding assistants, LLM APIs, or generative AI tools to accelerate engineering delivery and product innovation.
  • Proficiency with Claude Code (Anthropic’s agentic CLI coding tool) or equivalent AI-powered development environments (GitHub Copilot, Cursor, etc.) for daily development tasks, code generation, refactoring, and debugging.
  • Familiarity with prompt engineering, RAG (Retrieval-Augmented Generation), and integrating LLM APIs (Anthropic Claude, OpenAI, etc.) into data pipelines or analytical workflows is a strong plus.

CTC: Decent Hike on Current CTC

Job Category: Engineering
Job Type: Full Time
Job Location: Bangalore

Apply for this position

Allowed Type(s): .pdf, .doc, .docx