Senior Data Engineer / SSE
listed 42 days ago
Nirdisha removes a role 90 days after it was listed.
- Sector
- HR tech
- City
- Bengaluru
- Area
- Koramangala and Outer Ring Road
- Experience
- 3 to 5 years
- Role family
- Data and ML
- Salary
- Not disclosed
- Posted
- 27 Jul 2026 · 1 month ago
- Last checked
- 7 Sept 2026
About the role
Company: Apna
Role: Senior Data Engineer / SSE
Team: Data Platform / Engineering
Location: Work from Office - Domlur, Bangalore (5 days / week)
Experience : 4-6 Years of Experience
About the Role
Apna is looking for a Senior Software Engineer to build and scale our core data platform. This role will work on large-scale data pipelines, lakehouse architecture, query platforms, workflow orchestration, and data reliability systems that power analytics, product intelligence, machine learning, business dashboards, experimentation, and operational decision-making across Apna.
We are looking for someone who can think deeply about data architecture, design reliable pipelines, improve data quality, and help build a platform that can scale with Apna’s growth.
Requirements
What You’ll Own
You will be responsible for designing, building, and operating critical parts of Apna’s data platform, including:
- Building scalable batch and near-real-time data pipelines across product, business, growth, and ML use cases.
- Designing and improving our lakehouse architecture using technologies likeApache Hudi.
- Working with query engines such asPresto / Trinofor large-scale analytical workloads.
- Building and maintaining orchestration workflows usingApache Airflow.
- Creating reusable data models, curated datasets, and reliable data marts for analytics and product teams.
- Improving data platform reliability, observability, SLA tracking, lineage, and data quality checks.
- Optimizing storage, compute, query performance, and pipeline costs.
- Partnering with product, analytics, ML, and backend engineering teams to understand data needs and convert them into scalable platform solutions.
- Driving engineering standards around data modeling, schema evolution, partitioning, deduplication, backfills, replayability, and pipeline ownership.
- Mentoring data engineers and influencing architecture decisions across teams.
What We’re Looking For
Must Have
- Strong experience indata engineering, preferably at scale.
- Hands-on experience withApache Airflowor similar orchestration systems.
- Strong knowledge ofPresto / Trinoor other distributed query engines.
- Good understanding ofApache Hudiconcepts such as:
- Copy-on-write vs merge-on-read
- Upserts and deletes
- Incremental reads
- Compaction
- Clustering
- Timeline and commits
- Schema evolution
- Partitioning strategy
- Data warehouse
- Data lake
- Lakehouse
- Lambda architecture
- Kappa architecture
- Medallion architecture
- Event-driven data architecture
- Experience with Kafka, Spark, Flink, Hive, Iceberg, Delta Lake, or BigQuery.
- Experience building internal data platforms or self-serve data infrastructure.
- Experience with data quality frameworks such as Great Expectations, Deequ, Soda, or custom validation systems.
- Exposure to ML feature pipelines or feature stores.
- Experience with metadata management, data catalogs, lineage, and governance.
- Experience with cloud infrastructure such as AWS, GCP, or Azure.
- Understanding of privacy, compliance, PII handling, and access control in data systems.
- Critical business and product datasets are reliable, discoverable, and trusted.
- Pipelines are observable, recoverable, and have clear SLAs.
- Query performance improves across major analytical workloads.
- Data freshness and quality issues reduce significantly.
- Teams can build on top of the data platform faster without reinventing pipelines.
- The platform can scale with Apna’s user, job, employer, and engagement data.