Short Description
To design, build, and continuously optimize our core data infrastructure and pipelines. This role will take ownership of end-to-end data architectures, ensure data reliability and governance, and collaborate closely with Analytics, Data Science, and cross functional teams to enable process automation and data-driven decision-making across the organization.
Experience
Role & Responsibility (R&R)
- Architecture & Pipeline Design: Design, build, scale, and maintain robust batch and real-time data pipelines (ETL/ELT) to ingest data from diverse sources.
- Data Modeling & Storage: Architect scalable data warehouses, data lakes, and lakehouses using modern schema modeling techniques (e.g., Star/Snowflake schema).
- Performance & Optimization: Continuously monitor, troubleshoot, and optimize data infrastructure for query speed, cost efficiency, and reliability.
- Data Quality & Governance: Implement automated data validation, testing frameworks, CI/CD pipelines, and data governance policies (access control, metadata management, lineage).
- Cross-Functional Collaboration: Partner with Data Scientists, ML Engineers, and Business Analysts to transform raw data into analytics-ready datasets and feature stores.
- Technical Leadership: Conduct code reviews, and establish best practices for engineering excellence across the team.
Knowledge/ Ability
- Strong analytic skills related to working with unstructured datasets.
- Experience with data warehouse & processing : Microsoft Fabric, Redshift.
- Experience with big data tools: Hadoop, Spark, Kafka, etc.
- Experience with data pipeline and workflow management tools: Azkaban, Luigi, Airflow, etc.
- Proficiency in scripting languages: SQL, Python, Java, Scala.
- Data Modeling & Transformation: Deep understanding of dimensional modeling and experience with modern transformation frameworks.
Education
Computer Science, Information Technology, Data Engineering or related fields.