Job Description
Our Vision!
Founded in 2006 with 650+ engineers & global presence (8 delivery centers in Europe & North America) we strive to become a leading East-European technology partner for growing organizations in need of digital transformation of their products and services!
What you’ll do
– Design, develop, and maintain production-grade ETL/ELT pipelines using Apache Spark on Databricks
– Build robust data ingestion workflows from multiple sources including databases, APIs, streaming platforms, and cloud storage
– Implement and optimize Delta Lake architectures following medallion architecture patterns (Bronze, Silver, Gold)
– Develop real-time streaming data pipelines using Structured Streaming and Delta Live Tables
– Design and implement scalable data lakehouse solutions on Databricks
– Optimize Spark jobs for performance, cost-efficiency, and reliability
– Configure and manage Databricks workspaces, clusters, and compute resources
– Implement data partitioning, indexing, and caching strategies for optimal query performance
– Implement CI/CD pipelines for Databricks notebooks, jobs, and workflows
– Collaborate with data scientists to productionize ML models using MLflow
– Implement data governance policies using Unity Catalog
– Document technical designs, data lineage, and operational procedures
– Participate in code reviews and mentor junior team members
– Ensure compliance with data security, privacy, and regulatory requirements
What you need to be successful
– Proven track record of building production-scale data pipelines processing large datasets
– Experience with at least one major cloud platform (Azure, AWS, or GCP)
– Strong proficiency in Python and SQL; Scala experience is a plus
– Databricks: Expert knowledge of Apache Spark (PySpark, Spark SQL), Delta Lake, Delta Live Tables
– Data Processing: Experience with batch and streaming data processing patterns
– Cloud Platforms: Proficiency with Azure Data Lake Storage, AWS S3, or GCP Cloud Storage
– Version Control: Git/GitHub for code collaboration and version management
– Orchestration: Experience with workflow orchestration tools (Databricks Workflows, Airflow, Azure Data Factory)
– Strong understanding of distributed computing concepts and data architecture
– Experience with data modeling, dimensional modeling, and data warehouse design
– Knowledge of data quality frameworks and testing methodologies
– Ability to write efficient, maintainable, and well-documented code
– Strong problem-solving skills and attention to detail
Next Steps for you!
– Apply
– CV screening
– HR Interview
– Technical Interview
– Offer presented by our CEO