Description
We’re looking for a senior-minded Nearshore engineer who can turn SAS-based logic into robust, governed Databricks pipelines using Python, SQL, and modern CI/CD practices—while maintaining strict delivery discipline.
Required experience
- 4+ years of professional data or software engineering experience.
- Strong Python and SQL skills.
- Hands-on Databricks experience with Unity Catalog, Workflows, and Databricks Asset Bundles.
- Proficiency with GitLab CI/CD (pipelines, merge request workflows, and automated testing) and disciplined Git branching and code reviews.
- Clear written and spoken English for client-facing collaboration.
- Demonstrated ownership: scoping work, delivering outcomes, and proactively flagging risks without being prompted.
Focus areas
- Pipeline reliability, validation, and delivery of converted code.
- Parity checks between SAS and Databricks outputs.
- Deployment and governance using DABs + GitLab CI/CD + Unity Catalog.
Additional experience
- Spark and Delta Lake performance tuning.
- Data validation and reconciliation experience.
- Infrastructure-as-code or DAB-based deployment experience.
- SAS reading ability and exposure to healthcare data (plus experience with Azure) are valued.
How we work: We value clarity, accountability, and continuous improvement. We’ll expect you to communicate trade-offs, confirm assumptions early, and build trust through predictable delivery, thoughtful reviews, and transparent risk management.
Funciones
We’ll rely on you to drive the end-to-end conversion delivery from SAS inventories to validated Python/SQL outputs on Databricks, with a strong focus on pipeline reliability and data validation.
- Pipeline engineering: Build and run pipelines that process SAS inventories and produce converted outputs.
- Quality and parity validation: Validate converted code for parity against SAS outputs (e.g., row counts, checksums, schema, and data types).
- Deployment ownership: Own deployments through Databricks Asset Bundles (DABs) and GitLab CI/CD, ensuring repeatable releases.
- Databricks governance: Manage Unity Catalog objects, permissions, and promotion across environments.
- Operational excellence: Troubleshoot job failures and performance issues, and take preventive actions to improve pipeline stability.
- Performance tuning: Apply tuning techniques for Apache Spark and Delta Lake to meet reliability and execution-time expectations.
- Data reconciliation: Use data validation and reconciliation practices to ensure correctness and consistency.
- Infrastructure-as-code mindset: Implement DAB-based deployment patterns and support automated, testable delivery workflows.
- Client collaboration: Communicate progress, risks, and technical decisions clearly with client and partner stakeholders.
Deseable
- Experience converting SAS workflows to Python/SQL in production environments.
- Healthcare data exposure and familiarity with typical data quality and privacy expectations.
- Azure exposure and understanding of how cloud services fit into end-to-end delivery and operations.
- Deep experience optimizing Spark jobs (partitioning, caching strategies, skew handling) and Delta Lake (file sizing, compaction patterns).
Beneficios
- Contrato a largo plazo.
- 100% Remoto.
- Vacaciones y PTOs
- Posibilidad de recibir 2 bonos al año.
- 2 revisiones salariales al año.
- Clases de inglés.
- Equipamiento Apple.
- Plataforma de cursos en linea
- Budget para compra de libros.
- Budget para compra de materiales de trabajo
- mucho mas..