Logging-in

AWS Databricks Data Engineer

RemoteUnited StatesContract
$40 - $45 hourly
About the Job

Job Title: AWS Databricks Data Engineer

Location: Remote

Duration: Long Term


An AWS Databricks Data Engineer to design, build, and maintains scalable data pipelines and lakehouse architectures using Apache Spark, Python, and SQL. Optimize cloud storage, ensure data quality, and integrate seamlessly with other AWS services.

Core Responsibilities:

  • Pipeline Development: Design and build batch and streaming data pipelines using PySpark, Delta Lake, Autoloader and Delta Live Tables (DLT) to ingest data sets into Databricks.
  • AWS Integration: Build cloud-native data solutions leveraging AWS services (e.g., S3 for storage, IAM for access management, and AWS Glue or Lambda for serverless tasks
  • Data Governance & Security: Configure and manage access controls using Databricks Unity Catalog to ensure compliance and monitor lineage
  • Orchestration & CI/CD: Automate pipeline deployments using Databricks Workflows, Apache Airflow, and CI/CD tools (e.g., GitHub, GitLab). [123]
  • Cross-Functional Collaboration: Partner closely with stakeholders to build robust feature stores and prepare datasets for various consumption needs
 
Typical Qualifications & Technical Skills:

  • Industry: Healthcare Payer industry experience. Have worked on MMIS data sets – claims, provider, member enrollment and similar data sets
  • Experience: 3-5 years of hands-on data engineering experience, specifically on Databricks.
  • Programming: High proficiency in Python (specifically PySpark) and advanced SQL.
  • Big Data & Cloud: Strong understanding of Apache Spark, Data Lakehouse architecture, and working within a production AWS environment.
  • Databricks Ecosystem: Familiarity with the Databricks platform ecosystem, including notebooks, Delta Lake, and Unity Catalog.