Data Engineer
About the Role
Job Title: Data Engineer Job Type: Full-time Location: Remote We are looking for a Data Engineer to build and scale the data infrastructure that powers AI-driven products and research initiatives. In this role, you will develop distributed data pipelines, manage large-scale datasets across cloud environments, and design reliable data systems that support data processing, experimentation, and model development at scale.
What You'll Do
- Design, build, and maintain scalable data pipelines to ingest, process, and transform large-scale datasets from multiple sources.
- Develop and optimize distributed data processing workflows using Spark and cloud-native technologies.
- Build and maintain data storage solutions across SQL and NoSQL systems, ensuring scalability, performance, and reliability.
- Design and implement data architectures on AWS to support high-volume data ingestion, processing, and distribution.
- Write efficient Python and SQL code to extract, transform, validate, and analyze large datasets.
- Ensure data quality, integrity, monitoring, and operational reliability across data pipelines and storage layers.
- Collaborate with AI researchers, data scientists, and engineering teams to support data-intensive applications and experimentation.
- Implement automation, orchestration, and monitoring workflows to support scalable and efficient data operations.
Required Skills and Qualifications
- Strong proficiency in Python, SQL, and distributed data processing frameworks such as Apache Spark.
- Hands-on experience with AWS data services and cloud-native data architectures.
- Experience working with both SQL and NoSQL databases.
- Experience managing and processing large-scale datasets in distributed environments.
- Strong understanding of data partitioning, performance optimization, and scalable data architectures
Requirements
- Exposure to AI/ML workflows or research environments.
- Experience with data visualization tools such as Matplotlib, Seaborn, or Plotly.
- Familiarity with LLM-related data workflows (datasets for training, evaluation, or prompt experimentation).
Compensation & Logistics
The national pay range for this full-time position is base salary of $100,000 –$150,000 USD. All employees are eligible for equity compensation, and employees may also receive performance-based bonuses, dependent on role and subject to company policies. micro1 provides a comprehensive benefits package, including up to 100% reimbursement for health-insurance premiums, paid time off, a 401(K) plan with a company match, and additional benefits designed to support a high-performing, remote-first workforce.
Disclaimer
See Pay Min/Max columns for the hourly rate range.
Related roles
CAM Programmer
CAM Programmer — remote, paid AI training/evaluation work (other) ($80-$120/hr).
Senior Database Reliability Engineer
Senior Database Reliability Engineer — remote, full-time salaried role ($220,000-$260,000/yr).
Computer Systems Analyst
Computer Systems Analyst — remote, paid AI training/evaluation work (data-analysis) ($60-$120/hr).
