← Back

Member of Technical Staff (Data Intelligence)

Reka • Engineering

About this role

Join Reka's Data Intelligence team to ensure our data is high quality and can be produced at petabyte scale reliably and efficiently. Work closely with model researchers, data infrastructure engineers, and cross-functional partners to define "good data" and make smart choices about data that show up in model behavior. Reka is a globally distributed foundation model startup headquartered in San Francisco building useful multimodal AI. Our mission: build multimodal AI and use it to empower organizations and businesses. Our founding team and members contributed to breakthroughs in AI over the past decade from Google DeepMind, Facebook AI Research (FAIR), and successful startups. Key responsibilities: Work with model researchers to define "good data" (quality metrics, validation checks, thresholds), explore open-source datasets and create internal ones for fundamental World Models, build algorithms for automated data quality assessment and domain adaptation from synthetic to real data, track datasets, metadata, provenance, and versions for reproducible experiments, own CI/CD and development tooling for data stack (GitHub, Python, PyTorch), track and optimize throughput, storage, and compute utilization. Required: Strong ML and deep learning fundamentals with experience building and operating large-scale data/compute systems. Comfortable moving between research questions and production engineering. Demonstrated research experience with data compositions, quality, and dataset releases. Ability to design and execute experiments with unbiased outcomes. Practical experience with distributed processing and orchestration (Spark, Ray, Airflow). Solid Python skills, familiarity with modern model training workflows (datasets, checkpoints, experiment tracking). Strong data quality instincts—measuring, monitoring, preventing regressions at scale. Fast-moving environment comfort, prioritization ability, clear communication with researchers and engineers. Bonus: Large video dataset experience, dataset curation for training, or building internal tooling for ML evaluation/analysis. Benefits: 5 weeks paid leave, comprehensive healthcare (vision, dental), visa support including H1B and OPT transfers.
Apply now →

Job details