Senior Data Scientist, AI Training Data

Cognichip Inc.
Cognichip Inc.

Software Engineering, Data Science

Redwood City, CA, USA

Posted on Aug 21, 2026
**Job Title** Senior Data Scientist, AI Training Data **About Cognichip** We build AI-native tools for semiconductor engineering, combining large proprietary models, agentic workflows, and domain-specific intelligence to help engineers design, verify, and optimize chips faster. **About the Role** We're looking for a Senior Data Scientist to own the data our models learn from. Semiconductor engineering data is specialized, scarce, and often tightly licensed — very different from the general text and code used to train most large models. Turning it into training-ready, evaluation-ready datasets is one of the highest-leverage inputs to our model quality. In this role, you'll design the curation, synthetic data generation, and quality-modeling work that turns raw technical material into usable datasets, running at scale on our internal data infrastructure (managed by a dedicated platform team, so you can focus on the data itself). You'll work closely with domain engineers to figure out what the models actually need, and with our AI team to connect dataset improvements to measurable model performance gains. **Key Responsibilities** - Curate licensed and open-source technical datasets — collection, cleaning, annotation, and integration across the engineering lifecycle- Build automated pipelines for sourcing, license classification, and normalization of public data- Design synthetic and augmented data generation workflows to keep pace with model training demand- Develop quality-modeling approaches: deduplication, contamination/leakage detection, license and PII screening, difficulty/diversity scoring, and dataset-to-eval attribution- Write large-scale distributed data processing jobs, partnering with a platform team on infrastructure needs (throughput, versioning, lineage, reproducibility)- Translate observed model weaknesses and feedback into targeted, well-sourced datasets- Build and maintain retrieval/embedding datasets that support product features- Run exploratory analysis and produce insights that guide modeling, product, and go-to-market decisions- Collaborate across engineering, AI research, product, and business teams to turn ambiguous needs into concrete datasets- Establish data governance practices: license provenance, documentation, retention, and compliance **Required Qualifications** - MS or PhD in Computer Science, Data Science, Statistics, or related field- 5–10 years of hands-on experience in data science or ML data work, with ownership of production datasets used by other teams- Expert Python and strong SQL, with experience processing large datasets using distributed computing frameworks (e.g., Spark)- Practical experience preparing text or code corpora for LLM training, fine-tuning, or evaluation- Solid applied statistics and ML foundations, with experience in a major ML framework (PyTorch, TensorFlow, or scikit-learn)- Familiarity with modern data orchestration and versioned storage systems- Working knowledge of data governance practices — licensing, provenance tracking, handling of confidential/contractual data- Strong ability to work with domain experts and convert ambiguous requests into delivered datasets **Preferred Qualifications** The following items are not required but are great bonuses: - Exposure to hardware or engineering domain data (e.g., specialized design/verification formats and workflows). You don't need deep prior expertise — just genuine interest in learning a technical domain deeply- Experience building retrieval systems: chunking strategies for technical documents, embedding models, vector databases, and evaluation of retrieval-augmented systems- Experience with synthetic data generation using LLMs, including agentic pipelines built with modern orchestration frameworks- Familiarity with annotation tooling and workflows for expert-labeled data, including inter-annotator agreement and active learning- Contributions to open-source data, hardware, or AI projects; or published research in data-centric AI, dataset curation, or code models- Prior experience at a startup or early-stage team where you helped define a data function rather than inherited an existing one **What It’s Like Here** - We’re a fast-moving AI startup with a collaborative, high-trust culture- We value technical excellence, ownership, and the freedom to experiment- Our best work happens when our builders and innovators work closely together to turn ambitious ideas into category defining products- We operate on a hybrid schedule with four days in office, one day remote- If you’re excited to build cutting edge tools that empower semiconductor engineers and reshape how chips are designed, you’ll feel right at home