Product Engineer (Member of Technical Staff, Product)
Company Overview
Collinear is a research-focused AI company dedicated to building systems that enhance the reliability, alignment, and usefulness of intelligent models in real-world settings. The company operates across various domains including evaluation, simulation, post-training, and reinforcement learning (RL) to assist enterprise teams in deploying trustworthy AI systems. Their platform supports critical functions such as automated AI judges, red teaming, high-signal data curation for fine-tuning, and Reinforcement Learning with Human Feedback (RLHF). Collinear emphasizes a fast-paced, autonomous work environment where engineers own outcomes end-to-end, with minimal bureaucracy.
Job Summary
The Product Engineer (Member of Technical Staff, Product) will be responsible for developing and scaling the quality infrastructure of Collinear’s evaluation platform. This role is pivotal in ensuring the trustworthiness of large language models (LLMs) and agentic systems through the creation of evaluation frameworks, automated verification systems, and quality metrics. The candidate will operate at the intersection of product engineering and applied machine learning evaluation, contributing to core product code and defining evaluation methodologies that handle probabilistic correctness.
Responsibilities
- Design and develop automated testing and evaluation infrastructure for LLM-based products, including judges, simulations, and agentic pipelines.
- Build regression and quality gating systems to detect model and pipeline drift before deployment to customers.
- Define and refine evaluation methodologies such as rubric design, LLM calibration as judges, golden datasets, and adversarial/red-team test sets in collaboration with research and product teams.
- Instrument and analyze data pipelines to identify quality issues in synthetic data generation, labeling, and fine-tuning workflows.
- Partner with product engineers, researchers, and customers to translate real-world failure modes into repeatable, automated tests.
- Own and communicate quality metrics and reporting, transforming evaluation results into actionable signals for engineering and product decisions.
- Contribute to core product codebases, not just testing infrastructure, ensuring robust and reliable system development.
Qualifications
- 3–5 years of experience as a software or product engineer with a proven track record of shipping production systems.
- Experience with LLM evaluation, data quality assessment, or evaluation frameworks, including LLM-as-judge pipelines, benchmark suites, or training/data pipelines.
- Strong software engineering fundamentals with the ability to own entire services (design, code, test, deploy, monitor).
- Practical proficiency in Python and modern data/ML tooling.
- Experience with probabilistic systems and quality engineering for LLMs, understanding the nuances of evaluation in uncertain environments.
- Excellent communication skills, capable of translating evaluation results and quality trade-offs to both technical and non-technical stakeholders.
- Educational background in Computer Science, Software Engineering, Data Science, or related fields.
Preferred Skills
- Experience with reinforcement learning environments, agentic evaluations, or multi-turn simulation testing.
- Familiarity with red-teaming and adversarial testing for LLMs.
- Knowledge of synthetic data generation or curation for fine-tuning and RLHF.
- Prior experience working in fast-paced startup environments with minimal oversight.
Environment
This role is suited for a dynamic, fast-moving environment, potentially involving remote, hybrid, or in-office work settings. The work involves collaboration across cross-functional teams including research, product, and engineering, with a focus on building reliable, scalable evaluation systems for AI models.
Salary
Not specified.
GrowthOpportunities
Not specified.
Benefits
Not specified.