HumanBit Logo

Product Engineer | Scrabble & Jigsaw

full-time
Posted on 15-09-2026

Job Description

Product Engineer (MTS – Product)

About the Company

A fast-growing AI company building reinforcement learning environments and infrastructure used by frontier AI labs to train and evaluate the next generation of AI agents.

They create high-fidelity environments that replicate complex real-world software products and enables AI agents to interact with these environments, perform tasks, and be evaluated against measurable outcomes.

As AI agents become increasingly capable, the next challenge is enabling them to perform realistic, judgment-heavy software work. The company is building the infrastructure, evaluation systems, and quality frameworks required to solve this challenge.

About the Role

We are looking for a Product Engineer (MTS – Product) to own and scale the quality backbone of the evaluation platform.

This is not a traditional QA role. You will be a product engineer who specializes in quality, building evaluation frameworks, test harnesses, automated verification systems, regression suites, and quality gates that enable teams and customers to trust AI evaluation results.

You will work at the intersection of product engineering, AI evaluation, and applied ML, writing production code while also designing benchmarks, rubrics, and testing methodologies for systems where correctness is often probabilistic rather than deterministic.

What You'll Do

  • Design and build quality testing infrastructure for LLM evaluations, including judges, simulations, and agentic pipelines.

  • Build regression and quality-gating systems to identify model and pipeline drift before it impacts customers.

  • Define and evolve evaluation methodologies, including:

    • Rubric design

    • LLM-as-judge calibration

    • Golden datasets

    • Benchmark suites

    • Adversarial and red-team test sets

  • Instrument and analyze data pipelines to identify quality issues across synthetic data generation, labeling, and fine-tuning workflows.

  • Partner with product engineers, researchers, and customers to convert real-world AI failure modes into repeatable and automated tests.

  • Own quality metrics and reporting, translating evaluation results into actionable insights for engineering and product decisions.

  • Contribute directly to core product development — this is a build-first engineering role, not a role focused solely on test infrastructure.

  • Help establish engineering and quality practices as the evaluation platform scales.

What We're Looking For

  • 3–5 years of experience as a software or product engineer with experience shipping production systems.

  • Hands-on experience with LLM evaluation, data quality, or AI/ML evaluation, such as:

    • Evaluation frameworks

    • LLM-as-judge pipelines

    • Benchmarking systems

    • Model evaluation

    • Training or data pipelines

  • Strong software engineering fundamentals with the ability to own services end-to-end — from design and development to testing, deployment, and monitoring.

  • Strong hands-on experience with Python and modern data/ML tooling.

  • Ability to work effectively in ambiguous environments and help define quality methodologies for emerging AI systems.

  • Strong product instincts with an understanding of quality from the end customer's perspective, beyond traditional test coverage.

  • Excellent communication skills and the ability to explain evaluation results, quality trade-offs, and technical findings to both technical and non-technical stakeholders.

Nice to Have

  • Experience with RL environments, agentic evaluations, or multi-turn simulation testing.

  • Familiarity with LLM red-teaming or adversarial testing.

  • Experience with synthetic data generation, data curation, RLHF, or fine-tuning workflows.

  • Experience working in a fast-moving startup environment with high ownership and minimal oversight.

  • Exposure to AI agents, evaluation platforms, benchmarking, or developer tooling.

Why Join?

  • Work at the frontier of how advanced AI agents are trained and evaluated.

  • Solve challenging infrastructure and evaluation problems that don't yet have established playbooks.

  • Work closely with product engineers, researchers, and AI teams.

  • Take significant ownership in a small, high-caliber team.

  • Build systems that directly influence the quality and capabilities of the next generation of AI agents.

  • High ownership, low process, and an environment where your work ships and has meaningful impact.

Location: Bengaluru
Experience: 3–5 years
Role: Product Engineer (MTS – Product)

Powered by
HumanBit Logo