Skip to main content
HomeAnnotation
Confidential — prepared for technical reviewJuly 2026
AI Training Data · Benchmarks · Frontier Evaluation

Technical capability portfolio.

Three years of continuous engagement in AI data annotation, benchmark engineering, model evaluation, and adversarial red-teaming — measuring reasoning, tool use, computer interaction, terminal operations, planning, instruction-following, and real-world task completion for model families at the current frontier.

At a glance

Continuous delivery
3 yrs
Named frontier organizations
12+
Distinct engagements
21
Major benchmark systems
4
/ Executive brief

Built for labs and platforms that cannot afford shallow data.

The same annotate–review–QA system that built Terminal Bench from Linux fundamentals through autonomous long-horizon agents also produced the mathematics stress-testing behind Mixtral 8x7B, the adversarial terminal benchmarks that separate unreliable elicitation from absent capability, and PhD-level power-electronics problems engineered against current SOTA models.

This is not a generalist labeling operation applying one shallow process to every task. It is a single delivery pipeline that flexes from foundational benchmark construction, to frontier red-teaming, to specialist SME work — without lowering the quality bar.

Audience 01

Frontier labs & foundation-model teams

Adversarial evaluation, reliability gaps, and training signal that shapes the next release — not commodity labels.

Audience 02

Enterprise AI / agent platforms

Production-scale SFT, tool-calling, long-horizon workflows, and safety data under NDA-ready delivery.

Audience 03

Research & evaluation programs

Benchmark construction, Harbor-validated coding agents, GUI/computer-use suites, and process-level execution review.

/ Selected proof points

Evidence, not claims.

Three representative outcomes from the portfolio — model contribution, reliability measurement, and SME depth.

01Mistral → Mixtral

Model contribution

Stress-testing that shaped Mixtral 8x7B

Adversarial mathematics and logic tasks engineered against Mistral’s then-current enterprise model. The resulting data fed directly into the creation of Mixtral 8x7B — evaluation that became training signal for the next model.

0216-run protocol

Reliability signal

The gap between “can solve” and “reliably solves”

Adversarial Terminal Bench tasks where golden solutions pass every test, while top commercial models — including Claude and Qwen — pass only a fraction across repeated runs (e.g. a subset of 16 attempts). That gap isolates unreliable elicitation from absent capability.

03PhD-level SME

SME depth

PhD-level problems built to break SOTA

Complete research problems in Power Electronics, Embedded Systems, Chip Design, and Power Systems — constructed to break Gemini, Claude, and GPT. Not templated questions: work that requires genuine electrical-engineering depth to author.

/ Engagement matrix

23 engagements. Six tiers.

Ordered from the most technically demanding work to the most foundational. Filter by tier. Confidential rows are marked — not guessed.

12+

Named frontier organizations

21

Distinct engagements

4

Major benchmark systems

  • Tier 01 · Frontier stress-testing

    Soros — Electrical Engineering

    Gemini, Claude, GPT (SOTA)

    PhD-level Power Electronics, Embedded Systems, Chip Design & Power Systems problems

  • Tier 01 · Frontier stress-testing

    Soros — PCB Design

    Emerging PCB-design LLMs

    Golden-solution PCB designs with full component specs; near-zero prior training signal

  • Tier 01 · Frontier stress-testing

    Mistral Math & Logic Stress-Test

    Mistral → Mixtral 8x7B

    Adversarial math/logic design that fed directly into the resulting model

  • Tier 01 · Frontier stress-testing

    Gemini Spreadsheet Stress-Test

    Gemini 3 Pro

    7–8-class spreadsheet taxonomy + adversarial structured-data tasks

  • Tier 01 · Frontier stress-testing

    Qwen Time-Reasoning Stress-Test

    Qwen

    Interdependent, multi-timezone scheduling to stress temporal logic

  • Tier 01 · Frontier stress-testing

    Alibaba/Tencent Adversarial Terminal Bench

    Claude, Qwen (commercial)

    Golden solutions pass all tests; SOTA models pass only a fraction across 16 runs

  • Tier 02 · Benchmarks & coding agents

    Google GDM SWE-Bench

    Google DeepMind

    Harbor validation, Agentic Vet panel, patch & bug-fix remediation

  • Tier 02 · Benchmarks & coding agents

    Google Antigravity

    Google Antigravity

    Live GitHub PR resolution on 1,000+-star repos via remote SSH

  • Tier 02 · Benchmarks & coding agents

    Meta Agentic Code Generation

    Meta

    GitHub-issue-based golden-response construction

  • Tier 02 · Benchmarks & coding agents

    Terminal Bench 1.0 → 3.0

    Alibaba, Meta, Google, xAI, Reflection AI, DataHub, TML

    Flagship Linux/agentic benchmark: fundamentals → workflows → long-horizon agents

  • Tier 02 · Benchmarks & coding agents

    OSWorld (+ GUI scripting)

    Multi-lab

    PyAutoGUI VM capture across Chrome, LibreOffice, Terminal, VS Code

  • Tier 02 · Benchmarks & coding agents

    Video / Execution Review

    Multi-lab

    Full-process scoring of recorded agent runs — not final outputs alone

  • Tier 02 · Benchmarks & coding agents

    Meta RLHF Code Evaluation

    Meta

    Python/JS/SQL execution validation + scoring for training feedback

  • Tier 03 · Agentic & tool-use SFT

    Meta OpenClaw

    Meta

    Swarm-agent VM tasks via OpenClaw/Maton + MCP; 3,000+ samples

  • Tier 03 · Agentic & tool-use SFT

    Apple Agentic AI

    Apple (Siri)

    Voice-style SFT prompting written as users actually speak

  • Tier 03 · Agentic & tool-use SFT

    ServiceNow Agentic AI

    ServiceNow

    Long-horizon, multi-step, end-to-end workflow design

  • Tier 03 · Agentic & tool-use SFT

    NVIDIA SFT

    NVIDIA

    100-API workflow tool-calling with script / LLM-judge verification

  • Tier 03 · Agentic & tool-use SFTConfidential

    Penguin SFT & Tool Calling

    Confidential engagement

    API-call accuracy; hallucination-free JSON handling

  • Tier 04 · Safety & personalizationConfidential

    Penguin Safety SFT

    Confidential engagement

    Jailbreak resistance; sensitive-information leakage prevention

  • Tier 04 · Safety & personalization

    Google DeepMind Personalization

    Gemini Pro 3.1 + beta models

    Cross-service memory evaluation — Gmail, Calendar, search history

  • Tier 05 · SME & verticalsConfidential

    Dove Bookkeeping

    Confidential engagement

    Finance-domain golden trajectories from real invoices, receipts & transactions

  • Tier 05 · SME & verticals

    Electrical Engineering & PCB (Soros)

    Gemini, Claude, GPT (SOTA)

    Covered in Tier 1 — highest technical demand in the portfolio

  • Tier 06 · Multimodal

    Video-Generation Frame Labeling

    Video-generation models

    Frame-to-frame instruction labeling for coherent video synthesis

Showing 23 of 23 engagements · Ordered most technically elite first within each tier

/ Flagship system

Terminal Bench 1.0 → 3.0

Flagship benchmark system

Evaluates whether an agent can think, plan, execute commands, debug errors, and solve real software-engineering and systems-administration problems inside a live Linux environment — rather than answering questions about them. Built for organizations including Alibaba, Meta, Google, xAI, Reflection AI, DataHub, and TML.

1.0Fundamentals

Directory navigation, file manipulation, log search, software installation, shell utilities, configuration editing, script execution, archive handling, and permissions.

2.0Realistic workflows

Git operations, Python debugging, dependency management, Docker, package troubleshooting, multi-step scripting, and environment setup — reasoning over memorization.

3.0Autonomous agents

Multi-stage reasoning, autonomous planning, code modification, repository-level understanding, test execution, root-cause analysis, failure recovery, and adaptation after unexpected failures.

/ Quality system

Annotate → Review → QA.

01

Annotate

High-quality benchmark and training tasks: realistic Linux and GUI scenarios, multi-step and long-horizon engineering work. Verify expected outputs. Test reproducibility from clean environments. Design for diversity and realistic interaction — not convenience.

02

Review

Independent validation: remove ambiguous instructions, hidden hints and shortcuts, information leakage. Ensure fairness, clarity, and edge-case coverage. Nothing ships on the annotator’s judgment alone.

03

QA

Production layer: annotation verification, policy compliance, consistency detection, reproducibility standards, structured feedback to annotators, and formal sign-off before delivery.

Delivery stack

PythonGitDockerPyAutoGUIVM capture environmentsMCP pipelinesRemote SSHHarbor validationHard-coded script checksLLM-judge verification

Enterprise posture

  • NDA-ready delivery

    Engagements under internal project codes when end-client names cannot be disclosed. Confidential work is flagged — never invented.

  • Independent review layer

    Nothing ships on a single annotator’s judgment. Ambiguity, leakage, and hidden shortcuts are removed before QA.

  • Reproducible environments

    Tasks verified from clean environments with deterministic execution — script checks and LLM-judge verification at scale.

  • Production sign-off

    Policy compliance, consistency checks, structured feedback, and formal approval before a dataset is production-ready.

Technical reach

Fed into training & evaluation pipelines across the frontier stack

Collectively contributing to LLMs, coding assistants, autonomous software-engineering agents, computer-use agents, and evaluation frameworks themselves.

AlibabaTencentMetaGoogle DeepMindGoogle AntigravityGDMxAIReflection AIDataHubTMLAppleServiceNowNVIDIAMistral

Named where disclosable · confidential engagements flagged in matrix

/ How we engage

From confidential brief to production delivery.

Designed for procurement, research leads, and evaluation programs that need a clear path from NDA to signed-off data.

  • AI data annotation
  • Benchmark engineering
  • Model evaluation
  • Adversarial red-teaming
  • Quality assurance & review
  • Terminal & Linux agents
  • GUI / computer-use agents
  • Software engineering workflows
  • Safety & alignment SFT
  • SME / domain-expert data
  • Multimodal generation labels
  • Agent execution process review
  1. 01

    Scope under NDA

    Define task families, success criteria, volume, and confidentiality boundary in a technical scoping call.

  2. 02

    Pilot cell

    Small, reviewable batch through the full annotate–review–QA pipeline — calibrated to your quality bar before scale.

  3. 03

    Production delivery

    Scaled throughput with reproducible environments, automated checks, and formal sign-off on every release.

/ FAQ

Before the briefing

Who is this portfolio for?

Frontier labs, foundation-model teams, enterprise agent platforms, and research/evaluation programs that need training data and benchmarks at production quality — not crowd-sourced labels.

How do you keep quality consistent across six tiers?

Every engagement runs through the same annotate → review → QA pipeline, with independent review, reproducibility checks, and formal sign-off before a dataset is production-ready.

Can engagements stay confidential?

Yes. Several portfolio rows were delivered under internal project codenames. We mark those as confidential rather than inventing or guessing end-client names.

What is Terminal Bench?

A progressive Linux/agentic benchmark built in three stages — fundamentals, realistic multi-skill workflows, and autonomous long-horizon agents — used across multiple frontier organizations.

Confidential technical briefing

The next benchmark. The next stress test. The next frontier release.

Foundational rigor, frontier-grade red-teaming, and specialist domain depth — inside one pipeline. Request a confidential technical briefing.