Tier 01 · Frontier stress-testing
Soros — Electrical Engineering
Gemini, Claude, GPT (SOTA)
PhD-level Power Electronics, Embedded Systems, Chip Design & Power Systems problems
Tier 01 · Frontier stress-testing
Soros — PCB Design
Emerging PCB-design LLMs
Golden-solution PCB designs with full component specs; near-zero prior training signal
Tier 01 · Frontier stress-testing
Mistral Math & Logic Stress-Test
Mistral → Mixtral 8x7B
Adversarial math/logic design that fed directly into the resulting model
Tier 01 · Frontier stress-testing
Gemini Spreadsheet Stress-Test
Gemini 3 Pro
7–8-class spreadsheet taxonomy + adversarial structured-data tasks
Tier 01 · Frontier stress-testing
Qwen Time-Reasoning Stress-Test
Qwen
Interdependent, multi-timezone scheduling to stress temporal logic
Tier 01 · Frontier stress-testing
Alibaba/Tencent Adversarial Terminal Bench
Claude, Qwen (commercial)
Golden solutions pass all tests; SOTA models pass only a fraction across 16 runs
Tier 02 · Benchmarks & coding agents
Google GDM SWE-Bench
Google DeepMind
Harbor validation, Agentic Vet panel, patch & bug-fix remediation
Tier 02 · Benchmarks & coding agents
Google Antigravity
Google Antigravity
Live GitHub PR resolution on 1,000+-star repos via remote SSH
Tier 02 · Benchmarks & coding agents
Meta Agentic Code Generation
Meta
GitHub-issue-based golden-response construction
Tier 02 · Benchmarks & coding agents
Terminal Bench 1.0 → 3.0
Alibaba, Meta, Google, xAI, Reflection AI, DataHub, TML
Flagship Linux/agentic benchmark: fundamentals → workflows → long-horizon agents
Tier 02 · Benchmarks & coding agents
OSWorld (+ GUI scripting)
Multi-lab
PyAutoGUI VM capture across Chrome, LibreOffice, Terminal, VS Code
Tier 02 · Benchmarks & coding agents
Video / Execution Review
Multi-lab
Full-process scoring of recorded agent runs — not final outputs alone
Tier 02 · Benchmarks & coding agents
Meta RLHF Code Evaluation
Meta
Python/JS/SQL execution validation + scoring for training feedback
Tier 03 · Agentic & tool-use SFT
Meta OpenClaw
Meta
Swarm-agent VM tasks via OpenClaw/Maton + MCP; 3,000+ samples
Tier 03 · Agentic & tool-use SFT
Apple Agentic AI
Apple (Siri)
Voice-style SFT prompting written as users actually speak
Tier 03 · Agentic & tool-use SFT
ServiceNow Agentic AI
ServiceNow
Long-horizon, multi-step, end-to-end workflow design
Tier 03 · Agentic & tool-use SFT
NVIDIA SFT
NVIDIA
100-API workflow tool-calling with script / LLM-judge verification
Tier 03 · Agentic & tool-use SFTConfidential
Penguin SFT & Tool Calling
Confidential engagement
API-call accuracy; hallucination-free JSON handling
Tier 04 · Safety & personalizationConfidential
Penguin Safety SFT
Confidential engagement
Jailbreak resistance; sensitive-information leakage prevention
Tier 04 · Safety & personalization
Google DeepMind Personalization
Gemini Pro 3.1 + beta models
Cross-service memory evaluation — Gmail, Calendar, search history
Tier 05 · SME & verticalsConfidential
Dove Bookkeeping
Confidential engagement
Finance-domain golden trajectories from real invoices, receipts & transactions
Tier 05 · SME & verticals
Electrical Engineering & PCB (Soros)
Gemini, Claude, GPT (SOTA)
Covered in Tier 1 — highest technical demand in the portfolio
Tier 06 · Multimodal
Video-Generation Frame Labeling
Video-generation models
Frame-to-frame instruction labeling for coherent video synthesis
Showing 23 of 23 engagements · Ordered most technically elite first within each tier