AI Coding Evaluation & LLM Benchmark Opportunities
Review code generated by frontier reasoning models, design failing test suites, verify algorithmic complexity, and craft golden responses across Python, Java, TypeScript, and SQL.
Benchmark Tip
Prepare for 60-minute algorithmic challenges (LeetCode medium complexity)
Benchmark Tip
Focus on explaining edge cases and suboptimal time complexity
Benchmark Tip
Demonstrate clean modular documentation standards
Verified AI Coding Opportunities (4)
Each listing displays sourced compensation rates and verified country eligibility.
AI Software Engineer & Research Contributor
Build, evaluate, and benchmark cutting-edge frontier AI models and LLM agent architectures. Mercor conducts automated AI-driven technical interviews to place top engineers with elite AI labs and Silicon Valley startups at industry-leading compensation.
AI Coding Evaluation & Code Review
Evaluate complex code responses produced by cutting-edge frontier LLMs. You will test solutions, craft failing test cases, verify algorithmic efficiency, and rewrite suboptimal code across Java, Python, and SQL.
Full-Stack AI Code Generation Evaluator
Review and evaluate coding assistants solving frontend, backend, and database queries. Assess functional accuracy, adherence to user prompts, clean code formatting, and security vulnerabilities.
Senior Software Engineer LLM Benchmarking
High-bar engineering evaluation for advanced models. Review generated multi-file repositories, assess architectural trade-offs, identify concurrency bugs, and formulate adversarial prompts.
Get Matched to AI Coding Gigs
Upload your resume or build your candidate profile to receive deterministic skill fit scores and real-time country eligibility verification.