CUDA Engineering Expert
Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. Applicant pay shown: $80 - $100 / hour.
- Platform
- Mercor
- Fit category
- Software engineering
- Remote/location
- Remote
- Pay
- Firm hourly pay: $80 - $100 / hour
- Eligibility
- Remote; check the work-authorization wording
- Difficulty
- Selective
- Listing/source checked
- Jul 1, 2026
- Inventory presence checked
- Aug 12, 2026
- Apply link checked
- Jun 27, 2026
Before you apply
A concise view of the reviewed listing. Confirm current terms and eligibility on the platform before applying.
Why this may be worth checking
- Experienced technical specialists
- Applicants who can show practical project evidence
What to prepare
- Available to work at least 20 hrs/wk
- Fluent in core C++ features through C++17
- Prepare concise evidence of the specialized experience the role asks for before starting the assessment flow.
Reasons to pause
- You need H1-B or STEM OPT support and the live listing says that is not supported
- You need guaranteed acceptance, guaranteed pay, or guaranteed hours
What still needs checking
- Review the official Mercor listing before applying. Requirements, screening, pay, hours, and project availability can change.
Decide on this role
Keep it for later, compare it with other current roles, or add it to your application tracker before you apply.
Save, Not for me, Compare, and application tracking remain in this browser.
Current application
Check the current Mercor listing
Apply on MercorYou'll continue on Mercor's site. This Apply link may be a referral link, and Specialist AI Work may be paid if the platform credits the referral. This does not change the role's advertised pay or how roles are ordered here.
What this role involves
Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance using profiler-guided analysis. You'll help evaluate, optimize, and reason about GPU kernels across modern hardware environments. This is a contract-based opportunity for specialists who enjoy squeezing performance out of modern GPU architectures.
Requirements
- Working knowledge of Python and Git
- Fluent in at least one GPU programming model, such as CUDA, HIP, Slang, HLSL, GLSL, or related kernel programming
- At least 1 year of professional or graduate-level research experience working with GPUs
- Strong understanding of GPU profiler performance metrics and how to use them to optimize kernels
- Ability to optimize GPU kernels without needing deep prior context on every algorithm
- Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization is a plus
- Experience optimizing kernels for NVIDIA Blackwell hardware is a plus
- Familiarity with NSight Compute is a plus
- Prior experience with GPU hardware organizations such as NVIDIA, AMD, or Qualcomm is a plus
- Open-source contributions related to GPU kernel optimization are a plus
Best for
- Engineers comfortable with role-specific assessments
Skip this if...
- You are not willing to complete Mercor's application or assessment process
Application tips
- Submit your resume or relevant technical background to get started
- Qualified applicants may be asked to complete a brief technical assessment or submit additional information
- Verify the live Mercor listing details before applying because availability, pay, and requirements can change.
- Review the application steps shown in Mercor so you know which resume, interview, form, or work-authorization items may be reused across roles.