Machine Learning Engineer, Inference & Serving (Speech LLM)

Posted 2 hours ago
Jobright.ai

This role is part of the Jobright TNT - the private hiring network connecting top talent with top AI startups like Perplexity, Mercor, Cresta, Suno and 150 more.


This is not a mass job posting. Only select, high-signal candidates are invited to Jobright TNT and recommended directly to hiring teams


Hiring Company: Plaud


One-liner: A hardware and software company making AI-powered voice recorders for note-taking.


Salary: $195K/yr - $365K/yr


Why Join Us:


• The best-selling AI note-taking device globally, 700K units shipped across 170+ countries

• Bootstrapped to $180M ARR with 10x YoY growth for two consecutive years, no outside funding

• 400+ person team led by serial entrepreneurs, Offices in San Francisco, Tokyo, Singapore

• Competitive compensation and benefits, with flexibility for strong candidates


Role Responsibilities


Qualifications


Required


• Have hands-on experience building and deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models

• Understand the intricate tradeoffs between latency, throughput, and Time-To-First-Token (or Time-To-First-Audio) in real-time streaming environments

• Have practical experience with continuous batching, KV cache management (e.g., PagedAttention), and stateful connections necessary for real-time conversational AI

• Possess a deep understanding of GPU architectures (NVIDIA Ampere/Hopper) and the memory hierarchy, allowing you to identify and eliminate hardware bottlenecks

• Communicate clearly and collaborate effectively, as you will sit at the critical intersection between the core ML training team and the backend infrastructure team

• Thrive in fast-moving environments and genuinely enjoy the systems-engineering challenge of squeezing every last drop of performance out of a cluster of GPUs

• Are obsessed with building AI systems that natively understand and generate speech, ultimately creating a hardware-software AI companion that amplifies human productivity


Preferred


• Frontier Serving Frameworks: Deep, under-the-hood familiarity with modern LLM serving frameworks like vLLM, TensorRT-LLM, SGLang, or NVIDIA Triton Inference Server (bonus points for active open-source contributions to these repositories)

• Real-Time Audio Streaming: Experience handling continuous audio streams over WebSockets or WebRTC, deploying neural audio codecs, and managing chunked audio generation to minimize conversational latency

• Advanced Inference Techniques: Implementing cutting-edge generation algorithms such as speculative decoding, lookahead decoding, or chunked prefill

• Model Compression & Quantization: Hands-on experience with post-training quantization (PTQ), deploying models in FP8, INT8, AWQ, or GPTQ, without degrading audio naturalness or ASR accuracy

• Large-Scale Distributed Systems: Deploying multi-GPU (Tensor Parallelism) and multi-node inference pipelines, and managing autoscaling infrastructure using Kubernetes


How can I join Jobright TNT:


If this is your first time applying to a Jobright TNT role, the process works as follows:


1. Apply to your first Jobright TNT role

2. We review your background to determine if you meet the TNT quality bar

3. If qualified, your application is directly recommended to the employer

4. Once accepted into TNT, you may be:

- Invited to apply for other exclusive TNT-only roles

- Invited to private, invite-only hiring events with top startups


You will be notified of your TNT selection result.


PS: All Jobright TNT roles are 100% real, directly hired by top AI startups we partner with, and come with priority review and higher response rates than the normal application queue.

Login to Apply Now

Recommended Jobs