The Opportunity
Cephable is building the future of privacy-first, on-device AI that helps people control, create, and automate across software on-device agents, speech, computer use, and more. Our technology runs locally—offline, secure, and fast—across consumer and enterprise environments, including productivity software and games.
We work at the intersection of speech recognition, multimodal ML, generative and reasoning models, and real-time systems, optimized for modern CPUs, GPUs, and NPUs.
Role Overview
We are seeking a Lead Machine Learning Engineer to own and advance Cephable’s core ML systems. This role is highly hands-on and technical, with responsibilities spanning model development, fine-tuning, optimization, quantization, testing, and deployment across on-device environments.
You will lead the design and implementation of ML models for speech recognition, natural language understanding, generative and reasoning tasks, and multimodal inference—ensuring they run efficiently, reliably, and privately on end-user devices.
Key Responsibilities
Model Development & Research
- Design, train, fine-tune, and evaluate ML models for speech recognition, generative and reasoning models, and multimodal inference
- Adapt open-source and foundation models using Hugging Face and related tooling
- Translate research ideas into production-ready systems
Optimization & On-Device Deployment
- Optimize models for low-latency, low-power, offline execution
- Perform quantization, pruning, and distillation
- Deploy models via ONNX Runtime and OpenVINO targeting CPU, GPU, and NPU backends
Systems & Infrastructure
- Build pipelines for training, evaluation, benchmarking, and regression testing
- Define and improve accuracy, latency, and resource metrics
- Partner with application and platform engineers to ensure seamless ML integration
- Communicate model performance, architectural decisions, and technical tradeoffs clearly to both technical and non-technical stakeholders
Technical Leadership
- Own Cephable’s ML architecture
- Set best practices and mentor team members
- Evaluate new tools, frameworks, and hardware
- Mentor engineers across the team on ML concepts and practices as the org grows
Required Qualifications
- 4+ years of experience in machine learning or ML systems
- Strong PyTorch experience
- Hands-on experience with Hugging Face
- Production deployment using ONNX Runtime and/or OpenVINO
- Experience with acceleration frameworks like CUDA and GPU workflows
- Strong software engineering skills (Python, C++, or systems-level experience)
- Excellent communication skills — able to explain complex ML concepts, tradeoffs, and decisions clearly to engineers, product stakeholders, and non-technical partners alike
Preferred Qualifications
- Speech recognition or voice assistant experience
- LLMs, SLMs, or reasoning models
- Multimodal ML experience
- Edge or on-device AI background
- Experience with QNN, WinML, and CoreML
About the Team
You'll be joining a small, highly collaborative engineering team of engineers. You will be the dedicated ML lead — there is significant greenfield opportunity here to shape systems, practices, and architecture from the ground up. Close partnership with application and platform engineers is a core part of the role.
Why Cephable
- Mission-driven impact: Your models run on real devices for real users — many of whom depend on Cephable as a primary way to interact with technology
- Greenfield ML ownership: Shape Cephable's ML architecture from the ground up with the trust and autonomy to do it right
- Cutting-edge stack: On-device inference, multimodal input, NPU targeting, and privacy-first AI
- Small team, high trust: Work directly with senior leadership in a low-bureaucracy environment
- Seed-stage momentum: Backed by top investors with enterprise partnerships at scale
What we offer:
Equity
- Meaningful equity grants (options) in a company with existing revenue and clear growth trajectory
- Standard 4-year vesting with 1-year cliff
Health & Wellness
- Health insurance
- medical
- dental
- vision
Time Off
120 hours per year accrued per pay period
Up to 40 hours carry over year to year (we want you taking vacation)
8 additional hours per year of tenure
15 paid Holidays
11 Federal Holidays, plus,
- Fridays before Labor and Memorial Day
- Extra day for July 4th
- Wed before and Friday after Thanksgiving
Location: Remote (Boston area preferred)