Machine Learning Researcher Audio
Protege
Remote
Europe
entry level to mid-level
June 1, 2026
$100k+
Want to apply for this job?
Subscribe to access the application link and 15,000+ more jobs
Job Description
Company Overview: We are building Protege to solve the biggest unmet need in AI — getting access to the right training data.
The process today is time intensive incredibly expensive and often ends in failure.
The Protege platform facilitates the secure efficient and privacy-centric exchange of AI training data.
Solving AI’s data problem is a generational opportunity.
We’re backed by world-class investors and already powering partnerships with some of the most ambitious teams in AI.
The company that succeeds will be one of the largest in AI — and in tech.
We’re a lean fast-moving high-trust team of builders who are obsessed with velocity and impact.
Our culture is built for people who thrive on ambiguity own outcomes and want to shape the future of data and AI.
ROLE OVERVIEW Data is the foundation of AI performance and we believe model quality starts with data quality.
For speech and audio models in particular the bar for signal fidelity consistency and quality control is exceptionally high.
We’re seeking a Machine Learning Researcher focused on audio data quality ML data evaluation and quality control to lead the evaluation and optimization of large-scale speech datasets used to train audio speech and multimodal models.
This role will be responsible not only for applying existing audio quality metrics but also for researching how audio data quality should be evaluated for machine learning systems and developing new methods benchmarks and evaluation frameworks that better predict downstream model performance.
You will help define what “high-quality audio data” means in the context of modern ML training.
That includes studying how different forms of acoustic degradation dataset inconsistency recording conditions speaker variation labeling quality segmentation quality and signal artifacts affect model behavior across ASR TTS speaker modeling representation learning and multimodal systems.
A core part of this role will be original research and method development: designing new approaches for measuring audio data quality validating those approaches against downstream model outcomes and translating research insights into practical evaluation tools filtering rules and quality standards used across Protege’s data platform.
This is an ideal role for someone deeply obsessed with audio data quality and signal understanding comfortable operating in both research and hands-on implementation modes and excited to help Protege become the ubiquitous platform for high-quality AI training data.
WHAT YOU’LL DO Research audio data quality for machine learning Investigate how audio quality signal properties dataset composition and localized acoustic issues affect downstream model training evaluation and deployment.
Develop new metrics benchmarks diagnostics and evaluation frameworks for measuring audio data quality in ways that are predictive of ML model performance.
Speech dataset characterization and metrics Analyze and summarize Protege’s audio catalog and maintain clear up-to-date quality scorecards and metrics for key speech datasets.
Develop methods to measure true acoustic properties directly from the waveform including effective bandwidth spectral energy distribution high-frequency roll-off noise clipping reverberation distortion and codec artifacts.
Segment-level quality evaluation Build workflows that evaluate diarized or segmented speech regions surfacing localized degradation that file-level averages may miss.
Apply multiple complementary quality metrics to detect bandwidth mismatches resampling artifacts clipping reverberation codec distortion and other forms of degradation.
Model and data evaluation Design and run targeted evaluations connecting audio quality issues to downstream model behavior including ASR performance speaker embedding stability learned speech representations and synthesis quality.
Test which audio quality metrics meaningfully correlate with model outcomes identify failure modes of existing metrics and design better alternatives when current approaches are insufficient.
Deterministic filtering and evaluation infrastructure Translate research findings into reproducible filtering rules quality gates and dataset selection strategies that improve dataset consistency across training runs.
Build scalable tools and pipelines for applying audio quality analyses across large datasets tracking results over time and making quality signals accessible to researchers engineers and data teams.
Cross-functional collaboration Work closely with ML researchers data engineers data operations and external partners to define measure and communicate the value of Protege’s audio data assets.
WHAT SUCCESS LOOKS LIKE Near-term: establish a trustworthy audio-quality baseline Create a trustworthy view of the quality consistency signal fidelity and training-readiness of Protege’s speech and audio datasets supported by metrics and scorecards the team can operationalize.
Then use targeted evaluations ablations and downstream model analysis to connect audio-quality issues to concrete dataset improvements and clearer prioritization over time.
WHAT YOU BRING PhD or equivalent Master’s degree + 4+ years industry experience in machine learning audio signal processing speech technology computer science statistics engineering or a related quantitative field.
Proven experience designing and running data evaluations audio analyses benchmarks ablations or slice-based analyses.
Strong understanding of speech/audio data and signal properties including sampling rates codecs bandwidth spectrograms reverberation clipping noise and perceptual quality.
Experience developing or critically evaluating metrics benchmarks or measurement frameworks for ML systems data quality speech technology or audio signal analysis.
Ability to connect low-level signal properties to downstream machine learning behavior including model accuracy robustness representation quality speaker consistency or synthesis quality.
Comfortable moving between research exploration and production implementation: you can formulate hypotheses run experiments analyze results and turn findings into scalable tools or decision rules.
Excellent written and verbal communicator; able to write concise technical docs and explain empirical results clearly.
High ownership and bias toward action; you independently scope questions design experiments and drive them to decisions.
BONUS SIGNALS - Experience with ASR TTS speaker modeling self-supervised speech models diarization or multimodal audio models. - Experience developing evaluation frameworks or performance metrics for training data. - Experience inventing adapting or validating audio quality metrics for ML training datasets. - Experience studying the relationship between dataset quality and downstream model performance. - Publications or open-source contributions in speech audio ML data-centric AI ML evaluation or related areas. - Cross-functional collaboration with product infrastructure data operations or partnership teams. - Experience collaborating with industry or academic labs on speech/audio research or data projects.
PROTEGE VALUES Pass the Loved Ones’ Test We act with integrity and do the right thing — especially when it’s hard and no one is watching.
Always Find a Way We are resourceful resilient builders who solve hard problems and push through obstacles.
Go Fast and Grow Fast Velocity matters.
We move with urgency learn quickly and continuously improve as individuals and as a company.
Practice Kindness and Candor We communicate directly and respectfully building trust through honest feedback and genuine care for one another.
Deliver Together We win as one team.
Collaboration accountability and shared ownership drive our success.
Own the Outcome.
Hone the Craft.
We take pride in our work sweat the details and continuously raise the bar for excellence.
More Jobs You Might Like
Helpful Resources
Salary & Savings Calculator
Compare salaries across European cities and calculate your potential savings. Understand cost of living and take-home pay for tech jobs in Europe.
Career Guides
Expert advice on landing high-paying tech jobs in Europe. Tips on interviews, salary negotiation, and career growth from The European Engineer.
Access 15,000+ High-Paying Tech Jobs
Get unlimited access to our full database of 15,000+ jobs with advanced filters, salary comparisons, and exclusive career guides from The European Engineer.