Role description
Within Personalization, the Speak Team owns the development of Spotify's state-of-the-art speech models, contributing to speech recognition, speech synthesis, and speech-to-speech models. We craft voice models that match human-level emotional expressiveness, so we can deeply engage our listeners and support creators at scale. Our groundbreaking work on speech synthesis relies on state-of-the-art deep learning methods and evaluation techniques, highly efficient data processing and model serving, and capturing audio of outstanding quality from our voice talent pool.
We're looking for a senior applied research scientist with experience in developing novel ML techniques and architectures and with a strong interest in working across a full production pipeline to produce state-of-the-art generative conversational speech-to-speech models. You'll collaborate with our engineering teams to help develop our production pipelines, explore new ideas and methods to improve quality, understanding and realism, as well as push the frontiers of what is possible with our speech technology.