verified listingSign up to apply with your verified profile — no re-entering experience or references.
source · wttj·req · jb_7fae05aabc·listed 11h ago

Software Engineer (Data Infrastructure & Acquisition)

Speechify·Cardiff, United Kingdom·Hybrid·Full-time
Sourced listing · wttjNo salary disclosed
Posted
21 July 2026
via wttj
Type
Full-time
Arrangement
Hybrid
United Kingdom
Deadline
20 August 2026
closes in 30d
compensation · not disclosed
Salary not shared
Sign up to see our estimate based on role, location, and seniority.
source · estimate pending

Summary

the pitch

Join Speechify, a leading AI company, as a Software Engineer focused on Data Infrastructure & Acquisition. In this role, you will be responsible for all aspects of data collection to support our model training operations. You will work closely with our scientists to deliver richer data at a larger scale and lower cost, and collaborate with the AI team and Speechify leadership to craft the dataset roadmap for our next-generation products. Candidates should have a BS/MS/PhD in Computer Science or a related field, 5+ years of industry experience in software development, and proficiency in Docker, Infrastructure-as-Code, and at least one major cloud provider (GCP).

Role

posted by company
  • We are looking for a skilled Software Engineer to join us
  • BS/MS/PhD in Computer Science or a related field
  • Experience with web crawlers, large-scale data processing workflows is a plus
  • Ability to handle multiple tasks and adapt to changing priorities
  • 5+ years of industry experience in software development
  • Proficiency in Docker and Infrastructure-as-Code concepts and professional experience with at least one major Cloud Provider (we use GCP)
  • Strong communication skills, both written and verbal
  • Proficiency with bash/Python scripting in Linux environments

Key responsibilities

  • Responsible for all aspects of data collection to support model training operations, including finding new sources of audio data and integrating them into the ingestion pipeline.
  • Operate and extend the cloud infrastructure for the ingestion pipeline, currently running on GCP and managed with Terraform.
  • Collaborate closely with scientists and other team members to enhance data quality and throughput, and to craft the AI Team’s dataset roadmap.