Validio is the agentic data management platform that automates data quality, lineage, and cataloging — so enterprises can finally trust the data behind their analytics, AI, and reporting. As a Data Scientist on our Algorithm team, you build the models at the heart of the product — the Anomaly Detection that separates good data from bad.
A validator watches a metric over time; an ML model decides, for each new value as it arrives, whether it's anomalous and worth waking someone up for. These models run against tens of thousands of time series at once — unattended, unlabeled, with nobody tuning them per customer. You'll own that end to end: reading a metric from a customer's warehouse, running the algorithm online in our pipeline, and persisting the state it carries between runs, fast and reliably at scale.
What you'll do
Build the algorithmic core — anomaly detection, plus data profiling, SQL generation, lineage parsing, and change point detection
Take a method from a research paper to a tested prototype, then to a production model running online in the engine
Work out why a validator over-alerts on a metric — say one with a strong daily and weekly cycle — and reshape how it forms its expected range
Tell whether a model change actually improved detection across customers, when there's no labelled set to score against
Keep a validator's metrics correct and fast in SQL across multiple warehouses
Work with Customer Success, Solutions Engineering, and Support — on the call when the customer describes the problem
What you'll bring
ML models you've built and deployed to production, serving real predictions to users
A strong grounding in statistical modelling — the assumptions a method makes, how it estimates uncertainty, and how it fails when the data breaks them
Fluent Python, and either some Rust already or the appetite to learn it here
SQL well beyond SELECT — window functions, query plans, dialect differences, and why a query that flies on one warehouse is unusable on another
An instinct for performance — reasoning about CPU, memory, and bottlenecks when you implement a new algorithm
The range to carry a problem from a vague business need through prototype to measured impact, and to direct AI tools well — knowing what they get subtly wrong
Bonus points for
Time series, signal processing, anomaly detection, or unsupervised methods
Inference engineering — optimising runtimes and model throughput on general-purpose hardware
Our stack, and why Rust for most of the engine and backend. Python for exploration and prototyping in the Algorithm team, where modelling usually starts, and for parts of the service like the SDK and IaC tooling. SQL in every dialect that matters — we push computation down to the warehouse when we can and stream through our own pipeline when we can't. Postgres and Redis for state, Kubernetes and Helm to deploy, which lets us run Validio both as SaaS and fully air-gapped. Claude, Gemini, and Codex in daily use.
Why Validio Founded in Sweden in 2019, we're trusted by data-driven enterprises like Nordea, Canva, Truecaller, Traveloka, and Deutsche Glasfaser. We raised a $30M Series A this year and grew ARR 800% in twelve months as enterprises got serious about making their data AI-ready. At around 40 people, you'll own initiatives end to end and make the call when you're closest to the problem — strict about testing and documentation, and fast about decisions.
Come as you are. We hire for what you'll build with us, and we welcome applicants of every background.