§ 01Snapshot
- Levelsenior data / ML engineer with more than a decade in data engineering, analytics and machine learning; several years of production FastAPI + Next.js work. Currently a solo founding engineer on their own product.
- Roles heldhead of an analytics department (built from scratch) · senior data scientist / analyst · senior data engineer · senior data / LLM engineer · founding engineer (solo, full-stack)
- Scaleset up the data function and hired the first team at three different companies · solo-built a full-stack data product that aggregates and refreshes hundreds of thousands of records from many public sources
- Industriesprogrammatic advertising · freight / logistics marketplace · telecom · betting and online entertainment · marketing-heavy consumer businesses · enterprise IT integration · consumer information products
- Company typeslarge telecom operators · large online marketplace · mid-size adtech · growing tech/product companies (tens to low hundreds of people) · early-stage solo product
- Specialtiesdata warehouse design from source ingestion to BI-ready marts · ETL / ELT pipeline reliability · RAG and LLM integration in products · fine-tuning open-weight LLMs · fraud detection and forecasting ML · AI-assisted solo development from prototype to production
- Best forsolo founder · seed startup · scale-up (especially the ~50-person "our numbers don't match" stage)
- Core stackPython (FastAPI, Django, Flask, asyncio, SQLAlchemy), TypeScript / Next.js / React, PostgreSQL, ClickHouse, MySQL, Vertica, Oracle, MS SQL, Airflow, dbt, Spark / PySpark, Hadoop / Hive, Kafka, RabbitMQ, Redis, PyTorch, Hugging Face, LoRA / QLoRA, LangChain, Weaviate, Tableau, Qlik Sense, Metabase, Superset, Docker, Jenkins, Sentry, AWS (EC2, S3), bare-metal hosting
- Time zonesEuropean time zones
§ 02Services
| ✓ | PR/MR code review | Python / FastAPI, Next.js, SQL, data pipelines |
| ✓ | Architecture review | data platform, DWH, LLM / RAG |
| ✓ | Architecture design | DWH, ETL, ML / LLM services |
| ✓ | Audit / due diligence | data stack and pipelines |
| ✓ | Vibe-code rescue | Python / FastAPI + Next.js MVPs |
| — | Team mentoring | Not offered by this expert |
§ 03Track record
Building an analytics department from scratch in programmatic advertising
An adtech company had data but no analytics function to turn it into reporting and decisions.
One of our experts was hired to lead data processing and analytics, built the department from scratch, hired and managed the team, and set up BI dashboards and analytical reporting on a Hadoop / Spark stack.
The company got a working analytics department and regular reporting where there had been none.
Fraud-detection algorithms for programmatic ad traffic
Fraudulent traffic distorted results and cost real money in programmatic campaigns.
One of our experts developed fraud-detection algorithms on the company's big-data stack, alongside the analytics reporting.
Suspicious traffic could be flagged algorithmically instead of found by hand after the fact.
DWH architecture and management dashboards for a large freight marketplace
The analytics team of a large freight marketplace needed one reliable warehouse, and management needed fast visibility during a sudden market shock.
One of our experts designed the DWH architecture serving the analytics team, deployed a BI platform and built dashboards, OLAP cubes and data marts for management.
Management had current numbers in dashboards when conditions were changing week to week.
Freight price forecasting with ensemble models
A freight marketplace wanted to predict cargo transportation prices.
One of our experts built an ensemble regression model for price forecasting and helped set up the company's data science function, including hiring.
The company got a working price-forecasting model and the start of an in-house data science team.
Churn prediction on marketplace data
The marketplace found out about lost customers only after they had already left.
One of our experts delivered a churn-prediction ML project on top of the warehouse they had designed.
The business could see churn risk ahead of time instead of afterwards.
An ML system that allocates ad budget across channels and countries
Marketing spent across many channels and countries without a clear way to compare their effectiveness.
One of our experts built an ML/AI system to analyse advertising channel effectiveness and combined macroeconomic indicators with ML models to recommend budget allocation by channel and country, alongside data marts and BI dashboards.
Budget decisions were based on modelled channel effectiveness instead of habit.
A company data warehouse from scratch
Data was scattered across multiple heterogeneous sources and there was no single place to report from.
One of our experts designed and implemented the company's data warehouse from scratch: ingestion pipelines from every source, BI-ready marts, Tableau reporting, plus hiring and onboarding data analysts.
One warehouse and one set of numbers the company could report from.
Fine-tuning open-weight LLMs for product use cases
Product features needed language models tuned to specific tasks, served reliably.
One of our experts fine-tuned several open-weight model families (LLaMA, Mistral, Gemma, GPT-J) with LoRA / QLoRA and built FastAPI backend services to serve the ML/LLM products, with training orchestrated through Airflow and Jenkins.
Fine-tuned models running in production use cases behind the company's own APIs.
A solo-built data product with RAG at its core
A consumer product needed scraping, cleaning and comparison of hundreds of thousands of public listings from many sources, plus LLM features, and there was only one engineer.
One of our experts, as the sole founding engineer, built it end to end: FastAPI backend, Next.js frontend, regularly refreshed scraping/ETL into aggregated statistics, RAG with vector-store retrieval in core flows, deployment on bare-metal plus cloud. They used AI coding assistants heavily but owned the architecture, code and deploy themselves.
A live full-stack product with LLM features, built and run by one person.
LLM-based data-quality checks in production pipelines
Production analytics and ML pipelines pulled from many sources, and bad data was caught late.
One of our experts builds and supports the pipelines and automated data-quality monitoring, including LLM-based checks, and productionises data-science outputs for downstream use.
Data-quality issues are flagged automatically before they reach dashboards and models.
Data-lake ETL for an enterprise IT integrator
Data lived in many source systems with little documentation.
One of our experts sourced and documented data across the organisation's systems, built ETL integrations into the data lake, and developed data marts on a big-data stack.
Source systems feeding a documented data lake with marts on top.
§ 04What you can book
Data audit for a company at the "~50 people, numbers don't match" stage
Reports are built by hand in spreadsheets, departments disagree on the same metric, and one person holds the whole analytics setup in their head.
Inventory the sources and existing reports, check metric definitions and pipeline reliability, and look at bus-factor and access risks.
A risk map, a target data architecture sized to the company, and a phased plan the in-house team can run.
Making a vibe-coded FastAPI + Next.js MVP production-ready
A founder built an MVP with AI tools; it works on a laptop but not reliably in production.
Database schema and migrations, async and background work in FastAPI, auth and secrets, Next.js data fetching, deployment and error monitoring.
A prioritised fix list, reviewed PRs for the critical items, and a simple deploy-and-monitor setup.
RAG pipeline review before launch
A team added RAG to its product and answers are inconsistent, slow or expensive.
Ingestion and chunking, vector-store setup and retrieval quality, prompt and context assembly, evaluation, cost per answer, and whether a fine-tuned or self-hosted model fits better.
A written review with concrete changes, plus a small evaluation set to measure them.