§ 01Snapshot
- Levelsenior individual contributor; 5+ years of backend engineering, recently moved into AI engineering on production agent systems.
- Roles heldsenior software engineer at an engineering consultancy (client work for large European companies), senior AI engineer on a production multi-agent platform, independent builder of automation and AI systems.
- Scaleindividual contributor; led one AI product end to end on a client engagement.
- Industriesindustrial / engineering (client work), HR and recruiting workflows, B2B lead generation, compliance tooling (e-invoicing, whistleblower reporting, PII protection).
- Company typeslarge enterprises via a consultancy, a mid-size services company, small businesses buying automation.
- Specialtiesmulti-agent reliability (guardrails, retries, fallbacks), LLM evaluation design (fixed thresholds, gold sets, shadow-mode A/B), data-pipeline correctness, PII / GDPR handling in AI pipelines, agent-readiness of websites against agent-facing web standards, business-process automation.
- Best forseed startups and small companies putting LLM agents into production; teams whose agent works in demos but fails in production; small businesses automating lead or intake workflows; companies that want their website usable by AI agents.
- Core stackTypeScript, Python, Go; Azure messaging (Service Bus), Azure SQL, Cosmos DB; LLM APIs, LangGraph, n8n, Supabase, local LLMs; Postgres, DuckDB, dbt, Dagster, Delta Lake; LLM evaluation and fine-tuning (LoRA, constrained output grammars).
- Time zonesAsia-Pacific; async-friendly.
§ 02Services
| ✓ | PR/MR code review | TypeScript / Python / Go, agent and data-pipeline code |
| ✓ | Architecture review | agent systems, data pipelines |
| — | Architecture design | Not offered by this expert |
| ✓ | Audit / due diligence | agent reliability, agent-readiness, data quality |
| ✓ | Vibe-code rescue | AI / automation builds |
| — | Team mentoring | Not offered by this expert |
§ 03Track record
Agents for a production multi-agent workflow platform
A production multi-agent platform automating recruiting workflows has to keep working when a model times out, returns junk, or a downstream system is slow. One of our experts builds agents for this platform with guardrails, retries and fallbacks, on a cloud message-bus architecture.
End-to-end AI transcription and summarization product with PII protection
A large European engineering group needed AI transcription and summarization, without leaking personal data. Working through a consultancy, one of our experts led the product end to end: speech-to-text, custom PII-detection models, a migration of the backend from PHP to Go, and CI/CD.
EU-compliant e-invoicing and whistleblower-reporting tools
European companies face mandatory e-invoicing and whistleblower-channel rules. One of our experts built EU-compliant e-invoicing and whistleblower-reporting tools as part of consultancy work for European enterprise clients.
Auditing a live lead pipeline that lost data while every test passed
A pipeline scraping and enriching thousands of businesses twice a day looked healthy: builds green. One of our experts audited it and found a fallback dedup key wrongly merging several hundred distinct businesses, an LLM output limit truncating more than half of enrichment results, a join multiplying a fact table many times over, and a data contract badly undercounting the most valuable leads. They rebuilt it as a medallion lakehouse with grain tests.
Fine-tuned small models vs a prompted generalist: a controlled study
Is it worth fine-tuning small models for a classification task? One of our experts ran a controlled study on a single consumer GPU: a prompted mid-size generalist against fine-tuned small language models and a classifier, with thresholds fixed before training, a human gold set, time-based splits and shadow-mode A/B. Outcome: accuracy was a tie, which the report states plainly; fine-tuning bought an order-of-magnitude speedup, deterministic results and fully valid structured output.
Lead-acquisition and voice-intake automation for small businesses
Small businesses lose leads to slow follow-up and missed calls. One of our experts builds automation for this: a daily pipeline that delivers qualified prospects each morning, and inbound voice agents that handle intake, booking and reporting, one of them built in under two weeks.
§ 04What you can book
Multi-agent reliability review
An agent system works in demos but fails unpredictably in production.
Map every model and tool call, check output validation, retry limits, timeouts, fallbacks and guardrails, and look for loops and silent failures.
A failure-path map, a ranked fix list, and suggested evaluation checks to keep it honest.
Agent-readiness audit of a website
More traffic will come from AI agents acting for users, and most sites aren't built for them.
Test the site against agent-facing web standards with an inspector tool, check whether agents can discover and call key actions, and compare with similar sites.
A readiness score, a gap list, and the changes that matter most.
Business-process automation with a concise proposal
An owner knows a process eats hours but not what automating it would take.
Map the process, pick the steps where automation or an agent pays off, and write a short, few-page proposal with scope, risks and how success will be measured, before building anything.
The proposal, then an automation built on n8n or a code-based agent.
LLM evaluation setup before a model or prompt change
A team changes prompts or models based on vibes.
Build a human-labelled gold set, fix the pass thresholds before testing, use time-based splits, and run changes in shadow mode before switching.
An evaluation harness and a written decision rule for the next change.
PII and GDPR review of an LLM pipeline
Personal data flows into third-party LLMs without a clear record.
Trace where PII enters, check detection and redaction before model calls, look at logging and retention, and consider air-gapped or local models where data can't leave.
A data-flow map, risk list and redaction / audit-layer design.
Rescue of an automation or agent build that stopped working
A no-code or AI-assisted automation grew piece by piece and now drops records or fails without anyone noticing.
Add monitoring and row-count checks, find where data is silently lost (dedup, truncation, joins), and move the fragile parts to tested code.
A findings list, fixes on the critical path, and basic data-quality tests.