Experience
Backend developer, then Big Data engineer, then the person responsible for a data platform.
Fifteen years, nine companies. Details of the Audiense work are on the projects page; this is the timeline.
I'm responsible for the data platform end to end: its architecture, a good part of its implementation, and helping product teams use it.
- Introduced Trino as the lake's query engine and Apache Iceberg as its table format; I maintain the 15 TB+ lake behind audience enrichment and the 20+ streaming consumers on it. →
- Finished the migration of profile enrichment from batch to real time.
- Cut Trino report queries by 28–51% at the median and internal network per query by 93%; failures under 0.2%. →
- Designed VEGA, the asynchronous query API product teams read the lake through, replacing per-team database copies and the ETL and Spark jobs that fed them. →
- Built the interest-classification pipeline — LoRA fine-tuning, cross-encoder reranking, model-agreement rule, vLLM on Kubernetes — taking human-judged strict precision from 36% to 52%. →
- Mentor the team on using AI in day-to-day engineering work.
Designed and built a data platform for fraud analysis in motor insurance claims, and took it to production. Scala, Kafka, Akka, Cassandra.
Batch and streaming systems for clients; research on real-time processing architectures. Built ML model deployment pipelines on Kafka Streams, with monitoring on the ELK stack.
Node.js/Express APIs, OAuth2 for internal and third-party authentication, and planning of the data-analysis stack (Elasticsearch, Spark, Kafka).
SmartData: real-time footfall counting from WiFi device detection with Spark Streaming and Cassandra. Also IoT monitoring over WebSocket and end-to-end encrypted live video streaming.
Medical research data-collection platform with IoT biometric device integration.
Stack
What I work with
In production, not from a course. Roughly in order of how much I use each one today.