I’m a software engineer working on agentic systems, observability, and infrastructure. Recently that has meant OpenTelemetry integrations, a metrics pipeline moving more than 50MB/s into a data lake, and a global Kubernetes platform running in over 50 datacenters. Before software, I spent a decade in trading and investments. The two fields have more in common than they might appear to.

The through line

Investment research is a measurement discipline. You have an idea, and the idea is worthless until you can say how it performed on data it never saw. I spent years learning that a strategy which looks excellent in sample usually looks that way because you fit it to the sample, and that separating a real edge from a flattering backtest is most of the job.

That habit came with me into software, where it needed almost no translation. A service that passes its tests is not the same as a service doing what you intended in production. You need evidence from the running system. A prompt that works on the ten examples you had in mind when you wrote it is an in sample result. An agent pipeline without evals is a strategy without a backtest.

Both fields ask the same two things of you, and the two pull against each other. You need enough imagination to generate candidates worth testing, whether strategy ideas then or eval scenarios and synthetic queries now, and enough discipline to disbelieve every one of them until a measurement that could have proven them wrong fails to. Observability, measurement, specification, and verification are four facets of the discipline, and I’ve spent my career on some version of it.

How I got here

I started building software in 1997, writing tools to research and implement strategies for investment selection and portfolio management. From 2001 to 2003 I took small contracts on the side while studying: stock selection research for equities funds, backtesting for asset managers. I then moved into managing money for ultra high net worth investors, eventually with responsibility for over $2B. In 2013 I left to become a proprietary trader, owning several short time horizon equities strategies end-to-end.

In 2015 I joined Honest Dollar, a seed stage fintech startup, responsible for investment algorithms and portfolio management across hundreds of accounts. I learned software testing as distinct from testing a stock strategy. I learned to deploy and operate what I’d written, and to build SaaS infrastructure. Distributed systems became part of the work and have stayed part of it ever since. Goldman Sachs acquired the company, and I stayed two years to see the migration through to completion.

After that I contracted for a while, building two trading systems and infrastructure for a stock exchange alongside a number of projects I took mostly because they paid. Then datacenter operations and automation at a large banking software company, and then data pipelines and analytics for a rideshare analysis tool whose market data reports were sold to asset managers. That last one settled the question of what I wanted to do next. I wanted to know the latency of queries, see which queries were used most, and understand the throughput of the pipelines. Coming from a field where you are expected to defend every number, I found it hard to accept less.

I moved into observability deliberately, joining Lightstep — a tracing company expanding into metrics and logs — shortly after ServiceNow acquired it. I managed a team of contractors building integrations across the metrics ecosystem, covering the major infrastructure components that emit metrics: Kafka, Cassandra, RabbitMQ and others. Two of those integrations were receivers for the OpenTelemetry Collector, and one of them was the initial SSH receiver.

The acquisition was still in its early stages when I arrived, so the job combined migration work with expanding the company’s capabilities in observability. That was my second time in that position, once from inside the company being acquired and once joining after the deal. Re-platforming a system against standards someone else set is work I like. You can’t stop serving traffic while you do it, so you need continuous evidence that nothing has drifted.

From there I moved to a platform team within the organization, where I operated a global Kubernetes platform spanning more than 50 datacenters, supported workload teams from their first Helm chart through the life of their services, ran the logs and metrics stack, and built and operated the high volume metrics pipeline feeding our data lake.

What I work on outside the office

Since 2015 I’ve set aside time for whatever is newest, and I’ve never regretted it. It started with deep learning, training image recognition models and assembling a personal pipeline for image recognition, OCR, and transcription. Sentence transformers came next, for enriching textual data. When LLMs arrived and added modalities, I began testing them as replacements for pieces of those pipelines.

Well before the models were good at it, I was having them generate code and tests so I could measure code generation error rates using software tests as an oracle, including the harder case where generated tests are themselves the oracle. Code generation has been a thread throughout, starting with spec to code tools like Lean and moving to LLMs as they became capable enough.

Today I’m building agents and the tooling around them for tracing and evals. The deeper statistical side of evaluation is something I’m still learning. Most of what I build now performs ops on my homelab infrastructure, manages personal documents, and curates unstructured data. This keeps the work in domains where I have the expertise to label and judge agent performance without the benefit of a separate product team.

What I’m looking for

I’m looking for my next role. I want to be responsible for agent pipelines end-to-end: the pipeline itself, the eval datasets, the measurement of performance against both synthetic queries I generate and real queries extracted from production, and the observability that shows what any of it is actually doing.

Responsible for the whole of it does not mean doing it alone, and the parts I would not want to do alone are the interesting ones. Which dimensions of a product are worth measuring is usually clearest to the people who own the product, and I would rather take that from a product team than guess at it. Statistical rigor in how evaluations are designed and read is somewhere an ML or data science team will improve on anything I would build by myself. I like the enabling side of the work as much as the building: helping other teams get their pipelines deployed on Kubernetes and instrumented well enough that they can say what their systems are doing. Not long ago the lead of a team I had supported got in touch to say their workload had reached the platform because of that work. I would rather hear that than almost anything about a system of my own.

The teams I’d be most at home on are the ones building agentic observability and evaluation itself; that is where my last several years and my own projects converge. Past that, the problem is the same anywhere someone has to prove that an agentic system behaves the way they claim. Fintech is where I bring something extra, having worked under the compliance and evidentiary requirements that turn that proof into a regulatory obligation.

What I believe

If you don’t have evidence of what your code is doing in production, it’s probably not doing what you intended. The hedge is deliberate. Probably is a low enough bar that it should worry us more than it does, and it is most of why I keep ending up in observability.

If you aren’t allocating time to experiment with cutting edge projects, you’re falling behind. This has always been somewhat true. In an age where AI and big data change on a scale of weeks, the interval has gotten short enough that the time has to be scheduled.

Contact

Feel free to connect with me, but please include a note with the request: