← Writing

The site that answers for me

· 5 min read · Davit Hayrapetyan

Most personal sites are a CV with better typography. Mine answers questions.

You can ask it whether I have worked with Kafka, what I did at Synopsys, or whether I can legally work in the EU, and it answers from a curated set of files about me. Under every answer there is a diagram of the pipeline that produced it, showing which model tier responded and which sat idle. That diagram is not decoration. It is the reason I built the thing.

Why not just use an API key

The obvious build is one call to a hosted model. I wanted the site to be an argument about how I build systems, and "one call to someone else's API" is not an argument.

So the answer chain has four tiers, tried in order:

  1. A large model running on a Mac in my flat.
  2. A hosted model, when the Mac is asleep or away.
  3. A small local model, as the last generator.
  4. A regex classifier, for intent only, when every model fails.

The first tier is the interesting one, and not because running a model at home is clever. It is interesting because it is usually unavailable. A laptop sleeps. It leaves the house. Designing around a dependency that is absent half the time forces every decision that matters in a distributed system: how long you wait, what you do instead, how you notice, and what you tell the person waiting.

The twenty-second bug that was not a bug

An absent machine does not refuse your connection. It drops the packet, and your HTTP client then sits through the operating system's retry schedule — twenty seconds or more — before failing. Put a frequently-absent machine first in a fallback chain and you have added twenty seconds to every request made while it is away.

The fix is a sub-second reachability probe before the real call, and a circuit breaker that eventually elides even the probe. Two breakers, actually, with different personalities: the home model gets a twitchy one that opens after two failures and stays open for five minutes, because nothing about "the laptop went to the office" resolves in sixty seconds. The hosted tier gets the ordinary one.

None of this is novel. All of it is the difference between a demo and something you can leave running.

Making the model faster by asking it to think less

Intent classification — deciding whether a question is about my career, my projects or my taste in music — was taking four to twelve seconds. I assumed the model was slow.

It was not. It is a reasoning model, and it was spending 525 tokens thinking before emitting a single word. For a task whose entire output is one label from a fixed list.

Turning reasoning off for that one call took it to 12 tokens and about 1.1 seconds, with accuracy unchanged at 44 out of 45 on my test set. Median latency across the set went from 4.6 seconds to 1.1.

I also tried capping the output length, which seemed obviously correct and was completely wrong: a reasoning model spends the budget thinking, hits the cap, and returns nothing at all. Every request silently fell through to the next tier. The test set caught it the same day. It is now a comment in the code telling me never to do that again.

What the assistant is not allowed to do

The model never holds a fact about me.

When the site generates a CV tailored to a job posting, the model chooses which of my experience is relevant and rephrases it in that employer's vocabulary — but company names, job titles, dates and degrees have no field in its output schema at all. It returns identifiers, and my own text is looked up against them. A model that wanted to invent an employer has nowhere to put it.

Everything it writes then goes through a verifier that checks each number against the source it came from. A figure that is not in my profile gets reverted to my original sentence. The first real run wrote a summary quoting a throughput number that was true and in my profile — just from a different job — and the check caught it.

This is the part I would keep if I threw away everything else. The value is not that the output is good. It is that the output does not need to be read line by line before I can send it.

The thing I got wrong

Every section of this site opens in a dialog when you click a card. Good interaction, and it hid the entire site from search engines: the HTML a crawler receives mentioned three of my past employers exactly zero times. Thirteen years of work, reachable only by clicking, next to an assistant that answers only when asked. Crawlers do neither.

The fix was a plain page with all of it in HTML. It is not a clever fix. It is the reminder that a system can be correct, fast, well instrumented and still fail at its actual job, because nobody checked the thing it exists to do.

Is it worth it

For a personal site? Honestly, no — not if the goal is a personal site.

It was worth it as a place to make decisions I would otherwise only argue about in design reviews: how to fail over, what to do when the expensive dependency is gone, where to put the boundary between a model and the code that trusts it, how to prove a generated document is not lying. Those decisions are all visible in this one, and it runs every day whether or not anybody visits.

Ask it something. If the Mac is awake you will see it answer first.