<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
<channel>
  <title>Jan Matus — Blog</title>
  <link>https://janmatus.com/blog/</link>
  <atom:link href="https://janmatus.com/blog/rss.xml" rel="self" type="application/rss+xml"/>
  <description>Field notes on system architecture, engineering leadership and shipping real systems.</description>
  <language>en</language>
  <lastBuildDate>Thu, 13 Aug 2026 08:00:00 +0200</lastBuildDate>
  <item>
    <title>Building paniAnetka: medical dictation that never leaves the room</title>
    <link>https://janmatus.com/blog/building-panianetka-on-device-medical-dictation/</link>
    <guid isPermaLink="true">https://janmatus.com/blog/building-panianetka-on-device-medical-dictation/</guid>
    <pubDate>Thu, 13 Aug 2026 08:00:00 +0200</pubDate>
    <description>A voice assistant for ultrasound physicians that runs entirely on the doctor&#x27;s own computer. What a hard privacy constraint does to your architecture — and why it was the right constraint to accept.</description>
    <content:encoded><![CDATA[<p>A radiologist finishes an ultrasound examination and starts talking: organ, findings, measurements, conclusion. Somebody has to turn that speech into a structured, printable report. For years the choices were a human typist or a cloud transcription service — and for medical audio, "cloud" is doing a lot of quiet work in that sentence. The recording of a patient examination leaves the practice, crosses a border or two, and lands on infrastructure the physician has never seen and could not audit if they tried.</p>
<p>paniAnetka is my answer to a simple question: what if it just… didn't? What if the entire pipeline — speech recognition, medical language processing, report structuring — ran on the computer already sitting in the examination room?</p>
<h2 id="the-constraint-is-the-product">The constraint is the product</h2>
<p>"Fully on-device" started as a compliance requirement and became the defining architectural decision. Nothing to host, nothing to breach, no data processing agreement gymnastics: audio never leaves the machine, so the GDPR story collapses from a legal project into a sentence. For a solo physician's practice — the people who have no legal department — that difference decides whether the tool is usable at all.</p>
<p>But a hard constraint like this bills you elsewhere, and it billed me in hardware. Cloud vendors solve their performance problems with a bigger fleet; on-device, you have to choose which devices to bet on. I standardised on Apple Silicon — an M-series MacBook as the recommended machine — for reasons as much operational as technical: a handful of configurations to test instead of the endless PC matrix, unified memory that lets the speech models and a 14B-class language model share one machine sensibly, and enough performance for real-time transcription in an examination room. A CPU-only Lite mode covers ordinary Windows laptops, but the full product quality is an Apple Silicon story — and betting on one platform for initial adoption is a trade I accepted deliberately. A practice can buy one specific, well-understood laptop today; hardware progress does the rest, and in two or three years the same workload will run on whatever commodity machine is already on the desk. Within that envelope is where most of the engineering lives: model selection and quantisation, keeping the pipeline deterministic, and treating every megabyte of memory as contested territory.</p>
<h2 id="polish-medical-speech-is-its-own-problem">Polish medical speech is its own problem</h2>
<p>General-purpose speech models are trained on the internet's idea of language. An ultrasound report is not that. It is dense with abbreviations, Latin, measurements dictated at speed, and the particular phrasing conventions of Polish radiology — inflected, compressed, and unforgiving of transcription errors, because a wrong number in a report is not a typo, it is a clinical fact that somebody may act on. A large part of the system is a correction and normalisation layer that treats the raw transcript as a hypothesis, not an answer: numbers are cross-checked, units normalised, anatomical terms resolved against the vocabulary actually used in reports.</p>
<h2 id="the-bottleneck-was-never-typing">The bottleneck was never typing</h2>
<p>The most valuable design insight did not come from the ML side at all. Watching how an examination room actually works made it obvious that the report is a workflow problem before it is a transcription problem: patients queue, the assistant preps the next examination while the physician finishes the previous one, and paper comes out of a printer at a specific desk. So paniAnetka grew a second face — the assistant's own day board: a live patient queue, one-click hand-off of the prepared patient to the physician, and finished reports printing at the assistant's station while the doctor keeps scanning. The dictation engine is the headline; the choreography is why the tool stays in the room.</p>
<h2 id="what-shipping-it-taught-me">What shipping it taught me</h2>
<p>Shipping software to a non-technical user on their own hardware is a masterclass in humility. There is no ops team and no SSH access — the product must install itself, diagnose itself, and recover itself, or it dies quietly in a place you will never see. Every failure mode I did not design for became a phone call. The disciplines that made it work are the unglamorous ones: deterministic pipelines you can replay, aggressive preflight checks, and update channels boring enough to trust.</p>
<p>paniAnetka is live — in Polish, for the Polish market — at <a href="https://panianetka.pl/" target="_blank" rel="noopener">panianetka.pl</a> (there is an <a href="https://panianetka.pl/en/" target="_blank" rel="noopener">English summary</a> too). And if your product has a "this cannot leave the device" constraint hiding in it — that is exactly the kind of system I <a href="/contact/">design and build</a>.</p>]]></content:encoded>
  </item>
  <item>
    <title>Sorry, your AI is only as good as you</title>
    <link>https://janmatus.com/blog/your-ai-is-only-as-good-as-you/</link>
    <guid isPermaLink="true">https://janmatus.com/blog/your-ai-is-only-as-good-as-you/</guid>
    <pubDate>Thu, 13 Aug 2026 08:00:00 +0200</pubDate>
    <description>On why the demo comes free, the product doesn&#x27;t, and what an AI agent has in common with a development team.</description>
    <content:encoded><![CDATA[<p>Every AI coding demo looks the same. Two sentences of prompt, thirty seconds of streaming text, and a working app appears. The audience is impressed — and to be fair, the demo is real. I've lived one myself. The first prototype of <a href="https://panianetka.pl/en/" target="_blank" rel="noopener">my medical voice-dictation product</a> — local-first speech recognition for Polish doctors, noise suppression, a working UI — came together astonishingly fast. Roughly a weekend of prompting.</p>
<p>Then I spent months turning that demo into a product. It is still teaching me lessons. And somewhere along the way I wrote down a sentence I keep coming back to: <strong>your AI is only as good as you.</strong></p>
<p>That sounds like a criticism of the models. It isn't. It's a criticism of the mental model most of us bring to the table.</p>
<h2 id="the-vending-machine-fallacy">The vending machine fallacy</h2>
<p>The popular picture of AI-assisted development is a vending machine: insert prompt, receive software. It works beautifully for demos and collapses the moment you need a product, for one simple reason: <strong>you cannot compress a product into a prompt.</strong></p>
<p>However hard you try, there will always be assumptions. Which error states matter. What "fast enough" means. Whether that config value can change at runtime. How the new module should relate to the three that already exist. You can't enumerate all of it up front — partly because it's too much, and partly because you don't know all of it yet either.</p>
<p>An AI model, faced with a gap, does not stop and raise a blocker. It fills the gap with something statistically plausible. Multiply that by a few hundred small decisions and you get software that works — in a way known only to the model.</p>
<h2 id="a-better-mental-model-you-just-hired-a-team">A better mental model: you just hired a team</h2>
<p>Here is the frame that actually changed how I work: treat the AI not as a tool, but as a development team that just showed up at your virtual office. Absurdly fast, never tired, broadly read, surprisingly skilled. But still — a team.</p>
<p>And every engineering lead knows the uncomfortable truth about teams: unless you've somehow hired a group of supernaturals, <strong>the team's output can only be as good as your input.</strong> Even excellent engineers, with the best intentions, cannot read your mind. What they need is exactly what an AI needs.</p>
<p><strong>Requirements — functional and non-functional.</strong> Nobody writes "must be debuggable" or "must degrade gracefully when the network drops" into a two-line prompt, and then everybody is surprised that the result is neither. The non-functional ones are the first casualties, because they are invisible in a demo and painfully visible in production.</p>
<p><strong>Context.</strong> What already exists, what to reuse, what must not be touched, which conventions the codebase follows. A new hire who starts committing on day one without reading the repo produces exactly the kind of mess you'd expect. So does a model.</p>
<p><strong>Constraints and hints.</strong> Where the traps are. Which trade-off matters here and which doesn't. Which use cases are real and which are imaginary. "Keep it simple, we don't need multi-tenancy" is one sentence that saves a week of speculative architecture — and a lot of duplication and accidental complexity.</p>
<p><strong>Attendance.</strong> This is the part people skip. A real lead answers the team's questions, asks their own, challenges a suspicious design, discusses the pros and cons, sometimes throws in the missing idea. AI-assisted development is the same loop, just compressed from weeks to minutes. Modern systems are genuinely complex; intuition alone will not carry anyone through the trade-offs — not a human team, and not an agent.</p>
<p>Skip all of that and an unattended team will still produce <em>something</em>. Something compiles. Something demos well. But UX, predictability, debuggability, ease of maintenance, documentation — those don't happen by accident in human teams, and they don't happen by accident with AI either.</p>
<h2 id="verification-is-where-the-product-is-made">Verification is where the product is made</h2>
<p>Even with great input, you don't ship a team's work unread. You review it. The same applies here, and it's not bureaucracy — it's the actual job. The model does not know your definition of done. It doesn't feel the difference between "passes the happy path" and "I'd bet my invoice on it."</p>
<p>So you read the diff. You question the abstraction that appeared out of nowhere. You ask why a dependency was added and whether the retry logic actually retries. And when something smells wrong, you say so — because, much like a good team, the model usually fixes things quickly once it's told <em>what</em> is wrong, not just that something is.</p>
<p>Skip verification, and you're not shipping your product. You're shipping the model's assumptions with your name on them.</p>
<h2 id="the-ceiling">The ceiling</h2>
<p>And now the loop closes, because these two threads — "treat it like a team" and "it's only as good as you" — are really one thought:</p>
<p><strong>You can't specify what you don't understand. You can't verify what you can't judge.</strong></p>
<p>If non-functional requirements aren't part of your vocabulary, they won't be part of your prompt. If you can't smell a bad abstraction, you'll approve it. If you've never debugged a production incident at 3 a.m., you won't know why observability belongs in the instructions. The model will follow you off any cliff — confidently, in fluent prose.</p>
<p>Agents don't work on their own. They don't invent on their own. They don't innovate on their own. At minimum, they need a prompt — and that prompt is an instruction to a development team, polished only as well as you can polish it. It carries everything you know. It also carries everything you don't.</p>
<p>That's why AI is not a replacement for engineering judgment. It's a multiplier of it. Multiply solid architecture instincts, clear requirements, and a habit of honest review, and the throughput is genuinely remarkable — I would not have built my product solo without it. Multiply zero…</p>
<p>Well. The math does the rest.</p>]]></content:encoded>
  </item>
  <item>
    <title>Run a technical debt registry like a risk register</title>
    <link>https://janmatus.com/blog/run-a-technical-debt-registry/</link>
    <guid isPermaLink="true">https://janmatus.com/blog/run-a-technical-debt-registry/</guid>
    <pubDate>Wed, 12 Aug 2026 08:00:00 +0200</pubDate>
    <description>A backlog is where debt goes to be forgotten; a registry is where it goes to be decided. What a working registry entry actually contains.</description>
    <content:encoded><![CDATA[<p>Most organisations believe they track technical debt because their issue tracker has a <code>tech-debt</code> label. Look inside that label and you will find the same thing everywhere: hundreds of tickets, no owners, no costs, sorted by age, read by nobody. A backlog is where debt goes to be forgotten.</p>
<p>A registry is a different instrument. Companies already know how to run one — every risk department does. A risk register works because each entry is forced through the same questions: what is it, what does it cost us, who owns it, and what did we <em>decide</em> to do about it. Technical debt deserves exactly that treatment, for exactly the same reason: it is a liability the organisation is carrying whether or not anyone writes it down.</p>
<h2 id="what-a-registry-entry-contains">What a registry entry contains</h2>
<p>Six fields do the work. Everything else is decoration.</p>
<p><strong>Description — in system terms, not code terms.</strong> "The scheduler and the transport layer share state, so they cannot be released independently" is a registry entry. "Refactor the scheduler" is a wish. The test: someone outside the team should understand what is constrained and why it matters.</p>
<p><strong>Cost.</strong> The estimated cost to close the item — an honest range, not wishful thinking. This number is allowed to be large; if it's small — why not fix it immediately?</p>
<p><strong>Interest.</strong> What the item costs while it stands: where delivery is slower, which defect class keeps recurring, what every integration pays again. This is the field organisations skip, and it is the only one that turns the registry into a management tool. Interest is paid every single day, in the currency that is actually being tracked — team velocity, time to resolution.</p>
<p><strong>Trigger conditions.</strong> The future event that changes the decision: "if we add a third product variant, this becomes blocking", "if this module needs certification, the shortcut fails audit". Triggers are what make accepted debt safe to accept — they define when <em>accept</em> expires.</p>
<p><strong>Owner.</strong> A name. Not a team, not a guild. Unowned debt is unmanaged debt.</p>
<p><strong>Decision.</strong> The standing verdict, dated: <em>pay down</em> (scheduled against the budget), <em>accept</em> (with triggers), or <em>watch</em> (interest unclear — measure it). The decision field is the entire point of the registry. Every item has one, which means nobody can later claim the debt was invisible. It was on the ledger, and the organisation chose.</p>
<p>Ok, there may be a 7th field — <strong>Reason</strong> — why the actual shortcut has been taken. It's helpful for tracking the psychology of the organisation, but it does require a mature approach to managing technical debt, or reverse-engineering of decision making.</p>
<h2 id="cadence-portfolio-review-not-confession">Cadence: portfolio review, not confession</h2>
<p>A registry that is written once is a backlog with more data. The registry that is being worked on a cadence lives — once a quarter, align the registry review with roadmap planning. The review is portfolio management: has any interest estimate changed, has any trigger fired, does the pay-down budget go to the highest-interest items, does the work down the road trigger any registry items. Thirty minutes with the right five people, not an all-hands lament.</p>
<h2 id="anti-patterns-worth-naming">Anti-patterns worth naming</h2>
<p><strong>The amnesty sprint.</strong> A one-off purge applied to a recurring liability; covered in <a href="/blog/technical-debt-is-a-governance-problem/">the previous post</a>.</p>
<p><strong>No owner.</strong> See above. This kills more registries than any tooling choice.</p>
<p><strong>Treating all debt as bad.</strong> Some debt is deliberate leverage — taken to ship, with a recorded trade-off and a trigger. The registry is not a wall of shame; it is the record that separates leverage from loss.</p>
<h2 id="start-with-ten-items">Start with ten items</h2>
<p>Do not inventory the whole system — a complete registry is a stalled registry. Take the ten items engineers complain about most, force each through the six fields, and hold one review. The first review usually retires two items nobody still cared about, prices three that turn out to be expensive, and produces the first genuinely informed decision the organisation has made. That is more debt management than most companies have ever done.</p>
<p>The tooling does not matter — a spreadsheet with six columns and a review log is enough. The cadence is what does the work.</p>
<hr>
<p><em>The registry is one of three moves — visible, priced, decided — from my webinar on managing technical debt: <a href="/talks/technical-debt/">watch it here</a>.</em></p>]]></content:encoded>
  </item>
  <item>
    <title>Technical debt is a governance problem, not a code problem</title>
    <link>https://janmatus.com/blog/technical-debt-is-a-governance-problem/</link>
    <guid isPermaLink="true">https://janmatus.com/blog/technical-debt-is-a-governance-problem/</guid>
    <pubDate>Tue, 11 Aug 2026 08:00:00 +0200</pubDate>
    <description>The dangerous debt is the debt nobody decided to take. Why refactoring sprints keep failing, and what recording decisions changes.</description>
    <content:encoded><![CDATA[<p>Every engineering organisation has a version of the same meeting. The engineers say the codebase is slowing them down. Management hears a request for time that produces no feature. After some negotiation a "refactoring sprint" is granted. It happens once. Six months later the meeting repeats, with the same slides and worse numbers.</p>
<p>Nobody in that room is wrong. The engineers are right that delivery is degrading. Management is right that "the code is bad" is not an investable proposition. The meeting fails because both sides are discussing a symptom.</p>
<h2 id="the-debt-is-not-in-the-code">The debt is not in the code</h2>
<p>Technical debt is usually defined as shortcuts in the code. That definition points at the wrong layer. The code is where debt <em>shows up</em> — the debt itself is the sum of architecture decisions your organisation made implicitly.</p>
<p>Schedule pressure is the standard villain in debt stories, but pressure does not create debt directly. What pressure removes is the moment where a trade-off gets recorded. A team under deadline couples two components that should have stayed separate. That may even be the correct call — shipping matters. The damage is not the coupling; it is that nobody wrote down that the coupling was chosen, what it costs, and what future condition should trigger revisiting it.</p>
<p>Eighteen months later, nobody remembers a decision was made at all. The coupling is just "how the system is". New work routes around it, which is how one implicit decision breeds five more. That is the compounding mechanism — and it is a records problem, not a craftsmanship problem.</p>
<h2 id="you-are-already-paying-the-interest">You are already paying the interest</h2>
<p>Debt has a principal — what it would cost to fix — and interest: what it costs you every month you do not. Interest shows up as slower delivery on anything touching the affected area, a recurring class of defects, integration pain, and the onboarding drag of a system nobody can explain.</p>
<p>Most debt arguments are about the principal, which management correctly reads as an expensive proposal with an uncertain return. The stronger argument runs the other way: measure the interest, because you are paying it <em>right now</em>, in currency management already tracks — lead time, defect rates, integration effort. A liability with a measured monthly cost is an investment decision. "The code is bad" is a complaint.</p>
<h2 id="why-refactoring-sprints-keep-failing">Why refactoring sprints keep failing</h2>
<p>The refactoring sprint is a structural mismatch: a one-off amnesty applied to a recurring liability. It treats debt as an anomaly to be purged instead of a balance to be serviced. The sprint pays down whatever principal fits in two weeks, changes nothing about how new debt is taken, and leaves no record — so the balance quietly rebuilds and the credibility of the next request erodes with it.</p>
<p>Compounding is easiest to see in product families. In one radio product family I worked on, a platform strategy held roughly 95% of the software common across more than ten shipped variants. That number was not hygiene or luck — it was the direct output of debt being taken <em>deliberately</em>: boundaries chosen on purpose, exceptions argued and agreed, trade-offs revisited when conditions changed. Full disclosure: little of it was written down — it lived in the heads of a genuinely good team. That is the expensive version of the discipline, and it works exactly as long as those heads stay in the room. Writing it down is how the same discipline survives growth, turnover and acquisitions. The same portfolio without that discipline forks a little with every variant, and the divergence itself becomes the debt — multiplied by the roadmap.</p>
<h2 id="what-governance-changes">What governance changes</h2>
<p>If the dangerous debt is the debt nobody decided to take, the fix is to make deciding cheap and visible:</p>
<p><strong>Record decisions when they are made.</strong> A lightweight architecture decision record — a page, not a ceremony — capturing the trade-off, the cost, and the condition under which it should be revisited. Debt taken this way is leverage. It is how products ship on time. The record is what makes it leverage instead of loss.</p>
<p><strong>Keep a registry, not a backlog label.</strong> Debt items need owners, a priced interest estimate, and an explicit standing decision: pay down, accept, or watch. A backlog is where debt goes to be forgotten; a registry is where it goes to be decided. (That registry deserves its own post — coming next.)</p>
<p><strong>Budget the interest.</strong> A standing allocation for debt service, reviewed on the same cadence as roadmap planning, replaces the hero sprint with something management recognises: portfolio maintenance.</p>
<p>None of this requires better engineers. It requires the same thing every other liability in the company already has — a ledger, an owner, and a review date.</p>
<p>If your debt conversation is still engineers versus management, you do not have a debt problem yet. You have a visibility problem, and that one is cheaper to fix.</p>
<hr>
<p><em>I gave a full webinar on this — making debt visible, pricing it, and paying it down deliberately: <a href="/talks/technical-debt/">watch it here</a>.</em></p>]]></content:encoded>
  </item>
</channel>
</rss>
