Case study
RefiningJarvis
Originally Mark Liv's self-hosted AI assistant. I forked it and changed it substantially around the parts that matter for how I'd actually use it: voice in and out, memory that persists between sessions, a browser it can drive and hand over, and an approval gate in front of anything with real-world consequences. JavaScript for the agent and browser layer, Python for the model and voice plumbing. Currently in a refining stage rather than a scaffolding one.
Credit where it is due
This is a fork of Mark Liv's Jarvis, not something I started from an empty directory. The original project is the foundation; the work here is the rework described below. If you are reading this to judge what I can build, it is fair to read the original project first and treat this as the delta.
What I changed
Four things, in rough order of how much they matter to me using it daily. Voice in and out, so it feels like an assistant rather than a text box you have to already be holding. Memory that persists between sessions, so it knows context without re-explaining. A shared Chromium instance it can drive and then hand over, so it can show me what it found instead of describing it. And an approval gate in front of anything with consequences, because an assistant that can act on the world without asking is worse than one that only talks.
Why the approval gate mattered most
Everything else is convenience. The gate is the difference between an assistant you trust with your machine and one you do not. It has to be reliable rather than merely present — a gate that occasionally fails open is worse than no gate, because you stop checking.
How it is built
Roughly 80% JavaScript for the agent loop and browser layer, 15% Python for the model and voice plumbing, plus a Dockerfile and shell for packaging. Self-hosted rather than dependent on someone else's API. The live build runs on Hugging Face Spaces and is password-gated, which is deliberate for something that can drive a browser — but the source is public.
Where it actually is
Not a fresh scaffold. The core is built and running; what it is in now is a refining stage — tightening the parts where it acts on the world rather than just talking about it, and making that approval boundary dependable. See the limitations below for what is not finished.
What I learned
- An approval gate is a design problem, not a feature to bolt on. It has to be trustworthy or it is worse than having none.
- Sharing one browser between the agent and the person removes the worst part of browser automation — having no idea what it is looking at.
What it doesn't do yet
Listed deliberately — a project's limits are as useful to a reviewer as its features.
- A fork. The original author is credited above and the work here is the rework, not the original.
- Transcripts are stored but are not searchable yet.
- Plugins are not hot-loadable at runtime — the original required a restart and that is not fixed.
- Still refining — the action layer is the least settled part.
- The live build is password-gated, so there is no public demo you can just open.