Fixed rules first, AI second
2 min read
My background is in operations, not software. Earlier this year I spent two months using AI coding tools to build Fluent Lens, an iOS app for learning languages with YouTube and Netflix videos. You watch with subtitles in two languages, and when you tap a word, the app explains it in the context of the sentence. I tested it through TestFlight, but it isn’t on the App Store.
The design question that kept coming back was which decisions to leave to the AI model.
German showed me why. If you tap “zurückkehren” in “Teilung wird zurückkehren”, what you need explained is “wird zurückkehren”, the whole verb. A model can work that out, but not every time, and it sounds just as sure when it gets it wrong. So in Fluent Lens that choice is made by rules in ordinary code, and several of the difficult cases are covered by tests. One of those tests feeds the backend a wrong answer on purpose, one that reads the “ihr” in “Meister Oogway, ihr habt mich gerufen” as plural and claims 98 per cent confidence. In that line, “ihr” is a formal way of addressing a single master. The test only passes if the rule corrects the answer.
The model still does what it’s good at, which is explaining. But if the app can’t verify something, such as the grammar of a word, it leaves the field empty instead of showing “Unknown” or “0%”. I would rather a learner saw a blank than a confident guess.
The app also has a live mode, which translates speech from the microphone. Its translation model is billed for the length of the audio, silence included. So a session only opens while you hold the button or have live mode switched on. It closes shortly after you let go, or as soon as you leave the screen or lock the phone, and twelve seconds of silence end it automatically.
I also got one safeguard wrong. At first, every session was capped at two minutes to protect credits. In practice it mostly cut off real conversations, and to the user it looked like an error. Idle sessions were already being caught much earlier, so the cap went up to five minutes.
I recognised the split between rules and the model from operations work, where anything that has to happen the same way every time gets a written procedure and people’s judgement is kept for everything else.