Suwali LogoBook a demo
← All posts

We Built a Chatbot That Answers With Its Feet

Why large language models hallucinate, and how Suwali is built like a library that answers from the shelves rather than from memory.

LLM hallucinations

When I was a little girl, I always thought that quicksand and Fata Morganas were going to occupy a big portion of my adult life. Even though I was wrong about the quicksand, the hallucinatory part somehow found its way back through a door I did not expect – my job. I build AI systems, and it turns out the machines I help make hallucinate for a living.

A hallucination, in the standard psychological textbook sense, is a perception that occurs without any external stimulus. In other words, you see something, and there is nothing there. Mirages, like Fata Morganas, do not count: they are illusions, real perceptions distorted by the atmosphere (which is why my childhood cartoons technically lied to me about the desert). The word itself entered English a long time ago, in 1646, from Latin alucinari, “to wander in the mind”. Besides explaining my brain on Mondays, the term has been around in computer vision for more than two decades, originally describing models that detected objects that were not there.

The term stayed niche until large language models started doing it loudly and in public. Suddenly, everyone knew the word and the consequences. A chatbot tells you the Mona Lisa was painted in 1754; it invents a financial figure, a court case, a list of medical references that sound plausible and exist nowhere. And the term itself is a terrible misnomer since a hallucination is still a form of perception, and large language models have no perception to fail: we are just dressing up statistical processes in fancy but still psychological clothing. When I discussed this with the neuroscientist Anil Seth, he said he preferred “confabulation”, but I think that term carries the same anthropomorphic baggage, since confabulation is something humans do when they unconsciously fill in missing memories with plausible inventions, again implying a psychology that these systems simply do not have. I told him my own preferred term was botshit, which I am told is not appropriate for every occasion. This may be one of those occasions.

So, where do hallucinations come from? Despite the overwhelming number of LinkedIn experts, we do not fully know. There is still no empirically validated and coherent theory of why exactly they happen. But we do know the contributing factors: training data that contradicts itself, divergence between the source and the target, overfitting to spurious correlations, errors in how knowledge gets encoded into billions of parameters and decoded back out, and many others. It is not surprising that a model hallucinates about a second-order modal justification logic, on which it does not have much data compared to the latest Hollywood blockbuster.

In talks and in classes, my favorite starting slide is just to ask “what is the single purpose of a large language model” in an anonymous survey. I often get back results like “to understand humans”, “to generate knowledge”, “to mimic human reasoning” and similar profound answers, but the truth is that we built them for a single purpose, and that is “to complete a sentence”. In more technical terms, to predict the next most probable token. This process optimizes fluency rather than the truth or the factual groundedness of data. When the most probable continuation happens to be true, wonderful! When it is false, you get confident nonsense delivered in exactly the same tone. There is even a built-in engineering tradeoff: turn the randomness down, and the model becomes accurate but dull, and turn it up, and it becomes creative, human-like, and more likely to invent things (as we do). Just based on the setup, you cannot surgically remove one and keep the other – you are either Dr. Jekyll or Mr. Hyde.

That brings me to the question we as engineers get asked most often, usually by clients or someone holding a procurement budget hostage: “Can you guarantee 100% accuracy?” No, and anyone who guarantees it is selling you something. (Ironically, I am also selling you something. It is called Suwali, and we will get to it shortly.)

Does that mean you should not implement a chatbot? Also no. We do not refuse to work with humans because they misremember, exaggerate, mix things up, and occasionally insist that it works on their machine (guilty as charged). Instead, we build review processes and audit trails around them, so the same logic applies here. We are not trying to make a perfect system since such a thing does not exist, but the real question is whether we have managed to engineer the system so its failures are rare, traceable, detectable, and, in the end, cheap.

The bots often fail at the edges – precise numbers, exact dates, citations, names of obscure entities, often anything that happened recently, anything in a low-resource language (pretty much everything but English), and any question phrased so that a plausible-sounding answer exists whether or not a true one does. That is, the model would sooner hand you a wrong year than admit it has no idea (and we all know people like that). However, bots are genuinely good at conversations where the source of truth sits right next to the model: summarizing a document you provide, finding the information in your collection of documents, translating texts, and answering questions grounded in a curated knowledge base, with sources attached.

The philosophy behind Suwali, which we are building at Meedan, is that of a library. A plain chatbot is the person who read everything once, years ago, and now answers from memory fluently and confidently, but sometimes tells you that Ryan Gosling is the mayor of Chicago. Suwali is not allowed to answer from memory at all – we only borrow its reasoning power. Asked a question, it has to walk to the shelves – a knowledge graph stocked with content curated by our partners – and pick up the actual books. When I was a kid, we always had an old librarian who knew exactly where everything was. Mine wore reading glasses and floral dresses, and on good days, a real flower pinned to her cardigan. And she was a big woman who moved through the shelves at an unhurried pace, right up until a question excited her, at which point she could outrun every child in the building. She seemed to carry the entire library in her head, but she never answered with her head – she answered with her feet. She would walk me to the shelf, pull the book, open it to the page and show me how to finish my biology essay fast. If the library did not have what I needed, she simply said so, without making up that lions mate like jellyfish, and no amount of pleading could make her invent a second floor that did not exist.

Suwali is my childhood librarian who finds books the way she did, both by the catalog and by knowing what a book is about – full-text and semantic search combined – as it must answer exclusively from the pages in front of it. Every source Suwali cites is reconciled against the documents it actually pulled, so citing a book your library does not hold is structurally impossible. But we still have to fight the strong overconfidence of the bot, so before an answer reaches you, we have a team of additional librarians – or LLMs as judges – whose entire job is to check the previous one. The system goes back to the shelves with a rephrased question, maybe even with a hypothetical answer to be verified, or simply to check whether this is about biology after all. Some questions do not belong on the shelves at all, and for those, it walks you to the front desk, where a human being works. And because misinformation has never waited politely for an English translation, our shelves are multilingual from the start, from Arabic to Hindi.

All of these measures, of course, do not make hallucinations impossible, but they do make them rare and accountable. But my childhood handed me two images, and only one of them was a lie. The other one wore flowers and answered with her feet. Given the choice of what to build my adulthood around, I picked the librarian.

Funded by:

Sida Press Forward
Patrick J. McGovern Foundation McNulty Foundation National Philanthropic Trust Battery Powered