Symbols Travel, Meaning Is Rebuilt

This blog is from the “deleted scenes” of my book, Semantic Webs of Meaning, Technics Publications, 2026 (mostly Chapter 5).

TL;DR (Sort of)

Since this TL;DR is relatively long for a TL;DR, here is the elevator pitch:

Humans and LLMs live largely in rich, subsymbolic internal worlds. Knowledge graphs live in an explicit, serialized world. Intelligence increasingly depends on moving between the two—and continually checking that what was reconstructed still corresponds to what we meant.

Knowledge graphs and LLMs are not competing ways of storing the same thing. They represent knowledge in fundamentally different forms.

A knowledge graph serializes selected knowledge into explicit symbols—identities, relationships, assertions, provenance, rules, constraints, and other things we deliberately want to preserve and inspect. An LLM holds what it has learned subsymbolically, distributed through an enormous learned function that we can examine mechanically but cannot simply read. A human brain appears to have much the same problem: we can inspect neurons and synapses, but we cannot point to the place where dog isA mammal is written down.

I like to think of that subsymbolic world as a “dreamworld version of the Library of Congress“. The books are no longer sitting on shelves. They have been dissolved into the structure of the place itself. A symbol entering that world does not retrieve a little stored definition; it sets an enormous learned landscape into motion.

That is what I mean in this article by deserialization. It is not merely parsing. An LLM tokenizer can identify the incoming tokens; the important part comes afterward, when those tokens interact with context and billions of learned parameters to produce a much richer internal state. Humans have an even harder version of the problem because our incoming stream is multimodal—words, tone, gestures, facial expressions, surroundings, memories, expectations, and physical experience all interact.

Communication therefore does not transmit meaning directly. Symbols travel; meaning has to be reconstructed by the receiver. Humans do this with other humans. Humans do it with LLMs. LLMs can do it with other LLMs. And either humans or LLMs can serialize portions of what they know into a knowledge graph so that the knowledge becomes persistent, inspectable, shareable, and available to many other intelligences at once.

This also explains the symbiosis between LLMs and KGs:

  • KGs give LLMs explicit reference points. They can ground the model in identities, relationships, provenance, enterprise definitions, rules, and current assertions that should not depend entirely on what happens to be implicit in the model.
  • LLMs give KGs access to their enormous subsymbolic world. They can interpret sparse symbolic structures, relate them to broader learned knowledge, explain them, generate hypotheses, and help humans author and maintain the graph.
  • LLMs can serialize too. Prose, SQL, JSON, RDF, Python, and even an entire knowledge graph can be outputs of the same subsymbolic system. A KG therefore does not need to represent only what a human believes; one AI could explicitly publish its own beliefs, assumptions, preferences, constraints, or intentions in a KG for another AI to inspect.

None of this means the serialization is automatically true. Serialization is a representational question; epistemology is a trust question. Humans and LLMs can both serialize something wrong. Evidence, provenance, confidence, review, and governance remain separate issues.

And serialization is inherently lossy. My words do not contain my brain state. Reconstruction of those words is made using a different brain and a different lifetime of experience. That is why conversation is an iterative synchronization process: No, that isn’t quite what I meant. Oh, now I understand. LLM conversations work much the same way.

That same generative reconstruction also provides an intuitive way to think about hallucination. An LLM is very good at filling in what was not explicitly supplied. Usually that is an enormous strength. But when it fills in something unsupported, the newly generated token becomes part of the context for the next round of interpretation. An invented serialization can become the premise for the next deserialization, allowing a plausible mistake to grow coherently.

A KG provides something stable to return to.

Comparing Knowledge Graphs and Large Language Models

We often talk about knowledge graphs (KGs) and large language models (LLMs) as if they are competing ways of encoding and expressing knowledge. Do we need knowledge graphs when most, if not all, of that knowledge is in a frontier LLM? LLMs are relatively automatic to train, but KGs are tedious to author and maintain.

Rather, they are two sides of the same coin—serialized and subsymbolic, respectively. Rather than competing, they can work together in a symbiotic relationship. In fact, I’ve written about the symbiotic relationship between KGs and LLMs in Enterprise Intelligence (Chapter 4, page 94), Semantic Webs of Meaning, and many blogs on this site.

That symbiosis is:

  • AI tends to hallucinate, and like all intelligences, is subject to imperfect information. The explicit nature of KGs grounds AIs into “reality” or at least our explicit intentions.
  • KGs are very tedious to author and maintain. AI assistance can substantially mitigate that tedium.

It is really the “roux” of Enterprise Intelligence.

Those two benefits are the practical expression of a deeper architectural complementarity. A KG helps an LLM because it externalizes selected knowledge into an explicit, persistent form. An LLM helps a KG because it can interpret, extend, and help author that explicit form using a much richer learned internal representation. Seen this way, the symbiosis is not just a collection of useful techniques—RAG on one side and AI-assisted KG authoring on the other. It is really about moving knowledge back and forth between symbolic and subsymbolic forms.

And that is where I run into something that has always bothered me about the symbolic side.

What has always bothered me about ontology—and taxonomy, epistemology, and similar machinery—is that it is still symbolic. We explain one symbol with other symbols. We define Employee using more words, classes, properties, and relationships. We describe the provenance of a statement (which are sequences of symbols) with still more statements. Even reification, which lets us distinguish the assertion from claims about the assertion, does not escape that symbolic world—it simply gives us a more precise way to talk about the symbols and their provenance.

That is not a criticism of KGs. Symbolism is really their strength. Serialization makes knowledge explicit, inspectable, queryable, governable, and transferable. But serialization is also necessarily selective. The symbol is not the entire meaning that produced it.

A person who writes KahiliGinger→invasiveIn→Hawaii may have years of experience behind those few symbols: memories of hiking through dense stands of ginger, knowledge of native plants, conversations with conservationists, photographs, smells, places, consequences, and countless relationships that were never encoded into the triple. The KG preserves what we chose to make explicit. It does not somehow contain the entire mental world from which that assertion emerged.

This is also why I have become interested in what I call rich properties. Some knowledge is better preserved in its original form—a photograph, schematic, recording, CAD model, workflow, mathematical model, or simulation—rather than flattened into thousands of triples. The KG can identify the artifact, connect it to entities, record provenance, and expose the pieces needed for querying and reasoning without pretending that every bit of meaning has to be converted into ontology.

That starts to look naturally neuro-symbolic. A neural system is good at interpreting patterns that are difficult to specify completely in advance; a KG is good at giving selected meanings explicit identity, relationships, provenance, and structure. Neither representation needs to swallow the other.

That distinction is central to what follows. A KG is a serialized representation of selected knowledge. A human brain or LLM contains something much richer, but unfortunately, largely unreadable from the outside. Communication between them requires repeatedly crossing that boundary: meaning becomes symbols, and symbols become meaning again.

Note that “symbols” for brains include things like words, whereas they are “tokens” for LLMs.

Serialization and Deserialization of Human Conversation

In a nutshell, this article is about describing how highly intelligent LLMs (and I’d say LLM-centric AI is highly intelligent today, even if we’re only arguably near AGI) are largely subsymbolic and interact with each other through symbolic interfaces. That starts with person to person, extends to person to AI, and even AI to AI.

A little digression into more intuition about the meaning of “subsymbolic”:

The subsymbolic world of a brain or LLM is like a dreamworld Library of Congress. The books have been dissolved into the structure of the place itself. Ideas that repeatedly appeared together have changed its terrain; paths have formed between some regions, while others remain distant. A symbol entering that world does not retrieve a book from a shelf. It sets the whole learned landscape into motion. What emerges is a new internal state shaped by everything the system has learned.

Serialization then performs the remarkable opposite trick: it compresses that enormously rich, high-dimensional state back into a tiny sequence of symbols that can leave the system and enter another one.

Symbols travel. Meaning doesn’t. Meaning has to be reconstructed inside the dreamworld of the receiver.

Figure 1 illustrates two people, each with brains wired from different sets of experiences. But the experiences are shaped through varying mixes and degrees of cultural contexts—times of growing up, countries, socio-economic group, different sets of experiences, where every experience represents a “flap of a butterfly’s wing”, leading to unpredictable future states.

Those experiences continuously shape the brain through perception and learning, in a form we still do not know how to directly decipher. When people communicate portions of that internal world, thoughts are serialized into sequences of movements—speech, writing, gestures, and expressions—which another person must perceive and deserialize.

The symbols expressed by the first person are sensed and deserialized by other people into their brains. But because the brains of the 2nd party(ies) are crafted from different experiences, there is going to be some level of different interpretations, whatever translations are required to make it make sense in the current context, and reactions.

Figure 1 – Two people talking to each other.

Figure 2 is roughly the same as Figure 1, except the person on the left is having a conversation with an LLM—through a chat interface.

Figure 2 – How humans and LLM interact.

In an informal conversation, Figure 2 works very well. A person might want to learn about the history of Hawaiian-style plate lunches and their similarities with the American South’s “meat and three”. The conversation leans towards the characteristics that is casual, single topic, happens at the human’s convenience, proceeds at a comfortable pace, the subject is known, and is mostly inconsequential.

Communication at Enterprise Scale

At the hustle and bustle pace of our professional work lives, conversations are different.

  • Formal — because the answers may drive action. A casual answer can be interesting. An enterprise answer may affect money, customers, operations, compliance, or strategy. It needs to be clear, defensible, and based on information appropriate to the decision.
  • Of varying topics — because enterprise work crosses domains. A single problem may involve customers, products, finance, inventory, logistics, contracts, geography, and time. The conversation can move quickly from one subject to another, often requiring relationships among them rather than isolated facts.
  • About real-time problems — because the world does not wait for the conversation. A late shipment, declining KPI, production issue, pricing anomaly, or customer problem may already be unfolding. The information used in the discussion must be current enough to support action.
  • Data-intensive — because important answers often depend on more than language. The answer may require counts, trends, measures, comparisons, correlations, thresholds, and current operational state. Language alone is not enough; the conversation must reach into enterprise data.
  • Context-dependent — because the same words can mean different things in different parts of the business. “Customer,” “order,” “revenue,” “active,” or “risk” may have precise meanings that depend on domain, policy, system, geography, or time. The enterprise needs those meanings to be explicit.
  • Multi-user — because many people must work from the same understanding. A useful answer cannot depend entirely on what happens to be inside one person’s head or one LLM conversation. Important knowledge must be shareable, inspectable, and reusable across people and systems.
  • Persistent — because the conversation does not end when the chat window closes. Decisions, definitions, relationships, and conclusions may need to be used again tomorrow, next quarter, or by someone who was not part of the original exchange.
  • Auditable — because people may later ask why an answer was given. It may be necessary to trace an assertion to a source, a calculation, a rule, a document, or a person. “The model said so” is not sufficient provenance.
  • Governed — because not every person or system should see or change everything. Enterprise knowledge has ownership, security, privacy, versioning, and authorization requirements. Those controls must survive beyond an individual conversation.

This is where the human-to-LLM picture of Figure 2 becomes insufficient by itself. A human can serialize some portion of what is in their head into words, gestures, documents, or diagrams. An LLM can deserialize those symbols into its own internal representation and respond. That is enough for a great deal of useful conversation.

But an enterprise cannot depend on every important fact being freshly serialized by a human every time it is needed. There is only finite time and we can’t be in more than one place at a time.

People therefore encode what they know into more persistent forms: documents, rules, schemas, databases, semantic models, and knowledge graphs. This lets them be in many places at the same time—at least salient aspects of them. A KG is especially interesting because it preserves not merely facts, but explicit relationships among things, in a succinct manner. It is a serialized form of knowledge that can exist outside both the human brain and the LLM for simultaneous consumption by many brains and LLMs.

That is the point of Figure 3. It depicts a person authoring a KG with the assistance of an LLM, which an LLM can used by other humans and LLMs. The human is still the source of meaning. But instead of repeatedly explaining that meaning from scratch, some of it is deliberately encoded into a structure that can be inspected, queried, governed, shared, and reused.

The KG does not replace the human’s internal understanding, nor does it capture even a small fraction of the knowledge of that person. It preserves a selected, explicit portion of it.

And it does not replace the LLM’s internalized knowledge either. It gives the LLM something its own learned representation does not naturally provide: a persistent, explicit, enterprise-specific account of what this organization means, knows, and currently believes to be true.

Figure 3 – Humans building a knowledge graph assisted by LLMs.

The symbiotic relationship between KGs and LLMs follow a similar pattern. A KG is a very good method for encoding and caching serialized, symbolic knowledge, from people purposefully authoring them. An LLM is very good at deserializing symbols back into meaning and formulating a response.

Figure 4 illustrates how an AI RAG process leverages a KG authored by a human SME.

Figure 4 – The knowledge graph as a stand-in, a “reasonable facsimile”, of human knowledge workers.

Symbols Are Not Meaning

Consider the word:

dog

Those three letters typed in the line above are not a dog. They do not bark. They do not have fur. They do not chase squirrels and clearly want to be petted as I encounter them on my walks. They do not make some people happy and other people nervous.

The word is a symbol. When a human reads it, something much larger happens. The symbol is deserialized by the brain into a network of associations built over a lifetime.

Depending on the person, dog may activate memories of a childhood pet, the shape of four-legged animals, barking, wolves, veterinarians, bites, loyalty, breeds, smells, emotions, or a thousand other things. The symbol is tiny. The flood of associations it activates is enormous.

This is what human language has always done. We serialize pieces of our internal worlds into symbols so they can be transmitted to other people. Other brains deserialize those symbols using everything they have previously learned. Of course, it’s possible there are people in the world that have no concept of a dog, or at least that particular three-letter symbol if they don’t speak English. In that case, other processes are activated (confusion, questions, etc.)

Writing is therefore a remarkable compression scheme. The sentence, “The dog waited by the door.“, contains very little information physically. But another human can construct an entire scene from it. That scene was never contained in the ink.

Knowledge Graphs Are More Succinct than Language

Language is highly flexible. In the complex world in which we live, there is so much new phenomenon that appears that we don’t have single word name for it. There are so many unique properties that we can’t fully describe it. So, language has many ways to say something, and we employ a process of clarification between the people in the conversation.

But if we were to draw a box around a set of “the world“, we could remove the vagueness of real language and express things much more compactly, succinctly, explicitly. Instead of saying:

Alice works for Acme.

we can represent something resembling:

Alice → worksFor → Acme

Now the components have identifiers. The relationship has an explicit meaning. Alice can be identified as a Person. Acme can be identified as an Organization. worksFor can be related to other properties in an ontology.

We can go considerably further. We can say where the assertion came from, when it was true, who made the assertion, what evidence supports it, and whether another source disagrees. That is powerful precisely because the meaning has been serialized into explicit structure.

A KG can preserve distinctions that otherwise disappear in prose, like blending into a big crowd.

It can distinguish each of these:

  • Person → isA → Employee
  • Reviewer → assesses → Person
  • Court → determined → LegalStatus
  • Commentator → believes → Proposition

The relationships are explicit.

  • We can inspect them.
  • We can govern them.
  • We can query them.
  • We can reason over them.
  • We can disagree about them without losing track of exactly what we are disagreeing about.

But there is also a weakness. The machine sees symbols. ex:Dog is still a symbol. rdfs:subClassOf is still a symbol.

An IRI may be globally unique, impeccably governed, linked to an ontology, and surrounded by thousands of perfectly constructed triples. But something still has to interpret all of that.

The Machine Symbol Deserializer

This is where I think LLMs become particularly interesting. An LLM has absorbed a vast amount of human symbolic output.

Words, sentences, descriptions, arguments, code, classifications, explanations, analogies, documentation, and innumerable relationships expressed indirectly through language have influenced its learned parameters.

So when an LLM encounters the symbol, dog, it does not merely process three characters. That symbol participates in a huge learned relational space.

Dogs relate to mammals, pets, wolves, barking, leashes, veterinarians, breeds, homes, loyalty, bites, animal shelters, and innumerable other concepts. Not as a neat collection of explicit triples. It may even relate to childhood since many people think of their childhood dogs, or sleepless nights and subsequent tough days at work from the neighbor’s barking dog.

The relationships are distributed, implicit, approximate, and enormously rich.

In that sense, the LLM acts as a machine symbol deserializer. Symbols are deserialized into a subsymbolic body. Give it a compact symbolic representation and it can expand that representation into a much larger field of learned relationships. That is something traditional software is generally terrible at.

Parsing Is Not Deserialization

I need to clarify parsing from what I am calling deserialization. Parsing identifies and structures the incoming signal. Deserialization is what happens when that parsed input is run through the enormously rich learned function that gives it meaning.

For a text-based LLM, the parsing side is relatively straightforward because we designed it. Text is broken into tokens. Those tokens are converted into numeric representations and presented to the neural network. Multimodal models complicate this with images, audio, and other inputs, but those interfaces are still engineered systems whose mechanisms we largely understand.

The difficult part begins after that.

A token does not map to one neatly stored meaning. It interacts with the surrounding context and with billions of learned parameters. The result is a distributed internal state reflecting many possible associations, relationships, expectations, and interpretations.

That is what I mean here by deserialization.

It is not:

symbol → definition

It is more like:

parsed symbols → massively learned function → internalized meaning

Humans have an even more complicated front end.

There is no obvious human equivalent of a tokenizer. We hear words while simultaneously seeing facial expressions, gestures, posture, objects, and surroundings. Tone, timing, memory, expectation, emotion, and prior experience can all affect what we think we heard or saw. Even the distinction between perception and interpretation is not especially clean.

So for a person, parsing itself is multimodal and deeply intertwined with deserialization.

A spoken sentence may arrive together with a raised eyebrow, a hand pointing toward an object, a pause before a particular word, and years of shared history between the people speaking. The brain must somehow integrate all of those inputs into a usable internal state.

That makes the contrast roughly:

LLM: engineered parsing → subsymbolic interpretation
Human: multimodal biological perception ↔ subsymbolic interpretation

The double arrow in the human case is intentional. Parsing and meaning do not appear to be cleanly separable stages.

LLM Deserialization Is Also Iterative

There is another complication on the LLM side. An autoregressive LLM does not deserialize a prompt once, construct an entire answer internally, and then serialize that completed answer.

It works incrementally.

Conceptually, the loop is:

serialized tokens → internal state → next-token probabilities → selected token → expanded serialized context → revised internal state → next token → …

Each generated token becomes part of the serialized context used to produce the next one.

Implementations normally reuse previous computation rather than recomputing everything from scratch, but conceptually the important point remains: serialization and deserialization repeatedly feed each other during generation.

This helps explain why an LLM response can evolve as it is produced. A token selected early in the answer changes the context from which later tokens are generated. Two otherwise identical runs that select different tokens can therefore begin to diverge.

So deserialization is not merely the decoding of symbols. It is an ongoing interaction between the serialized stream and a massively learned subsymbolic function.

And the same basic distinction exists in human conversation. We do not simply parse a sentence and stop. Each new word, expression, gesture, and response changes the context in which the next signal will be interpreted.

The conversation itself becomes part of the deserializer.

Why the Knowledge Graph and LLM Need Each Other

The strengths and weaknesses line up almost suspiciously well.

A KG can say exactly: KahiliGinger → invasiveIn → Hawaii

An LLM may know an enormous amount about kahili ginger, invasive plants, Hawaii, ecosystems, gardening, land management, and related species. But the LLM’s knowledge is not necessarily explicit, current, attributable, or even correct.

The KG can provide the explicit assertion. The LLM can deserialize it into a much broader context.

  • The KG says: Here is what we know, expressed precisely.
  • The LLM says: Here is what these symbols may mean in the larger world I have learned.

That makes the relationship much richer than the familiar statement that a KG can “ground” an LLM. Grounding is part of it. But grounding describes only one direction. The graph constrains the LLM with explicit knowledge.

The LLM also expands the meaning of the graph. A triple that is trivial to a graph engine may become the starting point for explanation, analogy, hypothesis generation, classification, integration, or another query when interpreted by an LLM.

The relationship is symbiotic.

Humans Already Do This

Humans, more specifically, our brain, are also symbol deserializers. But humans have one enormous advantage. We learn inside a live, physical, fully immersive world.

Suppose someone tells a child: The stove is hot.

Those are symbols. But the child may also see the glowing burner, feel heat radiating from it, watch steam rise from a pan, see a parent pull a hand away, and perhaps make the unfortunate decision to touch something hot.

The statement becomes integrated with perception, action, memory, consequence, emotion, and the behavior of other people. The world tests the model immediately. That is very different from learning only from serialized descriptions of somebody else’s experience.

Humans receive far less information than modern machine-learning systems can process. But our information arrives through a remarkable integration mechanism: vision, sound, touch, proprioception, language, action, social interaction, memory, prediction, and consequence are continuously reconciled against the same world.

The human brain spends decades constructing that integrated model. An LLM can train far faster and absorb vastly more serialized human experience. But most of that experience has already passed through somebody else’s brain before reaching it. It is largely experience that has already been serialized.

Not Everything Is a Symbol

There is another complication. Much of what intelligence works with is not merely symbolic.

Consider:

Revenue = $14.7 million

The characters $14.7 million are symbols.

But what they represent is a measure. Fourteen million is greater than seven million. The difference between 14 and 7 has meaning.

We can add them, compare them, divide them, aggregate them, calculate rates of change, test thresholds, and determine whether something is unusually high or low. That is fundamentally different from the relationship between arbitrary labels such as Customer and Supplier.

So there is another representational world:

measures.

Business intelligence has lived largely in this world for decades.

Then there are vectors. Vectors consist of measures, but collectively they describe location in a learned relational space. They can tell us that two things are similar without explicitly telling us the relationship between them.

That is extraordinarily useful. But similarity is not the same thing as integration.

Knowing that two passages are close in vector space does not tell us:

this customer owns that account,

or:

this component caused that failure,

or:

this policy superseded that policy on January 1.

Those explicit relationships belong to a different representational mechanism.

And beneath an LLM are weights. Weights are different again. A weight is not a fact in anything resembling the sense of an RDF triple.

No particular weight means: Boise is in Idaho.

Instead, billions or trillions of learned values collectively determine how activation propagates through the model. They encode dispositions, patterns, and learned tendencies.

What emerges from those weights is an enormous implicit relational capability. So perhaps we should stop trying to force all machine knowledge into one representational category.

We have at least:

  • Symbols — explicit things and relationships.
  • Measures — quantities whose magnitudes and mathematical relationships matter.
  • Vectors — positions in learned spaces of similarity and association.
  • Weights — learned parameters that determine how a neural system responds.

Each does something the others do poorly.

The Interesting System Is the One That Uses All of Them

This is where the discussion becomes architectural. The future intelligent enterprise is probably not a giant knowledge graph. It is probably not an LLM with everything stuffed into its context window. It is not a vector database. And it certainly isn’t just another dashboard.

The interesting system is the one that moves intelligently among representations.

  • A semantic layer may efficiently produce measures.
  • A KG can explicitly connect those measures to products, customers, business processes, organizational concepts, policies, external knowledge, and other domains.
  • Vector representations can discover things that resemble each other even when nobody explicitly modeled the relationship.
  • Machine-learning models can recognize patterns and make predictions.
  • Functions can transform one representation into another.

And the LLM can sit among them, repeatedly converting compact symbolic representations into a much larger implicit context and then serializing useful pieces of that context back into language, queries, hypotheses, classifications, or new candidate relationships.

Seen this way, the LLM is not the knowledge base. It is not the ontology. And it is not a replacement for the semantic layer.

It is something different. It is an extraordinarily capable deserializer of human symbols.

Serialization Goes Both Ways

Of course, the LLMs do serialize too. Ask it:

  • a question and it converts its internal state back into words.
  • to extract entities and relationships and it may produce triples.
  • to write SQL and it translates an intention into another symbolic language.
  • to propose an ontology and it can turn implicit associations into explicit classes and relationships.

So the process really forms a loop:

explicit symbols → implicit relational meaning → new explicit symbols

The KG is particularly valuable on the explicit side of that loop. The LLM is particularly valuable on the implicit side. And the interface between the two may be more important than either technology considered by itself.

Perhaps That Is the Real Symbiosis

There has been a tendency to ask whether KGs still matter now that we have LLMs. I think that asks the wrong question. Brains did not make language unnecessary. Symbolic language is an interface. Like with any software functions, domain to domain handoffs, etc., communication needs to be controlled.

Language became valuable precisely because brains existed to interpret it. Books did not need to contain everything a reader knows. They only needed enough symbols to activate the appropriate structures already inside the reader.

Perhaps a knowledge graph plays a similar role for machine intelligence. It does not need to encode everything an LLM knows. That would be absurd, as you’d know if you tried encoding even one domain in a KG.

It needs to encode the things we want to make explicit: identities, relationships, definitions, provenance, constraints, measurements, rules, and assertions that matter to us.

The LLM supplies the enormous implicit context surrounding those symbols. The KG supplies the explicit structure the LLM lacks. One is not the primitive version of the other. They are different representations with different strengths.

And once I started thinking of the relationship in terms of serialization and deserialization, the reason they belong together became much clearer:

  • Knowledge graphs serialize knowledge so machines and humans can inspect it.
  • LLMs deserialize those symbols into a much larger learned world of relationships.

The interesting intelligence emerges when we keep moving between the two.

The Part We Still Cannot Read

There is a last interesting similarity—and an important difference—between a human brain and an LLM. In both cases, the internalized knowledge is much harder to decipher than the physical machinery that contains it. But the interfaces are different. We engineered the LLM’s path from symbols into the network and back out again. The human brain evolved its own multimodal serializer and deserializer, and those processes remain part of what we are still trying to understand.

For a brain, we can trace neurons and synapses, measure electrical activity, and increasingly map connections. For an LLM, we can inspect every parameter, activation, layer, and weight. Yet we won’t pinpoint neurons/synapses or weights that explicitly hold the information: dog → isA → mammal

The knowledge is distributed through the system. We can examine the pieces, but we usually cannot look at a particular synapse or weight and say, “That is where this fact is stored.” In that sense, both brains and LLMs contain knowledge in an internalized, largely undecipherable form.

There is, however, an important difference. With an LLM, we built the machinery that crosses the boundary between symbols and that internal form. We know how text is broken into tokens, how those tokens are converted into numeric representations, how they pass through the neural network, and how the resulting state is converted into probabilities over possible output tokens. We may not understand what all of the internal patterns mean, but we understand the mechanism that gets symbols in and symbols back out.

With a human brain, the serializer and deserializer are themselves part of the mystery. A person hears words, sees gestures, reads writing, and somehow turns those signals into an internal state of understanding. In the other direction, an internal thought somehow becomes a sentence, facial expression, drawing, or gesture. Neuroscience understands pieces of these processes, but we certainly did not design them, and we do not yet have anything approaching a complete description of how meaning becomes symbols and symbols become meaning.

So there are really two mysteries in the human case:

  1. How is knowledge internally represented?
  2. How does the brain translate between that representation and external symbols?

For an LLM, the first mystery remains substantial. The second is largely engineered.

Short Term Memory and Chat Sessions

There is another useful distinction: learning versus immediate interpretation.

During LLM training, enormous amounts of serialized material are processed and gradually alter the model’s weights. That is the long-term form of learning. During an individual conversation, the model is not normally rewriting those weights. Instead, the current prompt and conversation create temporary internal activations and context that influence what happens next.

Humans have a rough analogue. Years of experience alter long-term memory and neural structure, while the current conversation, what happened five minutes ago, and what we are presently paying attention to form a much more temporary working context.

So both systems have something resembling long-term internalization + short-term context, although the biological and artificial mechanisms are very different.

Exact LLMs and Exact Brains Doesn’t Result in the Same Answer

One final complication is that serialization is not necessarily a perfectly reversible process.

Suppose two identical copies of an LLM begin with the same learned weights and receive the same conversation. Under ordinary probabilistic generation, they can produce different next words and therefore quickly diverge into different conversations. The model produces a distribution of possible next tokens; the generation process may select among them probabilistically.

That does not mean the internal knowledge itself randomly changes each time. And if generation is configured deterministically and the computation itself is deterministic, identical models should produce identical results. The variability comes largely from how one possible symbolic output is selected from the model’s possible outputs.

With humans, no brains are exactly alike in the way we can work with exact copies of LLMs. Two people can hear exactly the same sentence and construct somewhat different meanings from it because they bring different lifetimes of experience to the deserialization process.

That may be the deepest reason serialization matters:

The serialized message is not the meaning. It is a compact signal from which another intelligence reconstructs meaning using what it already contains.

A KG is unusual because it deliberately keeps much more of that knowledge outside the intelligence, in an explicit serialized form that other humans and machines can inspect directly. That is precisely why its symbolic information complements both subsymbolic brains and LLMs so well.

LLM Generated Knowledge Graphs

A KG is not inherently a human serialization format. It is simply one possible external, explicit serialization of an internal state.

A human can go:

subsymbolic/internalized knowledge → serialize → speech, writing, diagram, RDF, knowledge graph

An LLM can do essentially the same thing:

subsymbolic/internalized state → serialize → prose, JSON, SQL, RDF, knowledge graph

Figure 5 illustrates how even LLMs can benefit from externalizing serialized information for the benefit of other LLMs.

Figure 5 – LLMs interacting through their respective, explicitly expressed beliefs and assertions.

Both are serializations. The KG version simply chooses a much more explicit structure.

The LLM is not somehow “moving knowledge into a graph”. It is serializing a selected portion of its internal state into KG form, exactly as the human is serializing a selected portion of their internal understanding into the same graph.

The crucial difference is epistemological, not representational. An LLM can produce a perfectly valid-looking triple that is wrong. Believe it or not, a human can too. So producing the serialization does not establish its truth. Provenance, evidence, review, confidence, rules, and human scrutiny remain separate questions.

That gives us a clean distinction:

Serialization answers: “How do we make some of this internal knowledge explicit?”
Epistemology answers: “Why should anyone believe what we just made explicit?”

And now the overall architecture becomes wonderfully symmetric:

Human ↔ symbols ↔ LLM

while either side can also produce:

internalized knowledge → serialized knowledge graph

And the KG can then become input again:

serialized knowledge graph → human or LLM → new internalized meaning

So the KG isn’t really a third kind of intelligence in this picture. It is a persistent serialized medium through which intelligences can externalize, inspect, exchange, preserve, challenge, and reconstruct portions of what they know.

Communication Is an Iterative Synchronization

The main theme of this blog is: Meaning is not transmitted directly. What moves between two intelligences is a serialization of meaning: words, gestures, writing, images, triples, diagrams, or some other external representation. The receiving intelligence must reconstruct meaning from that serialization using its own internalized knowledge, history, context, and expectations.

That reconstruction is necessarily imperfect.

When I say something to another person, I do not transfer the state of my brain into theirs. I reduce some portion of what I mean into speech, expression, gesture, or writing. It’s made in the context of my assumptions about what the other person knows and how I think he’s currently interpreting our conversation. The other person then rebuilds meaning using a completely different brain shaped by a completely different lifetime of experience.

It would actually be surprising if that worked perfectly on the first try. Instead, it’s going to be an iterative process. That is why human conversation contains so much repair:

  • Person 2: “No, that’s not quite what I meant.”
  • Person 1: “Oh, I thought you meant…”
  • Person 2: “Yes, but only in this case.”

We repeatedly serialize, deserialize, compare, correct, and try again. Over the course of a conversation, our internal models become sufficiently synchronized for the purpose at hand.

Not identical. Just synchronized enough.

The same issue appears when humans communicate with LLMs. A prompt does not somehow inject my intended meaning directly into an LLM. The model receives symbols, converts them into its own internal state, and reconstructs what those symbols probably mean using what it learned during training and what is currently present in the conversation.

It then serializes some portion of that internal state back into tokens for me. I deserialize those tokens into my subsymbolic brain. And the cycle begins again.

This helps explain why misunderstandings are not an unusual failure of either humans or LLMs. They are almost an expected consequence of communication itself. In a massively complex world such as the one we live in, even an ASI will not always “get it” on the first try.

Serialization is lossy. The receiver must reconstruct what was never fully transmitted.

When Reconstruction Turns Into Imagination

This also offers an intuitive way to think about so-called LLM hallucination.

One of an LLM’s great strengths is its ability to fill in what was not explicitly stated. That is essential for explanation, analogy, deduction, summarization, creativity, and ordinary conversation. But the same capability can fill in something that is not true.

When the available evidence is strong, the model may reconstruct a very useful interpretation. When the evidence is ambiguous, incomplete, or weak, it may construct something merely plausible. In ordinary language, this begins to resemble imagination.

The problem becomes more interesting because LLM generation is iterative. The model does not construct an entire answer internally and then simply print it. It generates a token, adds that token to the growing serialized context, and then uses the expanded context to generate the next token.

So if the model introduces an unsupported idea early in the response, that invention becomes part of the context from which subsequent meaning is reconstructed.

In other words: an invented serialization can become the premise for the next round of deserialization.

The model can then continue very coherently from something that was never grounded in the first place. That is one route by which a small misunderstanding can grow into a convincing hallucination.

Humans are not entirely foreign to this process either. We mishear things. We infer missing details. We remember imperfectly. We construct explanations. We occasionally become increasingly confident in something that began as an assumption. Conversation repairs much of this through continued synchronization as mentioned above.

Reality repairs even more. We ask again. We check a document. We look at the data. We ask another person. We compare what we think with what actually happened.

The Knowledge Graph as a Reference Point

This gives the KG another role in the symbiosis between KGs and LLMs. A KG is not merely a collection of facts. It is a persistent serialized reference point.

Human and LLM internal states are enormously rich, dynamic, and largely inaccessible from the outside. The KG preserves selected knowledge in an explicit form that can be inspected, shared, queried, challenged, versioned, and revisited.

That gives repeated acts of deserialization something stable to return to. Without such a reference point, communication can drift:

internal meaning → lossy serialization → mistaken reconstruction → generated assumption → assumption becomes new context → further drift

With explicit grounding, we gain another possibility:

internal meaning → serialization → reconstruction → check against explicit knowledge → correction → better synchronization

That may be the most important point of all. Humans, LLMs, and KG do not contain knowledge in the same form.

  1. Humans and LLMs operate through rich internalized, subsymbolic states.
  2. Knowledge graphs preserve selected knowledge in serialized, explicit form.
  3. Communication is the repeated crossing of that boundary.

And because every crossing is imperfect, intelligence depends not only on knowing things, but on continually synchronizing what we think we know.

Leave a comment