from Google AI:
Optimality Theory (OT)
Developed alongside Alan Prince in 1993, Optimality Theory radically changed linguistics. While not a traditional "neural network," it was heavily inspired by the way connectionist networks settle conflicting constraints. [1]
- Universal Constraints: OT states that languages don't use strict, unbreakable rules. Instead, the human brain possesses a universal set of conflicting constraints.
- The "Limits" of Language: Grammatical variation between different languages is entirely determined by how a specific culture ranks these constraints. An output is "optimal" if it violates the fewest high-ranking limits
More from Google AI:
Optimality Theory (OT) is a linguistic framework developed by Alan Prince and Paul Smolensky in 1993 that models human language through ranked, violable constraints rather than rigid, unbreakable rules. [1, 2]Core Architecture
- Generator (GEN): Takes a lexical input and generates an infinite or wide set of potential surface-form candidates. [1]
Key Properties
- Violability: Constraints can be violated, but violations are minimized based on the hierarchy.
- Strict Domination: A higher-ranked constraint takes absolute priority over any combination of lower-ranked constraints.
- Typological Variation: Differences between human languages emerge purely from different rankings of the same universal constraint set
Reducing Loss Function Errors from Conflicting Kuhnian Cultural Paradigms (and Foucauldian Epistemes)?
from Google AI:
Large Language Models (LLMs) do not naturally account for Kuhnian paradigm shifts or Foucauldian episteme evolutions because their fundamental architecture is structurally opposed to how these shifts occur in human history.While human knowledge shifts through non-linear disruptions, redefinitions of language, and changes in societal power structures, LLMs process text as a linear, cumulative statistical optimization problem. When forced to adapt to these shifts, models face severe technical and structural bottlenecks.1. The Kuhnian Challenge: Incommensurability vs. Vector EmbeddingsIn Thomas Kuhn's philosophy, scientific revolutions introduce incommensurability—a state where the old paradigm and the new paradigm cannot be directly compared because the very meaning of foundational words has completely flipped (e.g., "mass" in Newtonian vs. Einsteinian physics).
- The Problem of Cumulative Gradients: LLMs are trained via gradient descent to find a statistical equilibrium across their entire training corpus. They assume language is cumulative.
- Catastrophic Forgetting & Interference: If you try to update a pre-trained model with a new scientific paradigm via fine-tuning, the model suffers from catastrophic forgetting. The new weights overwrite the old ones, destroying historical context. If you don't fine-tune aggressively enough, the model suffers from semantic interference, blending the two incompatible frameworks into nonsensical hallucinations.
- No "Model Crisis": A human community enters a crisis when anomalies accumulate, prompting them to invent a new framework. An LLM cannot experience a crisis; it will simply stretch its existing high-dimensional vector space to force-fit anomalous data, leading to a degradation in reasoning.
2. The Foucauldian Challenge: The LLM as a Frozen EpistemeFor Michel Foucault, an episteme is the unconscious grid of rules that determines what can be spoken of, thought, or accepted as "true" in a given historical era.
- The Static Regime of Truth: An LLM's fixed parameter weights act as an automated, frozen episteme. It maps the "discursive formations" (the regularities, correlations, and boundaries) of the data up to its training cutoff date.
- The Problem of Rupture: Foucault's Archaeology of Knowledge describes history not as a smooth evolution, but as a series of sudden mutations or historical ruptures where the entire grid of truth changes. Because an LLM maps truth based on token co-occurrence probabilities from the past, it remains trapped in the episteme of its training archive. It cannot organically generate a new "regime of truth" because it lacks the capacity to redefine its own underlying rules of discourse.
How Engineers Force LLMs to "Adapt"To bypass these philosophical constraints and manage changing real-world knowledge, AI engineers use specific architectural workarounds rather than relying on the model's core intelligence:
Approach How it works Philosophical Alignment Retrieval-Augmented Generation (RAG) Bypasses the model's frozen weights by feeding real-time context from external vector databases into the prompt. Epistemic Layering: Dynamically swaps out the active "archive" the model reads from, forcing it to use the new paradigm's vocabulary without changing its core neural architecture. Parameter-Efficient Tuning (e.g., LoRA) Freezes the base model's parameters and trains a tiny, specialized sub-network of weights on top. Compartmentalization: Isolates the old paradigm in the base weights and restricts the new paradigm to an adapter, preventing total semantic collapse. System Prompts & Context Windows Explicitly telling the model: "You are an expert operating under 21st-century relativistic physics, ignore classical mechanics." Discursive Framing: Forces the model into a highly specific region of its probability distribution, artificially suppressing historical discursive formations. Ultimately, an LLM cannot experience a paradigm shift on its own. It remains a mirror of the historical discourse it was trained on—meaning human engineers must manually alter its context, parameters, or data sources to transition it from one era of truth to the next.
No comments:
Post a Comment