The Elusive Quest to Measure Neural Similarity: Beyond Numbers to Understanding
What does it mean for two brains—or even a brain and an AI model—to be alike? It’s a question that sounds deceptively simple but quickly unravels into a labyrinth of complexity. As someone who’s spent years grappling with this, I can tell you: we’re not just talking about comparing apples to oranges; we’re trying to compare apples to abstract concepts, and the measuring tape keeps changing.
The Comparative Impulse: From Darwin to Neurons
Let’s start with the big picture. Comparative analysis is the backbone of biology. Darwin didn’t just observe differences between species; he used them to build a theory of evolution. What’s fascinating here is how this principle has trickled down to neuroscience. We’ve long compared brain structures across species—think of the hippocampus and its link to spatial navigation. But now, with advances in recording technology, we’re diving deeper, comparing neural activity across populations of neurons.
What makes this particularly fascinating is the shift in scale. We’re no longer just comparing organs or species; we’re comparing the intricate firing patterns of neurons. It’s like moving from studying entire cities to analyzing the conversations happening in every coffee shop. But here’s the catch: we’re still figuring out what it means for these conversations to be ‘similar.’
The AI Twist: When Biology Meets Silicon
And then there’s AI. Modern artificial neural networks are often compared to biological brains, but the parallels are messy. Sure, both rely on distributed computation, but the differences—spike-based vs. analog communication, for instance—are glaring. This raises a deeper question: Can we even use the same metrics to compare biological and artificial systems?
High-profile projects like Brain-Score and the Algonauts Project are trying to bridge this gap, but it’s like comparing a symphony orchestra to a rock band. Both make music, but the instruments, rhythms, and goals are wildly different. Personally, I think this is where the real challenge lies: not in finding similarities, but in understanding what those similarities mean.
The Metric Maze: Too Many Tools, Too Little Clarity
Here’s where things get messy. The computational literature is flooded with methods to quantify neural similarity—geometric approaches like RSA and CKA, predictive models, and hybrid methods. A recent review counted over 30 such methods. On one hand, this is a treasure trove of tools. On the other, it’s a recipe for confusion.
One thing that immediately stands out is how many of these methods are closely related, often collapsing into a few underlying frameworks. For instance, RSA and CKA are essentially equivalent with a slight tweak. What many people don’t realize is that this redundancy isn’t a flaw—it’s a clue. It suggests that we’re circling around a few core principles, even if we haven’t fully articulated them yet.
But here’s the rub: predictivity and similarity are often conflated. A model might predict neural activity well, but that doesn’t mean it’s similar to the brain. Predictive accuracy is asymmetric; similarity, ideally, should be symmetric. This distinction is crucial, yet it’s often overlooked. If you take a step back and think about it, we’re not just measuring numbers; we’re trying to map entire systems of thought and computation.
Metrics vs. Maps: The Difference That Matters
This brings me to a detail that I find especially interesting: the difference between a similarity score and a proper metric. A score is just a number; a metric is a map. When a measure obeys the triangle inequality—meaning it’s symmetric and coherent—it becomes a tool for navigation. We can embed brain regions or AI models into a shared space, cluster them, and explore relationships in a meaningful way.
But here’s the kicker: brains are too complex for a one-size-fits-all metric. What this really suggests is that we need a toolkit, not a single tool. Neuroscientists should report multiple metrics to capture the richness of neural computation. It’s hard work, but it’s the only way to move beyond surface-level comparisons to deeper mechanistic understanding.
The Leaderboard Trap: Why Scores Aren’t Enough
This leads to a broader critique of how we approach these problems. The field has a habit of ranking models on leaderboards, as if the highest score wins. But what does that score mean? In the kidney example, medullary thickness mattered because it pointed to a mechanism—countercurrent multiplication. Similarly, neural similarity scores only matter if they lead us to computational principles.
What this really suggests is that we’ve been prioritizing novelty over understanding. New metrics are minted constantly, but many are just incremental tweaks. We need to step back and ask: What are we missing? Are there aspects of neural computation that none of these metrics capture?
Where Do We Go From Here?
Personally, I’m optimistic. The fact that we’re having these conversations means we’re on the right track. We need to refine existing metrics, unify fragmented approaches, and develop new tools that capture overlooked aspects of neural computation.
But more importantly, we need to shift our mindset. Comparative analysis isn’t about finding the ‘best’ metric; it’s about building a deeper understanding of how brains—and AI—work. It’s about asking the right questions, not just getting the right answers.
So, the next time someone asks if two brains are alike, I’ll say: It depends on what you’re measuring—and why. Because in the end, similarity isn’t just a number; it’s a window into the mechanisms that make us who we are.