Why Do Neural Networks Become Modular?
Reflections on Our Latest Nature Machine Intelligence Paper
Years later, when both our research abilities and our interests have changed considerably, we often find ourselves returning to the questions that troubled us when we first entered the field.
I began working in network neuroscience around 2014. At the time, it was an exciting interdisciplinary field: researchers were eager to develop and test new network measures, relate them to behavioral performance, and use those relationships to explain cognitive mechanisms. Studies had already revealed statistical associations between the modular organization of brain networks, their dynamic reconfiguration, and learning performance.[1]
These findings were exciting, but there was an “elephant in the room” that I could never quite set aside. How much of these seemingly reasonable mechanistic explanations came from an understanding of the actual computations, and how much was a story added to a statistical association? Could some of the associations themselves even be “statistical windfalls,” found after repeatedly trying different measures and analytical paths?
For me, retrospective explanations of observed network structures in terms of information transmission efficiency, connection costs, or functional specialization could not fully answer the question I cared about most: Under what conditions is modularity useful, and why does it emerge during learning?
Later work on artificial neural networks deepened this question. For example, Saining Xie and colleagues showed that connectivity generated from random graphs could also achieve competitive performance on image recognition tasks.[2] This made me even more curious: how do the topological features so often discussed in biological networks relate to specific computational demands and learning processes?
To answer these questions, we needed more than additional correlations. We needed an experimental platform that would let us systematically vary conditions, track learning, and test hypotheses about mechanisms.
About three years ago, we began exploring this direction by extending approaches from the study of biological and social networks to artificial neural networks. This offered two immediate advantages. First, we could finally experiment relatively freely: change the tasks, adjust network capacity, intervene on connections, and observe what happened. Second, we could distinguish more clearly between network properties that emerge gradually during training and those already built into the architecture or initialization.
We pursued two studies along these lines. In our 2024 paper in Science Advances, we used continual familiarity detection to investigate how a task’s information-encoding demands relate to the formation and reconfiguration of network modules, as well as learning performance.[3]
In our latest paper in Nature Machine Intelligence, we asked a further question: when a network has limited representational capacity but needs to learn multiple tasks, how do those demands shape its structure? Our experiments showed that multitask learning can promote modularity, particularly when the task load puts pressure on network capacity. Progressively adding connections during training can further strengthen this modular organization.[4]
For me, the most important connection between these two studies is that they move the question forward: from “What kinds of network structure are associated with better performance?” to “What kinds of learning demands give rise to those structures?”
We chose relatively simple models so that the key factors could be isolated, controlled, and tested, and we carried out mathematical analyses under explicit assumptions. This line of inquiry is not limited to small models. Recent work has also begun to identify modular organization associated with different cognitive tasks in large language models.[5] This gives us an opportunity to continue investigating the relationship between network organization and intelligence across artificial systems of different scales and with different tasks.
Of course, cognition involves control as well as representation: which information is routed, which is stored, which is suppressed, and which is ultimately used to guide action. In my view, this is also an aspect of current analogies between LLMs and memory in the brain that deserves closer examination. Understanding what a system “remembers” is not enough to explain how it selects and uses those memories in light of its current goals. These are among the questions we hope to explore next.
What I wanted to share here is the origin of this research and the thinking that developed along the way. For the specific models, experiments, and results, please see our WeChat article. For me, this paper is also another attempt to address a question from more than a decade ago: to bring network dynamics and network representations together within a model, rather than relying on conjecture and interpretation.
Finally, I would like to thank Yuhang Wu, Shikuang Deng, Kangrui Du, and the other students for their careful experimental analyses, and all our collaborators for their contributions and support.
State Key Laboratory of Brain Machine Intelligence on WeChat: Read the paper overview (in Chinese).