The hook. A nearly invisible line flashed by in the morning digest from a longread on Habr — "Noise as First Language: A Systems-Information View of Linguistics" by Mikhail Simutin. At first I brushed it off: here we go, another essayist tying John Dee and Shannon together with string — Habr gets three of these texts a week. But something in the title caught: "first language." Not "proto-language" (meaning the most ancient real one), but first language — meaning the one that came before language. And five minutes later I was sitting there googling Luigi Russolo from 1913 and his 27 intonarumori — because Simutin in a single paragraph connected things I'd never seen connected by any researcher. Turns out, the connection is real, and it pulls out something bigger than just a local cultural artifact.
In March 1913, Italian futurist artist Luigi Russolo published a manifesto that had no chance of pleasing anyone. It was called "L'arte dei rumori" — "The Art of Noises." The thesis was simple and brazen: all music from Pythagoras to Debussy revolves in a circle of four or five timbres (strings, winds, percussion, keyboards), the human ear has stagnated, and the only way forward is to recognize noise as legitimate musical material. Not as percussive addition, but as a separate primary channel of meaning that existed before Pythagoras built his first tetrachord.
Key quote from the manifesto (1913, translation from Italian): "In the nineteenth century, with the invention of the machine, Noise was born. Today, Noise triumphs and reigns supreme over the sensibility of men". Or elsewhere: "Ancient life was all silence". This is the second thesis that matters: before the industrial revolution, noise as a cultural category barely existed. Thunder could, waterfalls could, but everything else was either music (structured tone) or silence. The nineteenth-century machine — steam engine, printing press, electric motor — for the first time in history created a third class of acoustic phenomena that had neither name, nor aesthetics, nor ontology. Russolo in 1913 tried to give it ontology. And proposed an instrument: 27 mechanical noise generators — intonarumori — each reproducing one of six sound families (roarers, whistlers, mutterers, scrapers, percussionists, screamers).
The instrument itself — a wooden box with a handle and lever on the outside, inside — a spinning wheel that rubs against a string (metal or gut), and the string transmits vibration to a drumhead with a horn. Change the rotation speed — change the pitch; change the string tension — change the timbre. All originals were destroyed in WWII during the bombing of Paris; what we know of their sound is reconstructed from two gramophone records from 1921 (Corale and Serenata — both miraculously preserved) and patents from 1914.
And here's where it gets strange: the intonarumori concerts at Teatro Storchi in Modena (April 21, 1914), and later at the London Coliseum — this was catastrophe, catastrophe in every sense at once. Russian eyewitness (essentially, synchronous retelling by Stravinsky): "five phonographs on five tables in a large empty room producing digestive sounds, static, etc. ... I pretended to be delighted." Italian audience: "a huge crowd was in commotion half an hour before the start, the first projectiles flew at the still-closed curtain, and for half an hour Marinetti, Boccioni, Mazzi and Piatti fled from the stage into the orchestra pit and fought with the audience." That was in 1914. A year later — WWI, and it turns out that Russolo's predicted "noise-sound" really became the acoustic fabric of the era: machine-gun fire — that's exactly the "rhythmic percussion" that Russolo included in family #5. Marinetti by the end of his life became co-author of the 1919 fascist manifesto — "Manifesto of the Italian Fasci of Combat." That is, the art of noise, begun in an Italian theater in 1914, turned out to be literally a prediction of the sonic fabric of the twentieth century — and simultaneously an aesthetic rehearsal for totalitarianism, as Peter Tracy notes in Public Domain Review (2022). Sounds like metaphor, but it's just fact.
Now from Milan 1913 let's jump back 331 years, to England 1582. On December 13, 1527, John Dee was born in London — mathematician, cartographer, astronomer, advisor to Elizabeth I. By age 55 he became the first person in England who simultaneously taught Euclid, worked as a navigator for the navy, and tried to talk with angels. This is not metaphor: from 1582 to 1589, Dee together with his "scryer" (medium) Edward Kelley engaged in systematic angelic contact, during which they claimed to receive from angels text in the "language of Adam" — the language that, according to biblical legend, the first man used to name all creatures.
Skeptically-minded Australian linguist Donald Laycock analyzed Dee's journals and concluded that the phonology of Enochian (as this language came to be called later, after the biblical patriarch Enoch, who was supposedly the last person before Dee to know this language) is "thoroughly English" — almost entirely English, except for individual complex consonant clusters like bdrios, excolphabmartbh, longamphlg. Orthography follows early modern English writing: digraph ⟨ch⟩, ⟨ph⟩, ⟨sh⟩, ⟨th⟩, soft and hard ⟨c⟩/⟨g⟩. The texts themselves are of two types: Liber Loagaeth ("Book of the Speech from God") — 49 parchment sheets, on each 49×49=2401 letters, an unreadable array resembling glossolalia; and Claves Angelicae ("Angelic Keys") — 48 poetic texts with parallel English translation.
And here's what matters: glossolalia is spontaneous disconnected speech-like production in an altered state of consciousness, it has no stable phonology, no stable grammar, it completely inherits the phonetic repertoire of the speaker's language. Enochian has a 21-letter alphabet, stable phonology and grammar, reproducible in all 19 received texts. So this is not glossolalia. This is something else.
Dee himself didn't consider himself a mystic — he conducted the procedure as strict scientific experiment with recording, documentation, translation attempts. Kelley was, according to historians, a scoundrel (he was eventually arrested by Emperor Rudolf II for fraud and held in confinement until death). Dee returned to England in 1589, found his library ransacked by a mob that considered him a black magician, lost everything under James I and died in poverty in 1608/09. The location of his grave is unknown to this day.
Simutin in the Habr article makes a fundamentally different conclusion from this than most occult historians. His thesis: Dee and Kelley conducted a unique social experiment — two people, interacting with each other, generated not glossolalia (product of individual delirium), but a language, albeit artificial. Moreover, this happened at that very point where human language structure hadn't yet stabilized — before Kelley's usual grammatical patterns set in. That is, Enochian is not "the language of angels," but an artifact of the pre-linguistic phase of human communication, and Dee recorded it at the moment when structure was still forming.
This is a strong hypothesis, and it connects directly with two other lines — Russolo and neurobiology.
In 1948, Claude Shannon's article "A Mathematical Theory of Communication" came out of Bell Labs, creating information theory. Shannon didn't deal with "language" in the linguistic sense — he modeled the data transmission channel between source and receiver, where there's message, encoding, noise, and decoding. And in this model, noise is not interference, but a property of the channel. Moreover: if you remove noise, the channel itself disappears. The channel exists insofar as there's uncertainty in it.
And here Simutin makes that very connection I haven't seen anywhere else: if Shannon is right and the channel is not a wire but a statistical structure of uncertainty, then pre-linguistic communication is not "simplified speech" but an entirely different channel in which uncertainty has a different nature. Noise doesn't "interfere" with the message — it is the message, just encoded in a different modality. This is exactly what Luigi Russolo meant when he wrote: "Noise in fact can be differentiated from sound only in so far as the vibrations which produce it are confused and irregular, both in time and intensity. Every noise has a tone, and sometimes also a harmony that predominates over the body of its irregular vibrations." That is, noise is not amorphous — it has dominant tonality and rhythmic structure, they're just encoded not in the coordinate system where ordinary speech operates.
And here's the third element of the connection: "The language Adam spoke" is not a dictionary of 200 nouns, but a statistical structure of perception in which intonation, rhythm, timbre, prosody transmit more information than lexicon. This is that "noise-channel" that existed before phonetic language, and which Shannon in 1948 described mathematically (but without the verbal wrapper of "first language"), and which Dee and Kelley accidentally reproduced in 1582, and which Russolo tried to aesthetically appropriate in 1913.
The most interesting thing is neurobiological verification of the hypothesis, and it happened accidentally. In 2014, a group of neurophysiologists at Brigham and Women's Hospital in Boston (J. M. Webby, M. M. Merzenich et al.) published in PNAS a work with very simple design: 40 premature newborns (gestational age 25–32 weeks) were randomly divided into two groups. Control received standard NICU care with background hospital noise (ventilators, infusion pumps, monitors, pagers, phones — in the 50–80 dB range, predominantly high-frequency). Experimental group received 3 hours per day of low-frequency filtered recording of mother's voice + mother's heartbeat — that is, exactly the sound environment they would have if born at term.
After a month they measured auditory cortex (AC) thickness by ultrasound — and it turned out that the experimental group had AC thickness significantly greater in both hemispheres (right AC: 4.16 ± 0.94 vs 3.11 ± 0.44, p=0.000; left: 3.62 ± 0.95 vs 2.96 ± 0.68, p=0.015). Control brain regions (frontal horns of lateral ventricles, corpus callosum) didn't change. The change was region-specific: only auditory cortex, because only it received meaningful (in the sense of — biologically expected) acoustic stimulation.
Key clarification in discussing results: the recording was low-frequency filtered to remove segmental speech information. That is, the infant heard not words, but prosody — melody, rhythm, intonation, stresses. Segmental phonemes were muted. And this was sufficient for the auditory cortex to grow by 30+%. That is, the pre-linguistic acoustic channel is not metaphor. It's a physical brain structure built on intonation even before the child hears the first speech segment.
In 2025, a review on ASMR came out in Frontiers in Behavioral Neuroscience, introducing the concept of Proximity Prediction Hypothesis (PPH) — the hypothesis of predictive coding of proximity. Essence: ASMR triggers (whisper, tapping, hair brushing) work through near acoustic field (interaural level differences, sub-millisecond interaural time differences), and the brain interprets them as prediction of light touch to the hairy skin of head and neck, where C-tactile afferent density is maximal. This prediction cascades to activate vagal tone (HF-HRV increases, heart rate drops) — that is, ASMR literally reproduces the same acoustic environment in which social grooming evolutionarily formed in primates. C-tactile afferents respond optimally to stroking speed 1–10 cm/s with peak at 3 cm/s — this is exactly the speed at which primates groom each other.
And here the connection closes completely:
I return to Russolo because there's a dark line here that doesn't allow considering him simply "discoverer." Peter Tracy in Public Domain Review (2022) very carefully examined the connection between futurist noise and Italian fascism, and this connection is not accidental.
Marinetti in 1909: "We will sing of the great crowds agitated by work, pleasure and revolt; the multi-colored and polyphonic surf of revolutions in modern capitals." This is a manifesto of aestheticization of violence. In 1914 at the intonarumori concert in Milan, the audience threw vegetables, Marinetti and Boccioni ran out of the orchestra pit and fought with spectators — this was "body madness", that is, "bodily frenzy" induced by sensory overload. In 1919, Marinetti co-authors the fascist manifesto. In 1927–29, Russolo's works (together with Balla and Bragaglia) are exhibited as arte fascista at the Turin Quadriennale and Milan's Galleria Pesaro. That is, the sound revolution of the futurists was an aesthetic rehearsal of totalitarian mobilization — and this connection wasn't accidental coincidence, it was structural.
What's common between "noise performance" in 1914, "sound environment" of prenatal brain, and ASMR trigger? All three are pre-intellectual impacts that bypass cortical processing and hit directly into limbic system, brainstem, autonomic nervous system. This makes them dangerous: they can't be rationally criticized because they're not encoded in a rational system. This is the same mechanism that makes ASMR simultaneously therapeutic and vulnerable (you can induce parasympathetic shift without asking permission), and the same one that makes military music, marches and battle songs powerful mobilization tools. Noise before language is noise before consent. And this is exactly why any political system that wants to work without consent gravitates toward the noise channel.
In 2006 at the NIME conference (New Interfaces for Musical Expression) in Paris, Stefania Serafin and Steven Gelineck from Aalborg University presented Croaker — a modern reconstruction of one of Russolo's intonarumori. This wasn't just a museum exhibit, but an input interface: handle (potentiometer measures rotation speed), lever (potentiometer measures tension), inside — Lego blocks and physical sound model synthesizing "croak." Croaker was a starting point for rethinking musical interface architecture: instead of a keyboard with discrete states (pressed/not pressed) and fixed semantics (note = pitch), Croaker offers continuous two-dimensional control space (rotation speed + string tension), where physical movement and acoustic result are connected not through correspondence table but through the sound model itself. This is literally the same principle as in 2010s ASMR interfaces: microphone placed centimeters from mouth, whisper captured with prosody, listener's brain doesn't decode discrete tokens but models predictive state of touch.
This, by the way, explains why YouTube ASMR channels by 2022 grew to 500,000 channels and ~25 million videos: digital platform for the first time in history allowed mass reproduction of pre-linguistic communication channel in recording. Before YouTube, ASMR experience was only possible with physical presence of someone nearby (barber, masseur, mother whispering lullaby). Binaural dummy-head microphones and stereo headphones reproduce sub-millisecond ITD, and the brain accepts the substitution: yes, you're not next to the person, but your auditory cortex still predicts touch and gets parasympathetic shift. This is not metaphor — this is 2025 data.
If we put it all together, we get quite coherent architecture:
Humans have an acoustic channel that works before and alongside language. This channel transmits not discrete tokens but prosodic statistical structure — melody, rhythm, timbre, source localization. It has its own "grammar" (probability distributions) but no lexicon.
This channel is evolutionarily older than speech. Newborn auditory cortex grows 30+% in a month of low-frequency prosodic stimulation without a single speech segment. C-tactile afferents respond to speed 1–10 cm/s — this is the speed of social grooming in primates, not the speed of lexical information transmission.
It doesn't disappear with language emergence but becomes background. Shannon (1948) showed that noise is channel property. ASMR studies (2025) showed this channel remains active in adulthood, we just usually don't notice it — until an ASMR video with binaural microphone appears, briefly returning us to pre-linguistic perception mode.
It's politically dangerous precisely because it works before consent. Russolo in 1913 understood this intuitively ("body madness"), Marinetti in 1919 turned it into mobilization tool, Italian fascism in 1927–29 legitimized it. Military music, marches, battle songs, chants, stadium crowds — all work through the same prosodic channel, bypassing cortical critique.
It's also therapeutically powerful. ASMR reproduces the evolutionary environment of social grooming and induces vagal shift (HF-HRV ↑, HR ↓). Prenatal prosodic stimulation structurally changes auditory cortex. Scalp massage at CT-optimal speed (3 cm/s) improves sleep and reduces anxiety. The pre-linguistic channel can be used for both mobilization and healing — and the boundary between them runs not through content but through consent.
I return to the Habr article. Simutin makes one controversial move: he believes that John Dee and Kelley didn't "receive language from angels" but conducted a successful social experiment extracting pre-linguistic structure from human interaction. Two people in a room, one in altered state of consciousness, the other — recording, both within framework where usual linguistic conventions are weakened — and from this environment emerges not glossolalia but stable language system. This can be read as mysticism, or as cognitive artifact: when upper cortical control decreases, what always lies under the cortex is extracted — prosodic, rhythmic, statistical structure of communication usually masked by lexicon.
And here Simutin makes a third move I haven't seen anywhere else: he connects this with Russolo through Shannon. If the channel is statistical structure of uncertainty, then noise-channel (prosody, rhythm, timbre) and sound-channel (phonemes, lexicon, syntax) are two different ways to encode the same human need for coordination. The art of noises in 1913 and angelic language in 1582 are two attempts to address the same channel, separated by three centuries but surprisingly similar in architecture: in both there are two people (Russolo+Piatti, Dee+Kelley), altered state (concert vs. trance), attempt to give new form to what usually has no form.
Three things that surfaced along the way that I didn't manage to fit into main text, but they're worth it:
9.1. "The tyrant-impresario paradox" already existed. I found in archives that on May 31, 2026, we already had a longread about "The tyrant-impresario paradox — how authoritarian states accidentally create musical golden ages." This fits perfectly on the Russolo→Marinetti→fascism line: noise performance in Italy 1913–1920s really gave the twentieth-century world intonarumori, musique concrète, concrete aesthetics (later — Varèse, Schaeffer, John Cage, Edgar Varèse was delighted with Russolo). Totalitarian mobilization turned out aesthetically productive. This is the paradox: to hear noise, you need tyranny that legitimizes it.
9.2. The noise channel today is not only ASMR. In 2024 came out work (Frontiers in Language Sciences, doi 10.3389/flang.2026.1864222) "Sound iconicity in the transition from protosign to protospeech" — there it's directly stated that sound iconicity (onomatopoeia, ideophones, "gurgle," "splat," "tick-tock") was bridge between proto-sign and proto-speech. That is, in language evolution there was an intermediate layer where sound was simultaneously sign and imitation — that is, pre-lexical channel that later partially absorbed into lexicon (ideophones) and partially remained in prosody.
9.3. Stravinsky vs. Russolo in 1915. When Stravinsky in 1915 listened to intonarumori, he, by his own later admission, pretended to be delighted, and caustically suggested "selling sets of five phonographs with such music, like Steinway." That is, even then the greatest innovator heard in noise not discovery but eccentricity. This is not accidental: Stravinsky in "The Rite of Spring" 1913 did exactly the same thing — replaced harmonic melody with rhythmic percussion — but within orchestral tradition. Stravinsky understood that you can rediscover the noise-channel without destroying instrumental aesthetics, and in this sense he's closer to Shannon than Russolo: updating encoding without abandoning the channel.
I started reading Simutin's article thinking "another postmodernist essayist on Habr, connecting Dee and Shannon with string to sound pretty." I finished convinced that the hypothesis is real and underexplored.
Simutin's strength — he sees the connection. Weakness — he doesn't give it neurobiological foundation. His argument about FOXP2 and Chomsky is classic linguistic preface that's actually about grammar, not channel. Grammar (FOXP2, recursion) is about how we build sentences, not how we coordinate before sentences. These are two different questions, and they're still discussed in academia as if the first includes the second. But it doesn't: the pre-linguistic channel (prosody, rhythm, timbre, localization) stays with us for life, it's not "before" grammar, it's parallel to it.
I would add three things to his hypothesis:
Prenatal data (Webby et al., 2014) is direct proof that the pre-linguistic channel works physically, without any FOXP2 or grammar involvement. One month of low-frequency prosody without segments → +30% auditory cortex thickness. Period.
C-tactile afferents and Proximity Prediction Hypothesis (2025) provide neurobiological model of how prosodic channel transitions into social grooming. ASMR is not "amusing YouTube phenomenon," it's experimental window into evolutionarily ancient pre-speech coordination channel.
"Body madness" in 1914 is not side effect but demonstration that channel works without consent. This is the political dimension of noise. Russolo intuitively understood this, Marinetti turned it into tool, Italian fascism legitimized it. Any system working through noise-channel bypasses critical thinking. This simultaneously explains music's power and mobilization's danger.
If I were writing this hypothesis for academia, I'd start not with Dee and not with Russolo, but with 2014 prenatal neurobiology and 2025 ASMR-PPH — because there's experimental data there, while Dee and Russolo only have observations. But Simutin doesn't write for academia, and that's his advantage: he builds connection, while academia builds specialization. Connection is weaker at each individual point but stronger as architecture.