How the Octave and Diatonic Scale Were Born: From Pythagorean Ratios to Do-Re-Mi

Start on middle C on a piano and play only the white keys going up, and by the eighth note you land on another note also called “C.” Yet that final C is a physically different sound from the first. It vibrates the air exactly twice as fast. The frequency — how many times the air oscillates per second — has doubled, and yet our ears insist these two notes are somehow “the same.” Why does an interval like C to D sound completely distinct, while C and the C an octave above even share a name? In fact, there are only seven distinct notes — do © through ti (B) — and the eighth is simply the first note returning an octave higher. That is exactly why the span from the first note to the eighth is called an “octave,” from the Latin for “eighth.” Starting from that question unravels a bigger one: the do-re-mi-fa-sol-la-ti-do scale — seven distinct notes, even though it’s popularly counted as eight — is not a law of nature. It’s the product of thousands of years of human choices and compromises.

Why an Octave Sounds Like the Same Note

Two notes whose frequencies differ by exactly a factor of two are said to be an octave apart. Press a guitar string down at exactly its midpoint, and the resulting note sits precisely one octave above the open string. The phenomenon in which two notes related by this 2:1 ratio get called by the same name is known in music theory as “octave equivalence.”

This isn’t merely a cultural convention — it reflects how human hearing actually works. Researchers describe it through a perceptual property called “pitch chroma.” Sounds whose frequencies relate by a factor of 2, 4, or 8 — octave multiples — get grouped by the brain into the same perceptual category.[1] Brain-response studies on newborns have found traces of this same octave recognition, adding to evidence that it’s less a learned convention and more a built-in feature of the auditory system itself.[1] The “saptak” of Indian classical music and the “pitch class” of Western music theory are simply two different vocabularies describing the same underlying perceptual fact.[1]

If the octave is music’s most basic frame, how many notes to divide it into was an entirely separate question. And the person most often credited with first tackling that question mathematically is the ancient Greek philosopher Pythagoras.

Pythagoras, String Ratios, and the Blacksmith Legend

The story goes like this: Pythagoras was walking past a blacksmith’s shop when he noticed that several hammers striking an anvil produced sounds — some pleasing together, some not. He supposedly went inside, weighed the hammers, and discovered that the ratios of their weights held the secret to consonance — the quality of multiple notes sounding pleasant together.

It’s a captivating story, but it doesn’t hold up physically. The pitch a hammer produces when it strikes an anvil depends not on the hammer’s weight but on far more complicated factors — the anvil’s vibrational properties, the way the blow is struck, and so on. Combining hammers of different weights in neat ratios simply doesn’t produce notes matching those ratios.[2] The oldest surviving text recording this legend was written centuries after Pythagoras’s death, and scholars generally regard it as a dramatic invention by a later school seeking to mystify Platonic philosophy.[2] In other words, the blacksmith tale isn’t history — it’s later generations dressing up the Pythagorean worldview that “the universe is built from mathematical ratios” into a memorable story.

Illustration of Pythagoras and bells from Franchino Gaffurio's treatise
A woodcut from Franchino Gaffurio’s Theorica musicae (c. 1492), depicting Pythagoras experimenting with intervals using bells and various instruments. The blacksmith-hammer story is a later fabrication, but the enduring popularity of images like this one shows just how widely that legend was believed for centuries. Source: Wikimedia Commons (Public Domain)

The instrument the Pythagoreans are actually believed to have used was far simpler: a single-stringed device called a monochord. By adjusting the length of the string and comparing the resulting sounds, they found that certain ratios produced particularly pleasant, well-matched tones. Press the string down to exactly half its length (a 2:1 ratio) and you get a note exactly one octave higher. Press it to two-thirds length (3:2) and you get what we now call a perfect fifth; press it to three-quarters length (4:3) and you get a perfect fourth.[2] Why these clean whole-number ratios sound especially pleasing to the human ear remains an active research question in psychoacoustics even today — but the ancient Greeks had already arrived at the answer empirically.

Dividing the Octave Was a History of Compromise

This is where the trouble starts. Stacking perfect fifths (the 3:2 ratio) on top of one another to fill out the notes within an octave is known as Pythagorean tuning. Stack twelve perfect fifths starting from C, and in theory, after passing through seven octaves, you should land back exactly on the note you started with. But when you actually do the math, it doesn’t quite work out. Multiplying 3 by itself twelve times and multiplying 2 by itself some whole number of times can never produce exactly equal results.[3] A small but stubbornly irreducible gap remains between the two numbers — known as the “Pythagorean comma.” Think of it like trying to complete a circle by laying identical bricks end to end: the last piece always ends up slightly too long or slightly too short.[3]

That tiny gap became a problem running through the entire history of music. Keep the perfect fifths pure, and the octave doesn’t quite line up; make the octave line up exactly, and some fifth, somewhere, ends up slightly out of tune. The tuning system that emerged later, known as just intonation, used not only the 3:2 ratio but other simple whole-number ratios like 5:4 to make harmonies sound purer — but it introduced a new problem: transposing a piece to a different key made the intervals sound wrong.[4]

The practical solution European music eventually settled on was equal temperament: dividing a single octave into twelve semitones of mathematically identical size, so the spacing between every adjacent pair of notes is exactly the same. This means that no matter which note you start on when transposing, the same relationships between notes are preserved. The tradeoff is that nearly every interval, including the perfect fifth, ends up slightly detuned from its pure whole-number ratio. In effect, musicians gave up a handful of perfectly beautiful intervals in exchange for a system that works everywhere.[4]

Twelve-Tone Equal Temperament: A Ming Dynasty Calculation That Beat Europe to It

Mention equal temperament and most people think of Johann Sebastian Bach and his keyboard collection, The Well-Tempered Clavier. But the first person to work out the precise mathematics of dividing a semitone into twelve mathematically equal parts wasn’t European at all.

In 1584, Zhu Zaiyu (朱載堉), a Ming dynasty prince and scholar, laid out in his monumental treatise Lülü Jingyi (律呂精義) a mathematically exact method for dividing an octave into twelve equal intervals. His method involved repeatedly dividing the length of a string or pipe by the twelfth root of two (approximately 1.059463) — a principle identical to the standard equal-temperament calculation used today.[5] Zhu Zaiyu didn’t stop at theory: he actually built a twelve-string tuning apparatus based on these calculations.[6]

In Europe, just a year later, in 1585, the Dutch mathematician Simon Stevin independently arrived at a similar equal-temperament calculation. That the two men reached essentially the same conclusion at almost the same time, with no contact between them, is often cited as a striking historical coincidence.[5] Zhu Zaiyu’s calculations, however, were more precise and led directly to a working, practical instrument, which has led music historians to reassess him as a genuine pioneer of equal-temperament theory.[7] Bach himself wasn’t active until the 18th century — more than a hundred years later — which means equal temperament wasn’t a Western invention at all, but a mathematical conclusion reached almost simultaneously at opposite ends of the Eurasian continent.

Diagram of pitch calculations from Zhu Zaiyu's Lülü Jingyi
A pitch-calculation diagram from Zhu Zaiyu’s Lülü Jingyi (c. 1584). This primary source shows that a Ming dynasty scholar worked out the mathematics of twelve-tone equal temperament with precision, ahead of Europe. Source: Wikimedia Commons (Public Domain, 17th-century Ming-era document)

Where Do-Re-Mi Actually Came From

The question of how many notes to divide the octave into was settled through a long history of mathematics and compromise. But the names we give those notes came from a completely different path.

Guido d’Arezzo, an 11th-century Italian monk and music theorist, noticed his choir singers struggling to memorize unfamiliar chants and devised a new teaching method to help. What caught his attention was the text of a hymn honoring John the Baptist. The hymn had an unusual structure: each successive line began exactly one note higher than the one before it. Guido took the first syllable of each line and used it as the name for that note. Ut, Re, Mi, Fa, Sol, La — six note names were born.[8]

A seventh note, Si, was added to this six-note system a bit later. In Italy, the awkward-to-sing “Ut” was eventually replaced by the easier “Do,” giving us the do-re-mi-fa-sol-la-ti system we recognize today.[9] In the English-speaking world, 19th-century music educator Sarah Glover changed “Si” to “Ti” so that every syllable would start with a different consonant — which is why continental Europe still uses “Si” while English speakers use “Ti” to this day.[9]

Illustration of the Guidonian Hand, an 11th-century solfège memorization device
The Guidonian Hand. A medieval teaching tool that assigned notes to points on the joints of the hand to help singers memorize solfège syllables, shown here as depicted in a 15th-century manuscript. Source: Wikimedia Commons (Public Domain)

What’s striking is that Guido never set out to invent a “scale” at all — he was simply building a memory aid. The solfège system used in music education worldwide today turns out to have started as a practical shortcut for easing the memorization burden on medieval choir singers.

For most English-speaking readers, though, these syllables aren’t primarily known from an 11th-century Italian monastery at all — they’re known from a Salzburg hillside. In “Do-Re-Mi,” the 1959 Rodgers and Hammerstein show tune from The Sound of Music, Maria teaches the Von Trapp children to sing by pairing each syllable with an English homophone: “Doe, a deer, a female deer,” “Ray, a drop of golden sun,” and so on through “Tea, a drink with jam and bread.”[13] The song, later carried into millions of American, British, Australian, and Canadian living rooms by the 1965 film, did more to popularize Guido’s naming system across the English-speaking world than any classroom ever could — even if almost nobody singing along realizes they’re reciting a nine-hundred-year-old mnemonic.

The Seven-Note Scale Isn’t a Universal Answer

Following the story this far, it might seem like do-re-mi-fa-sol-la-ti-do is the natural, inevitable answer for music. But it’s just one option among many the world’s musical traditions have arrived at.

Indonesian gamelan music, from Southeast Asia, uses two scales entirely different from the Western seven-note system: slendro, built from five notes, and pelog, built from seven. Both systems are structured with interval spacing fundamentally different from Western equal temperament, and neither can be reproduced on a standard piano keyboard or guitar fretboard.[10]

Slendro-tuned gamelan ensemble from the Kraton Kasepuhan in Indonesia
A set of Indonesian gamelan instruments tuned to the slendro scale. The distinctive overtone structure of bronze percussion instruments produces a sonic aesthetic entirely different from that of Western harmony. Source: Wikimedia Commons (CC BY-SA 3.0)

In Indian classical music, an octave is divided into twenty-two microtonal units called shrutis, and each raga — a melodic framework — selects its own distinct subset of notes from within that space.[11] Arab maqam music divides the octave into twenty-four quarter tones, capturing pitch distinctions half the size of a Western semitone.[12] Many traditional music cultures across Africa widely use five- or six-note scales, and distinct pentatonic traditions persist across various East Asian musical cultures as well.[10]

What these examples show isn’t a hierarchy in which one system is more “natural” than another. The octave, as a physical framework, arises from a shared feature of human hearing. But how many pieces to divide that framework into, and at what intervals, was a decision each culture reached independently, shaped by its own aesthetics and needs.

Conclusion: The Scale We Take for Granted Was One Choice Among Many

Do-re-mi-fa-sol-la-ti-do has long been treated as the default standard of Western music, but pull apart how it actually came together and the story turns out to be far messier than that. The fact that an octave sounds like “the same note” is a universal feature of human hearing. But dividing it into seven notes, and choosing this particular set of intervals, is closer to an accident — the tangled outcome of Greek string-ratio experiments, precise mathematical calculations by a Ming dynasty scholar, and a memory trick invented by an Italian monk. Musical traditions that divide the very same octave in entirely different ways are still very much alive around the world today. Do-re-mi-fa-sol-la-ti-do was never the one true path music had to take — it was simply one answer, among several, that humanity arrived at while working with the same physical raw material, and one that happened to spread widely.


References

[1]: Attention, Perception, & Psychophysics, “Pitch chroma discrimination, generalization, and transfer tests of octave equivalence in humans” (peer-reviewed journal; https://link.springer.com/article/10.3758/s13414-012-0364-2); Developmental Science, “The Neural Reality of Pitch Chroma in Early Infancy” (peer-reviewed journal; https://onlinelibrary.wiley.com/doi/full/10.1111/desc.70037)

[2]: Wikipedia, “Pythagorean hammers” (CC BY-SA 4.0; https://en.wikipedia.org/wiki/Pythagorean_hammers); Wikipedia, “Pythagorean tuning” (CC BY-SA 4.0; https://en.wikipedia.org/wiki/Pythagorean_tuning)

[3]: Wikipedia, “Pythagorean comma” (CC BY-SA 4.0; https://en.wikipedia.org/wiki/Pythagorean_comma)

[4]: Wikipedia, “Just intonation” (CC BY-SA 4.0; https://en.wikipedia.org/wiki/Just_intonation); Wikipedia, “Equal temperament” (CC BY-SA 4.0; https://en.wikipedia.org/wiki/Equal_temperament)

[5]: Wikipedia, “Equal temperament” (CC BY-SA 4.0; https://en.wikipedia.org/wiki/Equal_temperament)

[6]: Wikipedia, “Zhu Zaiyu” (CC BY-SA 4.0; https://en.wikipedia.org/wiki/Zhu_Zaiyu)

[7]: Music Theory Spectrum (Oxford Academic), Rehding, A., “Fine-Tuning a Global History of Music Theory: Divergences, Zhu Zaiyu, and Music-Theoretical Instruments” (peer-reviewed journal; https://academic.oup.com/mts/article-abstract/44/2/260/6609877)

[8]: Wikipedia, “Ut queant laxis” (CC BY-SA 4.0; https://en.wikipedia.org/wiki/Ut_queant_laxis)

[9]: Wikipedia, “Solfège” (CC BY-SA 4.0; https://en.wikipedia.org/wiki/Solfège)

[10]: Wikipedia, “Slendro” (CC BY-SA 4.0; https://en.wikipedia.org/wiki/Slendro); Wikipedia, “Pelog” (CC BY-SA 4.0; https://en.wikipedia.org/wiki/Pelog)

[11]: Wikipedia, “Shruti (music)” (CC BY-SA 4.0; https://en.wikipedia.org/wiki/Shruti_(music))

[12]: Wikipedia, “Arab tone system” (CC BY-SA 4.0; https://en.wikipedia.org/wiki/Arab_tone_system)

[13]: Wikipedia, “Do-Re-Mi” (CC BY-SA 4.0; https://en.wikipedia.org/wiki/Do-Re-Mi)

You Might Also Like