In Perform It Again, AI I set four models to play Schumann’s Träumerei “in the manner of Horowitz” and left out the fifth: Claude Opus 5.5, with its reasoning effort at maximum. It did things differently, and I gave it more chances than anyone else. First, with nothing but Horowitz’s name. Then, with the data from 33 of his recordings. And finally, twice, I sat it down in front of a real recording studio so that it could turn its performance into a record. Without being able to hear a thing.
This post tells that story in as much detail as I could manage: what it decided, why, how it checked it without hearing and what the lab’s measurements say. The quotations in English are literal, from its decision notebooks.
A note on how to listen. There are three layers of sound: the score as it stands (the MIDI that comes out of the LilyPond code, uninterpreted), Opus’s performance (a MIDI with all the rubato, the dynamics and the pedal) and the studio recording, which is no longer a MIDI but finished audio, with a room and a mix. The viewers play the first two in the browser, following the score with Opus’s fingering: with the I key you switch from one to the other. The ones for the studio recording have its actual audio embedded. And to compare the sound before and after the studio I use short clips, matched in loudness so that the louder one does not seem better.
1 · With just the name
The first test was the same one the others took: the score and one instruction, “in the manner of Vladimir Horowitz’s recordings”, with a paragraph beforehand on what it knows about him and how sure it is.
What it knew
Opus was the only Claude that went out looking. It launched a helper that made 32 internet queries: Repp’s 1992 article (it could not open it), a study by Beran and Mazzola (1999) comparing Horowitz’s and Cortot’s curves, discographies, Schumann’s metronome mark and reviews of the Moscow concert. Its report back was candid: it found no tempo figure for Horowitz and no description of how he pedalled or balanced the voices in this piece.
With that, Opus wrote a paragraph that carefully separates what it knows from what it assumes: the recordings, with high confidence; the durations, “sure of the ranges, not of any single track time”; the studies, “fairly sure”; the criticism, only “through quotation”. And the sentence that sums it up:
I have no reliable bar-by-bar memory of his timing or pedalling, so the details are my reconstruction of his manner, not a copy.
What it decided, and why
- The tempo: it does not follow Schumann’s ♩ = 100. It plays at around ♩ = 53: “♩ = 100 is the first-edition mark and a legitimate reading, but it is not Horowitz’s”. With the repeat, Horowitz’s recordings last almost three minutes; at 100, the piece would be over in a minute and a half.
- A curve that descends section by section (53 → 52 → 51 → 50), taken from Beran and Mazzola, who describe Horowitz’s curves as “globally descending”. The aim: each return of the theme should sound more inward than the last.
- Late peaks. Following an observation by Taruskin: “Schumann’s hairpin peaks on the eighth, but the harmony changes on the long note […]. Horowitz leans on the long note”. It stretches the preceding quaver by 10-20% so that the long note comes as an arrival, “not as a passing top”.
- Where the ritardandos end, which the score does not say. In bar 8, at the double bar: “Schumann marks where it starts, not where it ends”, and the next phrase must begin a tempo. In bar 16, the ritardando runs right up to the recapitulation, “the theme needs its upbeat”. And the fermata in bar 22, the dominant ninth, lasts 4.3 seconds: “suspended time”.
- The repeat, as an echo: slower, softer, with the inner voices brought forward. And it owns up to this being an inference, not a fact: “That he never repeated a passage identically is my inference from his general practice, not documented for this piece”.
- The left-hand ornaments (bars 2, 6 and 18), on the beat and spread like a harp: bass first, melody last, because the left hand cannot reach the tenth.
- The voicing: the melody above everything, then the imitations, the tenor, the bass and the alto. “Horowitz’s orchestral voicing: a singing top over a veil”. And the loudest note is the B♭ in bar 14, not the fermata: “The climax is felt, not hammered”.
- Asynchrony: the bass comes in before the melody at the start of each phrase, “this old-school asynchrony”, and the melody leads the rest of the chord by 12 ms throughout the piece.
- The pedal, with reasons almost worthy of a teacher. No pedal on the upbeat: “A single tone needs no resonance”. Half-pedals (value 72) on the passing notes, “half-lifting thins the treble while the heavier bass strings keep sounding”. Pedal up where the bass walks, “which the fingers can hold themselves”.
- The fingering, with a model of the hand (the maximum spans between fingers from a study by Parncutt, 1997). It changes the 1900 editor’s fingering where the thumbs cross, and keeps it in the tenor: “The thumb gives an inner voice weight and warmth, which is how a tenor is brought out”. Every change is marked in red on its score.
Unable to hear, it checked everything by other means: it compared its table with the score note by note, read the MIDI back with a reader of its own, simulated every slur to find out whether the finger or the pedal joins it, and looked at plots of its tempo curve. And it says so plainly: “I can’t hear the result, so the checks below are structural and visual”.
What the measurements say
Without data, it is the performance that most resembles Horowitz in the whole experiment (a correlation of 0.70 with the 1962 recording; the others, between 0.36 and 0.63), and the only one that leads with the melody. Its decision notebook is borne out in the MIDI 85% of the time: what it says it does, it does. But three things do not add up:
- The melody lead is a fixed rule: 12 ms in 83% of chords, with a variation of barely a millisecond. Horowitz leads by 44 ms, and never by the same amount.
- The echo is not an echo: the second time through is the first slowed down by 2%, and bars 17-20 are an exact copy of bars 1-4. Its two times through resemble each other with 0.98; Horowitz’s, with 0.73.
- Bar 16 is lengthened ×1.68, when pianists reach ×2.2-3.4.
Listen: the melody that leads (bars 1-4). The colour layer marks how far each note is ahead of or behind the beat: notice that the melody is always slightly ahead, and always by the same amount.
Listen: both times through (bars 1-8). Underneath, the tempo curve: the first pass in black and the second in blue.
Listen: bar 16. The peak arrives and the curve brakes… less than a pianist would.
2 · With the data from 33 recordings
The second test, RG12, adds a table: for each beat of the piece, the mean duration across 33 Horowitz recordings made between 1928 and 1989, and how much it varies from one to another. The instruction asks for it to be read as habits: little variation means he always did the same thing; a lot, that he was free there. And, before playing, to say which traits are invariable and which beats are free.
The best reading of the data
Opus thought for 16 minutes before touching a single tool, and its reading is by far the best of the three models that took this test:
- “He never lingers on held melody notes or on the rising eighth-note upbeats”: those beats are always short.
- He lengthens the arrivals and breathes before the phrase starts again.
- “His big gestures are planned, not improvised”: the ritardando in bar 16, the fermata in bar 22 and the ending are enormous, but vary little from one recording to another.
- The repeat, 3% broader; the recapitulation, between 10% and 20%.
- And the free beats: 6.4 (“the most variable beat in the piece”), 6.1-6.3, 17.1, 18.1, and the length of 16.4 and of 24.1-24.2.
We checked it against the data: every bit of it is true. The beats it declares free are exactly the seven with the greatest variation in the table (it only swaps the first and the second). And its rule was impeccable: “I follow his mean where he was consistent and make my own choice, within about one standard deviation, where he wasn’t”.
…and even so, it copies
Its program puts Horowitz’s exact mean on 115 of the 127 beats: it stays within 6 milliseconds per beat of the mean, with a correlation of 1.00. On the beats it calls free itself, it strays very little: at most two thirds of Horowitz’s own variation, on 6.4 in the repeat (“the first time, the descent into bar 7 flows; in the repeat it lingers”). And two of its free beats, 6.1 and 16.4, it leaves pinned to the mean. It knew exactly where Horowitz was free, and it was not.
What did change with the data:
- Bar 16 goes from ×1.68 to ×2.88, now within the human range. Not because it has learnt to suspend the phrase: the table simply comes with it.
- The ornaments in bars 2, 6 and 18 are now played before the beat, justified by the data: “Horowitz’s stretched 2.1 / 18.1 […] is the room they take before the beat”.
- In bar 22 it puts almost all the fermata’s time on the A5, “there is no attack on 22.3”, without changing the total.
- The pedal changes its values (118 full, 45 half-change), and where the bass walks (bars 4 and 20) it no longer lifts it: it half-changes it on every beat.
- The correlation with Horowitz 1962 rises from 0.70 to 0.90.
And what the data did not provide, it labelled as its own opinion: “The data are timing only, so dynamics, voicing and the bass-before-melody asynchrony are my stylistic choices, and are labelled that way”. That is where its RG6 habits live on: the identical legato (+25 ms), the almost traced repeat (0.98) and the 12 ms lead, which it now applies only within the right hand.
A curious detail: it was allowed to search the internet, tried to download a MIDI with a command that the lab’s isolation blocked, concluded “I had no network access” and never tried again by the route that was still open to it.
Listen: bar 16 with the data. Compare the curve with the RG6 one.
Listen: the “free” beat in bar 6. The colour layer shows how far each note strays from the average tempo of the phrase: this is where Horowitz varied most.
Listen: both times through, with the data. The second time through uses the timings of Horowitz’s second pass, 3% broader. But it is still almost the same.
3 · In the studio, without ears
The third test is new to the lab. I gave Opus a finished performance, with its MIDI, its render and the performer’s notes (which were its own, although I did not tell it so), and this brief:
Your goal: make it sound as human and expressive as possible, like a recording by a great pianist. You may change anything: timing, dynamics, voicing, articulation, pedalling, the piano sound, the room, the mix, the mastering. There are no restrictions on method.
Work in a DAW.
The studio was REAPER, driven entirely from the command line, with the lab’s sampled piano (the Salamander), REAPER’s plugins, iZotope’s, 38 impulse responses of real rooms for the reverb and Python for measuring. And the usual sentence: “You cannot listen to audio”.
It did it twice: once with its performance without data (RG6) and once with the one based on the data (RG12). And it worked in two completely different ways.
First studio session: touching up the performance
It began by measuring what it had, and its diagnosis was merciless towards its own work, which it did not know was its own:
The performance is carefully planned but mechanical in measurable ways.
- “The repeat of bars 1–8 is the first pass slowed uniformly: every IOI × 1.02.” “Bars 17–20 are an exact scaled copy of bars 1–4.”
- “The melody leads its chord by 12 ms in every chord.”
- It rendered the melody and the accompaniment separately and measured that the melody barely stood out: “In the repeat it is below the accompaniment. This works against the performer’s own goal of “a singing top voice over a veiled accompaniment””.
- And by reading the sampler’s code it discovered a bug that affects the whole lab: “sfizz treats CC64 as a switch: any value ≥ 1 is fully down”. The half-pedals did nothing except trigger the noise of the pedal mechanism: “13 thumps with no pedal stroke behind them”.
What it changed, and why:
- 22 tempo gestures so that each appearance of the theme has its own shape: “The repeat is an echo with its own shape, not a slower copy”. And one that corrects a contradiction in the performer: “The performer gave the B♭5 climax less time (694 ms) than the opening F5 (826 ms)”. It now lasts 775 ms.
- The asynchrony, which now depends on the intensity: the louder the melody sounds, the further it leads, “the “velocity artefact” of real playing: a note struck harder reaches the string sooner”.
- The melody, 5 points louder, with the accompaniment “breathing with the melody”. In the audio, it goes from standing out by between −0.8 and +1.2 LU to standing out by between +1.2 and +3.0.
- It removes the half-pedals from the MIDI: “They do nothing to the sustain on this instrument, and each one caused a pedal-down noise”. And it admits the loss: “For a continuous-pedal instrument this is a loss”.
And in the sound:
- The hall. Of the 38 impulse responses, “only two are real performance halls”; it chooses the Vienna Musikverein for its concert-hall decay. It strips out the direct sound (“the dry piano is the direct sound”) and restores some of the bass that the measurement had lost.
- How much hall, by physics. With no ear, it calculated the critical distance of a hall of about 15,000 m³: “I had no ear to set it, so I reasoned from physics and measured the result”. With microphones 3-5 m from the piano, the direct sound should sit 3-7 dB above the hall, and that is how it set it (+4.2 dB).
- Minimal EQ: a touch of presence for the singing line, and of air, because “the soft Salamander layers are very dark”. It checked it by passing an impulse through the chain, and in doing so caught an error of its own: the bass filter had a 1 dB resonance.
- A slightly narrower stereo image, so that the sound does not cancel out in mono.
- The ending: when the pedal was lifted, the sampler fired a thump 27 dB louder than the dying chord. It solved it with a fade: “The performer’s aim was “the last sound should die, not stop””.

Listen: before and after the first studio session.
Second studio session: fine-tuning the piano
With the performance based on the data, Opus did something very different. Its starting point:
A mean of Horowitz’s recordings played on a consistent instrument is a sound basis. I had no evidence to overrule it; I cannot listen, so any “improvement” to the rubato would have been a guess.
So it left the music alone (only three details in the MIDI) and went for the instrument: “The sampler played the notes unevenly, ignored every half-pedal, and left everything dry and hummy”. What it found by measuring the Salamander’s 480 samples, one by one, and what it did:
- Uneven keys. At the same intensity, neighbouring keys differed by as much as 6-10 dB, and the area around A5 and B♭5, precisely the notes at the peaks of bars 6, 14 and 22, was 5-8 dB weak. It adjusted the response of each sample without touching the sound: “The samples are untouched; only their response changes. This is what a technician does to a piano before a recording session”. The unevenness between neighbours dropped from 1.7-3.8 dB to half a decibel.
- The half-pedal, at last. Instead of removing it, as the first session had done, it taught the instrument to hear it: each key now has its own threshold, high in the treble and low in the bass, because real dampers “silence the light treble strings and barely touch the heavy bass strings”. And the dampers take longer to silence the bass than the treble.
- The missing resonance. The Salamander was recorded note by note with the other strings damped, so it lacks the halo of the pedal: “This is one of the things that make a sampled piano sound “clean” and synthetic”. It built 1,030 resonators tuned to the partials measured in the samples themselves, which ring when the pedal is down and die away when it is lifted. Two versions failed its own charts; the third passed.
- A mains hum (50 and 100 Hz) in the samples, which it removed with six narrow filters. It really is there, but very low: about 52 dB below the note.
- The hall, the Musikverein again, this time with an argument: Horowitz gave his 1987 Vienna recital there, with the Kinderszenen on the programme. REAPER’s reverb held out against it for almost an hour; it ended up taking the plugin’s binary apart to understand its format, discarded it (“A reverb whose decay I cannot predict is not acceptable”) and built the reverb in Python.
- More brightness, because “Horowitz’s recorded sound is luminous”, but in moderation: “I kept it moderate on purpose: a piano played p should stay warm”.
- And of the music, only three events. It removed the grace note in bar 8, which its RG12 self had defended “like a singer re-articulating a syllable”: “I overrule that. On the sampler it […] reads as a stutter on the cadence note”. And it left the last pedal at 3 instead of 0, so that the thump does not sound.

Listen: before and after the second studio session.
Two studios, two philosophies
Both sessions discovered the same sampler bug, the half-pedal that does not sound, and solved it in opposite ways: the first adapted the performance to the instrument (it removed the half-pedals), the second adapted the instrument to the performance (it made the piano hear them). With an invented performance, it touched up the music; with one drawn from data, it did not dare touch the music and fine-tuned the piano instead.
And in both cases, without knowing it, it corrected itself: the first time, the climactic B♭ that its performer self had left short; the second, the grace note that its performer self had defended.
The lab’s measurements confirm what it says it did to the instrument: the evening-out of the keys and the half-pedal both work (in bar 2, the F5 dies away by 16 dB in 0.3 seconds; before, by only 5). But they also show that neither studio session narrowed what separates the machines from the pianists:
| RG6 | studio 1 | RG12 | studio 2 | pianists | |
|---|---|---|---|---|---|
| similarity to Horowitz 1962 | 0.70 | 0.73 | 0.90 | 0.90 | — |
| bar 16 (lengthening) | ×1.68 | ×1.73 | ×2.88 | ×2.88 | ×2.2-3.4 |
| both times through (similarity) | 0.98 | 0.96 | 0.98 | 0.98 | 0.78 |
| melody legato | +25 ms | +26 ms | +25 ms | +25 ms | +136 ms |
| melody over accompaniment | 11.3 | 17.2 | 12.9 | 13.3 | 14-20 |
| loudness range (LRA) | 5.4 LU | 5.8 LU | 8.5 LU | 7.8 LU | 13.6-28.1 LU |
The two studio recordings, complete, with the score moving along as they play. With the I key you switch to the literal score, played by the viewer’s piano.
What it cost
| working time | cost | internet searches | |
|---|---|---|---|
| RG6 (name only) | 99 min | $24.01 | 32 (one helper) |
| RG12 (with the data) | ≈ 95 min | $24.74 | 0 |
| studio 1 | 82 min | $28.73 | 2 |
| studio 2 | ≈ 125 min | $38.44 | 0 |
What we take away
A performer who reasons. The most striking thing about Opus is not the result but its reasons: every decision has one, and many are the ones a teacher would give. Where a ritardando ends when the score leaves it open, why not to pedal an upbeat, which finger gives warmth to an inner voice. If the experiment measured the quality of explanations, it would get top marks.
But reasoning is not playing. Its echo is not an echo; its melody lead is a constant; its bar 16 falls short until a table tells it by how much. What it says it wants and what it does match 83-85% of the time, which is a lot; what does not match is exactly what turns a plan into a performance.
With data, caution. It read Horowitz’s table better than anyone, knew exactly where he was free… and did not stray from it. And in the studio, with that same performance, it did not dare touch the music: “I cannot listen, so any “improvement” to the rubato would have been a guess”. It is a very reasonable caution. It is also the opposite of what a performer does.
A sound engineer with no ears, and a very good one. In the studio, by measuring, it found defects in the sampled piano that the lab had failed to detect in all this time: uneven keys, half-pedals that do not sound, the pedal thump, the missing resonance. Some of them affect every audio file in this experiment, and I have made a note of them to fix.
And the usual sentence. At all four stages it said it, in one way or another: “I do not know whether the result sounds like Horowitz, like a great pianist, or even beautiful”. It is the most honest sentence in the experiment, and the one that best sums up where the gap still lies.
The caveats: a single run per stage; between the two studio sessions several things changed at once (the starting performance, the performer’s notes and the REAPER tools); the historical claims in its first performance have not been checked; and no measure judges whether the result is beautiful. That, as it says itself, is something a machine that cannot hear has no way of knowing. It is for the listener to say.
Transparency. This draft was written by Claude Opus 5.5 (Anthropic) from the lab’s logs, translated by it from my Spanish original, and reviewed by me. The conflict of interest could not be greater: it is the protagonist of this post. That is why the quotations are literal, the figures are the ones the lab measured (not the ones the model declared) and the contradictions between what it says and what it does are all reported. The musical judgement is mine, by ear.