samhan.ioMUSIC BOXHow it works
A LITTLE MACHINE THAT COMPOSES

Music box.

Pick a composer and press play. Every piece is new.

Chorale in D minorReady

A melody waiting to happen.
Idea A Its answer BWatch them return in new phrases

Press Listen. “Another piece” shuffles key, tempo and sound.

The beginning
Shape the music Sound, tempo & composition

Measured interval patterns guide the melody and answering voice.

WHILE IT PLAYS

How this piece was made.

01Choose a world

02Try drafts

03Give it a journey
ThemeRecallTurnHome

One theme returns through changing harmony. A brief variation leads back to its familiar shape.

INSIDE THIS COMPOSER

First, give it something to remember.

Two bars establish a theme. It returns, changes shape, and eventually comes home.

Read the theory behind this moment

Learn about harmony and cadences ↗
A note is a choice
Why did this note win?

Each legal pitch receives these scores. For the invention and minuet, the displayed note belongs to the best complete bar: its immediate score may be lower than an alternative that leads to a worse continuation. The chorale still chooses the highest note score. The searched closing run shows note choices; tonic arrivals are marked separately.

Consideration

Freeze the explanation to inspect a decision while the music continues. This is a replay of the recorded choices, not live training.

How it worksThe composers it learned from, and what it learned.+
THREE COMPOSER STUDIES

Different habits. Different worlds.

Bach keeps the established thematic composer. Mozart adds balanced eight-bar phrase plans and broken-chord accompaniment. Vivaldi contrasts accented, short-bow ensemble returns with lighter solo episodes. Selective stepwise flourishes decorate the solo line, while the cadence notes remain protected. These articulation and ornament rules are authored choices, not learned from MIDI velocities. These are small, original studies inspired by each composer, not recovered compositions or a full sonata or concerto model.

Separate score-derived interval models guide note choices. For the new profiles, duration frequencies also help choose a reusable opening rhythm. Form, harmony, accompaniment and scoring weights remain authored rules. The new profiles keep up to 24 cadence paths, checking that the last two notes can reach a legal arrival; the original Bach search remains unchanged. The Bach suspension model stays in the Bach profile. The form and accompaniment choices draw on Vivaldi’s refrain–episode contrast and Classical broken-chord textures.

Does more training data help?

Training and validation loss are average surprise in bits per melodic interval; lower predicts these score streams better. Each curve uses nested training subsets with a fixed validation set. Smoothing is chosen on validation works, then evaluated once on separate test works. This is a learning curve, not a fitted scaling law or a beauty score.

Sources, methods and limits

Mozart: 54 movements from 18 piano sonatas, from Hentschel, Neuwirth and Rohrmeier’s Annotated Mozart Sonatas, CC BY-NC-SA 4.0. Vivaldi: openly licensed Mutopia editions, including six Op. 3 concertos, the Four Seasons, a double concerto and arias. Part-only exports and exact normalized duplicates are excluded. All movements and alternate exports of a work stay in the same split.

We extract outer-note streams, not expert-separated melodies. Mozart uses the upper staff and outer notes of the lower staff; Vivaldi track roles are inferred by mean pitch. Gaps break interval context. Written repeats and MIDI rendering differ, and the genres differ, so these losses do not rank composers or measure imitation quality. Small Vivaldi holdouts are particularly uncertain. No human listening study has validated the new profiles.

The interval-context idea is informed by Pearce and Wiggins on musical expectation; this small fixed-order model is not the full IDyOM system. The piano and string sounds now use a compact selection of actual instrument recordings from VSCO 2 Community Edition (CC0). Soft and strong dynamic layers, short and sustained bow samples, smoothed sustain loops and controlled releases shape playback. Featured violin lines use the same violin-section recordings as the opening, with a warmer string balance and more bass support. Sample-level trims reduce loudness jumps between recordings. Mozart’s upper parts sit an octave lower, with the piano brought forward against softer strings. The importance of articulation and connected notes is discussed by Jaffe and Smith. This is sample-based playback, not a physical bow/string simulation. Other keyboard colours retain the original synthesis. The 2.9 MB recording bank loads when needed; decoded samples are reused during the session. Sample credits, source URLs and transformations ↗

Download parameters, splits and measured losses ↗ · Edition credits and licenses

PHRASES, NOT JUST NOTES

Give an idea somewhere to go.

One two-bar theme now anchors the piece. Its pitch targets and rhythmic pattern return throughout. For inventions and minuets, the seed chooses one of three opening contours and one of three pulses. The second bar derives its targets from the first by repetition or a one-scale-step shift, preserving the first bar’s rhythm. The middle section lifts or dips the same idea before it returns; a held note or displaced accent stays recognisable on later appearances. The search still decides the actual notes using harmony, melodic motion and, if enabled, learned interval tendencies. Independently sampled corpus fragments are no longer used to build the opening.

The earlier harmonic plans give these returns a familiar setting, with each complete internal phrase now resolving to the tonic before the next begins. Theme resemblance has a stronger weight on later notes, and fresh random variation in the returning melody is reduced to 45% of its opening strength. A separate penalty discourages large jumps between bars. The melody keeps the opening rhythm while the accompaniment changes pace: a spacious second phrase, a more active third phrase, then a balanced return. The chorale closing gesture borrows a rhythmic value and, when the join allows, the direction of the theme’s last move. In the invention and minuet, the melody and inner voice now enter the final chord together, with no extra inner note after the arrival. Expressive playback eases toward the arrival and softens the landing. Each internal phrase ends with a coordinated dominant-to-tonic cadence: the melody settles on the tonic or its third, the bass reaches the tonic root, and a leading tone in the melody or inner voice resolves upward. These essential notes are present even with ambient pads off. The chorale has a settled close, a gentler answer, and a small turning figure. In inventions and minuets, the first three phrases vary between four, five and six bars. A longer phrase repeats a remembered cell; a shorter one reaches its cadence sooner. The last bar keeps the preceding rhythmic cell moving into the dominant chord. A short search follows that pulse into the tonic, considering melodic motion, chord fit, the arrival and the next phrase’s opening. The approach keeps moving in half-beat steps into a beat-aligned arrival. The search can allow an extra bar when a longer continuation fits better. The resolving leading tone may be in the melody or the inner voice. The following melody entrance stays within five semitones of the note just reached. The chorale keeps its earlier phrase plan. A small, fixed change to the last part of the motif gives later returns a distinct answer without changing their rhythmic identity. The resulting 19–23-bar piece (19 bars for chorales) gives the idea time to finish before it settles. The chorale compares six drafts using the opening theme (65%) and melodic continuity (35%). The invention and minuet compare twelve drafts using the opening theme (48%), continuity (28%) and phrase fit (24%), including stalled notes and jumps into cadences. These are hand-chosen proxies and cannot measure beauty.

A little musical program.

For inventions and minuets, the opening follows a compact recipe: invent one bar, repeat or shift it to make an answer, then recall the pair. The closing run inherits the last bar’s pulse and searches for a connected arrival. We checked the notation of Bach’s first invention as a design reference for continuing rhythmic motion through cadences. This is one example, not a measured model of all Bach endings. Later phrases reuse these operations under changing harmony. This is an experiment inspired by minimum description length: fewer independent decisions can make a pattern easier to recognise. We constrain the construction to this small grammar; we do not yet search all programs or measure their length in bits. Chord fit and voice-leading can alter the target pitches.

Learn when to bend a rule.

We reanalysed the same training scores for prepared dissonances: a note that fits the bass is held across a bass change, briefly clashes, then resolves downward within one beat. We count how often this happens among eligible opportunities, separated by voice, beat position, and the resulting interval. This is a measurable proxy, not a full harmonic analysis or a count of every kind of rule-breaking.

For example, among inner-voice downbeat opportunities that would create a fourth above the new bass, our detector found 111 such gestures in 822 opportunities (13.5%). The smoothed sampling probability is 13.4%. For a major seventh it found 15 in 763 (2.0%). The generator samples from the matching context, then samples the observed resolution delay and downward step. It also checks chord fit at the resolution, voice spacing and phrase rests. These extra safeguards and the different generated contexts mean the overall frequency need not match the corpus.

The live explanation shows the counts and probability when a learned exception occurs. Switching off learned tendencies disables this pass. Inspect counts, detector method and held-out results ↗

Does the exception model generalise?

On 79 held-out chorales, the detector found 217 successes in 4,172 eligible opportunities. Predicting with voice-role averages costs 0.271 bits per opportunity; conditioning on interval and beat position reduces this to 0.257. This is a modest prediction improvement for the detector’s labels, not proof that every detected event is a textbook suspension or that the music sounds better. Rates use 12 prior observations toward the role average; that smoothing choice is ours. We currently apply this model to the answering voice at bar changes, where the generated harmony changes.

Open Music Theory · preparation, suspension and resolution ↗

Gather, crest, release.

A shared phrase envelope gathers gradually, holds a brief crest late in the phrase, then releases more quickly at the resolving chord. Near the crest, the melody’s target lifts by one scale step while its remembered rhythm stays intact. Expressive playback adds a modest acceleration and crescendo, then softens and eases into the arrival. This is a hand-shaped expressive gesture, not a learned measure of emotion. Turning off expressive playback removes the timing and dynamic shaping; the composed pitches stay the same.

Listen for the space.

After the full phrase, the closing bar changes from the dominant chord to the tonic. In inventions and minuets, the melody keeps moving through the closing bar and its arrival lasts roughly one to one and a half beats. The internal phrases take three different breaths: a short comma, a middle pause and one deeper stop; the chorale keeps its quarter- or half-beat breaths. The tonic has enough time to sound without forcing every phrase into a long hold. The next melody entrance is slightly gentler, while the room tail carries across the pause. Ordinary bar lines keep flowing. In the quicker forms, expressive playback adds a smaller, graduated ease; the deeper pause occurs only once in the internal phrases. Reverb can linger across a breath, as it would in a room.

Ending profiles influence the closing register and colour. Inventions and minuets search their final approach using the same continuation process as internal cadences; chorales retain their small closing figures. The inner voice supplies the tonic third even when the melody finishes low. In a minor key the bright ending uses a major final chord. The live explanation follows this plan and distinguishes searched note choices from the coordinated tonic arrival.

Where the ideas come from

Triumph of the Cyborg Composer introduces Cope’s recombination and musical roles. US 7,696,426 describes EMI’s contextual grouping and protected signatures in its background, then a separate cadence-first, backward recombination method. We borrow the ideas of grouped gestures, context, and planning an arrival. We do not implement EMI, infer SPEAC labels, or run its backward segment-matching algorithm.

The experimental fragment pool is retained for inspection but is not used by the current generator. It contains training data only, with the source filenames retained. This new generator has not been validated in a listening study. The prediction results below apply only to the original interval model. Inspect the fragment pool and provenance ↗

ORIGINAL BACH EXPERIMENT · LEARNING FROM SCORES

Let Bach supply the tendencies.

We analysed Craig Stuart Sapp’s digital edition of 370 Bach chorales. After removing 11 exact transposition-normalised duplicates, we fitted the model on 280 pieces and reserved 79 pieces for evaluation. Entries with the same work catalogue number stay together.

For each voice, we count the distance between consecutive notes and how that distance depends on the previous interval. Frequent patterns earn a higher score. Small pseudocounts leave room for patterns absent from training.

Measured in the 280 training chorales
VoiceIntervals1–2 semitone stepsRepeated notes

Switch between learned and hand-written settings above, keeping the seed, key and form fixed. The learned model guides melody and answer pitches. In chorale mode, learned duration frequencies supply a two-bar rhythm that is reused throughout. Harmony, bass, form and thematic transformations still come from our rules.

Does it predict music it hasn’t seen?
Held-out interval prediction · bits per interval, lower is better
VoiceInterval frequenciesWith previous interval

The conditional model improves prediction on the held-out pieces. This checks the learned statistics, not the beauty of generated music. It is not a comparison against the hand-written composer, and closely related melodies may remain across the split despite catalogue grouping and exact duplicate removal.

What exactly did we learn?

Signed semitone interval probabilities, probabilities conditioned on the previous interval, note-duration counts, and a separate pool of recurring four-note gestures. Tied notes are merged; rests break melodic sequences; written repeats are not expanded. Interval statistics ignore absolute key. Major and minor pieces are pooled. The answering voice pools alto and tenor.

The model uses 0.5 pseudocount per possible interval (−24 to +24 semitones); conditional rows are smoothed toward the marginal with 20 pseudocounts. The composer blends its motion score with 4 × log(49 × learned probability). Those smoothing constants, the blend and the scoring scale are our design choices, not values recovered from C.P.U. Bach.

Learned chorale rhythms sample durations of ½, 1, 1½ or 2 beats, restricted to the space remaining in each bar. Inventions and minuets choose a flowing, measured or offbeat opening pulse from the seed. Subsequent phrases keep its rhythmic cells, with short phrase-end rests. Chorales are a different musical texture from inventions and minuets: transferring their pitch statistics is an experiment.

Download parameters, evaluation and score manifest ↗

Data: Craig Stuart Sapp’s Bach chorale edition. The edition and our derived parameter files are provided under CC BY-NC-SA 4.0. Source attribution.

WHAT WE CAN RECONSTRUCT

We have a blueprint.
Not the original machine.

Sid Meier and Jeff Briggs’s patent describes a hierarchy of musical sections and a “weighted exhaustive search”: score possible notes and rhythmic chunks, then choose the best. Hard rules reject options; softer tendencies favour others. Themes receive their own evaluation.

This is an independent, simplified implementation inspired by that architecture. It does not run the 3DO program or reproduce its exact output.

01 / PLAN

Give the piece a destination.

Our version arranges thematic cells and planned cadences into a 19–23-bar piece. The invention and minuet vary the internal phrase lengths; the chorale keeps three five-bar phrases and a four-bar conclusion. A dominant or subdominant chord leads into the final tonic, with several melodic shapes and voicings.

02 / CHOOSE

Make each note earn its place.

We rank pitches by chord fit, melodic motion, and similarity to a theme. The new learned option adds interval probabilities estimated from Bach scores. Seeded variation changes the scores. In the faster forms, a short beam search compares how each choice affects the rest of its bar. The panel above shows the chosen note alongside local alternatives.

03 / REMEMBER

Let an idea come back.

The chorale compares six deterministic drafts. The invention and minuet search four possible paths through each bar, then compare twelve completed drafts for theme, continuity and phrase fit. The winning draft’s first two bars supply pitch and rhythm cells. Later phrases mostly recall them, with a small lift or dip in the middle passage and a changed answer near two other returns. An answering voice alternates spacious and more active passages underneath. This is free imitation, not a strict fugue.

What is historical, and what is ours?

Documented: section generators, weighted candidate search, rules and tendencies, theme evaluation, and a separate performance stage.

Our choices: the chord plan, scoring weights, pitch ranges, phrase lengths and closing bars, seeded variation, simplified imitation, visualisation and synthesised sounds. The chorale still picks individual notes greedily. The invention and minuet now keep a four-path beam through each thematic bar and a twelve-path cadence continuation search before selecting among complete drafts. The scores and search limits are our choices. Our phrase grammar and cadence plans are hand-authored; they do not reproduce either patent’s full algorithms. Voice-leading is heuristic, not a guarantee of correct Bach counterpoint.

The patent discusses statistical analysis of existing music as earlier research in its background section; it does not establish that C.P.U. Bach learned its own parameters this way. Our corpus model is a new experiment, not a recovery of the original weights. No original source code was located in the sources searched. A patent explains a method; it is not a complete implementation.