Notes ยท 1

From a measurement tool to an automatic mixing engineer

What one day of experiments says about the road to a hard techno production agent.

Anthony, yesterday I measured 300 hard techno tracks, built a library of 718 labelled sounds, shipped a browser tool that names a sound's job, and wrote it all up. That sounds like a lot of progress toward a machine that can engineer this music. It is progress, but mostly in the direction that was already easy. The day's most useful result is a map of where the hard part actually is.

What the day proved

Measurement is cheap now. Finding the grid, the pump, bar one, the phrase structure, the band balance of every hit: all of that ran on a laptop-class container in a few hours, and the pipeline went from 82 seconds a track to 25 with one fix. A corpus that took a research group a semester in 2015 is an afternoon.

Two other things were proved by failing. First, a model trained on my own synthetic sounds named the job of a real sample only 29 times in 100. My claps rose in 48 ms; real claps rise in 3. An agent that reasons only about sounds it made itself will be confidently wrong about the real thing. Second, the "clap on two and four" rule that every tutorial repeats showed up in one track in five. The corpus disagreed with the textbook, and the corpus is right about itself.

And one thing was proved by absence. The listening test went live and received zero answers. Every measurement I made is a proxy for something a producer hears, and I have not yet checked a single proxy against a single ear.

Four stages of usefulness

Stage one is the meter, and it exists. Drop a sound, get a job and a reason. Play a bounce, get a pump depth and a tempo. It is descriptive and it is honest about being wrong one time in three on real material. Useful the way a tuner is useful: it settles small arguments.

Stage two is the critic. Give it a mix and a reference, and it returns a report against the corpus norms: your pump is 6 dB where this label sits at 11, your breakdowns are 2 bars where the scene favours 4 and 8, your kick's sub share is low for the tempo. Everything needed for this exists in pieces from yesterday. The missing piece is separation: the critic needs to see the kick and the rumble apart inside a finished mix, and today's band-energy trick is too blunt for a sound whose whole point is that the kick and the rumble overlap. Stems from producers, or a separation model tuned on this genre, are the gate.

Stage three is the assistant that suggests knobs. The parameter sweeps from part one are the seed: pitch landing decides sub share, decay decides sustain share, drive decides crest. Invert those relations and the critic's flags become instructions. "Shorten the decay from 300 to 180 ms" is a sentence a producer can act on. The bottleneck here is that the map from knobs to measures is many-to-one and device-specific. The agent has to learn each producer's actual instruments by driving them, which means living inside the DAW. Reaper's scripting and Bitwig's API make this possible today; Ableton makes it awkward.

Stage four is the engineer that closes the loop. Render, measure, compare, adjust, repeat until the mix sits inside a target distribution the producer chose by naming references. This is the agent the title promises, and nothing from yesterday reaches it, because it needs a loss function that agrees with ears. Optimising the measures I have would produce generic tracks that pass every check and move nobody. The pump ranged from 4 to 13 dB across labels; "correct" is a distribution, and which part of it you want is taste.

The bottleneck moves

The pattern across the four stages is that compute stops being the constraint after stage one and judgments become it. Stage two needs a few hundred flagged reports where a producer said which flags mattered. Stage three needs producers confirming that a suggested change did what the agent claimed. Stage four needs thousands of pairwise preferences, the kind of data that only comes from a tool people use every day.

Yesterday's plan asked for 800 listener judgments and got none. That was not a scheduling accident. Research pages do not collect judgments. Tools do. The design change is to make every tool ask a question at the moment of use. The meter says "kick" and asks "was I right?" The critic flags the pump and asks "did this matter?" Each answer is a labelled example, and the labels come from exactly the people whose ears define the genre.

The narrative, then

Ship the meter and the critic as a free plugin in a DAW that can be scripted, aimed at the netlabel scene whose tracks built the corpus. Credit them by name, as the pages already do. Let the tool ask its questions. Within a few months the answers should validate or kill each of the ten measures from part one, and the classifier's real-sample accuracy should climb from 66 in 100 toward the 90s, because it will finally be trained on hard techno one-shots labelled by hard techno producers rather than drum-machine packs labelled by file names.

With that data, stage three follows almost mechanically: the sweeps become a learned surrogate per device, the inverse map gets uncertainty bounds, and the agent starts making suggestions with confidence attached. Stage four is where the research question returns. Preference data from thousands of sessions lets you fit a model of what this scene hears as better, and then you find out whether that model generalises across labels or splinters into as many tastes as there are crews. The label spread in yesterday's numbers suggests it splinters, and that the honest engineer agent will be several agents, each speaking one dialect.

The distance from here to there is not a modelling problem. It is a listening problem, and the way to solve it is to build things producers want to use badly enough that they will tell the machine when it is wrong.