mp3→midi runs in your browser · nothing is uploaded

How Long Does Audio-to-MIDI Conversion Take?

Last updated 10 October 2026

Measurement scope: Figures below record historical tests of the GAME Vocal/Fast route in desktop Chrome. The original test date and complete device specifications were not recorded. These numbers do not predict Piano, Basic Pitch, Full Song Fast or HQ Local performance. Guidance reviewed 10 October 2026.

In a historical desktop Chrome test of GAME Vocal, conversion took roughly 1.4 seconds per second of audio after the model was ready. A 30-second clip took 44 seconds, one minute took 84 seconds, two minutes took 3 minutes 1 second, and three minutes took 4 minutes 41 seconds. The older build also spent 12 to 14 seconds starting its model. These figures are not a timing estimate for current Piano, Basic Pitch or HQ Local modes.

Those figures come from seven conversions run back to back in the same browser tab, on audio built to be exactly 5, 10, 20, 30, 60, 120 and 180 seconds long and to contain the same music at the same density throughout — two notes per second in every file. Timing starts the moment the file is handed to the page and stops when the download link appears.

The short version In the measured GAME Vocal route, inference dominated and used eight-second windows with context. A nine-second clip crossed a window boundary and cost more work than an eight-second clip. The roughly 12-second new-tab startup was measured on the older eager-loading build; current model loading starts when you first convert.

The measured table

Desktop Chrome, no other work on the machine, engine loaded and warm. "Audio fed" is what the engine's window planner actually sends to the model — it is the length of the file plus the context around each window.

Audio lengthWindowsAudio fedAnalysisTotal waitms per second of audio
5 s15.0 s5.6 s6.4 s1,126
10 s212.0 s13.6 s13.9 s1,358
20 s324.0 s28.4 s29.1 s1,421
30 s436.0 s43.4 s44.1 s1,446
60 s874.0 s83.7 s84.4 s1,395
120 s15148.0 s180.5 s181.4 s1,504
180 s23224.0 s280.1 s281.1 s1,556

Read the last two columns together, because they are the whole answer. Total wait rises faster than audio length — from 1,126 ms per second of audio on a 5-second clip to 1,556 ms on a three-minute one. But the number of seconds the model is actually asked to analyse rose from 5.0 to 224.0 over the same range, and the cost per second of that stayed between 1,126 and 1,250 ms the whole way. Nothing is slowing down as files get longer. The engine is simply being handed more audio than the file contains.

One repeat, to show how much of this is machine noise rather than structure: a second, independent three-minute conversion in a fresh tab took 4 minutes 20 seconds instead of 4 minutes 41. That is an 8% spread on identical input, so treat every number here as a bracket, not a stopwatch reading.

Where the time actually goes

The page reports four phases, and only one of them matters. These are the same seven conversions, broken down.

Decoding the file0.1 to 0.5 s. Turning your MP3, WAV or M4A into raw samples. Grows with file length, stays trivial.
Preparing the audio0.02 to 0.22 s. Downmixing to mono and resampling to the 44,100 Hz the model expects. Effectively free.
Analysis99.6% of the time on the three-minute file — 280.1 of 281.1 seconds. This is the neural network running on the CPU.
Writing the MIDI0 ms that the clock can resolve. A 359-note file is a few kilobytes; encoding it is instant.

Everything outside the model — reading the file, resampling it, writing the .mid — came to between 0.3 and 1.0 seconds per conversion, whatever the length. If you were hoping to speed things up by feeding the tool a smaller file format, that is where the hope dies: the decode step you would be optimising is 0.2% of the wait.

Historical GAME startup: a cost paid once per page

The older GAME build fetched and initialised its model before the first file. The table below records that historical behavior. The current site waits until a file is selected, and startup differs by mode and device.

SituationWait before the tool is ready
Browser that has never loaded the model13.9 s
New tab, same browser, model already cached11.6 – 11.9 s
Second file in a tab that has already converted something0 s

In that older build, the gap between the first two rows was about 2 seconds of downloading on the test connection. The remaining startup mainly initialised five ONNX sessions. Its 51,054,395 bytes of model weights could be reused from browser cache. Current cache behavior and model size depend on the selected mode and browser storage state.

The practical consequence is that the arithmetic people do in their heads is wrong at the short end. A 5-second clip does not take a twelfth of the time of a one-minute file. It takes 6.4 seconds against 84.4 — because the one-minute file amortises the same start-up over twelve times as much audio.

Reason two: the engine works in 8-second windows with context

This is the part that shows up in the numbers and nowhere else. The engine does not analyse your file as one signal. It cuts it into windows of 8 seconds, and for each window it sends 1 second of audio from before and after as context, so the model can hear what is coming. Only the middle 8 seconds of each window are kept as output.

Two things follow. First, the audio fed to the model is longer than your file. Second, the file length is rounded up to a whole number of windows, so crossing an 8-second boundary adds a full extra window's worth of work at once. Here is what the planner does to a range of lengths:

Audio lengthWindowsAudio fed to the modelFed ÷ your audio
8 s18.0 s1.000
9 s211.0 s1.222
16 s218.0 s1.125
24 s328.0 s1.167
32 s438.0 s1.188
64 s878.0 s1.219

Look at the first two rows. Going from an 8-second clip to a 9-second clip adds 12.5% more audio and 37.5% more model work, because the ninth second forces a second window and that second window arrives with its own context. A 9-second clip is relatively more expensive than a 16-second one. This is the only place in the whole pipeline where the cost is genuinely discontinuous, and it is a boundary you can see: at 120 BPM in 4/4, one window is exactly four bars, and the context is half a bar on each side.

The same effect explains the slow drift in the per-second price. A file that fits in one window pays no context at all — 1.000×. As files get longer, the proportion of audio that is context settles toward the 25% ceiling, so the ratio climbs from 1.000 through 1.200, 1.233 and 1.244 and stays there. Measured, on a clean set of single-file runs, the cost per second of audio fed was 1.09 s at 10 seconds, 1.16 s at 60 seconds and 1.15 s at 180 seconds — flat, while the cost per second of your audio rose from 1.31 s to 1.44 s over the same span. The drift is entirely accounting.

Reason three: your machine, not the file

Everything above was measured on one desktop. To see how much of the result is hardware, we ran two of the same files again with the CPU throttled to a quarter of its speed — an emulated slow device, not a physical phone, and it should be read as a direction rather than a specification.

Audio lengthFull speedCPU throttled to 25%Ratio
10 s13.9 s25.7 s1.85×
30 s44.1 s87.2 s1.98×

The throttled run came out about twice as slow, not four times — 1.85× on the 10-second file and 1.98× on the 30-second one. The shape of the curve was unchanged: the 30-second file cost 3.2 times the 10-second file at full speed and 3.4 times when throttled. So if you are on a slower machine, scale the whole table and keep the relationships.

What moved least was start-up. For the same cached-model case, the engine took 14.9 seconds to become ready with the CPU throttled against 11.6 to 11.9 seconds at full speed — about 25% slower rather than twice as slow. Most of that wait is fetching and creating five inference sessions rather than computing, so it is the part of the pipeline a slow CPU hurts least.

What the progress bar is really telling you

The bar moves once per window — 23 steps for a three-minute file, 4 for a 30-second one — and the time estimate next to it is derived from how long the work has taken against how far the bar has got. That method has one structural flaw: progress is reported just before a window runs, not after it finishes, so the estimate is always computed from one window less work than the bar claims.

Here is the three-minute run, panel against reality:

ProgressPanel saidActually remainingError
31%~2:313:5335% low
46%~2:333:0417% low
68%~1:411:519% low
87%~47s0:448% high
98%~12s0:05164% high

The estimate converges as the run goes on and then overshoots at the end, for the same reason in both directions: it assumes every window costs the same, and the first window costs more than the rest while the last one is usually short. On the 10-second file the flaw is at its worst — the panel said ~1s left at 55% while 13.6 seconds of work remained, because the estimate was computed from the six milliseconds of elapsed time between the analysis starting and the first window being handed off.

Two details of the display are worth knowing before you read anything into it. The bar never shows 0% or 100% during analysis; it is mapped across 12% to 98%. And any estimate under one second is printed as "~1s left" — a floor in the formatting, not a measurement. If you see that, you know the estimate has collapsed, not that the job is nearly done.

Practical answers

Frequently asked questions

How long does a three-minute song take to convert?

In the historical GAME Vocal desktop test, two runs of the same three-minute file took 4 minutes 41 seconds and 4 minutes 20 seconds with the model already loaded. The older build added 11.6 to 13.9 seconds of startup in a new tab. Current models and devices can take very different times; try a short clip first.

Why does a 10-second clip take 14 seconds?

Because the engine never processes your file in one piece. It splits the audio into 8-second windows and feeds each window with 1 second of context on either side, so a 10-second file becomes two windows carrying 12 seconds of audio. Measured: 13.6 seconds of analysis for a 10-second file, against 5.6 seconds for a 5-second file that fits in a single window. The doubling is not in your file, it is in the windowing.

Does conversion time grow in proportion to the length of the audio?

Only above about a minute, and even then not exactly. The cost per second of audio was 1.13 s at 5 seconds of audio, 1.36 s at 10 seconds, 1.40 s at 60 seconds and 1.56 s at 180 seconds. The reason the price per second rises is that the model is fed 1.2 to 1.24 times the audio you supplied, because every 8-second window carries a second of context on each side. The cost per second of audio the model actually sees stayed between 1.09 and 1.25 seconds across every file we ran.

Why does the progress bar move in steps instead of smoothly?

Because there is nothing to report between windows. The engine hands off one 8-second window at a time and reports progress once per window, so a three-minute file moves the bar in 23 steps and a 30-second file in 4. There is no finer-grained signal available — the inference inside a window is a single opaque call.

Why is the time estimate so wrong at the beginning?

Because progress is reported before the window it refers to has run, and the estimate divides elapsed time by progress. On a three-minute file the panel said ~2:31 left at 31% while 3 minutes 53 seconds of work actually remained — 35% low. On a 10-second file it said ~1s left while 13.6 seconds of work remained. The estimate becomes accurate at about 83% of the way through, then overshoots, because the final window is usually a short one: the last window of the three-minute file carried 4.5 seconds of audio against roughly 12 seconds for a full one. Any estimate below one second is displayed as ~1s left.

How long does it take on a phone?

We did not test a phone, and we are not going to invent a number for one. What we can offer is a directional measurement: with the CPU throttled to a quarter of its speed, the same 10-second file went from 13.9 to 25.7 seconds and the 30-second file from 44.1 to 87.2 seconds — roughly twice the wall clock, not four times. The shape of the curve was unchanged, so on slower hardware expect every number on this page to scale, with the one-time start-up staying closest to constant because most of it is downloading and initialising rather than computing.

Can I make a conversion finish faster?

Convert only the section you need, and keep the tab open for later files so the selected model can be reused. In the historical GAME Vocal test, eight-second windows made a nine-second clip more expensive than an eight-second clip. Current modes use different models and may have very different timing, so use a short representative clip to estimate the wait on your device.

Related: how long a file you can convert before the browser gives up, what local conversion costs you in speed and saves you in privacy, and whether a denser file takes longer.