Record to MIDI: What Works and What Does Not
Last updated 10 October 2026
Measurement scope: Figures below record historical tests of the GAME Vocal/Fast route in desktop Chrome. The original test date and complete device specifications were not recorded. These numbers do not predict Piano, Basic Pitch, Full Song Fast or HQ Local performance. Guidance reviewed 10 October 2026.
This tool has no live microphone recorder. Save a recording in another app, then choose or drop the audio file here. In a historical GAME Vocal test, an AAC voice memo and a browser Opus recording of the same performance produced matching 40-note MIDI files; that result is specific to those recordings.
Everything below is measured on this tool, in a browser, on the engine this site uses for vocals, humming and single melodies. The numbers are reproducible and the method is described so you can check the shape of it yourself.
First, the part that does not exist
Plenty of pages that convert audio let you sing into them. This one does not, and the useful thing to do is establish that as a fact rather than an impression, because a missing feature is otherwise easy to argue about.
For the original test, we opened the converter homepage in a headless browser and inspected its file input and capture-related JavaScript. The byte counts below belong to that historical build, not the current article page or current homepage.
| JavaScript in the historical homepage test | 984,704 bytes total — 91,867 bytes of the site's own code and 892,837 bytes of third-party code. This is not the current transfer size. |
| References to a capture API in that JavaScript | 0 — searched for getUserMedia, MediaRecorder, navigator.mediaDevices, enumerateDevices, createMediaStreamSource and AudioWorkletNode |
| Same check on the historical build | Its app.js was 74,028 bytes at the time; capture-API matches: 0 |
| The input element | <input type="file">, accept list of 17 extensions, capture attribute present? No |
| Buttons on the page | Download .mid · Play preview · Convert another · Re-run clean-up · Reset to defaults |
| Does the visible text mention recording? | No — the words "record" and "microphone" do not appear in the page's text at all |
| How audio actually gets in | Four ways, all of them files: the file picker, drag-and-drop anywhere on the page, pasting a file from the clipboard, and the "Convert another" button which reopens the picker |
That capture attribute is the small detail worth understanding. A file input with capture set tells a phone to open the camera or the microphone directly. Without it — which is the case here — the phone opens its file browser instead. So the absence of that one attribute is the mechanical reason this page will never ask for microphone permission: it has nothing to ask with.
The route that does work
Recording elsewhere and converting the file is not a workaround with a catch. The accept list on the input already covers what recorders write, so there is no conversion step and no re-export:
- Record. Any recorder will do — a phone voice-memo app, a browser recorder, or a desktop capture tool. Get as close to the source as you can and keep it to one line at a time.
- Save it as a file you can find. On a phone, that usually means sharing or exporting the memo out of the recorder app and onto the same device you are converting on.
- Drop the file into the converter. Pick it, drag it onto the page, or paste it. It is read locally, so nothing is uploaded while you wait.
- Check the piano-roll preview before downloading. The note count and the pitch range are printed on screen. If the range is not the range you sang, the recording is the problem, not the converter.
We ran that route end to end
The honest question is whether a recording made this way survives contact with the converter, or whether the compressed audio a recorder writes costs you something. So we took one performance — a 20-second hummed line, 40 quarter notes at 120 BPM spanning D4 to A4 — and saved it three ways: as the original uncompressed WAV, as an AAC recording at 128 kbps in an M4A container, and as an Opus recording at 96 kbps in a WebM container. The last two are what a phone voice memo and a browser recorder hand you. Each was converted once, in the same browser session, through the real file input.
| File | Size | Notes returned | Pitch range | Output .mid | Time |
|---|---|---|---|---|---|
| WAV, 44.1 kHz mono (baseline) | 1,764,044 B | 40 | D4–A4 | 432 B | 27,361 ms |
| AAC 128 kbps in M4A, 48 kHz mono (phone voice memo) | 324,604 B | 40 | D4–A4 | 432 B | 26,712 ms |
| Opus 96 kbps in WebM, 48 kHz mono (browser recorder) | 335,144 B | 40 | D4–A4 | 432 B | 26,539 ms |
Three files, one performance, 40 notes out of every one of them, the same pitch range, and the same 432-byte output. The recording is 5.4 times smaller than the WAV and the note content is unchanged: the voice memo is 18.4% of the WAV's bytes and the browser recording is 19.0%.
The timing is worth a sentence too. Twenty seconds of audio took between 26.5 and 27.4 seconds to convert on this machine, so the work runs at roughly 1.3 times the length of the recording. The compressed inputs were not measurably faster or slower than the uncompressed one — the cost tracks how long the audio is, not how big the file is.
The byte-level comparison, because "the same" is a claim
"Both produced 40 notes" is a weak statement. The stronger one is what the files actually contain, so we read all three back and compared them event by event. A .mid file here holds 88 events: 40 note-on messages, 40 note-off messages, 4 controller messages that declare the pitch-bend range, and 4 meta events — the tempo, the time signature, the track name and the end-of-track marker.
- The two recordings produced byte-identical files. The AAC output and the Opus output have the same 432 bytes and the same SHA-256 hash. Not similar — identical.
- Against the WAV, the difference is confined to the last two events in the file. The final note-off sits at tick 19,190 in the recorded versions and tick 19,200 in the WAV version — 10 ticks. At the 120 BPM written into the file, one tick is about 1.04 ms, so the gap is about 10.4 milliseconds.
- The other 86 events are identical — every note-on, every note-off, every controller message, the tempo, the time signature and the track name. The second of the two events that move is the end-of-track marker, which always sits immediately after the last note-off and so travels with it.
The cause is not mysterious: an AAC encoder adds a little padding at the start of the stream, so the decoded audio ends a fraction of a millisecond to a few milliseconds earlier than the original. The converter saw a marginally shorter buffer and wrote the last note-off marginally earlier. A ten-millisecond difference in the release of the final note is not audible and is not worth chasing.
There is a useful side effect of running this comparison: the WAV output in this run has the same hash as an earlier, independent run of the same WAV in a different test — which is how we know the pipeline is deterministic. Same input, same output, every time.
What each recorder hands you
The format you end up with depends on what you record with, and all of the common ones are already in the accept list, so none of them needs converting first.
| Recorder | What it writes | Does it go straight in? |
|---|---|---|
| iPhone or iPad voice memo | M4A holding AAC audio | Yes — tested, 40 notes out |
| Browser recorder (MediaRecorder) | WebM holding Opus audio | Yes — tested, byte-identical to the AAC result |
| Android recorder | Usually M4A/AAC; sometimes OGG holding Opus | Yes — both are in the accept list |
| Desktop capture tool, e.g. Audacity | WAV by default | Yes — uncompressed, the shortest path to analysis |
| Screen or system-audio capture | MP4, MOV or WebM | Yes — the audio track is read and the video ignored |
If you have a choice, WAV is the safest pick, but the reason is narrower than people expect. It is not that compressed audio damages the transcription — in our test the compressed recordings produced the same notes as the uncompressed one. It is that WAV removes one thing that could go wrong, and it is the format every recorder can produce.
Where recording, rather than converting, is what fails
None of the failures below are the converter's fault, and all of them are visible in the piano-roll preview before you download anything.
- A room, not a source. A microphone in a room records the room as well as the performance. Reflections and background noise are pitched information, and a model looking for a melody has to find it in that. Record close to the instrument or the mouth, in the quietest space you have, and record one line at a time.
- Too quiet, or clipping. Aim for a healthy level that does not touch the top of the meter. A recording so quiet that the melody is near the noise floor, or so loud that it distorts, is damaged before the converter ever sees it.
- Several things at once. The engine on this page follows a single line. A phone held up in front of a band records a mix, and a mix is the hard case for a melody engine. This site offers separate engines for piano, guitar, bass and other instruments, and a separate route for a full mix — but for a recording, the cleanest fix is to record one part at a time.
- Waiting on a long take. Conversion runs at roughly 1.3 times the length of the audio and the tab has to stay open throughout. Record a short excerpt first, check the roll, then commit to the whole thing.
What the file it writes actually is
Whatever route the audio took to get in, the writer behaves the same way, and we confirmed this by reading the three output files back byte by byte.
| Format | Standard MIDI File, format 0 — one track, every note in it |
| Resolution | 480 ticks per quarter note |
| Analysis sample rate | 44,100 Hz, mono, after decoding and down-mixing — a 48 kHz stereo recording is resampled and mixed down before the model sees it |
| Tempo | Measured from the note onsets and written into the file; 500,000 µs per quarter note (120 BPM) only when no steady pulse can be found |
| Time signature | 4/4, a constant |
| Track name | mp3 to midi |
| Pitch-bend range | ±2 semitones, declared with an RPN 0 message, centre 8192 |
| Velocity | clamp(amplitude × 127, 1, 127) — on all 120 notes we measured here it came out as 90, because the engine reported the same confidence value for every one of them |
| Note-off velocity | Fixed at 0x40 (64) on every note |
| Shortest note | 0.02 s — anything shorter is clamped up to it |
| Pitch bends written | None, on every file we measured here — the range is declared but no bend data is emitted |
| Audio content | None. The file carries note events only — 0 bytes of samples |
One number in that table is worth pulling out, because it is the one people are surprised by when they record themselves singing. Every note carried the same velocity — 90 — across all three files, because the engine reports the same confidence value (0.71) for each note it emits. There is no dynamic information in the output. If you recorded a phrase with a crescendo in it, the .mid will not contain the crescendo; you will draw that in your DAW afterwards.
Frequently asked questions
Can I record straight into this converter with my microphone?
No. The current converter accepts saved files through a picker, drag-and-drop or paste; it has no live microphone button. A historical audit of the converter homepage found no capture API calls, but its old JavaScript byte count is not a current performance measurement.
Why is there no record button?
Because the engine takes a decoded audio buffer, not a live stream. A microphone gives you a stream that has to be recorded, encoded into a container, saved, and then decoded before analysis can start — a separate pipeline with its own permission prompt, its own length limits and its own failure modes. This tool is file in, file out, and it says so. The input element carries an accept list of 17 extensions and no capture attribute, which is exactly why tapping the drop zone opens a file browser instead of the microphone. The page has five buttons and none of them records: Download .mid, Play preview, Convert another, Re-run clean-up and Reset to defaults.
I recorded a voice memo on my phone. Will it convert?
Yes. We made a mono 48 kHz AAC recording at 128 kbps in an M4A container — the shape a phone voice-memo app writes — from a 20-second hummed line of 40 quarter notes, and converted it. It came back as 40 notes, a pitch range of D4 to A4, and a 432-byte .mid file. A browser recorder's output in Opus at 96 kbps in a WebM container produced the identical result, byte for byte. Both are in the accept list of the file input, so both go straight in with no re-encoding step.
Which file does each recorder give me, and does it convert?
An iPhone voice memo writes an M4A holding AAC audio; a browser recorder using MediaRecorder writes WebM holding Opus; Android's recorder usually writes M4A, sometimes OGG holding Opus; a desktop capture tool such as Audacity writes WAV by default. All of those are in the accept list, and the ones we tested — AAC in M4A and Opus in WebM — both converted on the first try, with no error and no re-encoding step. WAV is the safest choice if you have the option, only because it is the shortest path to the analysis: it is already uncompressed.
The two recordings converted to identical bytes. Does recording quality not matter?
It matters, but not through the file format. The AAC voice memo and the Opus recorder file produced the same 432 bytes because the signal underneath them was the same clean line — the codecs had nothing to damage. What damages a recording is what a microphone picks up: room reflections, background noise, a source that is too quiet, clipping, or several instruments at once. Those change the audio the model sees, and no container choice compensates for them. Record close, record clean, record one line, and check the piano-roll preview before you import.
Why did the last note come out slightly short?
We compared the .mid files event by event. The output from the AAC recording and the output from the WAV of the same performance differ in exactly two of their 88 events, and both are the last two in the file: the final note-off sits at tick 19,190 instead of tick 19,200, and the end-of-track marker follows it. That is 10 ticks, which at the 120 BPM written into the file is about 10.4 milliseconds. Everything before those two events — all 86 of them — is identical. The cause is the small amount of padding an AAC encoder adds, which shifts the end of the decoded audio slightly. A ten-millisecond difference at the very end of a take is not something you will hear.
How long can a recording be?
There is no length limit imposed by the page, but there is a practical one, and it is set by your device rather than by the recorder. In our tests a 20-second file took between 26.5 and 27.4 seconds to convert, so the work runs at roughly 1.3 times the length of the audio — a ten-minute recording is a thirteen-minute wait, and the tab has to stay open for all of it. For a quick idea, record 20 or 30 seconds, convert that, and only then commit to the full take.
Related: what leaves your device when you convert, preview and inspect MIDI notes in the converter, and why a recording with drums in it comes out wrong.