mp3→midi runs in your browser · nothing is uploaded

Record to MIDI: What Works and What Does Not

Last updated 10 October 2026

Measurement scope: Figures below record historical tests of the GAME Vocal/Fast route in desktop Chrome. The original test date and complete device specifications were not recorded. These numbers do not predict Piano, Basic Pitch, Full Song Fast or HQ Local performance. Guidance reviewed 10 October 2026.

This tool has no live microphone recorder. Save a recording in another app, then choose or drop the audio file here. In a historical GAME Vocal test, an AAC voice memo and a browser Opus recording of the same performance produced matching 40-note MIDI files; that result is specific to those recordings.

Everything below is measured on this tool, in a browser, on the engine this site uses for vocals, humming and single melodies. The numbers are reproducible and the method is described so you can check the shape of it yourself.

First, the part that does not exist

Plenty of pages that convert audio let you sing into them. This one does not, and the useful thing to do is establish that as a fact rather than an impression, because a missing feature is otherwise easy to argue about.

For the original test, we opened the converter homepage in a headless browser and inspected its file input and capture-related JavaScript. The byte counts below belong to that historical build, not the current article page or current homepage.

JavaScript in the historical homepage test984,704 bytes total — 91,867 bytes of the site's own code and 892,837 bytes of third-party code. This is not the current transfer size.
References to a capture API in that JavaScript0 — searched for getUserMedia, MediaRecorder, navigator.mediaDevices, enumerateDevices, createMediaStreamSource and AudioWorkletNode
Same check on the historical buildIts app.js was 74,028 bytes at the time; capture-API matches: 0
The input element<input type="file">, accept list of 17 extensions, capture attribute present? No
Buttons on the pageDownload .mid · Play preview · Convert another · Re-run clean-up · Reset to defaults
Does the visible text mention recording?No — the words "record" and "microphone" do not appear in the page's text at all
How audio actually gets inFour ways, all of them files: the file picker, drag-and-drop anywhere on the page, pasting a file from the clipboard, and the "Convert another" button which reopens the picker

That capture attribute is the small detail worth understanding. A file input with capture set tells a phone to open the camera or the microphone directly. Without it — which is the case here — the phone opens its file browser instead. So the absence of that one attribute is the mechanical reason this page will never ask for microphone permission: it has nothing to ask with.

Why it is built that way The engine's input is a decoded audio buffer. A microphone hands you a live stream, which has to be recorded to a blob, encoded into a container, and decoded again before analysis can begin — a separate pipeline with its own permission prompt, its own length limits and its own ways to fail. This tool is file in, file out. Adding a recorder would mean adding that whole second pipeline, and a page that claims nothing leaves your device is a bad place to bolt one on casually.

The route that does work

Recording elsewhere and converting the file is not a workaround with a catch. The accept list on the input already covers what recorders write, so there is no conversion step and no re-export:

  1. Record. Any recorder will do — a phone voice-memo app, a browser recorder, or a desktop capture tool. Get as close to the source as you can and keep it to one line at a time.
  2. Save it as a file you can find. On a phone, that usually means sharing or exporting the memo out of the recorder app and onto the same device you are converting on.
  3. Drop the file into the converter. Pick it, drag it onto the page, or paste it. It is read locally, so nothing is uploaded while you wait.
  4. Check the piano-roll preview before downloading. The note count and the pitch range are printed on screen. If the range is not the range you sang, the recording is the problem, not the converter.

We ran that route end to end

The honest question is whether a recording made this way survives contact with the converter, or whether the compressed audio a recorder writes costs you something. So we took one performance — a 20-second hummed line, 40 quarter notes at 120 BPM spanning D4 to A4 — and saved it three ways: as the original uncompressed WAV, as an AAC recording at 128 kbps in an M4A container, and as an Opus recording at 96 kbps in a WebM container. The last two are what a phone voice memo and a browser recorder hand you. Each was converted once, in the same browser session, through the real file input.

FileSizeNotes returnedPitch rangeOutput .midTime
WAV, 44.1 kHz mono (baseline)1,764,044 B40D4–A4432 B27,361 ms
AAC 128 kbps in M4A, 48 kHz mono (phone voice memo)324,604 B40D4–A4432 B26,712 ms
Opus 96 kbps in WebM, 48 kHz mono (browser recorder)335,144 B40D4–A4432 B26,539 ms

Three files, one performance, 40 notes out of every one of them, the same pitch range, and the same 432-byte output. The recording is 5.4 times smaller than the WAV and the note content is unchanged: the voice memo is 18.4% of the WAV's bytes and the browser recording is 19.0%.

The timing is worth a sentence too. Twenty seconds of audio took between 26.5 and 27.4 seconds to convert on this machine, so the work runs at roughly 1.3 times the length of the recording. The compressed inputs were not measurably faster or slower than the uncompressed one — the cost tracks how long the audio is, not how big the file is.

The byte-level comparison, because "the same" is a claim

"Both produced 40 notes" is a weak statement. The stronger one is what the files actually contain, so we read all three back and compared them event by event. A .mid file here holds 88 events: 40 note-on messages, 40 note-off messages, 4 controller messages that declare the pitch-bend range, and 4 meta events — the tempo, the time signature, the track name and the end-of-track marker.

The cause is not mysterious: an AAC encoder adds a little padding at the start of the stream, so the decoded audio ends a fraction of a millisecond to a few milliseconds earlier than the original. The converter saw a marginally shorter buffer and wrote the last note-off marginally earlier. A ten-millisecond difference in the release of the final note is not audible and is not worth chasing.

There is a useful side effect of running this comparison: the WAV output in this run has the same hash as an earlier, independent run of the same WAV in a different test — which is how we know the pipeline is deterministic. Same input, same output, every time.

What each recorder hands you

The format you end up with depends on what you record with, and all of the common ones are already in the accept list, so none of them needs converting first.

RecorderWhat it writesDoes it go straight in?
iPhone or iPad voice memoM4A holding AAC audioYes — tested, 40 notes out
Browser recorder (MediaRecorder)WebM holding Opus audioYes — tested, byte-identical to the AAC result
Android recorderUsually M4A/AAC; sometimes OGG holding OpusYes — both are in the accept list
Desktop capture tool, e.g. AudacityWAV by defaultYes — uncompressed, the shortest path to analysis
Screen or system-audio captureMP4, MOV or WebMYes — the audio track is read and the video ignored

If you have a choice, WAV is the safest pick, but the reason is narrower than people expect. It is not that compressed audio damages the transcription — in our test the compressed recordings produced the same notes as the uncompressed one. It is that WAV removes one thing that could go wrong, and it is the format every recorder can produce.

Where recording, rather than converting, is what fails

None of the failures below are the converter's fault, and all of them are visible in the piano-roll preview before you download anything.

What the file it writes actually is

Whatever route the audio took to get in, the writer behaves the same way, and we confirmed this by reading the three output files back byte by byte.

FormatStandard MIDI File, format 0 — one track, every note in it
Resolution480 ticks per quarter note
Analysis sample rate44,100 Hz, mono, after decoding and down-mixing — a 48 kHz stereo recording is resampled and mixed down before the model sees it
TempoMeasured from the note onsets and written into the file; 500,000 µs per quarter note (120 BPM) only when no steady pulse can be found
Time signature4/4, a constant
Track namemp3 to midi
Pitch-bend range±2 semitones, declared with an RPN 0 message, centre 8192
Velocityclamp(amplitude × 127, 1, 127) — on all 120 notes we measured here it came out as 90, because the engine reported the same confidence value for every one of them
Note-off velocityFixed at 0x40 (64) on every note
Shortest note0.02 s — anything shorter is clamped up to it
Pitch bends writtenNone, on every file we measured here — the range is declared but no bend data is emitted
Audio contentNone. The file carries note events only — 0 bytes of samples

One number in that table is worth pulling out, because it is the one people are surprised by when they record themselves singing. Every note carried the same velocity — 90 — across all three files, because the engine reports the same confidence value (0.71) for each note it emits. There is no dynamic information in the output. If you recorded a phrase with a crescendo in it, the .mid will not contain the crescendo; you will draw that in your DAW afterwards.

Frequently asked questions

Can I record straight into this converter with my microphone?

No. The current converter accepts saved files through a picker, drag-and-drop or paste; it has no live microphone button. A historical audit of the converter homepage found no capture API calls, but its old JavaScript byte count is not a current performance measurement.

Why is there no record button?

Because the engine takes a decoded audio buffer, not a live stream. A microphone gives you a stream that has to be recorded, encoded into a container, saved, and then decoded before analysis can start — a separate pipeline with its own permission prompt, its own length limits and its own failure modes. This tool is file in, file out, and it says so. The input element carries an accept list of 17 extensions and no capture attribute, which is exactly why tapping the drop zone opens a file browser instead of the microphone. The page has five buttons and none of them records: Download .mid, Play preview, Convert another, Re-run clean-up and Reset to defaults.

I recorded a voice memo on my phone. Will it convert?

Yes. We made a mono 48 kHz AAC recording at 128 kbps in an M4A container — the shape a phone voice-memo app writes — from a 20-second hummed line of 40 quarter notes, and converted it. It came back as 40 notes, a pitch range of D4 to A4, and a 432-byte .mid file. A browser recorder's output in Opus at 96 kbps in a WebM container produced the identical result, byte for byte. Both are in the accept list of the file input, so both go straight in with no re-encoding step.

Which file does each recorder give me, and does it convert?

An iPhone voice memo writes an M4A holding AAC audio; a browser recorder using MediaRecorder writes WebM holding Opus; Android's recorder usually writes M4A, sometimes OGG holding Opus; a desktop capture tool such as Audacity writes WAV by default. All of those are in the accept list, and the ones we tested — AAC in M4A and Opus in WebM — both converted on the first try, with no error and no re-encoding step. WAV is the safest choice if you have the option, only because it is the shortest path to the analysis: it is already uncompressed.

The two recordings converted to identical bytes. Does recording quality not matter?

It matters, but not through the file format. The AAC voice memo and the Opus recorder file produced the same 432 bytes because the signal underneath them was the same clean line — the codecs had nothing to damage. What damages a recording is what a microphone picks up: room reflections, background noise, a source that is too quiet, clipping, or several instruments at once. Those change the audio the model sees, and no container choice compensates for them. Record close, record clean, record one line, and check the piano-roll preview before you import.

Why did the last note come out slightly short?

We compared the .mid files event by event. The output from the AAC recording and the output from the WAV of the same performance differ in exactly two of their 88 events, and both are the last two in the file: the final note-off sits at tick 19,190 instead of tick 19,200, and the end-of-track marker follows it. That is 10 ticks, which at the 120 BPM written into the file is about 10.4 milliseconds. Everything before those two events — all 86 of them — is identical. The cause is the small amount of padding an AAC encoder adds, which shifts the end of the decoded audio slightly. A ten-millisecond difference at the very end of a take is not something you will hear.

How long can a recording be?

There is no length limit imposed by the page, but there is a practical one, and it is set by your device rather than by the recorder. In our tests a 20-second file took between 26.5 and 27.4 seconds to convert, so the work runs at roughly 1.3 times the length of the audio — a ten-minute recording is a thirteen-minute wait, and the tab has to stay open for all of it. For a quick idea, record 20 or 30 seconds, convert that, and only then commit to the full take.

Related: what leaves your device when you convert, preview and inspect MIDI notes in the converter, and why a recording with drums in it comes out wrong.