MP3 Bitrate: 128 vs 192 vs 320 kbps Explained
By Mark Fulton · 2026-09-14 · 11 min read

MP3 bitrate is the number of bits the encoder is allowed to spend on each second of sound, and it sets file size exactly: an hour of constant bitrate audio is 57.6 MB at 128 kbps, 86.4 MB at 192 kbps and 144 MB at 320 kbps. Pick 128 kbps for speech, 192 kbps for anything with music in it, and 320 kbps only when the source is a lossless or very high quality recording. The rule that matters more than any tier: a bitrate can only preserve what the source already has. Extracting a 128 kbps soundtrack at 320 kbps gives you a file two and a half times bigger that sounds no better, and often slightly worse.
That is the whole decision. The rest of this page is the arithmetic behind it, the evidence on when a difference becomes audible, and the one mistake that quietly costs quality twice.
What does an audio bitrate describe?
A bitrate is a budget, not a quality score. "128 kbps" means 128,000 bits of data for every second of audio, and the encoder decides how to spend them. MP3 is a lossy format: the encoder analyses the sound, predicts which parts a listener is least likely to notice, and throws those away first. A bigger budget means it has to throw away less.
Three things follow from that, and each one clears up a common misunderstanding.
Bitrate is not the same as sample rate or bit depth. Sample rate (44.1 kHz, 48 kHz) is how many times per second the original waveform was measured. Bit depth (16 bit, 24 bit) is how precisely each measurement was stored. Bitrate is what is left after the encoder has compressed all of that down. You can have a 48 kHz file at 128 kbps or at 320 kbps.
The same bitrate does not mean the same quality across encoders. An MP3 from a good encoder and one from a poor encoder can share a bitrate and sound noticeably different. The encoder that has been refined the longest is LAME, which Hydrogenaudio's LAME knowledgebase entry describes as its recommended MP3 encoder, developed as open source since 1998. VidClip's extractor uses LAME, through FFmpeg's libmp3lame wrapper.
Constant and variable bitrate behave differently. Constant bitrate (CBR) gives every frame the same budget, so a silent pause costs as much as a crashing chorus. Variable bitrate (VBR) shifts bits toward the difficult passages and targets a quality level instead of a number. VBR is more efficient. CBR is predictable, which is why it is the mode that makes file size pure arithmetic. VidClip's MP3 tool encodes at constant bitrate, so the numbers below are exactly what you are choosing between.
How many megabytes is an hour at each tier?
At constant bitrate there is nothing to measure. Size is bitrate multiplied by time:
MB per hour = kbps × 3,600 ÷ 8 ÷ 1,000
Multiply kilobits per second by 3,600 seconds to get kilobits per hour, divide by 8 to get kilobytes, and divide by 1,000 to get megabytes. These figures use decimal megabytes, where 1 MB is 1,000,000 bytes. If your operating system reports sizes in binary units (MiB), divide by 1.048576, so 86.4 MB shows as roughly 82.4 MiB. All figures cover the audio stream only. Tags, embedded cover art and frame headers add a little on top.
The table is the part worth bookmarking. The last column is the one that decides whether a tier is worth its size.
| Bitrate | MB per hour | The content it suits | Can the source even justify it? |
|---|---|---|---|
| 32 kbps, mono | 14.4 | Speech headed to a transcription service | Yes, from almost anything with a voice in it. This is what VidClip's transcriber sends. |
| 64 kbps | 28.8 | Spoken word archives where size matters most | Yes for a single voice. Thin for music. |
| 96 kbps | 43.2 | Mono podcasts | Yes for speech recorded on a decent microphone. |
| 128 kbps | 57.6 | Lectures, interviews, meetings, talk podcasts | Yes from a screen recording, a video call, or a phone video of someone talking. A clip exported by VidClip's compressor carries 128 kbps AAC, so this is its natural ceiling. |
| 160 kbps | 72.0 | Stereo podcasts with music beds | Yes if the music was mixed in cleanly rather than picked up by a room microphone. |
| 192 kbps | 86.4 | Music for everyday listening, mixed soundtracks | Yes from a good camera, a concert clip with a direct feed, or a music video in a high quality file. |
| 256 kbps | 115.2 | Music you care about, stereo podcasts with heavy music | Only if the source audio is itself high bitrate or lossless. |
| 320 kbps | 144.0 | The ceiling of the MP3 format | Only from lossless or uncompressed audio, such as a recorder or camera that stores PCM. Almost never from a downloaded or shared video. |
| CD audio, uncompressed | 635.04 | The reference everything is compared against | Not an MP3 option. 44,100 samples × 16 bits × 2 channels = 1,411.2 kbps. |
The three rows in bold are the settings on VidClip's video to MP3 tool. The other rows are there so you can see where they sit.
Two patterns jump out. The jump from 128 to 320 kbps adds 86.4 MB to every hour, which is a whole extra 192 kbps file stacked on top. And CD audio is 635.04 MB an hour, so even a 320 kbps MP3 is keeping less than a quarter of the raw data. Lossy audio works because the part it discards is chosen carefully, not because the budget is generous. The same size equals bitrate times duration relationship drives video too, and the video compression arithmetic walks through it for the picture side of a file.
When can you hear the difference?
Less often than the size difference suggests, and the best evidence comes from blind testing rather than opinion.
The standard method in audio circles is the ABX test: you hear A, you hear B, then you hear X and have to say which one it was, many times over. If you cannot beat chance, you cannot hear a difference, whatever you believe. On that basis, Hydrogenaudio's LAME entry reports that LAME encodes typically reach transparency, meaning a listener cannot tell them from the original, at bitrates well below the maximum, and that "encoding with higher-bitrate settings will have no effect on the perceived quality." On 320 kbps CBR specifically, the same page says nobody has produced ABX results showing it ever sounds better than LAME's highest VBR presets.
The Hydrogenaudio page on transcoding gives a rough marker: for MP3 with LAME, transparency is usually reached around 192 kbps. That is where the default on the extractor comes from.
"Usually" is doing real work in that sentence. Transparency depends on the listener, the equipment and above all the material. When you go looking for a difference, listen where the encoder has the most to describe:
- Cymbals, hi hats and applause. Dense, noisy, high frequency sound.
- Hand claps and plucked strings. Sudden attacks that start from near silence.
- Reverb tails and room ambience. Quiet, complex detail at the edge of hearing.
- Wide stereo mixes. Lots of difference between the left and right channels.
One person talking into a microphone contains very little of that, which is why spoken word is routinely published at the lower tiers. Playback matters too. Laptop speakers, a car on the highway or earbuds on a train can easily mask details that a quiet room and good headphones would let you notice.
The test that settles it for your own file takes two minutes. Export the same clip at 128 and 192 kbps, find the busiest ten seconds, and play both at the same volume on the equipment you actually use. If you cannot pick the 128 out, you have your answer, and you certainly do not need 320.
Why does re-encoding compressed audio cost you twice?
This is the section that matters most if your audio is coming out of a video, and it is the part almost nobody mentions.
The audio inside a typical MP4 is not raw. It is already lossy, usually AAC, which means an encoder has already thrown information away once. To make an MP3, that AAC has to be decoded back to samples and handed to a second lossy encoder, which throws away a second set of information chosen by different rules. Hydrogenaudio puts it plainly: every time you encode with a lossy encoder, the quality will decrease, and there is no way to gain quality back even if you transcode a 128 kbps MP3 into a 320 kbps MP3.
So the two costs are these:
- The first encode already happened. Whatever was lost when the phone, camera or platform compressed the audio is gone for good. A higher output bitrate cannot rebuild it.
- The second encode adds its own damage. The MP3 encoder spends bits faithfully reproducing the first encoder's artefacts as if they were music, and then discards some real detail on top.
The practical rules come straight out of that:
- Match the source, do not exceed it. If the soundtrack is 128 kbps AAC, choosing 320 kbps buys file size and nothing else. Choosing 192 kbps gives the second encoder enough headroom that its own losses stay small.
- A lossy source makes low bitrates worse. Hydrogenaudio's transcoding page reports a listening test in which 256 kbps MP3 transcoded down to 128 kbps MP3 deteriorated very significantly compared with a 128 kbps MP3 made directly from the original. The same bitrate sounds worse when it starts from a copy.
- Go back to the best source you have. If you own the original recording, extract from that rather than from a copy that has already been shared, compressed and downloaded again.
- Skip the re-encode when you can. If the destination accepts
.m4a, copying the AAC stream out untouched costs nothing at all. The trade offs between copying and converting are covered in the guide to extracting audio from video.
To find out what a file actually contains, a tool like MediaInfo or FFmpeg's ffprobe will report the audio codec and bitrate before you choose. For example:
ffprobe -v error -select_streams a:0 -show_entries stream=codec_name,bit_rate,sample_rate,channels file.mp4
What should you pick for speech you'll transcribe?
Much lower than you would pick for listening. Transcription needs the words to be intelligible, not the room tone to be lifelike, so a music grade file mostly adds upload time.
VidClip's transcribe tool is a working example of that choice. Before anything leaves the browser, it strips the video and converts the audio to mono at 16 kHz and 32 kbps. At 32 kbps an hour of speech is 14.4 MB, so a 25 MB request fits about 104 minutes (25,000,000 bytes × 8 ÷ 32,000 bits per second is 6,250 seconds), which the tool rounds down to roughly 100 minutes per run. One honest note: transcription is the one VidClip tool that uploads. The compact speech file is sent to a recognition service. Every other tool, including the MP3 extractor, runs entirely in your browser.
If you are preparing a file for a different transcription service, the same thinking applies:
- Mono is enough for a single speaker. A second channel carries nothing the words depend on.
- Clean input beats high bitrate. Background music, echo and crosstalk hurt recognition far more than a modest bitrate does.
- Keep a listening copy separately. If you also want to publish the audio, export a 128 or 192 kbps MP3 for people and a small mono file for the machine.
Frequently asked questions
Is 320 kbps worth it?
Only when the source is lossless or uncompressed and you are listening on equipment good enough to reveal the difference. Hydrogenaudio's LAME documentation says no one has shown in ABX testing that 320 kbps CBR ever sounds better than LAME's best VBR settings, and those typically land well below 320. From a video's soundtrack, which is almost always already compressed, 320 kbps adds 57.6 MB per hour over 192 kbps and nothing you can hear.
What bitrate is CD quality?
Uncompressed CD audio is 1,411.2 kbps: 44,100 samples per second, 16 bits per sample, two channels. That works out to 635.04 MB per hour. No MP3 bitrate is literally CD quality, because MP3 is lossy and tops out at 320 kbps. What a good MP3 can be is transparent, meaning a listener cannot reliably tell it apart from the CD, and for LAME that usually happens around 192 kbps.
Does a higher bitrate fix bad audio?
No. Bitrate only controls how faithfully the encoder copies what it is given. Hiss, echo, clipping, a quiet recording or an already compressed source all come through intact at 320 kbps, just in a bigger file. Apple's podcast guidance notes that audio compression algorithms typically do not modify loudness, which is why levelling has to happen before encoding. Fix the recording or the mix first, then choose the smallest bitrate that sounds right.
What bitrate should a podcast be?
Apple Podcasts' audio requirements recommend 96 to 128 kbps for mono MP3 and 128 to 256 kbps for stereo MP3 at 44.1 or 48 kHz. For RSS feeds, Apple strongly recommends AAC over MP3. In practice, 128 kbps is the safe choice for a talk show, and 192 kbps is worth it if music beds or sound design carry a lot of the episode.
Hear the difference yourself
Numbers settle file size. Only your ears settle quality. Drop a video into the video to MP3 extractor, export the same clip at 128, 192 and 320 kbps, and compare the busiest ten seconds side by side. It all runs in your browser, so nothing uploads and audio exports never carry a watermark. The free tier handles files up to 200 MB, and Pro ($4/mo billed $24 every six months, or $79 once for life) removes that cap.