1. What bleed actually does to your audio
Put four people around a table with a microphone each and you do not have four recordings. You have four recordings of the whole room, each one merely favouring a different person. When one person speaks, all four mics capture it — one closely, three from further away.
Leave all four tracks open and two separate things go wrong.
The noise floor stacks up
Each open mic contributes its own hiss and its own share of room tone: air conditioning, traffic, computer fans, the building. These are largely uncorrelated between microphones, and uncorrelated noise adds up at roughly 3 dB per doubling of sources. Two open mics are about 3 dB noisier than one; four are about 6 dB noisier; eight about 9 dB.
That is why a recording can be clean on every individual track and still sound like it was made in a wind tunnel once everything is playing at once. Nothing is wrong with any one track.
Comb filtering makes voices sound hollow
This is the one people cannot name but always hear. The same voice reaches the near mic and the far mic at different times, because sound takes about 3 milliseconds to travel a metre. When you sum those tracks, some frequencies reinforce and others cancel.
The result is a series of notches through the frequency response — a comb filter. It sounds thin, phasey, slightly metallic, like the voice is coming through a tube. Editors usually reach for EQ at this point, which cannot help, because the problem is not tonal balance. It is the same sound arriving twice.
The distance is what matters, and the damage is not one notch but a whole series of them. A speaker one metre from their own mic and two metres from the next one produces roughly a 3 ms offset. That puts the first notch near 170 Hz and then repeats one every 333 Hz — 500, 833, 1167, 1500 and onwards — straight through the range that carries speech. It is the regular spacing that gives comb filtering its characteristic hollow ring, rather than any single missing frequency.
2. Why turning the other mics down does not work
The obvious response is to pull down the mics that are not being used. The problem is that which mics those are changes several times a minute. In a real conversation the floor passes back and forth constantly, with interruptions, overlaps, agreement noises and half‑finished sentences.
A static mix cannot express that. Every fader has to be up when its owner is speaking and down when they are not, and that has to be true at every moment of a two‑hour recording. There is no single set of levels that is correct for more than a few seconds at a time.
This is why the fix has to be dynamic. The only question is whether a human or a machine does the work.
3. The manual fix, and what it costs you
The manual method is straightforward and completely reliable. For each mic track, you cut it down to only the moments its owner is actually speaking, and remove everything else.
In Premiere that means, per track:
- Scrub through and find where that person starts and stops speaking.
- Cut at each boundary —
Cmd Kon macOS,Ctrl Kon Windows. - Delete the segments where they are silent.
- Add a short fade at every edit, or you will hear the room tone snap on and off. A few frames is usually enough.
That last step is not optional. An abrupt cut between room tone and silence is more audible than the bleed you removed — the ear is far more sensitive to a sudden change in noise than to a constant one.
The cost is time. A four‑person, ninety‑minute conversation typically contains somewhere between several hundred and a couple of thousand speech turns. At even a few seconds of editing per turn, this is comfortably a full working day, and it is the least creative day in the whole edit.
It is also the kind of work where attention degrades. The last twenty minutes of a two‑hour session rarely get the same care as the first twenty, and that is usually where the mistakes end up.
4. Why a noise gate usually makes it worse
The natural next thought is to automate it with a noise gate: set a threshold, and the track opens only when the signal is above it. Premiere ships with one. It very rarely works on conversation, for a reason worth understanding properly.
A gate keyed on absolute level cannot tell a loud neighbour from a quiet owner. If one person projects and sits close to their mic, their bleed into the mic beside them can be louder than the quiet person's own voice in their own mic. No single threshold separates those two cases, because they overlap.
Set the threshold high and you lose the quiet speaker's sentence beginnings and endings, which is the most damaging possible failure — clipped words read as bad editing rather than bad audio.
Set it low and every mic opens whenever anyone talks, which is precisely the situation you were trying to escape.
Gates also chatter. Around the threshold, ordinary speech dynamics push the signal above and below many times a second, and the gate opens and closes with it. Hold and release controls smooth this at the cost of accuracy in both directions.
Gates work well on a drum kit, where the source is loud, transient and close. Conversation is none of those things.
5. What actually separates speech from bleed
There are two signals that work, and neither is absolute loudness.
Level relative to that microphone's own speaking level
The useful question is not "is this loud?" but "is this loud for this microphone?" Every mic has its own characteristic level when its owner is speaking into it, set by their voice, their distance and the gain on the channel. Measure that level for each mic, and bleed sits well below it — typically far enough below to be unambiguous.
This works because it removes the comparison between different people. A quiet speaker at their own normal level is clearly speaking; a loud speaker bleeding into someone else's mic is clearly not at that mic's normal level, however loud they are in absolute terms.
Which microphone heard it first
Sound travels about a metre every 3 milliseconds. Whoever is closest to the speaker hears them first, and at 48 kHz a metre is around 140 samples — easily measurable if the tracks share a clock.
Arrival time is powerful because it is a fact about geometry rather than a judgement about volume. It does not care how loudly anyone speaks. It is most useful exactly where level comparisons struggle: two people talking at once, or one person much louder than another.
The catch: it only works if the recordings share a timebase. Mics into one interface or one recorder share a clock. Separate recorders that started independently do not, and their drift is larger than the difference you are trying to measure. This is the hard case, and anyone who tells you otherwise is selling something.
6. Reducing bleed while recording
Every minute spent on this at the recording stage saves considerably more in the edit, and none of it costs money.
- Get the mics closer. This is by far the biggest lever. Sound pressure falls about 6 dB every time the distance doubles, so halving the distance to your own mic improves the ratio between your voice and everyone else's by about 6 dB — more than any processing will recover later.
- Use directional microphones and aim them properly. A cardioid mic rejects sound from behind it. Point the back of each mic at the person most likely to bleed into it, not at the wall.
- Increase spacing between people where the room allows. Bleed falls with distance for the same reason proximity helps.
- Avoid recording into one shared room mic as well unless you genuinely need it. It is bleed by definition, and it will comb‑filter against everything.
- Record to one device where you can. A shared clock keeps arrival time usable, which preserves the strongest tool available for sorting this out afterwards.
- Soften the room. Curtains, rugs, bookshelves and bodies all reduce the reflections that make far‑mic bleed sound smeared.
None of this eliminates bleed. Four open mics in one room will always hear each other. The aim is to make the difference between near and far large enough that separating them afterwards is easy rather than marginal.
Doing it automatically
Hush is a Premiere Pro panel that does the manual method in section 3 — working out who has the floor moment by moment, muting the other microphones and fading every edit — using both signals from section 5 rather than a level threshold.
It builds a new sequence and leaves your original untouched, and it runs entirely on your own computer: no uploads, no cloud processing. The free version runs the complete detection on the first five minutes of your own recordings, which is the only test that tells you anything useful.
It is also honest about where it struggles. Separate recorders with microphones close together, as described above, are a genuinely hard case.
Try it on your own audio