Short answer: mixing a vocal is a chain of small steps, each doing one job. You clean the recording, high-pass the rumble, even out the dynamics with compression, tame the harsh sibilance with a de-esser, EQ for clarity, add space with reverb and delay on sends, then ride the level with automation so every word stays audible. Carve a little pocket for the voice in the instrumental and it sits without a fight.
The order matters more than any single setting, and the frequency figures here are approximate zones, not exact rules. Every voice and every song is a little different, so use your ears and treat the numbers as a starting point. None of this is tied to one product.
You give it a clean starting point, process it in a sensible order, and make room for it in the track. A vocal sits when its level is steady, its harsh moments are under control, its tone is clear, and nothing else is crowding the range it lives in. No single plugin does that on its own. It is the result of a few small steps stacked in the right order, plus a little space carved out for the voice in the arrangement.
The recording comes first. Get the cleanest take you can, tune it if the style calls for it, edit out distracting breaths and lip noise, and set a sensible, steady level before any processing. This is the least glamorous step and the one that matters most. Fixing a rough recording later is much harder than starting from a good one, and most of what people call a great vocal mix is really a great vocal that was then handled carefully.
Then the chain. Each step does one job. A high-pass filter clears rumble the voice does not use, a compressor evens out the dynamics, a de-esser tames sibilance, EQ shapes the tone for clarity, and reverb and delay place the voice in a space. We walk through the order in the next section, because the order is part of what makes it work.
And the track around it. A vocal that is processed well can still get buried if the instrumental is fighting it in the same range, so part of the job is carving a small pocket in the backing so the voice has somewhere to sit. More on that further down.
The best vocal move happens before you open a plugin. A close, clean, well performed take at a steady level gives every step below something good to work with. Time spent on the recording, the tuning and the edit almost always beats time spent trying to rescue a poor take with processing later.
A common order is: clean the recording, high-pass, compress, de-ess, EQ, then reverb and delay on sends, and automation last. Each step feeds the next, so the order is not arbitrary. You clear the rubbish before you compress, you compress before you shape the tone, and you add space and ride the level once the voice itself is under control. Here is the whole chain in order and a sensible place to start with each.
| Step | What it does | A place to start |
|---|---|---|
| Clean the recording | Tune it if the style needs it, edit out loud breaths and clicks, and set a steady level. The foundation everything else sits on. | Do this before any plugin |
| High-pass | Rolls off the low rumble and room noise the voice does not use, so it does not muddy the low end or trigger the compressor. | Around 80 to 100 Hz, by ear |
| Compression | Evens out the level so loud and soft words sit closer together. Fast-ish attack, moderate ratio. Two light stages often beat one heavy one. | 2:1 to 4:1, a few dB of gain reduction |
| De-essing | Turns down the harsh "s" and "t" sibilance when it spikes, without dulling the whole vocal. | Target the harsh band, a few dB only |
| EQ | Cuts mud in the low mids, lifts presence, adds air. Clarity, not a rebuild. | Cut around 200 to 400 Hz, gentle lift 3 to 6 kHz |
| Reverb and delay | Place the voice in a space and add depth. On send channels, not inserts, so you can control and share them. | Start subtle, high-pass the returns |
| Automation | Rides the level line by line so every word stays audible under the music. The finishing pass. | Last, after everything else |
The two compression stages in that table are worth a word. Rather than one compressor working very hard, two gentle ones in series often sound more natural: a fast one to catch the sharp peaks, then a slower one to hold the overall level steady. If compression is the step you are least sure about, our guide on what a compressor does and how to use it walks through threshold, ratio, attack and release in plain terms.
Treat the order as a strong convention rather than a law. Some engineers de-ess before the compressor so it does not clamp on the loud sibilants, and some make a broad tone move with EQ before compressing and a finer one after. Once you know what each step is doing, you can move them around with intent. The version above is a reliable default that works on most vocals.
Clean up the low mids first, then lift the top gently, and control the sibilance with a de-esser so the brightness does not turn into harshness. Clarity is mostly about removing what is in the way, and brightness is a small lift once the clutter is gone, not a big boost on a cluttered signal. Do it in that order and the voice opens up without getting brittle.
Cut the mud. Vocals carry their body in the low mids, roughly 200 to 400 Hz, and when that stacks up with the body of everything else in the track the voice turns boxy and thick. A gentle cut there opens it up. Sweep a narrow band to find the worst spot, then pull a few dB out rather than guessing. This is the same low mid buildup behind a cloudy mix in general, which we cover in why does my mix sound muddy.
Lift presence. A gentle boost somewhere around 3 to 6 kHz brings out the consonants and helps the words cut through the track, so a listener can follow the lyric. Keep it gentle. Too much here is exactly where a vocal starts to sound harsh and fatiguing, so a little usually goes further than you expect.
Add air. A soft high shelf up around 10 kHz and above adds openness and breath and makes a vocal feel expensive. This is tone rather than intelligibility, so a small amount does the job and more just adds hiss and edge.
Control the sibilance. Those presence and air lifts also raise the "s" and "t" sounds, which can get piercing. A de-esser turns down just the harsh sibilant band, and only when it spikes, so you keep the brightness without the sting. Set it to act on the offending sound alone, not the whole vocal, and use as little as clears the problem.
Cutting what is in the way usually beats boosting what is missing. Clear the mud around 200 to 400 Hz and a vocal often sounds brighter on its own, before you touch the top end. Then a small presence and air lift is all it needs, and because you are lifting a clean signal it stays clear instead of turning harsh.
Automation. After the compressor has done its broad leveling, ride the vocal fader line by line so every word sits right under the music, then make room for the voice in the track and keep your reverb and delay from washing over the words. Compression handles the fast, moment to moment jumps, but keeping a whole quiet line even against a loud chorus is a job you finish by hand.
Automate the level. A compressor cannot know that one line matters more than another, so it evens out the fast peaks but leaves whole phrases sitting too low or too high. Automation fixes that: draw the level up where a word disappears and down where one jumps out, so the lyric stays even from the first line to the last. It is the last step and often the real difference between an amateur and a professional vocal.
Carve a pocket. Even a well mixed vocal will struggle if the instrumental is loud in the same range. A small dip in the backing where the voice lives, often a gentle cut around the presence range on a busy synth or guitar bus, gives the vocal somewhere to sit without turning the whole track up. Some engineers do this with a sidechain from the vocal so the dip only happens while the vocal is actually singing.
Keep space on sends. Put reverb and delay on send channels rather than straight on the vocal, so you control how much wash sits behind the words, and high-pass the returns so the tails do not add mud. Too much space and the words blur into it. A little, placed well, adds depth without hiding the lyric.
Use parallel and bus tricks lightly. For a bigger, more consistent sound you can blend a heavily compressed parallel copy quietly underneath the main vocal, which lifts the soft detail without squashing the main take. And if you have several vocal tracks, a shared vocal bus with gentle compression and a touch of EQ can glue them so they read as one voice.
No amount of compression replaces automation for keeping every word present. Compression narrows the range automatically, but the final, deliberate pass of nudging lines up and down by hand is what keeps a listener able to follow every word from start to finish. Leave it until the end, once the tone and the space are settled.
Half of getting a vocal to sit is decided before you open a single plugin, in how much space the backing leaves for it. A busy, full-range arrangement fights the voice no matter how carefully you mix it. A focused one leaves a pocket the vocal drops straight into, and everything above gets easier.
We do not make a vocal plugin, so this is not that. What we make is instruments. The Collection is our three of them together, ARGISH, SILT and REHEAT, with three separate licence keys. They make clean, focused sounds that stay in their own part of the range, so the backing leaves a natural pocket for the voice and the vocal has room before you touch it. ARGISH is a self-playing drone synth, SILT is a tape-loop instrument, and REHEAT writes acid lines. Every one ships as AU, VST3, AAX and a standalone app on macOS, signed and notarized, and a VST3 on Windows.
You can hear one free in your browser first, no install and no account, then decide.
We write these when there is something worth writing down. One email when a new one lands or a new Tunary instrument ships. No newsletter, no schedule.