If you have spent any time with ACE-Step — the open-source music foundation model that writes a full song, vocals included, in about twenty seconds on hardware you already own — you know the gap between "this generation is actually good" and "this file can go to Spotify." Generation solves composition. It does not solve loudness targets, peak headroom, or the small artifacts diffusion models leave behind. This page is the workflow I use for ACE-Step output: what to export, which numbers your distributor will check, and where a per-track master fits in between.
What ACE-Step actually gives you (and does not)
For anyone new to it: ACE-Step is an open-source foundation model jointly developed by StepFun and ACE Studio. The v1.5 generation shipped on 28 January 2026 with a hybrid language-model-plus-diffusion architecture, and the team added the XL series (larger decoder) in April 2026. The properties that matter for this workflow:
- It runs locally. The v1.5 models run under 4 GB of VRAM; the team's own figures put a full song at under ten seconds on an RTX 3090. Batch mode can queue up to eight songs at once.
- The output is yours, commercially. The v1.5 weights are published under MIT (the original v1 model card carries Apache-2.0), and the project states its training data was licensed, royalty-free or synthetic — so generating music you intend to release is fine on the generation side. Your remaining obligations are the distributor's AI disclosure rules, not a model license.
- Range and control. 10-second loops up to ten-minute songs, lyrics in 50+ languages, reference-audio conditioning, cover/repaint/vocal-to-accompaniment modes, LoRA fine-tuning on your own style.
What it does not do is deliver a stream-ready master. Two things worth knowing about the audio itself:
- The technical report states an inherent quality ceiling: v1's pipeline builds on a mel-spectrogram autoencoder and a 32 kHz monophonic vocoder rather than end-to-end audio generation, which limits high-frequency fidelity. Community threads on 1.5 keep reporting the same family of issues — harshness or distortion up top, clipping in loud passages.
- Nothing in the generator is aiming at streaming loudness. It optimizes for "sounds right on a laptop," not for −14 LUFS integrated with true-peak headroom under the encoder.
Both problems are exactly the class of thing mastering exists to fix — and both can be measured before you spend a cent (more below).
Step 1 — Export from ACE-Step in the best format you have
Local runs (official repo, ComfyUI node packs) typically write WAV at your chosen sample rate; community guides routinely export 48 kHz. Cloud wrappers of the model may hand you MP3 or a fixed-rate WAV depending on the service.
Rules of thumb:
- Prefer lossless. If your run can emit WAV, give the masterer WAV — an already-compressed source gives any chain less to work with than a clean one.
- Export what you will actually release. Batch generation is cheap and fast here (eight songs at once), so treat early takes as audition material: keep three or four candidates per idea rather than re-rolling endlessly, then master only the keeper(s).
- If the model gave you stems (v1.5 supports track separation / multi-track layers), master the final stereo mix — stem mastering is a separate discipline this page does not pretend to cover.
Step 2 — Measure before you decide
The two numbers your distributor's pipeline cares about:
| What | Target for a stream-ready master (2026) | Why it matters |
|---|---|---|
| Integrated loudness | ≈ −14 LUFS integrated (the Spotify-family norm; Apple's Sound Check sits around −16) | Files louder than the pack get turned down to match — hitting target costs you nothing and buys encoder headroom. See what LUFS actually measures |
| True peak ceiling | ≤ −1 dBTP (Spotify's own guidance; use −2 dBTP if your master is louder than −14 LUFS) | AAC/Ogg encoding can add inter-sample peaks of a decibel or more. Without headroom that becomes audible clipping in the file listeners actually hear — details at what true peak is and why sample-peak lies |
You do not need a DAW for this check. Our free browser LUFS meter & normalizer computes integrated loudness plus true peak locally in the browser (nothing uploaded). Drag your ACE-Step export in and you will usually find one of three states: already near −14 with peaks under ceiling (you need tonal shaping, not a rescue); 3–6 LU quiet but clean (a gentle full master is all it wants); or loud-and-spikey with inter-sample overshoot (the classic "crackles on phones" case). Measure first — the free tool tells you which of the three you are in.
Step 3 — Master: your realistic options
| Option | Cost (verified 8 Sep 2026) | When it's the right one |
|---|---|---|
| BandLab mastering | Free, unlimited (3 presets) | Quick audition; no reference targets and no loudness spec shown — fine for demos, not for release files |
| DistroKid Mixea | $99/year for unlimited tracks; first mastered track free (distrokid.com/mixea). Separately, DistroKid's optional "Loudness Normalization" upload extra auto-adjusts audio to −14 dB LUFS / −1 dB true peak and is skipped if your file already meets target | You are all-in on DistroKid and release a lot. Note the master lives inside their ecosystem; see mastering before DistroKid for the trade-off |
| Per-track web mastering (this service) | $5 one-time per variant's 24-bit WAV, 3 free processing passes/day, up to 3 variants in parallel per upload. The polishing step specifically targets noise and AI artifacts before EQ/limiting runs | You generate locally and want a release file without installing a DAW chain; you A/B style targets on the same mix; your output is AI-generated and you don't want an ecosystem deciding your terms — LANDR's pricing page, for example, now states monthly limits apply to high volumes of AI-generated audio |
| Human indie engineer | Typically $50–150 per track (higher for specialists with credits) | A record that genuinely needs editorial fixes — wrong performance, structural mix problems — which no chain can invent into a file |
A practical note on "up to 3 variants in parallel": batch mode means you will keep several close retakes of the same song. Mastering all candidates before choosing is backwards; master the keeper once, but if two finalists are close, running both through the same style target in one queue pass keeps them sounding like a set rather than strangers — and costs $5 per WAV, not $100+ per engineer session.
Step 4 — Deliver: WAV first
- Send your distributor the 24-bit WAV master, never an MP3 re-upload of it.
- If you also use DistroKid's loudness-normalization extra, understand what it does in their own words: it adjusts level and headroom to "−14 dB integrated LUFS with −1 dB true peak maximum" — and costs nothing when the file already meets target. A master that sits at spec passes through untouched; a spikey one gets processed by their pipeline instead of one you auditioned first.
- Keep the free MP3 (256 kbps) for sharing drafts with people; keep the WAV as your archival delivery file.
Budget reality check: 4-track EP generated locally
| Step | Cost |
|---|---|
| ACE-Step generation, one month of batch runs on your own GPU | $0 (electricity aside — open-source weights, no subscription) |
| 3–5 retakes per song as audition material (free MP3 previews here, or just local listens) | ~$0 |
| Mastering: 4 keepers × $5 WAV variants | $20 |
| DistroKid Musician plan for distribution | ~$24.99/year |
| Total to have a stream-ready EP out the door | under $50, no recurring mastering subscription |
Compare that against a $15–39/month subscription you would keep paying between releases, or four tracks at typical indie engineer rates ($200+). The per-track math starts to lose only when your monthly output crosses roughly 8–10 paid masters — beyond that a subscription legitimately wins, and I will say so rather than oversell.
What mastering does NOT fix in ACE-Step songs
Honesty section, because "AI fixes everything" marketing is worth nothing:
- A broken vocal performance stays broken. Off-key hook or a take you don't believe — regenerate (or repaint the section; v1.5's local editing modes exist precisely for this), then master the good one.
- The 32 kHz ceiling is not recoverable. Mastering cannot add high-frequency detail that was never in the signal the model produced. If harshness in a specific take is generator-level, a different retake or an XL-variant run is the fix — polishing and EQ will make it tolerable at loudness target, not pristine.
- Structural mix problems are mix problems. Muddy 808s fighting the kick inside one generation: mastering polishes around them. ACE-Step's stem separation plus a DAW pass (or Vocal2BGM) is the real fix when you need it.
Frequently Asked Questions
Can I release ACE-Step music commercially? Is mastering it allowed anywhere?
On the model side, yes — v1.5 weights are MIT-licensed and the project documents its training data as licensed/royalty-free/synthetic. On the distribution side, each platform sets its own AI disclosure rules (DistroKid, Spotify et al.), so check current terms at upload time. Mastering in between changes nothing: our service has no separate restriction on generated audio — unusual among online mastering services in 2026 — which matters if you release more than a few tracks a month.
Should I export WAV or is MP3 from the web UI good enough?
Use WAV whenever your setup offers it (local runs and ComfyUI exports easily give you lossless at 48 kHz). If you only have an MP3 export, mastering still helps — loudness targeting, peak control and artifact cleanup all apply to compressed sources too — but a clean file gives every step of the chain more to work with. Rule of thumb: never re-compress a finished master back to MP3 before upload.
My ACE-Step songs sound fine on my headphones but get quieter or crackle on Spotify — why?
Quiet playback usually means your player is normalizing and hiding the real level; crinkle-on-phones means inter-sample peaks above 0 dBFS that only appear after the platform's encoder runs. Both are measurable in minutes with the free browser LUFS/true-peak tool — if true peak sits above −1 dBTP, a proper limiting pass is worth its weight even when everything else is fine (full explanation at /blog/true-peak).
Do I need Ozone or a DAW to master local AI tracks?
No. That was the old path: $55–499 of plugin, a metering chain you have to learn, and manual limiting per song while batch generation gives you candidates every few seconds — a mismatch in tempo as much as skill. The workflow this page describes (measure free → pick style target → A/B up to three variants → take the WAV) is built for exactly that pace: minutes per track, $5 when it's the one you ship.
Do I need different masters for Spotify vs Apple Music?
No — both converge on roughly −14 LUFS integrated with a −1 dBTP ceiling (Apple's Sound Check sits slightly lower but accepts the same master). One stream-ready master serves the major platforms; keep a separate loud take only if you also play club/DJ contexts. For the full platform table see what -9/-14/-16 LUFS presets mean, and if your output comes from Suno instead, the workflow is nearly identical — see mastering Suno output.