Open-source licenses & attribution
StemSwell is built on open-source models and libraries with clean licensing. Weights and code are verified separately before anything ships. This page is generated from our license register.
Attribution-required models in use
- Cinematic 3-stem (dialogue/music/FX): Apache-2.0 (Bandit-v2) · weights: CC-BY-SA-4.0 (DnR-v3 checkpoint; attribute; no proprietary fine-tunes)
- Declip / general repair: MIT (VoiceFixer) · weights: CC-BY-4.0 (attribute)
- MP3 / lossy repair: CC-BY-SA-4.0 (JusperLee/Apollo) · weights: CC-BY-SA-4.0 (attribute; no proprietary fine-tunes)
- "Describe the fix" restore/master: Apache-2.0 (SonicMaster) · weights: Apache-2.0 (amaai-lab/SonicMaster) + GATED Stable-Audio-Open VAE (Stability AI Community License; gate accepted for the deploy account 2026-07-22)
- Regenerate / continue a region (Stable Audio 3): MIT (Stable Audio 3) · weights: Stability AI Community License + Gemma Terms of Use (T5Gemma)
- Restyle a region (Stable Audio 3): MIT (Stable Audio 3) · weights: Stability AI Community License + Gemma Terms of Use (T5Gemma)
Licenses & attribution
VocalParse singing transcription (2026-10-07)
vocal-to-lyric-notes is public/commercial-OK with positive credits, but remains deploy-pending. VocalParse source licence is Apache-2.0. The pinned released checkpoint card explicitly declares Apache-2.0; the unauthenticated Hub API reports gated: false. Its underlying Qwen3-ASR code and Qwen3-ASR-1.7B weights also declare Apache-2.0. Verified directly on 2026-10-07. No gate was accepted.
The released VocalParse snapshot contains the extended tokenizer, audio processor and complete fine-tuned weights. The integration loads that snapshot directly, with the pinned Apache VocalParse forward patch and prompt helper; it does not download the base Qwen checkpoint or any separate aligner, separator, F0, beat tracker or dereverb weights. Upstream notices remain in the installed source and Volume snapshot. Transformers, qwen-asr and HF Hub are Apache-2.0; torch is BSD-style, numpy/soundfile BSD, librosa ISC and mido MIT. ffmpeg/SoX remain subprocess dependencies in the server image, never browser code.
The model is primarily trained on Mandarin. Its note values and BPM describe a symbolic score, not measured physical note durations or audio timestamps. Independent 10-second chunks avoid unsupported long-input inference; boundary words and melody continuity are uncertain. MIDI and JSON disclose these limits.
YuE2 personal-use integration (2026-10-06)
generate-yue2 is owner/superuser-only, zero credits and excluded from public surfaces. The pinned YuE2 0.1.6 source at 1647252d68b70cbead046ffd3f7caf036dde47ee explicitly grants Apache-2.0 for inference code; third-party notices preserve MIT grants for Oobleck/stable-audio-tools and NVIDIA SnakeBeta. This integration installs the SHA-256-pinned source archive rather than the ambiguous licence metadata in the 0.1.5 model-hosted wheel. The current YuE repository contains YuE2; the old YuE implementation is on its YuE-v1 branch and supplies no licensing evidence here.
It remains deploy-pending because the separately downloaded qwen.tiktoken vocabulary asset's grant is unverified: the checkpoint licence explicitly excludes tokenizer files. The Apache-licensed tokenizer wrapper does not by itself establish a grant for that model-hosted asset. Do not publish this route until the asset is cleared; the NC checkpoints keep it restricted regardless.
Both YuE2-3B and the required YuE2-Vae weights are CC-BY-NC-4.0, and both Hub APIs report gated=false. No terms click or access request was performed. No external F0, dereverb, SheetSage2, MERT2 or transcription weights are loaded. Users supply their own lyrics and optionally a melody/harmony ABC plan.
Sources: official model/API, pinned code licence, pinned dependency/version manifest, weight licence, third-party notices, required VAE, YuE2 paper. See tool runbook.
Binding source for the licensing policy is DESIGN.md §4. This file is the human-readable register that backs the public /legal/oss attribution page (DESIGN §4.5). Every model requires BOTH a verified code license AND a verified checkpoint license before it ships. The machine-readable copy lives per-tool in packages/shared/src/registry.ts (licenseFlags) and modal/registry.py.
Legend: ✅ clean · ⚠️ counsel-flag (ship with the note below) · 🔴 blocked.
MuLaCover restricted research integration (2026-10-06)
Required model attribution: MuLaCover by MuLa Labs, with the official repository. New job sidecars retain that attribution and source alongside the noncommercial output notice.
mulacover and mulacover-midi are owner/superuser only, zero credits and excluded from all public catalogs. MIDI and audio were measured successfully on L40S. Source: Apache-2.0 code. Official weights and generated outputs require noncommercial use (CC-BY-NC-4.0 plus custom terms). An output cannot be commercialized merely because this app charges zero credits for inference.
MIDI conditioning uses public HeartCodec-oss-20260123 and Qwen3-Embedding-0.6B (Apache-2.0 cards). It avoids audio extraction checkpoints. Audio conditioning also needs YourMT3 and five ChordNet checkpoints, each with separate terms. Upstream third-party notices explicitly report unresolved GPL-3.0 versus Apache-2.0 provenance for YourMT3. The entire inference path runs in a subprocess; this does not clear the checkpoint rights or authorize public distribution. Audio remains a research integration only. The weights, codec and encoder were publicly accessible without a license-click gate on 2026-10-06; no terms were accepted or access requested.
MOSS conversation transcription (2026-10-06)
moss-diarize is public, commercial-OK, and charges credits. Both the MOSS code and MOSS checkpoint are Apache-2.0. Cross-window speaker embeddings use the Apache-2.0 SpeechBrain ECAPA checkpoint and Apache-2.0 SpeechBrain code. Transformers is Apache-2.0; torch/torchaudio are BSD-style, numpy BSD, librosa ISC, soundfile BSD, and PyAV BSD. Audio decoding uses ffmpeg as a subprocess. No additional VAD, F0, tokenizer, dereverb, gated pyannote, or alignment weights are loaded. The tokenizer and Whisper/Qwen-derived architecture are part of the pinned MOSS checkpoint.
The public Hub APIs reported gated: false and license: apache-2.0 for both checkpoints on 2026-10-06. Neither an access request nor terms acceptance was performed. Upstream LICENSE/NOTICE files remain with the pinned packages and cached checkpoint snapshots. Speaker labels are anonymous estimates, not verified identities; mixed overlap is excluded from speaker-embedding excerpts.
UniPASE damaged-call repair (call-repair)
Commercial use is permitted. Code is MIT with inherited MIT (WavLM, WavTokenizer architecture) and Apache-2.0 (Cisco PASE, ESPnet) files. Their notices remain in the pinned checkout's licenses/ directory. Checkpoint card explicitly grants Apache-2.0; the Hub API reported gated: false on 2026-10-06. The runtime loads only DeWavLM-Omni.pt, Adapter.pt, Vocoder_DWO-L1.pt, and PostNet.pt from that revision. The adapted WavLM backbone is contained in DeWavLM-Omni; there is no separate pretrained WavLM download, tokenizer weight, or validation-only Vocoder_WavLM-L24.pt in this pipeline. ESPnet is installed without training extras; inference uses its Apache separator base and layer lookup. PyTorch/torchaudio, numpy, soundfile, einops, OmegaConf and Hugging Face Hub are permissive dependencies; ffmpeg runs as a subprocess. No NC preprocessing is used.
Public tier, positive credits, deployment pending. Author recognition-error and speaker-similarity results are motivation rather than evidence of superiority over StemSwell's installed tools. See the integration and benchmark.
Shipping models (Phase 2 set)
SoulX-Singer-SVC clean-vocal conversion (2026-10-06)
singing-conversion is a commercial/public integration. Its ephemeral Modal smoke passed; production has not been deployed. The Apache-2.0 grant covers the SoulX code and model weights; the SVC checkpoint is in the shared model repository. The pinned checkpoint repository's model card also declares Apache-2.0 and expressly covers model weights. Its files include model-svc.pt with no separate licence or non-commercial supplement. The model card's usage disclaimer describes intended uses and responsible conduct; it contains no non-commercial condition overriding that grant. The pinned GitHub README identifies the SVC release in this shared checkpoint repository and likewise grants Apache-2.0 for code and weights.
Its mandatory Whisper-base encoder is Apache-2.0 in the Transformers distribution (original Whisper code/weights are MIT). The vocoder is included in the SVC state dict and requires no separate vocoder checkpoint. Embedded Amphion flow-matching code retains its MIT notice.
The upstream preprocessing bundle is excluded, never downloaded or imported. It includes becruily karaoke separation (uncleared weights), anvuew dereverb (restricted in DESIGN §4.2), and an RMVPE checkpoint with no separate grant in the bundle's empty model card. The integration instead accepts clean isolated vocals and computes 50-Hz F0 with pyworld (MIT) and WORLD (BSD-3-Clause) harvest/stonemask DSP. There are no F0 weights, ASR/lyric/MIDI preprocessing, phonemizers, or bundled separation/dereverb dependencies. This clearance applies to this specific route, not the upstream bundle or synthesis pipeline. Both required HF repositories are ungated in the unauthenticated Hub API audit.
SoulX code revision 81aeb3ae772c70093c3de74dc23c92d983801ae4, SVC checkpoint revision 40493ad90286056c7a9095035164434a79daa8c9, and Whisper revision e37978b90ca9030d5170a5c07aadb050351a65bb are pinned. CPU weight bake is explicit; GPU jobs load local files with HF offline and cannot fetch auxiliary models. The image build requires exactly two upstream Whisper identifiers before patching them and compiles the patched source; missing/mismatched loads abort the build. The pinned cli.inference_svc imports only the SVC model and shared audio/config utilities, with no preprocessing import; the real smoke also succeeded with the entire upstream preprocess/ directory removed.
| Tool | Model / route | Code | Weights | Status | Attribution |
|---|---|---|---|---|---|
stem-split-2 | Mel-Band RoFormer "Kim" via audio-separator; htdemucs_ft fallback | MIT | Kim: MIT relicense reported; confirm in writing; htdemucs_ft: MUSDB18 provenance | ⚠️ | Demucs (MIT code) |
convert | ffmpeg (LGPL build, subprocess) | LGPL-2.1 | N/A | ✅ | N/A |
loudness-normalize | ffmpeg loudnorm two-pass (subprocess) | LGPL-2.1 | N/A | ✅ | N/A |
trim | ffmpeg trim/afade (subprocess) | LGPL-2.1 | N/A | ✅ | N/A |
ffmpeg must be an LGPL build: no --enable-gpl, no libfdk-aac; invoked strictly as a subprocess (DESIGN §3.9, §4.3, §4.4).
CPU catalog (Phase 3): belt, effects, time/pitch, master, analyze
All CPU tools; each records both licenses in modal/waveforge/registry.py + packages/shared/src/registry.ts. No weights unless noted; nothing here is on the do-not-ship register.
| Tool(s) | Route | Code | Weights | Status | Notes |
|---|---|---|---|---|---|
fade speed gain eq dynamics silence-remove channel-ops reverse extract-audio metadata-edit spectrogram-png waveform-image | ffmpeg filters (subprocess) | LGPL-2.1 | N/A | ✅ | LGPL-safe encoders only (libmp3lame/native aac/libopus/flac/pcm); never libfdk-aac |
limiter | ffmpeg alimiter + oversampled soxr | LGPL-2.1 | N/A | ✅ | true-peak safe |
loudness-report | ffmpeg ebur128 + astats (+ pyloudnorm) | LGPL-2.1 / MIT | N/A | ✅ | N/A |
time-stretch pitch-shift | signalsmith-stretch via python-stretch; pyworld formant mode; ffmpeg fallback | MIT (python-stretch + signalsmith-stretch), MIT wrapper of BSD WORLD (pyworld), LGPL (ffmpeg) | N/A | ✅ | Do NOT use Rubber Band (GPL-or-pay, §4.2); ffmpeg fallback flagged quality=basic |
reference-master | Matchering 2.0; isolated subprocess, never imported (waveforge.tools._matchering_run) | GPL-3 (process isolation) | N/A | ✅ | 2 inputs (target + reference); §4.3 |
tape-stop stutter-riser stutter-spindown gate-stutter | Waveforge own numpy DSP | MIT | N/A | ✅ | DESIGN §3.10 |
key-chords | librosa chroma + Krumhansl key; BTC chords (opt-in) | ISC (librosa) / MIT (BTC) | MIT/ISC | ✅ | chroma-template chord fallback flagged quality=basic |
bpm-beats | Beat This! | MIT | MIT | ✅ | CPU |
structure | MSAF | MIT | MIT | ✅ | allin1 optional; verify weights first |
audio-to-midi | basic-pitch | Apache-2.0 | Apache-2.0 | ✅ | CPU |
python-stretch (github.com/gregogiudici/python-stretch) is MIT, wrapping the MIT Signalsmith Stretch library, the sanctioned Rubber Band alternative (§3.8). Matchering's GPL-3 is quarantined to a subprocess, so no GPL symbols link into the MIT/Apache/LGPL app graph.
GPU wave (Phase 3): separation / clean & restore / enhance / analyze
Deploy column: ✅ in the deployed default (smoke-tested) · ⏳ implemented but deploy-flagged (weights/repo bake pending) · ✋ flagged OFF (checkpoint license unverified; never deployed). Full metadata per tool lives in modal/waveforge/registry.py + packages/shared/src/registry.ts.
| Tool | Model / route | Code | Weights | Status | Deploy | Attribution |
|---|---|---|---|---|---|---|
stem-split-4 stem-split-6 karaoke | Demucs htdemucs_ft / htdemucs_6s (+ opt-in Kim RoFormer) | MIT | MUSDB18 provenance | ⚠️ | ✅ | Demucs (MIT) |
denoise-music dereverb | audio-separator UVR VR-arch (UVR-DeNoise, UVR-DeEcho-DeReverb) | MIT | MIT | ✅ | ✅ | UVR / audio-separator |
denoise-voice | DeepFilterNet3 | MIT | MIT | ✅ | ✅ | DeepFilterNet |
superres-48k | AudioSR (basic) | MIT | MIT | ✅ | ✅ | AudioSR |
lyrics | faster-whisper (large-v3) + WhisperX align, on the vocal stem | MIT / BSD | MIT / BSD | ✅ | ✅ | Whisper / WhisperX |
tags-mood | CLAP (LAION larger_clap_music_and_speech) zero-shot | Apache-2.0 | CC0/Apache | ✅ | ✅ | LAION CLAP |
cinematic-3stem | Bandit-v2 (dialogue/music/FX) | Apache-2.0 | CC-BY-SA-4.0 (DnR-v3, Zenodo 12701995) | ✅ | ✅ | Bandit-v2 (CC-BY-SA-4.0; attribute) |
vocal-restore | Resemble-Enhance | MIT | MIT | ✅ | ✅ | Resemble-Enhance |
declip | VoiceFixer | MIT | CC-BY-4.0 | ✅ | ✅ | VoiceFixer (CC-BY-4.0; attribute) |
lossy-repair | Apollo (official checkpoint) | CC-BY-SA-4.0 | CC-BY-SA-4.0 | ✅ | ✅ | Apollo (CC-BY-SA-4.0; attribute) |
bwe-fast | AERO | MIT | MIT | ✅ | ✅ | AERO |
speech-bwe | AP-BWE | MIT | MIT (code + weights, explicit) | ✅ | ✅ | AP-BWE |
describe-fix | SonicMaster (text-prompted) | Apache-2.0 | Apache-2.0 ckpt ✅, but GATED Stable-Audio-Open VAE dependency | ⚠️ | ✋ | N/A (gated; see below) |
Licensing calls made this wave (wave-3 bake, 2026-07-21):
- Bandit-v2 (`cinematic-3stem`): ✅ DEPLOYED. Code Apache-2.0; the only weights
are the CC-BY-SA-4.0 DnR-v3 checkpoints (Zenodo record 12701995, license API-verified; DnR-v3 training data is also CC-BY-SA-4.0). We ship the multi checkpoint, built directly from the model modules (no ray/Hydra). Attribute Watcharasupat et al., "Remastering Divide and Remaster…" (arXiv 2407.07275) + the Zenodo DOI on /legal/oss. ShareAlike attaches to any redistributed *weights* derivative, not to a user's separated stems. (⚠️ Do NOT substitute the older BandIt-v1 "Plus" checkpoint with a different, unverified license.)
- Apollo (`lossy-repair`): ✅ DEPLOYED. Official
JusperLee/Apollo
pytorch_model.bin is CC-BY-SA-4.0 (HF cardData.license verified); repo LICENSE is also CC-BY-SA-4.0. Attribute "Apollo (Kai Li & Yi Luo), CC-BY-SA-4.0" on /legal/oss. Do NOT use community "Apollo universal / by Lew" fine-tunes.
- AERO (`bwe-fast`): ✅ DEPLOYED, MIT (code; weights MIT-by-inclusion). The
12→48 kHz checkpoint is Google-Drive-only upstream, mirrored to the weights Volume.
- AP-BWE (`speech-bwe`): ✅ DEPLOYED, MIT for BOTH code and weights (explicit
weights_LICENSE.txt). 16→48 kHz checkpoint mirrored to the weights Volume.
- SonicMaster (`describe-fix`): ✅ GA (gate accepted 2026-07-22). SonicMaster's own
checkpoint (amaai-lab/SonicMaster) is Apache-2.0 (verified); the flow-matching inference is fully wired (TangoFlux, faithful to infer_single.py). Its TangoFlux architecture decodes in the Stable-Audio-Open 1.0 Oobleck-VAE latent space, and stabilityai/stable-audio-open-1.0 is a GATED repo under the Stability AI Community License. The owner accepted the gate on the deploy HF account `dakxjn` (`davehugface@xaai.ch`) on 2026-07-22; the token now resolves the gated VAE files (HTTP 200, was 403). Promoted to GA at 6 credits (L40S): un-flagged in modal/waveforge/registry.py + packages/shared/src/registry.ts (removed from the DEPLOY_FLAGGED set), deployed on its own describefix L40S image (SonicMaster repo + diffusers AutoencoderOobleck), with the gated VAE fetched at runtime via the `waveforge-hf` Modal secret (HF_TOKEN) and cached on the weights Volume. ⚠️ ONGOING LAUNCH CONDITIONS (Stability AI Community License; track these): 1. Revenue cap: commercial use is permitted only below ~US$1M annual revenue; above that, a separate commercial agreement with Stability AI is required. Re-check this the moment StemSwell revenue approaches the cap. 2. Attribution: display the Stability AI attribution (and the Community License reference) on /legal/oss. The SonicMaster checkpoint itself stays Apache; only its Stable-Audio-Open VAE dependency carries the Stability condition. (No payment/company commitment was required; free registration under the Community License.)
- CC-BY / CC-BY-SA weights (
cinematic-3stemBandit-v2,declipVoiceFixer,
lossy-repair Apollo) require visible attribution on /legal/oss and for the SA/no-proprietary-finetune ones, no proprietary fine-tunes may ship.
- Demucs weights keep the MUSDB18 provenance counsel-flag (
stem-split-*,
karaoke, and the lyrics vocal-isolation pre-step). The do-not-ship community RoFormer fine-tunes (viperx ep_317, becruily, unwa, gabox, aufr33/Sucial) are refused at runtime by tools/separate.py even if a param slips one through.
- The
⏳ deploy-pendingtools are clean-licensed but not baked into a deployed
image yet (research-repo integration / conflicting dep pins), so they are deploy-flagged and /spawn refuses them until their bake is verified.
Wave-2: voice (§3.5) + generate (§3.6)
Deploy column: ✅ deployed (smoked) · ⏳ implemented, deploy-flagged (image/inference bake pending) · 🔴 blocked (checkpoint license). All checkpoint licenses below were verified 2026-07-21 against the HF model cards / repo LICENSE files (cited).
| Tool | Model / route | Code | Weights (checkpoint) | Deploy | Attribution |
|---|---|---|---|---|---|
pitch-correct | WORLD (pyworld) harvest F0 + resynth toward scale; torchcrepe/RMVPE opt-in | MIT/BSD | N/A (DSP; no weights) | ✅ | N/A |
tts | Kokoro-82M (espeak-ng phonemizer server-side) | Apache-2.0 | Apache-2.0 (hexgrad/Kokoro-82M) | ✅ | Kokoro (Apache) |
tts-clone | Chatterbox (Perth watermark on every output) | MIT | MIT (ResembleAI/chatterbox) | ⏳ | Chatterbox |
text-to-music | ACE-Step 1.5 2B base/turbo (DiT-only, L40S) | MIT | MIT (ACE-Step/Ace-Step1.5 bundle + ACE-Step/acestep-v15-base) | ✅ | ACE-Step 1.5 |
extend-inpaint | ACE-Step 1.5 repaint with bounded context | MIT | MIT, same bundle | ✅ | ACE-Step 1.5 |
vocal-to-accompaniment | ACE-Step 1.5 base complete | MIT | MIT, same bundle | ⏳ | ACE-Step 1.5 |
music-cover | ACE-Step 1.5 cover / base-only lego | MIT | MIT, same bundle | ⏳ | ACE-Step 1.5 |
voice-conversion | kNN-VC (bshall/knn-vc); replaces the NC-blocked Vevo | MIT | MIT (WavLM + HiFi-GAN, redistributed by the MIT repo) | ✅ | kNN-VC |
generate-fast | DiffRhythm (+ MuQ-MuLan style encoder) | Apache-2.0 (code + DiT) | 🔴 required MuQ-MuLan encoder is CC-BY-NC-4.0; VAE is Stable-Audio-Community | ✋ | N/A (NC-blocked; see below) |
The bundled Qwen3 text encoder also has an upstream Apache-2.0 grant, including its tokenizer. Retain those upstream notices alongside the ACE-Step bundle's MIT notice. Both permit commercial use; no non-commercial F0, separator or dereverb weights are used.
Licensing calls made this wave:
- `voice-conversion`: switched Amphion Vevo → kNN-VC (MIT). Amphion *code* is MIT, but
the *released Vevo/VevoSing checkpoints* are non-commercial (amphion/Vevo CC-BY-NC-4.0, amphion/Vevo1.5 CC-BY-NC-ND-4.0), so per DESIGN §4.1 (weights ≠ code) / §4.2 they cannot ship. Per counsel (team-lead + research-models), voice-conversion now ships kNN-VC (bshall/knn-vc, MIT code + weights: WavLM feature kNN-matching + HiFi-GAN), a permissive zero-shot VC that works for speech and singing. The Vevo NC checkpoints were added to the DESIGN §4.2 register (team-lead commit c63610a). FreeVC / OpenVoice V2 are the MIT alternatives if kNN-VC quality is insufficient.
- SonicMaster (`describe-fix`): ✅ license CLEARED. The HF card
amaai-lab/SonicMaster
declares `license: apache-2.0` (checkpoint), and the paper/repo (AMAAI-Lab/SonicMaster, arXiv 2508.03448) are Apache. The wave-1 blocker ("HF tag unverified → flagged OFF") is resolved. It is now deploy-pending (Apache is fine; the flow-matching inference wiring + L40S image bake are still to do), not flagged-off. Training data = the SonicMaster paired degraded/clean dataset (self-constructed via 19 simulated degradations).
- ACE-Step 1.5: public, commercial-OK. Code licence is MIT; the official bundle (including VAE and Qwen3 text encoder) and base checkpoint model cards identify MIT, ungated as inspected 2026-10-06. No external preprocessing or LM planner is used. The old installation used the original ACE-Step repository, whose code is Apache-2.0, so the former MIT-code comment was incorrect. Old weight rights are not reused as evidence for 1.5. For the new bundle,
the authors state training on a legally-compliant mix (licensed + public-domain/royalty-free + synthetic MIDI-to-audio). Runs in bf16 → L40S (Turing/T4 has no native bf16 and cuDNN can't find an engine for its bf16 ops there; verified failure, fixed by moving to L40S).
- Kokoro (`tts`): ✅ Apache-2.0 checkpoint; espeak-ng phonemizer stays server-side (§3.5).
- Chatterbox (`tts-clone`): ✅ MIT; embeds a Perth (perceptual) watermark in every output;
disclose this in the UI/ToS.
- DiffRhythm (`generate-fast`): 🔴 NC-BLOCKED (found during the wave-3 bake). The
DiffRhythm code and DiT weights (ASLP-lab/DiffRhythm-*) are Apache-2.0, but the pipeline's get_style_prompt() unconditionally loads the MuQ-MuLan style encoder (OpenMuQ/MuQ-MuLan-large, CC-BY-NC-4.0) to embed even a pure-text caption, and the DiT was trained against MuLan's embedding space (not swappable for a permissive encoder). MuQ is on the §4.2 do-not-ship register, so the whole tool is non-commercial as-is. (Its VAE, ASLP-lab/DiffRhythm-vae, is additionally Stable-Audio-Community, not Apache.) Stays deploy-flagged. Responsible-use note (per DESIGN §4.6 ToS): the DiffRhythm authors' responsible-use notice must be mirrored in our ToS *if* it ever ships. To un-block: license MuQ-MuLan commercially from Tencent, or retrain the DiT against a permissive style encoder. MuQ-MuLan is now recorded on the do-not-ship register.
Do-not-ship register (never deploy; DESIGN §4.2)
MUSDB18-HQ/MoisesDB-trained community checkpoints without explicit grants (viperx BS-RoFormer ep_317, becruily, unwa, gabox, aufr33/Sucial); Banquet (NC); LarsNet weights (NC); anvuew dereverb (GPL); Bandit-v1 (NC); original FlashSR (no license); Essentia (AGPL); madmom pretrained (NC); TempoCNN (AGPL); MERT/MuQ (NC; incl. MuQ-MuLan OpenMuQ/MuQ-MuLan-large CC-BY-NC-4.0, DiffRhythm's required style encoder); Seed-VC (GPL, archived); so-vits-svc upstream (AGPL); Diff-HierVC (NC); MusicGen/AudioCraft/JASCO weights (NC); Tango/TangoFlux (NC); Tencent LeVo/SongGeneration (NC); Muse (tainted provenance); F5-TTS (NC); Fish/OpenAudio S1 (NC); XTTS-v2 (NC/defunct); DeepAFx-ST (NC); Rubber Band (GPL-or-pay; use signalsmith-stretch); libfdk-aac (non-free).
The register remains binding for general availability. The only exception is the restricted tier below (DESIGN §4.8): owner / super-user only, 0 credits, never marketed/GA.
Restricted tier (DESIGN §4.8): owner / super-user only, 0 credits
These tools integrate a do-not-ship (or hosted-only copyleft) checkpoint behind the `restricted: true` gate, served ONLY to role in {owner, superuser}, always 0 credits, excluded from the public directory / SEO / SSR / /legal/oss / plan pages, never part of any paid flow. They relocate the owner's existing personal use of these checkpoints into his own app without commercializing NC-licensed inference (they do not sell, meter, or market NC weights). Server enforcement lives in createJob (+ pipelines validation); the registry license test asserts every register-listed model is absent from GA OR restricted: true with baseCredits == 0.
| Tool | Model / checkpoint | Code | Weights | Why restricted | Notes |
|---|---|---|---|---|---|
speech-repair-reuse | NVIDIA RE-USE, revision 022e920d727347a64d6c21fbf0628f2a5f37ad78 | NVIDIA NSCL v1; no independent permissive grant for bundled inference | [NVIDIA One-Way Noncommercial License (NSCL v1)](https://github.com/NVlabs/HMAR/blob/main/LICENSE) (linked by the model card) | Noncommercial research and education only; owner/superuser, 0 credits | Released 9.6M Mamba checkpoint differs from the paper. No gated access or auxiliary checkpoints. Mamba/causal-conv1d Apache-2.0; torch BSD; HF Hub/einops Apache/MIT; ffmpeg subprocess. Isolated pinned A10G image. |
stem-split-2-pro | BS-RoFormer model_bs_roformer_ep_317_sdr_12.9755 (viperx) via audio-separator | MIT | MUSDB18/MoisesDB-trained, no explicit grant (§4.2) | NC-style training-data provenance | The owner's own tools/modal-vocals/modal_vocals.py checkpoint (~12.98 SDR). Rides the deployed separation image. |
karaoke-pro | same BS-RoFormer ep_317 | MIT | same | same | Karaoke framing (instrumental primary). |
dereverb-pro | anvuew dereverb_mel_band_roformer_anvuew_sdr_19.1729 (+ optional aufr33 UVR-De-Echo-Normal) | MIT | anvuew: GPL weights; UVR-De-Echo: community (§4.2) | GPL weights + community checkpoint | The owner's proven reverb→echo chain (modal_cleanup.py). GPL weights are hosted-only-safe (SaaS inference does not distribute the software), so this one could go GA later on that basis; kept restricted for now per scope. |
drum-split | LarsNet (polimi-ispl/larsnet) | MIT (arch) | non-commercial (§4.2) | NC weights | Distinct from the Phase-2 GA drum-kit-split (retrain on StemGMD CC-BY-4.0, §4.7). |
voice-convert-pro | Seed-VC (Plachtaa/seed-vc @ 51383ef) | GPL-3 (archived); isolated subprocess (§4.3) | GPL (archived); hosted-only-safe | GPL, hosted-only | Ports modal_svc.py. GPL run only as a subprocess (inference.py), never imported. |
generate-musicgen | Meta MusicGen (medium) via audiocraft | MIT (code) | CC-BY-NC-4.0 (§4.2) | NC weights | audiocraft image. |
fish-s2pro | Fish Audio S2 Pro + bundled tokenizer + DAC codec | Fish Audio Research License | Fish Audio Research License | non-commercial, including output use | Restricted personal use only, 0 credits. No separately downloaded preprocessing model. Code license, weights license. |
generate-fast | DiffRhythm (+ MuQ-MuLan style encoder) | Apache-2.0 | DiT Apache, but required MuQ-MuLan encoder is CC-BY-NC-4.0 (§4.2) | transitive NC dependency | Reclassified from BLOCKED → restricted. GA still needs a commercial MuQ-MuLan license or a permissive-encoder retrain. |
sheetsage2 | SheetSage2 + MERT-v2-FullSong | Unresolved separate inference-code grant | CC-BY-NC-4.0 for both adapter and backbone | NC weights + unverified code rights | Owner/superuser only, 0 credits, deploy-pending. The HF repository has a CC-BY-NC-4.0 LICENSE; its README labels that grant as weights. The linked Apache-2.0 YuE repository does not contain the HF implementation. See SheetSage2 verification. |
muscriptor | MuScriptor small / medium (Kyutai and Mirelo) | MIT; optional MuseScore GPL-3 subprocess only | CC-BY-NC-4.0, gated contact-sharing and specific-use conditions | Non-commercial gated weights | Owner/superuser only, 0 credits, never public, deploy-pending. Small/medium T4 MIDI and small PDF/MusicXML smoked 2026-10-07. CPU prepare must seed the target Volume before local-only GPU loading. |
MuScriptor sources: upstream code and MIT licence, small gate, medium gate. The gate also requires agreement to share contact information, warrants necessary input/output rights, and includes indemnification conditions. StemSwell must never accept these terms or request access automatically. Each model size has its own gate. Code and checkpoint revisions are pinned in waveforge/tools/muscriptor.py. The audio path uses ffmpeg, the model's mel conditioner and in-repo tokenizer; no separator, F0, dereverb, pretrained beat tracker or soundfont is invoked. MuseScore Studio 4.7.5 is only an external renderer of PDF/MusicXML (GPL-3, subprocess isolation); it adds no checkpoint restrictions or browser dependency. Tempo detection is disabled so the imported Beat This package never downloads its optional checkpoint. MIDI retains seconds at a fixed tick tempo, with constant note velocity; MuseScore performs its own notation import quantization.
Deploy state: the original seven restricted tools were baked, smoked and un-flagged by 2026-07-23. SheetSage2 remains deploy-pending while its separate inference-code grant is unresolved. Before enabling it, prepare both pinned models on the production weight Volume and validate the serving path; see the pre-enable checklist. Restricted image builds for the older tools remain gated by WAVEFORGE_BUILD_RESTRICTED_IMAGES=1; SheetSage2's function and image are registered unconditionally.
Mega split (DESIGN §3.1) ships as a restricted built-in pipeline preset: the full v2 DAG fan-out (MEGA_SPLIT_PRESET_V2): 6-stem → (drum-split on drums) ‖ (karaoke on vocals) → dereverb the isolated lead ≈ 13 labeled stems. Seeded per-account via scripts/seed-restricted-presets.mjs. Runs once drum-split is baked.
GPL tools: process-isolated only (DESIGN §4.3)
Matchering (GPL-3), SoX, aubio, audiowaveform: CLI subprocess or a separate Modal function, files in/out, never import/linked.
Client bundle dependencies (must be MIT/BSD/Apache/ISC; DESIGN §4.4)
| Package | License |
|---|---|
| wavesurfer.js v7 | BSD-3-Clause |
| next, react, react-dom | MIT |
| tailwindcss, tailwindcss-animate | MIT |
| @radix-ui/\* | MIT |
| class-variance-authority, clsx, tailwind-merge | MIT |
| lucide-react | ISC |
| firebase (web SDK) | Apache-2.0 |
| zod | MIT |
No copyleft ships to the browser. @ffmpeg/core (GPL) is deferred; when client-side conversion lands it must be a custom audio-only LGPL build.
Counsel review queue (DESIGN §9)
- Demucs MUSDB provenance (facebook/demucs issue #327).
- Kim RoFormer written MIT-relicense grant.
- SonicMaster / Amphion Vevo checkpoint tags (future phases).
Stable Audio 3: conditional commercial hosting
Verified 2026-10-07. Public editing tools stable-audio-region and stable-audio-restyle, 5 base credits, with production weights prepared and deployed RPC/GCS jobs verified before clearing their deployment flags. StemSwell is below the stated revenue threshold (Kevin's supplied eligibility fact).
- Code: MIT,
pinned at 3a82c807b69cf4b7c5c05270011a5d5e47abac18. Preserve copyright and permission notices. Torch, Flash Attention, Transformers, Hub, NumPy, soundfile and einops have BSD/MIT/Apache grants; ffmpeg is subprocess-only.
- Weights: Small Music
at 8bdfb53c6b0fbc62113cd571fd37373efe15fa14, Medium at d4d3579395a8ac0df4ab13b5b1d7138ce7290f81: Stability AI Community License, with bundled T5Gemma encoder/tokenizer under Gemma terms. SAME is included; no separate Google/SAME or NC preprocessing weights are used.
- Commercial grant: the gates expressly link the July 5, 2024 agreement.
Section III expressly permits commercial products/services, including hosted/API use. Commercial registration with Stability is required. The grant terminates when the licensee/affiliates individually or in aggregate exceed US$1M annual revenue from any source; obtain Enterprise terms then. The licensing landing page advertises Enterprise to API providers, but the operative grant has no independent API-provider exclusion and expressly includes hosted/API use below the cap. This is an interpretation of the published grant, not a negotiated Enterprise clearance. Keep pending until production preparation and rollout verification are complete; gate acceptance is not registration evidence.
- Attribution: IV(a) requires agreement copies, NOTICE and prominent
attribution for products/services. Powered by Stability AI. This Stability AI Model is licensed under the Stability AI Community License, Copyright © Stability AI Ltd. All Rights Reserved. Copies are served at /licenses/stable-audio3/, linked from user terms; tool pages credit Stability and this register appears on /legal/oss. Prepare retains Stability, Gemma and upstream NOTICE files on the Volume; the MIT code notice is also served with the licence copies. Stability AUP, effective September 30, 2026, is incorporated. IV(b) also prohibits use of materials/outputs to create/improve other foundational generative models. Output ownership (IV(c)) is only to the extent permitted by law, with no guarantee of third-party rights. Terms are revocable on breach; no endorsement.
- Gemma: April 1, 2026 Terms list T5Gemma
in the Appendix. Sections 1.1(b), 2.2 and 3.1 permit hosted web/API distribution without an NC restriction. Hosted services must provide agreement copies and notice, and make Section 3.2 use restrictions enforceable in user terms. /legal/terms incorporates 3.2 and the Prohibited Use Policy, links copies and prohibits violating those restrictions. The NOTICE-file requirement in 3.1(d) exempts Hosted Services, but the bundled notice is retained. We do not modify weight files or redistribute derivatives; the documented in-memory configuration only redirects encoder loading to local prepared paths. Google claims no output rights (3.3); users remain responsible. No Google trademark or endorsement rights are granted. Breach terminates the grant.
- Gate: Kevin must accept conditions with HF account
dakxjn, including
contact sharing, Stability licence/privacy and Gemma 3.2. No automatic gate acceptance or access request, account creation or settings/secret changes.
After Kevin confirmed access, development CPU preparation and an audit of both actual pinned LICENSE.md, LICENSE_GEMMA.md and NOTICE files succeeded. Both models carry identical July 5, 2024 Stability and April 1, 2026 Gemma terms; the served copies now reproduce the pinned texts. The operator separately confirmed Stability commercial registration on 2026-10-07 for customer-facing commercial use. No personal registration details are recorded here.
The grants permit the proposed public SaaS under these conditions; they are not blanket permissive weights. Restricted zero-credit access does not cure unmet terms. Commercial registration is confirmed; continuing revenue eligibility must still be maintained. Re-audit licences when changing pinned revisions. See implementation/release steps.