MusGU+ Evaluation: Stable Audio Open Small

The Music-Generative Usable+ AI (MusGU+) framework is a musician-centered evaluation framework designed to assess how generative music models can be adapted, used, and controlled in real-world creative contexts. The framework evaluates models along three complementary dimensions, with each dimension addressing a key question from the musician's perspective:

πŸ“– Read the detailed evaluation criteria, return to the discovery tool or inspect the model's YAML source file.

Affiliation: Stablility AI

Architecture: latent diffusion

Musical applications

style transfer, text-to-music

Adaptability

40%

Hardware Requirements

~ Partially supported

Stable Audio Open Small research paper reports using institutional-scale hardware for fine-tuning (8 x H100 GPUS). However, the released source code supports adaptation on a single consumer-grade GPU, although with significant memory and storage demands. Adapting the model with CPU-only is impractical, and larger GPU setups will significantly improve training and fine-tuning feasibility.

Dataset Size

✘ Not supported

Stable Audio Open Small training relies on large-scale curated data (i.e., thousands of hours of audio). No documentation suggests that effective training or fine-tuning can be achieved with small personal datasets.

Adaptation Pathways

βœ”οΈŽ Fully supported

The released tooling supports LoRA, full fine-tuning, and training on custom datasets using pretrained model initialization.

Technical Barriers

✘ Not supported

Adaptation requires substantial technical expertise, including configuration of diffusion training pipelines, dataset preparation, and GPU-based training. No user-friendly interface for model adaptation is provided

Model Redistribution

~ Partially supported

Redistribution of adapted models is permitted under the Stability AI Community License, but subject to licensing conditions and commercial revenue thresholds.

Usability

50%

Interface Availability

~ Partially supported

The model can be run through a basic Gradio UI provided in the stable-audio-tools library, offering a simplified interactive interface. However, it requires installation and environment setup by the user.

Access Restrictions

~ Partially supported

Stable Audio Open Small can be run locally with open pretrained checkpoints and inference code without paywalls, usage limits, or subscriptions, subject only to license agreement acceptance.

Real-time Capabilities

~ Partially supported

Stable Audio Open Small can be interactive on high-end consumer GPU for short audio segments (e.g., achieving sub-200ms response time), but it not designed for low-latency or streaming use on CPU and is impractical for live performance on usual personal hardware.

Workflow Integration

✘ Not supported

Although optimized for low-latency generation on GPU, this model does not offer native integration with DAWs, live music environments, visual programming systems, or musical hardware. Usage is limited to file- or script-based generation via research tooling, with no direct embedding in existing creative workflows

Output Licensing

~ Partially supported

Generated outputs are owned by the user and may be used for both non-commercial and commercial purposes. However, commercial use is conditional on registration and subject to revenue thresholds, attribution requirements, and restrictions on downstream uses (e.g., training other foundational models). That is, output usage is permitted but not unrestricted.

Community Support

βœ”οΈŽ Fully supported

Support exists via a Discord channel (i.e., specific sub-channel β€œstable-audio”), GitHub issues and documentation.

Controllability

38%

Conditioning Inputs

~ Partially supported

Stable Audio Open Small supports text prompts and audio initialization. Initial audio can guide generation through voice-derived control, beat alignment, and audio-to-audio style transfer.

Time-Varying Control

~ Partially supported

The model enables time-localized influence through audio initialization (e.g., beat-aligned or voice-guided generation), but does not provide explicit or fine-grained time-varying control over musical attributes.

Feature Disentanglement

✘ Not supported

No disentangled or explicitly separable musical controls are provided. Musical attributes are implicitly entangled within prompt- and audio-based generation.

Control Parameters

~ Partially supported

Stable Audio Open Small provides several inference parameters, including duration, diffusion steps, seed, audio initialization strength, and sampling controls. These allow generation behavior to be adjusted but do not provide direct manipulation of internal representations or fine-grained controls for musical attributes.