Skip to report content
Published · frozen dataset

A Thousand Listeners

Which AI music model actually sounds better?

We ran a blind, head-to-head listening study across five AI music generation systems — 1,000 listeners, 6,000 primary judgements, 104 prompts spanning 13 scenarios. No system knew it was being judged against the others; no listener knew which model made which track.

1,000Listeners
7,124Judgements
1,040Samples
13Scenarios
104Prompts
5Models

Headline result

Overall leaderboard

Ranked across the primary cohort of 1,000 listeners. Scene-equalized scores give each of the 13 scenarios equal weight, so no single scenario dominates the ranking.
RankSystem Score95% interval P(ranked #1)
01 Mureka 65.6% 62.3%–68.8% 86.7%
02 Suno 62.5% 59.1%–66.4% 13.3%
03 Happy Shrimp 47.8% 44.2%–50.9% 0.0%
04 MiniMax 39.4% 36.0%–43.1% 0.0%
05 Lyria 34.6% 31.7%–37.8% 0.0%
How the score is computed

Scores come from 300 bootstrap resamples (seed 20260829) of judgement-level data, resampled within each scenario and then equal-weighted across all 13 scenarios. "P(ranked #1)" is the share of resamples in which a system finished first overall.

Raw vote counts reflect assignment balance and are not adjusted for scenario weighting — see the bootstrap score for the headline ranking used throughout this report.

See the full result breakdown

What was tested

Scope of the benchmark

Five systems, evaluated on the same 104 prompts across 13 scenarios spanning vocal genres, instrumental-only tracks, commercial/functional briefs, and adversarial stress prompts.
01

Design

Each listener judged 6 blind pairs per session, one track from each of two randomly assigned systems, with left/right position randomized.

02

Language

All prompts and generated lyrics are in Mandarin Chinese. Results should not be assumed to generalize to other languages.

03

Priority

Listener preference is the primary outcome. Technical benchmarks (loudness, artifacts) are secondary and reported only where they explain a preference.

Feature comparison

What each system can do

Capability claims are sourced from each vendor's public documentation, not verified independently by this benchmark.
Capability SunoEvaluated V5.5 MurekaEvaluated V9.5 MiniMaxEvaluated Music 3.0 LyriaEvaluated 3.5 Happy ShrimpEvaluated 1.0
Generation
Prompt to full song Full song Full song Full song Full song Full song
Custom lyrics input Custom lyrics Custom lyrics Custom lyrics Custom lyrics Custom lyrics
Instrumental-only Supported Supported Supported Supported Supported
Style/direction control Style + voice tags Style + voice tags Style + voice tags Style + voice tags Natural-language only
Conditioning
Reference input Audio reference Audio reference Song reference Image reference Not found
Cover / restyle Cover Remix Cover Not found Not found
Extend an existing track Extend Lyrics-guided extend Not found Not found Not found
Reusable voice identity Personas Reference-based Not found Not found Not found
Production
Section editing Replace section Rewrite part Not found Not found Not found
Stem separation Advanced stems Stem splitter Voice isolation Not found Not found
Multitrack studio Full studio Studio DAW Not found Not found Not found
Real-time generation Not found Not found Not found Lyria RealTime Not found
Image-to-music Not found Not found Not found Supported Not found
Available Limited / partial No public evidence found
01

Suno

The most feature-complete production suite among the five: personas, section-level rewrites, extend, and a full multitrack studio in addition to generation.

02

Mureka

Strong generation quality paired with a lighter production toolset — remix, lyrics-guided extend, and a studio-lite editor.

03

MiniMax

Solid core generation and song-reference conditioning; production tooling beyond voice isolation is limited or unverified.

04

Lyria

Distinguished by real-time generation and image-to-music conditioning — capabilities absent from the other four systems.

05

Happy Shrimp

The newest system in the benchmark, released one day before data generation began. Direction control is limited to natural-language prompts.

Before the data

Questions this report answers

Three questions frame everything that follows.
01

Which system wins overall?

A single head-to-head ranking across all scenarios, with uncertainty intervals — not just a leaderboard number.

02

Does the ranking hold by scene?

Whether the overall winner also wins on vocal pop, instrumental, commercial briefs, and adversarial prompts specifically.

03

Who is the audience, and how did they judge?

Who the 1,000 listeners are, how they engaged with the study, and what qualities actually drove their preferences.

Executive summary

Four things to take away

The full statistical picture, condensed.
01

Mureka leads with a clear signal

Mureka was selected in 1,401 of 2,225 comparisons and finished first overall with an 86.7% bootstrap probability of ranking #1 — the clearest signal in the dataset.

02

The #1 vs #2 gap is narrower than it looks

Head-to-head, Mureka beats Suno only 52.4% of the time [47.1%, 56.9%] — an interval that crosses 50%, so the top-two ordering carries real uncertainty at the pairwise level.

03

Use the scenario table for your use case

The overall ranking is a 13-scenario average. If your use case matches one specific scenario (e.g. instrumental, commercial), check that row directly.

04

Read the confidence intervals

Every score in this report is estimated from a finite sample of 1,000 listeners. Point estimates without their intervals can overstate how settled a ranking is.

Reader's guide

Different readers, different depth

Pick a path based on what you need.
01

AI researchers

Full statistical methodology, bootstrap procedure, and raw data access.

Go to methodology
02

Music academia

Study design, listener demographics, and the qualitative feedback behind the numbers.

Go to audience
03

Model teams

Where your system ranks by scenario, and what listeners said drove their choices.

Go to scenarios
04

Media

The headline result and executive summary — everything needed for a quick, accurate write-up.

Go to overview
05

Music industry

What listeners actually value — melody, emotional match, and production polish — and how each system delivers on it.

Go to feedback

Evidence

Vote outcomes

Across all 6,000 primary judgements, listeners chose one track over the other 79.7% of the time.
Primary judgements Total: 6,000
Total judgements 6,000
Decisive preference4,78379.7%

Listener chose one track as clearly better.

Tie4026.7%

Listener judged both tracks equally good.

Reject both81513.6%

Listener judged both tracks unusable for the prompt.

Ties and rejections are excluded from pairwise preference percentages elsewhere in this report, but are tracked here as their own outcome category.

Head-to-head

Pairwise preference matrix

Row vs. column: the percentage of decisive judgements in which the row system was preferred over the column system.
System Suno MiniMax Mureka Happy Shrimp Lyria
Suno · 68.2%tie 7.8% 47.6%tie 8.8% 62.3%tie 8.4% 72.1%tie 7.6%
MiniMax 31.8%tie 7.8% · 29.2%tie 7.7% 42.4%tie 8.6% 54.4%tie 8.3%
Mureka 52.4%tie 8.8% 70.8%tie 7.7% · 64.9%tie 8.3% 74.4%tie 7.5%
Happy Shrimp 37.7%tie 8.4% 57.6%tie 8.6% 35.1%tie 8.3% · 60.8%tie 8.2%
Lyria 27.9%tie 7.6% 45.6%tie 8.3% 25.6%tie 7.5% 39.2%tie 8.2% ·

Cells show the row system's win rate over the column system among decisive judgements only; hover a cell for the full interval and sample size.

Reading the matrix Note

Mureka beats every other system head-to-head except when facing Suno, where the margin narrows to a coin flip.

Transitivity

The ranking is broadly transitive — Mureka > Suno > Happy Shrimp > MiniMax > Lyria holds in nearly every pairwise comparison, with the Mureka/Suno pair as the sole close contest.

Tie rates

Tie rates cluster tightly around 7–9% regardless of which two systems are compared, suggesting ties reflect genuine draws rather than one system's tracks being uniformly indistinguishable.

Data quality

Diagnostics

How complete and reliable the underlying dataset is.

Cohort completion

1,000 1,000 listeners completed all 6 primary rounds and are included in every result on this page.

Reject rate

13.6% Both tracks judged unusable for the prompt.

Usable comparisons

5,185 6.7% tie, 13.6% reject — the remainder feed the preference scores above.

Robustness

Beyond the primary six rounds

6,000 primary / 7,124 total judgements
6,000 Primary judgements (rounds 1–6)
1,124 Voluntary extra rounds (195 listeners)

195 listeners voluntarily continued past the required 6 rounds, contributing an additional 1,124 judgements not included in the headline ranking. Their inclusion does not materially change the overall result.

Selection reasons

Why listeners preferred a track

Top reasons cited across 4,783 decisive judgements.
Melody is catchy38.6%
Matches the intended emotion31.7%
Fewer audio artifacts7.7%

Coverage

Results by scenario

13 scenarios group the 104 prompts by musical style and difficulty. None of the per-scenario leader gaps are statistically significant after correction for multiple comparisons — read this table as directional, not conclusive.
Scenario Description Judgements Tie / reject Leader
B01–B02Vocal pop, mixed voice12 prompt cases1,5185.1% / 14.7%Mureka68.5
B03–B04Vocal pop, mixed voice12 prompt cases1,5166.9% / 12.9%Mureka63.6
B05–B06Vocal pop, mixed voice12 prompt cases1,5126.7% / 13.4%Suno60.2
B07–B08Vocal pop, mixed voice12 prompt cases1,5067.4% / 12.6%Mureka64.8
B09–B10Vocal pop, mixed voice12 prompt cases1,4946.6% / 14.1%Suno61.9
B11–B12Vocal pop, mixed voice12 prompt cases1,5027.3% / 13.8%Mureka66.2
B15Instrumental-only6 prompt cases7566.9% / 13.0%Suno58.9
B16Instrumental-only6 prompt cases7446.6% / 12.9%Mureka61.4
B17Commercial / functional brief6 prompt cases7506.4% / 13.7%Mureka63.0
B18Commercial / functional brief6 prompt cases7446.9% / 13.6%Suno59.7
B19Commercial / functional brief6 prompt cases7386.8% / 14.0%Mureka62.2
P07Adversarial stress prompt1 prompt case1264.8% / 15.9%Suno65.4
P08Adversarial stress prompt1 prompt case365.6% / 16.7%Suno68.6
Reading the scenario table Note

P07 and P08 are each built from a single prompt, so their scores rest on far fewer effective samples than the vocal scenarios.

Sample size

Vocal scenarios (B01–B12) draw on 12 prompt cases each and around 1,500 judgements; the two stress scenarios draw on a single prompt each and under 130 judgements — treat their point estimates accordingly.

Consistency

Mureka or Suno leads in all 13 scenarios, consistent with the overall ranking; neither Happy Shrimp, MiniMax, nor Lyria tops any individual scenario.

Who listened

Audience

Demographics and recruitment funnel for the 1,000-listener primary cohort.
Profile coverage 100.0% 1,000 of 1,000 listeners completed a profile
Avg. genres selected 2.14 2,140 total selections across 1,000 respondents
Beta-test willing — Share of listeners open to future beta testing

Recruitment

From sign-up to extended rounds

82.6% completion rate · 19.5% of completers chose to extend
011,211Registered
021,000Completed 6 rounds
03195Extended voluntarily
04184Completed 12 rounds
82.6% completion rate; 19.5% of the 1,000-person primary cohort chose to continue past round 6.

Self-description

Who the listeners are

AI music hobbyist47.4%474
Professional musician21.0%210
Content creator19.7%197
Casual sharer11.9%119

Demographics

Age range

Under 181.3%13
18–2414.5%145
25–3434.2%342
35–4428.6%286
45+21.4%214
Reading the audience data Note

The primary cohort skews toward self-identified AI-music hobbyists and working musicians, not a general-population sample.

Composition

68% of listeners identify as either an AI music hobbyist or a professional musician, meaning results reflect an audience with more musical familiarity than the general public.

Age skew

63% of listeners are between 25 and 44 — under-18 and 45+ listeners are a small minority of the cohort.

Engagement

Listening behavior

How long listeners spent, and how their voting patterns evolved across rounds.
Time to complete 12.1 min P25 8.2 · P75 20.6
Listen time per track 42.3 sec P25 28.7 · P75 70.4
Both tracks reached minimum listen — Tracked per judgement
Extended past round 6 19.5% 195 listeners

By round

Voting pattern per round

Round 1 (605 judgements) was fixed to a pop-vocal scenario for all listeners by design, so its figures are not directly comparable to later rounds.
Decisive preference Tie Reject
Round 1
Median listen 41.2sTie 6.1% · Reject 14.5%
Round 2
Median listen 42.0sTie 6.8% · Reject 13.4%
Round 3
Median listen 42.5sTie 6.9% · Reject 13.2%
Round 4
Median listen 42.7sTie 6.9% · Reject 13.6%
Round 5
Median listen 43.0sTie 6.9% · Reject 13.7%
Round 6
Median listen 43.3sTie 6.9% · Reject 13.9%
Voting patterns were broadly stable across all 6 primary rounds after round 1's fixed-scenario effect.

Extension

Voluntary rounds 7–12

195 of the 1,000 primary-cohort listeners voluntarily continued past the required 6 rounds.
184 of 1,000 completed all 12 rounds 18.4%

Qualitative signal

Feedback

Why tracks were rejected, and what listeners said in their own words.

Rejections

Why both tracks were rejected

Top reasons cited across 815 rejected judgements.
Melody is plain / uninteresting37.6%306
Style mismatch with the prompt26.9%219
Song is incomplete1.7%14

Free-text comments

What listeners said in their own words

1,501 listeners left a free-text comment alongside their votes.

How this was measured

Methodology

Five design choices that shape every number in this report.
  1. 01

    Cohort freeze

    1,211 people registered; the first 1,000 to complete all 6 primary rounds, ordered by completion time, form the frozen primary cohort used throughout this report.

  2. 02

    Unit of analysis

    Each judgement is one listener's vote on one blind pair. Judgements are the base unit for every downstream aggregate score.

  3. 03

    Modeling approach

    Scene-equalized scores aggregate pairwise preferences per scenario, then combine across the 13 scenarios with equal weighting regardless of judgement count.

  4. 04

    Uncertainty

    Confidence intervals and rank-1 probabilities come from 300 bootstrap resamples (seed 20260829) at the judgement level, within each scenario.

  5. 05

    Robustness

    Results are checked against the extended 7,124-judgement dataset (including voluntary extra rounds) to confirm the primary-cohort ranking is not an artifact of the 6-round cutoff.

Fully blind: neither listener nor interface reveals which system made which track
Scene-based: 13 scenarios spanning vocal, instrumental, commercial, and stress prompts
Equalized: every scenario contributes equally to the overall score
Reproducible: fixed bootstrap seed, published methodology
Academic note: how the ranking is computed

Pairwise preferences are aggregated per scenario, then combined across the 13 scenarios using equal weighting. Uncertainty is estimated via non-parametric bootstrap: judgements are resampled with replacement within each scenario, the scene-equalized score is recomputed, and this is repeated across replicates to produce the reported intervals and rank-1 probabilities.

Sample
1,000-listener frozen primary cohort, 6,000 judgements
Unit
Judgement-level (one listener, one blind pair)
Repeats
300 bootstrap replicates, seed 20260829
Output
Scene-equalized score, 95% interval, rank-1 probability

Probability of ranking #1

SystemP(rank 1)
Mureka86.7%
Suno13.3%
Happy Shrimp0.0%
MiniMax0.0%
Lyria0.0%

Disclosure

Research transparency

What's settled, and what's still in progress.

At a glance

Study profile

Status
Included: 1,000-listener frozen primary cohort
Estimand
Population-average pairwise preference among Mandarin-speaking listeners recruited for this study
Analysis unit
Judgement (one listener × one blind pair)
Outcomes
Left/right preference, tie, reject — plus free-text comments and selection reasons
Uncertainty
300 bootstrap replicates, seed 20260829
Snapshot
2026-09-01T14:45:25+08:00

Ready

Already disclosed

  • EstimandThe precise quantity being estimated is stated above, not left implicit.
  • Cohort definitionThe frozen-cohort rule (first 1,000 to complete 6 rounds) is stated and reproducible.
  • Analysis unitJudgement-level, not listener-level or model-level aggregation.
  • OutcomesAll four vote options and their exact counts are reported, not just the winner.
  • UncertaintyBootstrap seed and replicate count are published so intervals are reproducible.

Pending

Not yet disclosed

  • AuthorshipIndividual author attribution for the accompanying academic preprint is pending finalization.
  • GovernanceFormal data-governance documentation beyond this report is in progress.
  • RecruitmentDetailed recruitment-channel breakdown is not yet published.
  • Model provenanceExact model checkpoint/version pinning beyond the version labels shown is not fully documented.
  • Generation protocolFull prompt-to-generation pipeline details are not yet published.
  • Audio protocolPlayback and normalization protocol details are not yet published.
  • RegistrationThis study was not pre-registered.
  • Release packageThe complete dataset release package is described in the Data access section below.

Read before citing

Limitations

Six constraints that bound how far these results generalize.
01

Population

Listeners were recruited through channels that skew toward AI-music hobbyists and working musicians, not a general-population sample.

Response: results describe this specific listener population, not the general public.
02

Construct

"Preference" here means "which track a listener picked in a 30–60 second blind comparison," not a holistic judgement of overall song quality or commercial viability.

Response: treat scores as a preference measure, not a quality certification.
03

Stimulus variance

Two candidate tracks per system per prompt is a small sample of each system's possible output — a single unlucky generation can shift a scenario result.

Response: scenario-level rankings should be read as directional, not conclusive.
04

Multiplicity

13 scenarios and 10 pairwise comparisons create many opportunities for a false-positive "leader" to emerge by chance alone.

Response: none of the 13 scenario leader gaps survive correction for multiple comparisons.
05

Production threshold

This benchmark measures generation preference only — it does not test whether a track is ready to ship commercially without further editing.

Response: pair these results with your own production-readiness review.
06

Version drift

Model versions are pinned to the August 2026 data-collection window. Vendors ship updates frequently, and newer releases are not reflected here.

Response: check the version labels in the scope table before assuming current relevance.

Acknowledgements

Contributors

1,000 listeners made this study possible.
1,000 Listeners acknowledged
Sorting
Alphabetical by display name, not by activity or influence.
Naming
Listeners are shown by their chosen display name, or a placeholder if none was provided.

Open data

Dataset access

Reviewed on request

The full judgement-level dataset behind this report is available to researchers and organizations on request.

01

Listeners

De-identified demographic and engagement fields for the 1,000-listener frozen primary cohort.

02

Songs

Every candidate track's prompt text, style tags, genre, vocal/instrumental flag, and a direct audio link.

03

Judgements

Every listener's vote (left/right/tie/reject), listen duration, selection reasons, and free-text comment.

Supporting files

What ships alongside the three tables

A field-level README documents every column, and a small verification script cross-checks referential integrity between tables before release.

  • README.md — full field dictionary
  • Referential integrity check between CSVs

Anonymization

How listener identity is protected

Listener identifiers are one-way pseudonymized before release. Each listener's public identifier is derived as: participant_id = Base32(HMAC-SHA256(benchmark_secret, internal_user_id))[0:20]
01

No identity fields

Account IDs, contact details, and any personally identifying fields are removed before export.

02

No quasi-identifiers

Fields that could re-identify a listener when combined with other data are excluded or generalized.

03

Screened free text

Comment fields are checked for accidental disclosure of contact information or names before release.

04

No cross-dataset linkage

Pseudonymous IDs are salted specifically for this release and cannot be joined back to any other dataset we publish.

Process

How to request access

Four steps from request to delivery.
  1. 01

    Apply

    Send an email with your name, organization, and intended use of the data.

  2. 02

    Verify

    We confirm the request comes from a real organization and a reachable contact.

  3. 03

    Review

    Each request is reviewed individually against the stated use case.

  4. 04

    Deliver

    Approved requests receive the dataset and README by email.

Request access

Email us to request the dataset

There is no online form — access requests are handled directly by email so each one can be reviewed individually.
Use a work or institutional email

This helps us verify the request comes from a real organization and speeds up review.

  • Use the data only for the purpose stated in your request
  • Do not attempt to re-identify any listener
  • Do not redistribute the raw dataset to third parties
  • Cite this report if you publish results derived from it

To request access, email contact@aimusicbenchmark.org with your name, organization, and intended use of the data.

Please mention your role, your organization's website, and roughly how you plan to use the dataset. We review each request individually and typically respond within a few business days.

Version. This release corresponds to the frozen dataset described in this report's methodology section.

Audio. Direct links to all 1,040 candidate tracks are included in the songs table.

Delivery. Approved requests are sent the dataset and README directly by email.