A Thousand Listeners
Which AI music model actually sounds better?
We ran a blind, head-to-head listening study across five AI music generation systems — 1,000 listeners, 6,000 primary judgements, 104 prompts spanning 13 scenarios. No system knew it was being judged against the others; no listener knew which model made which track.
Headline result
Overall leaderboard
| Rank | System | Score | 95% interval | P(ranked #1) |
|---|---|---|---|---|
| 01 | Mureka | 65.6% | 62.3%–68.8% | 86.7% |
| 02 | Suno | 62.5% | 59.1%–66.4% | 13.3% |
| 03 | Happy Shrimp | 47.8% | 44.2%–50.9% | 0.0% |
| 04 | MiniMax | 39.4% | 36.0%–43.1% | 0.0% |
| 05 | Lyria | 34.6% | 31.7%–37.8% | 0.0% |
Scores come from 300 bootstrap resamples (seed 20260829) of judgement-level data, resampled within each scenario and then equal-weighted across all 13 scenarios. "P(ranked #1)" is the share of resamples in which a system finished first overall.
| Rank | System | Selected | Comparisons | Selection rate |
|---|---|---|---|---|
| 01 | Mureka | 1,401 | 2,225 | 63.0% |
| 02 | Suno | 1,340 | 2,099 | 63.8% |
| 03 | Happy Shrimp | 870 | 1,971 | 44.1% |
| 04 | MiniMax | 608 | 1,858 | 32.7% |
| 05 | Lyria | 399 | 1,858 | 21.5% |
What was tested
Scope of the benchmark
Five systems, evaluated on the same 104 prompts across 13 scenarios spanning vocal genres, instrumental-only tracks, commercial/functional briefs, and adversarial stress prompts.Design
Each listener judged 6 blind pairs per session, one track from each of two randomly assigned systems, with left/right position randomized.
Language
All prompts and generated lyrics are in Mandarin Chinese. Results should not be assumed to generalize to other languages.
Priority
Listener preference is the primary outcome. Technical benchmarks (loudness, artifacts) are secondary and reported only where they explain a preference.
Feature comparison
What each system can do
| Capability | SunoEvaluated V5.5 | MurekaEvaluated V9.5 | MiniMaxEvaluated Music 3.0 | LyriaEvaluated 3.5 | Happy ShrimpEvaluated 1.0 |
|---|---|---|---|---|---|
| Generation | |||||
| Prompt to full song | Full song | Full song | Full song | Full song | Full song |
| Custom lyrics input | Custom lyrics | Custom lyrics | Custom lyrics | Custom lyrics | Custom lyrics |
| Instrumental-only | Supported | Supported | Supported | Supported | Supported |
| Style/direction control | Style + voice tags | Style + voice tags | Style + voice tags | Style + voice tags | Natural-language only |
| Conditioning | |||||
| Reference input | Audio reference | Audio reference | Song reference | Image reference | Not found |
| Cover / restyle | Cover | Remix | Cover | Not found | Not found |
| Extend an existing track | Extend | Lyrics-guided extend | Not found | Not found | Not found |
| Reusable voice identity | Personas | Reference-based | Not found | Not found | Not found |
| Production | |||||
| Section editing | Replace section | Rewrite part | Not found | Not found | Not found |
| Stem separation | Advanced stems | Stem splitter | Voice isolation | Not found | Not found |
| Multitrack studio | Full studio | Studio DAW | Not found | Not found | Not found |
| Real-time generation | Not found | Not found | Not found | Lyria RealTime | Not found |
| Image-to-music | Not found | Not found | Not found | Supported | Not found |
Suno
The most feature-complete production suite among the five: personas, section-level rewrites, extend, and a full multitrack studio in addition to generation.
Mureka
Strong generation quality paired with a lighter production toolset — remix, lyrics-guided extend, and a studio-lite editor.
MiniMax
Solid core generation and song-reference conditioning; production tooling beyond voice isolation is limited or unverified.
Lyria
Distinguished by real-time generation and image-to-music conditioning — capabilities absent from the other four systems.
Happy Shrimp
The newest system in the benchmark, released one day before data generation began. Direction control is limited to natural-language prompts.
Before the data
Questions this report answers
Three questions frame everything that follows.Which system wins overall?
A single head-to-head ranking across all scenarios, with uncertainty intervals — not just a leaderboard number.
Does the ranking hold by scene?
Whether the overall winner also wins on vocal pop, instrumental, commercial briefs, and adversarial prompts specifically.
Who is the audience, and how did they judge?
Who the 1,000 listeners are, how they engaged with the study, and what qualities actually drove their preferences.
Executive summary
Four things to take away
The full statistical picture, condensed.Mureka leads with a clear signal
Mureka was selected in 1,401 of 2,225 comparisons and finished first overall with an 86.7% bootstrap probability of ranking #1 — the clearest signal in the dataset.
The #1 vs #2 gap is narrower than it looks
Head-to-head, Mureka beats Suno only 52.4% of the time [47.1%, 56.9%] — an interval that crosses 50%, so the top-two ordering carries real uncertainty at the pairwise level.
Use the scenario table for your use case
The overall ranking is a 13-scenario average. If your use case matches one specific scenario (e.g. instrumental, commercial), check that row directly.
Read the confidence intervals
Every score in this report is estimated from a finite sample of 1,000 listeners. Point estimates without their intervals can overstate how settled a ranking is.
Reader's guide
Different readers, different depth
AI researchers
Full statistical methodology, bootstrap procedure, and raw data access.
Go to methodologyMusic academia
Study design, listener demographics, and the qualitative feedback behind the numbers.
Go to audienceModel teams
Where your system ranks by scenario, and what listeners said drove their choices.
Go to scenariosMedia
The headline result and executive summary — everything needed for a quick, accurate write-up.
Go to overviewMusic industry
What listeners actually value — melody, emotional match, and production polish — and how each system delivers on it.
Go to feedbackEvidence
Vote outcomes
Listener chose one track as clearly better.
Listener judged both tracks equally good.
Listener judged both tracks unusable for the prompt.
Ties and rejections are excluded from pairwise preference percentages elsewhere in this report, but are tracked here as their own outcome category.
Head-to-head
Pairwise preference matrix
| System | Suno | MiniMax | Mureka | Happy Shrimp | Lyria |
|---|---|---|---|---|---|
| Suno | · | 68.2%tie 7.8% | 47.6%tie 8.8% | 62.3%tie 8.4% | 72.1%tie 7.6% |
| MiniMax | 31.8%tie 7.8% | · | 29.2%tie 7.7% | 42.4%tie 8.6% | 54.4%tie 8.3% |
| Mureka | 52.4%tie 8.8% | 70.8%tie 7.7% | · | 64.9%tie 8.3% | 74.4%tie 7.5% |
| Happy Shrimp | 37.7%tie 8.4% | 57.6%tie 8.6% | 35.1%tie 8.3% | · | 60.8%tie 8.2% |
| Lyria | 27.9%tie 7.6% | 45.6%tie 8.3% | 25.6%tie 7.5% | 39.2%tie 8.2% | · |
Cells show the row system's win rate over the column system among decisive judgements only; hover a cell for the full interval and sample size.
Mureka beats every other system head-to-head except when facing Suno, where the margin narrows to a coin flip.
Transitivity
The ranking is broadly transitive — Mureka > Suno > Happy Shrimp > MiniMax > Lyria holds in nearly every pairwise comparison, with the Mureka/Suno pair as the sole close contest.
Tie rates
Tie rates cluster tightly around 7–9% regardless of which two systems are compared, suggesting ties reflect genuine draws rather than one system's tracks being uniformly indistinguishable.
Data quality
Diagnostics
Cohort completion
1,000 1,000 listeners completed all 6 primary rounds and are included in every result on this page.Reject rate
13.6% Both tracks judged unusable for the prompt.Usable comparisons
5,185 6.7% tie, 13.6% reject — the remainder feed the preference scores above.Robustness
Beyond the primary six rounds
195 listeners voluntarily continued past the required 6 rounds, contributing an additional 1,124 judgements not included in the headline ranking. Their inclusion does not materially change the overall result.
Selection reasons
Why listeners preferred a track
Top reasons cited across 4,783 decisive judgements.Coverage
Results by scenario
| Scenario | Description | Judgements | Tie / reject | Leader |
|---|---|---|---|---|
| B01–B02 | Vocal pop, mixed voice12 prompt cases | 1,518 | 5.1% / 14.7% | Mureka68.5 |
| B03–B04 | Vocal pop, mixed voice12 prompt cases | 1,516 | 6.9% / 12.9% | Mureka63.6 |
| B05–B06 | Vocal pop, mixed voice12 prompt cases | 1,512 | 6.7% / 13.4% | Suno60.2 |
| B07–B08 | Vocal pop, mixed voice12 prompt cases | 1,506 | 7.4% / 12.6% | Mureka64.8 |
| B09–B10 | Vocal pop, mixed voice12 prompt cases | 1,494 | 6.6% / 14.1% | Suno61.9 |
| B11–B12 | Vocal pop, mixed voice12 prompt cases | 1,502 | 7.3% / 13.8% | Mureka66.2 |
| B15 | Instrumental-only6 prompt cases | 756 | 6.9% / 13.0% | Suno58.9 |
| B16 | Instrumental-only6 prompt cases | 744 | 6.6% / 12.9% | Mureka61.4 |
| B17 | Commercial / functional brief6 prompt cases | 750 | 6.4% / 13.7% | Mureka63.0 |
| B18 | Commercial / functional brief6 prompt cases | 744 | 6.9% / 13.6% | Suno59.7 |
| B19 | Commercial / functional brief6 prompt cases | 738 | 6.8% / 14.0% | Mureka62.2 |
| P07 | Adversarial stress prompt1 prompt case | 126 | 4.8% / 15.9% | Suno65.4 |
| P08 | Adversarial stress prompt1 prompt case | 36 | 5.6% / 16.7% | Suno68.6 |
P07 and P08 are each built from a single prompt, so their scores rest on far fewer effective samples than the vocal scenarios.
Sample size
Vocal scenarios (B01–B12) draw on 12 prompt cases each and around 1,500 judgements; the two stress scenarios draw on a single prompt each and under 130 judgements — treat their point estimates accordingly.
Consistency
Mureka or Suno leads in all 13 scenarios, consistent with the overall ranking; neither Happy Shrimp, MiniMax, nor Lyria tops any individual scenario.
Who listened
Audience
Recruitment
From sign-up to extended rounds
Self-description
Who the listeners are
Demographics
Age range
The primary cohort skews toward self-identified AI-music hobbyists and working musicians, not a general-population sample.
Composition
68% of listeners identify as either an AI music hobbyist or a professional musician, meaning results reflect an audience with more musical familiarity than the general public.
Age skew
63% of listeners are between 25 and 44 — under-18 and 45+ listeners are a small minority of the cohort.
Engagement
Listening behavior
By round
Voting pattern per round
Extension
Voluntary rounds 7–12
Qualitative signal
Feedback
Rejections
Why both tracks were rejected
Top reasons cited across 815 rejected judgements.Free-text comments
What listeners said in their own words
1,501 listeners left a free-text comment alongside their votes.How this was measured
Methodology
Five design choices that shape every number in this report.-
01
Cohort freeze
1,211 people registered; the first 1,000 to complete all 6 primary rounds, ordered by completion time, form the frozen primary cohort used throughout this report.
-
02
Unit of analysis
Each judgement is one listener's vote on one blind pair. Judgements are the base unit for every downstream aggregate score.
-
03
Modeling approach
Scene-equalized scores aggregate pairwise preferences per scenario, then combine across the 13 scenarios with equal weighting regardless of judgement count.
-
04
Uncertainty
Confidence intervals and rank-1 probabilities come from 300 bootstrap resamples (seed 20260829) at the judgement level, within each scenario.
-
05
Robustness
Results are checked against the extended 7,124-judgement dataset (including voluntary extra rounds) to confirm the primary-cohort ranking is not an artifact of the 6-round cutoff.
Academic note: how the ranking is computed
Pairwise preferences are aggregated per scenario, then combined across the 13 scenarios using equal weighting. Uncertainty is estimated via non-parametric bootstrap: judgements are resampled with replacement within each scenario, the scene-equalized score is recomputed, and this is repeated across replicates to produce the reported intervals and rank-1 probabilities.
- Sample
- 1,000-listener frozen primary cohort, 6,000 judgements
- Unit
- Judgement-level (one listener, one blind pair)
- Repeats
- 300 bootstrap replicates, seed 20260829
- Output
- Scene-equalized score, 95% interval, rank-1 probability
Probability of ranking #1
| System | P(rank 1) |
|---|---|
| Mureka | 86.7% |
| Suno | 13.3% |
| Happy Shrimp | 0.0% |
| MiniMax | 0.0% |
| Lyria | 0.0% |
Disclosure
Research transparency
At a glance
Study profile
- Status
- Included: 1,000-listener frozen primary cohort
- Estimand
- Population-average pairwise preference among Mandarin-speaking listeners recruited for this study
- Analysis unit
- Judgement (one listener × one blind pair)
- Outcomes
- Left/right preference, tie, reject — plus free-text comments and selection reasons
- Uncertainty
- 300 bootstrap replicates, seed 20260829
- Snapshot
- 2026-09-01T14:45:25+08:00
Ready
Already disclosed
- EstimandThe precise quantity being estimated is stated above, not left implicit.
- Cohort definitionThe frozen-cohort rule (first 1,000 to complete 6 rounds) is stated and reproducible.
- Analysis unitJudgement-level, not listener-level or model-level aggregation.
- OutcomesAll four vote options and their exact counts are reported, not just the winner.
- UncertaintyBootstrap seed and replicate count are published so intervals are reproducible.
Pending
Not yet disclosed
- AuthorshipIndividual author attribution for the accompanying academic preprint is pending finalization.
- GovernanceFormal data-governance documentation beyond this report is in progress.
- RecruitmentDetailed recruitment-channel breakdown is not yet published.
- Model provenanceExact model checkpoint/version pinning beyond the version labels shown is not fully documented.
- Generation protocolFull prompt-to-generation pipeline details are not yet published.
- Audio protocolPlayback and normalization protocol details are not yet published.
- RegistrationThis study was not pre-registered.
- Release packageThe complete dataset release package is described in the Data access section below.
Acknowledgements
Contributors
- Sorting
- Alphabetical by display name, not by activity or influence.
- Naming
- Listeners are shown by their chosen display name, or a placeholder if none was provided.
Open data
Dataset access
The full judgement-level dataset behind this report is available to researchers and organizations on request.
Listeners
De-identified demographic and engagement fields for the 1,000-listener frozen primary cohort.
Songs
Every candidate track's prompt text, style tags, genre, vocal/instrumental flag, and a direct audio link.
Judgements
Every listener's vote (left/right/tie/reject), listen duration, selection reasons, and free-text comment.
Supporting files
What ships alongside the three tables
A field-level README documents every column, and a small verification script cross-checks referential integrity between tables before release.
- README.md — full field dictionary
- Referential integrity check between CSVs
Anonymization
How listener identity is protected
Listener identifiers are one-way pseudonymized before release. Each listener's public identifier is derived as:participant_id = Base32(HMAC-SHA256(benchmark_secret, internal_user_id))[0:20]
No identity fields
Account IDs, contact details, and any personally identifying fields are removed before export.
No quasi-identifiers
Fields that could re-identify a listener when combined with other data are excluded or generalized.
Screened free text
Comment fields are checked for accidental disclosure of contact information or names before release.
No cross-dataset linkage
Pseudonymous IDs are salted specifically for this release and cannot be joined back to any other dataset we publish.
Process
How to request access
Four steps from request to delivery.-
01
Apply
Send an email with your name, organization, and intended use of the data.
-
02
Verify
We confirm the request comes from a real organization and a reachable contact.
-
03
Review
Each request is reviewed individually against the stated use case.
-
04
Deliver
Approved requests receive the dataset and README by email.
Request access
Email us to request the dataset
There is no online form — access requests are handled directly by email so each one can be reviewed individually.This helps us verify the request comes from a real organization and speeds up review.
- Use the data only for the purpose stated in your request
- Do not attempt to re-identify any listener
- Do not redistribute the raw dataset to third parties
- Cite this report if you publish results derived from it
To request access, email contact@aimusicbenchmark.org with your name, organization, and intended use of the data.
Please mention your role, your organization's website, and roughly how you plan to use the dataset. We review each request individually and typically respond within a few business days.




