2026:Music Audio Generation
Task Captain: Yingxin Liang, Rythmix AI
Contents
Task Description
The MIREX 2026 Cover Song Generation task asks an AI music system to remake a full song according to a style or arrangement prompt.
The output can change the style, instruments, arrangement, timbre, or tempo, but it should still be recognizable as a version of the source song. The evaluation checks whether the melody is preserved, the harmony works, the key is coherent, the prompt is followed, and the arrangement sounds complete.
For each test item, participants will receive:
- one full-length source song
- one prompt
Participants must generate one complete cover-song audio file for every test item. The output should cover the full song structure rather than provide only an excerpt, preview, or isolated section. No fixed duration is imposed because tempo and arrangement may change.
Examples
The following examples use the same full-length source song with three target style prompts. For each prompt, results from two reference-conditioned systems are provided.
These examples are for explanation only and will not be used as test items.
Full audio player: Listen to all examples
| Source audio | ACE-Step 1.5 SFT XL | SongEcho | Exact prompt |
|---|---|---|---|
| Source song (MP3) | Audio (MP3) | Audio (MP3) | Prompt 01 — Classic pop ballad
|
| Audio (MP3) | Audio (MP3) | Prompt 02 — Indie rock
| |
| Audio (MP3) | Audio (MP3) | Prompt 03 — Soul folk acapella
|
The examples should make the input-output relationship clear: the system receives the source song and prompt, then generates a new full-song arrangement. They are not examples of voice conversion or ordinary human-performed covers.
Evaluation
The official evaluation will be based on expert listening. Three trained evaluators are planned. System identities will be hidden during evaluation.
Each generated cover will be rated on a five-point scale in the following dimensions:
- Melody - whether the main melody is preserved or plausibly adapted, without salient pitch or contour errors.
- Harmony - whether the chord progression and chord-melody relationship are musically coherent.
- Key Consistency - whether the output maintains a coherent tonal center, unless modulation is intentional.
- Style Consistency - whether the output follows the target style or arrangement prompt.
- Arrangement and Production - whether the instrumentation, texture, balance, timbre, and production support a complete musical result.
The rating scale is:
- 5: very good, with no salient problem;
- 4: acceptable, with only minor imperfections;
- 3: a noticeable issue, but still musically usable;
- 2: a severe issue affecting musical coherence or usability;
- 1: a substantial or persistent failure.
The official system score will be the average of the five dimension scores across test items and evaluators. Dimension-level results will also be reported.
Evaluators may provide short time-stamped notes for important errors. Inter-rater agreement will be reported with the final results.
Supporting Objective Features
The organizers may report diagnostic features such as large-leap ratio (LLR), pitch range (PR), pitch stability (PS), key confidence (KC), in-key note ratio (IKNR), key-change rate (KCR), loudness deviation, loudness range (LRA), and spectral contrast (SC).
KCR refers to the rate at which the estimated local key changes across adjacent analysis windows. These features are provided only as descriptive diagnostic evidence. They will not contribute to the official score, system ranking, or tie-breaking. In particular, key-related features do not measure chord-function quality and will not be used as automatic harmony scores. PR and PS are also descriptive pitch summaries rather than validated quality scores.
Dataset
The test set will focus on contemporary tonal popular music, such as pop, pop rock, synth-pop, dance-pop, acoustic pop, folk-pop, and pop ballads.
To keep the evaluation consistent, styles dominated by highly complex, non-functional, modal, or tonally ambiguous harmony will not be included.
All source songs will be provided by the task captain, who will secure the permissions needed to distribute and transform them for this evaluation. Participants may use the source songs only for this task and may not redistribute them.
Official test materials must not be used for training or fine-tuning. Participants may generate multiple candidates and select one for submission, but the selection process must be disclosed.
Submission
Both commercial and open-source systems are allowed, provided that participants have the right to use the system for this evaluation and disclose the system name, version, and any external services used.
Participants will run their systems themselves and submit generated audio and metadata. Model code and checkpoints are not required.
Each submission should include:
- one complete audio output for every official source-prompt pair;
- metadata linking each output to its test item;
- a short system description or MIREX extended abstract; and
- disclosure of training data, pretrained models, candidate selection, and any manual post-processing.
For each test item, participants must report the total number of candidate outputs generated before selecting the submitted version. For example, a submitted output selected from 20 generations should be reported as 1 selected from 20.
The final audio format, naming convention, and submission method will follow the instructions provided by the MIREX organizers.
Baseline
At least one reference-conditioned music generation system will be included as a baseline. A text-only music generation system may also be reported as a comparison because it does not use the source-song reference.
Before the test set is released, the organizers will run the reference-conditioned baseline on all test items to confirm that the inputs, prompts, audio requirements, and evaluation workflow function as intended.
Resources
- Diagnostic framework and code: https://github.com/TiaaL/songecho-cover-metrics
- Associated paper: A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features
Timeline
| Stage | Date |
|---|---|
| Test songs and prompts release | August 5, 2026 |
| Submission deadline | September 15, 2026 |
| Evaluation questionnaire preparation | September 16-20, 2026 |
| Subjective evaluation | September 21-October 5, 2026 |
| Result aggregation and analysis | October 6-12, 2026 |
| Final results ready | By October 15, 2026 |
Contact
If you have any questions or suggestions about the task, please contact:
- Yingxin Liang — Hangzhou Xiaoying Innovation Technology Co., Ltd. (Rythmix AI) - Email: yingxinliang.yxl@gmail.com yingxinliang.yxl@gmail.com