<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://music-ir.org/mirex/w/index.php?action=history&amp;feed=atom&amp;title=2026%3AMusic_Audio_Generation</id>
	<title>2026:Music Audio Generation - Revision history</title>
	<link rel="self" type="application/atom+xml" href="https://music-ir.org/mirex/w/index.php?action=history&amp;feed=atom&amp;title=2026%3AMusic_Audio_Generation"/>
	<link rel="alternate" type="text/html" href="https://music-ir.org/mirex/w/index.php?title=2026:Music_Audio_Generation&amp;action=history"/>
	<updated>2026-07-27T19:21:04Z</updated>
	<subtitle>Revision history for this page on the wiki</subtitle>
	<generator>MediaWiki 1.31.1</generator>
	<entry>
		<id>https://music-ir.org/mirex/w/index.php?title=2026:Music_Audio_Generation&amp;diff=15064&amp;oldid=prev</id>
		<title>Junyan: /* Contact */</title>
		<link rel="alternate" type="text/html" href="https://music-ir.org/mirex/w/index.php?title=2026:Music_Audio_Generation&amp;diff=15064&amp;oldid=prev"/>
		<updated>2026-07-26T16:17:56Z</updated>

		<summary type="html">&lt;p&gt;‎&lt;span dir=&quot;auto&quot;&gt;&lt;span class=&quot;autocomment&quot;&gt;Contact&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;table class=&quot;diff diff-contentalign-left&quot; data-mw=&quot;interface&quot;&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;tr class=&quot;diff-title&quot; lang=&quot;en&quot;&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #222; text-align: center;&quot;&gt;← Older revision&lt;/td&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #222; text-align: center;&quot;&gt;Revision as of 16:17, 26 July 2026&lt;/td&gt;
				&lt;/tr&gt;&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l153&quot; &gt;Line 153:&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;Line 153:&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class='diff-marker'&gt;&amp;#160;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #222; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;- Email: yingxinliang.yxl@gmail.com&lt;/div&gt;&lt;/td&gt;&lt;td class='diff-marker'&gt;&amp;#160;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #222; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;- Email: yingxinliang.yxl@gmail.com&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class='diff-marker'&gt;&amp;#160;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #222; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;[mailto:yingxinliang.yxl@gmail.com yingxinliang.yxl@gmail.com]&lt;/div&gt;&lt;/td&gt;&lt;td class='diff-marker'&gt;&amp;#160;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #222; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;[mailto:yingxinliang.yxl@gmail.com yingxinliang.yxl@gmail.com]&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class='diff-marker'&gt;−&lt;/td&gt;&lt;td style=&quot;color: #222; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;&lt;/del&gt;&lt;/div&gt;&lt;/td&gt;&lt;td colspan=&quot;2&quot;&gt;&amp;#160;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class='diff-marker'&gt;−&lt;/td&gt;&lt;td style=&quot;color: #222; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;[[Category:MIREX 2026]]&lt;/del&gt;&lt;/div&gt;&lt;/td&gt;&lt;td colspan=&quot;2&quot;&gt;&amp;#160;&lt;/td&gt;&lt;/tr&gt;
&lt;/table&gt;</summary>
		<author><name>Junyan</name></author>
		
	</entry>
	<entry>
		<id>https://music-ir.org/mirex/w/index.php?title=2026:Music_Audio_Generation&amp;diff=15062&amp;oldid=prev</id>
		<title>Junyan: Created page with &quot;'''Task Captain:''' Yingxin Liang, Rythmix AI  == Task Description ==  The MIREX 2026 Cover Song Generation task asks an AI music system to remake a full song according to a s...&quot;</title>
		<link rel="alternate" type="text/html" href="https://music-ir.org/mirex/w/index.php?title=2026:Music_Audio_Generation&amp;diff=15062&amp;oldid=prev"/>
		<updated>2026-07-26T16:17:25Z</updated>

		<summary type="html">&lt;p&gt;Created page with &amp;quot;&amp;#039;&amp;#039;&amp;#039;Task Captain:&amp;#039;&amp;#039;&amp;#039; Yingxin Liang, Rythmix AI  == Task Description ==  The MIREX 2026 Cover Song Generation task asks an AI music system to remake a full song according to a s...&amp;quot;&lt;/p&gt;
&lt;p&gt;&lt;b&gt;New page&lt;/b&gt;&lt;/p&gt;&lt;div&gt;'''Task Captain:''' Yingxin Liang, Rythmix AI&lt;br /&gt;
&lt;br /&gt;
== Task Description ==&lt;br /&gt;
&lt;br /&gt;
The MIREX 2026 Cover Song Generation task asks an AI music system to remake a full song according to a style or arrangement prompt.&lt;br /&gt;
&lt;br /&gt;
The output can change the style, instruments, arrangement, timbre, or tempo, but it should still be recognizable as a version of the source song. The evaluation checks whether the melody is preserved, the harmony works, the key is coherent, the prompt is followed, and the arrangement sounds complete.&lt;br /&gt;
&lt;br /&gt;
For each test item, participants will receive:&lt;br /&gt;
&lt;br /&gt;
* one full-length source song&lt;br /&gt;
* one prompt&lt;br /&gt;
&lt;br /&gt;
Participants must generate one complete cover-song audio file for every test item. The output should cover the full song structure rather than provide only an excerpt, preview, or isolated section. No fixed duration is imposed because tempo and arrangement may change.&lt;br /&gt;
&lt;br /&gt;
=== Examples ===&lt;br /&gt;
&lt;br /&gt;
The following examples use the same full-length source song with three target&lt;br /&gt;
style prompts. For each prompt, results from two reference-conditioned systems&lt;br /&gt;
are provided.&lt;br /&gt;
&lt;br /&gt;
These examples are for explanation only and will not be used as test items.&lt;br /&gt;
&lt;br /&gt;
'''Full audio player:''' [https://tiaal.github.io/songecho-cover-metrics/mirex-audio/ Listen to all examples]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Source audio&lt;br /&gt;
! ACE-Step 1.5 SFT XL&lt;br /&gt;
! SongEcho&lt;br /&gt;
! Exact prompt&lt;br /&gt;
|-&lt;br /&gt;
| rowspan=&amp;quot;3&amp;quot; | [https://tiaal.github.io/songecho-cover-metrics/mirex-audio/audio/source_original.mp3 Source song (MP3)]&lt;br /&gt;
| [https://tiaal.github.io/songecho-cover-metrics/mirex-audio/audio/ace_step_prompt_01_classic_pop_ballad.mp3 Audio (MP3)]&lt;br /&gt;
| [https://tiaal.github.io/songecho-cover-metrics/mirex-audio/audio/songecho_prompt_01_classic_pop_ballad.mp3 Audio (MP3)]&lt;br /&gt;
| '''Prompt 01 — Classic pop ballad'''&lt;br /&gt;
&amp;lt;br /&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Warm nostalgic classic pop ballad, expressive piano, acoustic guitar, soft live drums, melodic bass and elegant strings, intimate verses and emotional choruses, preserve the original vocal melody, lyrics and song structure.&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| [https://tiaal.github.io/songecho-cover-metrics/mirex-audio/audio/ace_step_prompt_02_indie_rock.mp3 Audio (MP3)]&lt;br /&gt;
| [https://tiaal.github.io/songecho-cover-metrics/mirex-audio/audio/songecho_prompt_02_indie_rock.mp3 Audio (MP3)]&lt;br /&gt;
| '''Prompt 02 — Indie rock'''&lt;br /&gt;
&amp;lt;br /&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Indie rock emotional female vocals, pop, folk, dreamy, acoustic guitar, electric guitar, strings, drums, bass, percussion&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| [https://tiaal.github.io/songecho-cover-metrics/mirex-audio/audio/ace_step_prompt_03_soul_folk_acapella.mp3 Audio (MP3)]&lt;br /&gt;
| [https://tiaal.github.io/songecho-cover-metrics/mirex-audio/audio/songecho_prompt_03_soul_folk_acapella.mp3 Audio (MP3)]&lt;br /&gt;
| '''Prompt 03 — Soul folk acapella'''&lt;br /&gt;
&amp;lt;br /&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;soul folk acapella, pop, folk, singer-songwriter, mellow&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
The examples should make the input-output relationship clear: the system receives the source song and prompt, then generates a new full-song arrangement. They are not examples of voice conversion or ordinary human-performed covers.&lt;br /&gt;
&lt;br /&gt;
== Evaluation ==&lt;br /&gt;
&lt;br /&gt;
The official evaluation will be based on expert listening. Three trained evaluators are planned. System identities will be hidden during evaluation.&lt;br /&gt;
&lt;br /&gt;
Each generated cover will be rated on a five-point scale in the following dimensions:&lt;br /&gt;
&lt;br /&gt;
# '''Melody''' - whether the main melody is preserved or plausibly adapted, without salient pitch or contour errors.&lt;br /&gt;
# '''Harmony''' - whether the chord progression and chord-melody relationship are musically coherent.&lt;br /&gt;
# '''Key Consistency''' - whether the output maintains a coherent tonal center, unless modulation is intentional.&lt;br /&gt;
# '''Style Consistency''' - whether the output follows the target style or arrangement prompt.&lt;br /&gt;
# '''Arrangement and Production''' - whether the instrumentation, texture, balance, timbre, and production support a complete musical result.&lt;br /&gt;
&lt;br /&gt;
The rating scale is:&lt;br /&gt;
&lt;br /&gt;
* '''5:''' very good, with no salient problem;&lt;br /&gt;
* '''4:''' acceptable, with only minor imperfections;&lt;br /&gt;
* '''3:''' a noticeable issue, but still musically usable;&lt;br /&gt;
* '''2:''' a severe issue affecting musical coherence or usability;&lt;br /&gt;
* '''1:''' a substantial or persistent failure.&lt;br /&gt;
&lt;br /&gt;
The official system score will be the average of the five dimension scores across test items and evaluators. Dimension-level results will also be reported.&lt;br /&gt;
&lt;br /&gt;
Evaluators may provide short time-stamped notes for important errors. Inter-rater agreement will be reported with the final results.&lt;br /&gt;
&lt;br /&gt;
== Supporting Objective Features ==&lt;br /&gt;
&lt;br /&gt;
The organizers may report diagnostic features such as large-leap ratio (LLR), pitch range (PR), pitch stability (PS), key confidence (KC), in-key note ratio (IKNR), key-change rate (KCR), loudness deviation, loudness range (LRA), and spectral contrast (SC).&lt;br /&gt;
&lt;br /&gt;
KCR refers to the rate at which the estimated local key changes across adjacent analysis windows. These features are provided only as descriptive diagnostic evidence. They will '''not''' contribute to the official score, system ranking, or tie-breaking. In particular, key-related features do not measure chord-function quality and will not be used as automatic harmony scores. PR and PS are also descriptive pitch summaries rather than validated quality scores.&lt;br /&gt;
&lt;br /&gt;
== Dataset ==&lt;br /&gt;
&lt;br /&gt;
The test set will focus on contemporary tonal popular music, such as pop, pop rock, synth-pop, dance-pop, acoustic pop, folk-pop, and pop ballads.&lt;br /&gt;
&lt;br /&gt;
To keep the evaluation consistent, styles dominated by highly complex, non-functional, modal, or tonally ambiguous harmony will not be included.&lt;br /&gt;
&lt;br /&gt;
All source songs will be provided by the task captain, who will secure the permissions needed to distribute and transform them for this evaluation. Participants may use the source songs only for this task and may not redistribute them.&lt;br /&gt;
&lt;br /&gt;
Official test materials must not be used for training or fine-tuning. Participants may generate multiple candidates and select one for submission, but the selection process must be disclosed.&lt;br /&gt;
&lt;br /&gt;
== Submission ==&lt;br /&gt;
&lt;br /&gt;
Both commercial and open-source systems are allowed, provided that participants have the right to use the system for this evaluation and disclose the system name, version, and any external services used.&lt;br /&gt;
&lt;br /&gt;
Participants will run their systems themselves and submit generated audio and metadata. Model code and checkpoints are not required.&lt;br /&gt;
&lt;br /&gt;
Each submission should include:&lt;br /&gt;
&lt;br /&gt;
* one complete audio output for every official source-prompt pair;&lt;br /&gt;
* metadata linking each output to its test item;&lt;br /&gt;
* a short system description or MIREX extended abstract; and&lt;br /&gt;
* disclosure of training data, pretrained models, candidate selection, and any manual post-processing.&lt;br /&gt;
&lt;br /&gt;
For each test item, participants must report the total number of candidate outputs generated before selecting the submitted version. For example, a submitted output selected from 20 generations should be reported as &amp;lt;code&amp;gt;1 selected from 20&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The final audio format, naming convention, and submission method will follow the instructions provided by the MIREX organizers.&lt;br /&gt;
&lt;br /&gt;
== Baseline ==&lt;br /&gt;
&lt;br /&gt;
At least one reference-conditioned music generation system will be included as a baseline. A text-only music generation system may also be reported as a comparison because it does not use the source-song reference.&lt;br /&gt;
&lt;br /&gt;
Before the test set is released, the organizers will run the reference-conditioned baseline on all test items to confirm that the inputs, prompts, audio requirements, and evaluation workflow function as intended.&lt;br /&gt;
&lt;br /&gt;
== Resources ==&lt;br /&gt;
&lt;br /&gt;
* Diagnostic framework and code: https://github.com/TiaaL/songecho-cover-metrics&lt;br /&gt;
* Associated paper: [https://arxiv.org/abs/2607.19688 A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features]&lt;br /&gt;
&lt;br /&gt;
== Timeline ==&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Stage&lt;br /&gt;
! Date&lt;br /&gt;
|-&lt;br /&gt;
| Test songs and prompts release&lt;br /&gt;
| August 5, 2026&lt;br /&gt;
|-&lt;br /&gt;
| Submission deadline&lt;br /&gt;
| September 15, 2026&lt;br /&gt;
|-&lt;br /&gt;
| Evaluation questionnaire preparation&lt;br /&gt;
| September 16-20, 2026&lt;br /&gt;
|-&lt;br /&gt;
| Subjective evaluation&lt;br /&gt;
| September 21-October 5, 2026&lt;br /&gt;
|-&lt;br /&gt;
| Result aggregation and analysis&lt;br /&gt;
| October 6-12, 2026&lt;br /&gt;
|-&lt;br /&gt;
| Final results ready&lt;br /&gt;
| By October 15, 2026&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Contact ==&lt;br /&gt;
&lt;br /&gt;
If you have any questions or suggestions about the task, please contact:&lt;br /&gt;
&lt;br /&gt;
- '''Yingxin Liang''' — Hangzhou Xiaoying Innovation Technology Co., Ltd. (Rythmix AI)&lt;br /&gt;
- Email: yingxinliang.yxl@gmail.com&lt;br /&gt;
[mailto:yingxinliang.yxl@gmail.com yingxinliang.yxl@gmail.com]&lt;br /&gt;
&lt;br /&gt;
[[Category:MIREX 2026]]&lt;/div&gt;</summary>
		<author><name>Junyan</name></author>
		
	</entry>
</feed>