Play as You Like: Timbre-Enhanced Multi-Modal Music Style Transfer

Chien-Yu Lu; Min-Xin Xue; Chia-Che Chang; Che-Rung Lee; Li Su

doi:10.1609/aaai.v33i01.33011061

Authors

Chien-Yu Lu National Tsing Hua University
Min-Xin Xue National Tsing Hua University
Chia-Che Chang National Tsing Hua University
Che-Rung Lee National Tsing Hua University
Li Su Academia Sinica

DOI:

https://doi.org/10.1609/aaai.v33i01.33011061

Abstract

Style transfer of polyphonic music recordings is a challenging task when considering the modeling of diverse, imaginative, and reasonable music pieces in the style different from their original one. To achieve this, learning stable multi-modal representations for both domain-variant (i.e., style) and domaininvariant (i.e., content) information of music in an unsupervised manner is critical. In this paper, we propose an unsupervised music style transfer method without the need for parallel data. Besides, to characterize the multi-modal distribution of music pieces, we employ the Multi-modal Unsupervised Image-to-Image Translation (MUNIT) framework in the proposed system. This allows one to generate diverse outputs from the learned latent distributions representing contents and styles. Moreover, to better capture the granularity of sound, such as the perceptual dimensions of timbre and the nuance in instrument-specific performance, cognitively plausible features including mel-frequency cepstral coefficients (MFCC), spectral difference, and spectral envelope, are combined with the widely-used mel-spectrogram into a timbreenhanced multi-channel input representation. The Relativistic average Generative Adversarial Networks (RaGAN) is also utilized to achieve fast convergence and high stability. We conduct experiments on bilateral style transfer tasks among three different genres, namely piano solo, guitar solo, and string quartet. Results demonstrate the advantages of the proposed method in music style transfer with improved sound quality and in allowing users to manipulate the output.

Play as You Like: Timbre-Enhanced Multi-Modal Music Style Transfer

Authors

DOI:

Abstract

Downloads

Published

How to Cite

Issue

Section

Information

Developed By

Subscription