A method, apparatus, device, and readable storage medium for transferring accompaniment styles.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-19
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]有鉴于此,本发明的目的在于提供一种伴奏风格迁移方法、装置、设备及可读存储介质,解决了现有技术中伴奏风格迁移效率较低的技术问题
[0044]可见,本发明通过对原始乐器信号中的单一乐器轨信号进行归一化处理,得到目标归一化乐器信号;利用风格提取模型对单一乐器轨信号进行风格向量提取,得到目标乐器风格向量;其中,风格提取模型为提取乐器轨信号风格特征的模型;基于原始乐器信号的目标归一化乐器信号和目标乐器风格向量,利用风格迁移模型进行风格迁移,得到目标乐器伴奏信号;其中,风格迁移模型为基于风格向量和归一化乐器信号改编乐器信号的模型。和当前通过人工试听原曲伴奏,分析记录下每种乐器的演奏方式,然后选择相似的乐器,按照原曲伴奏的方式进行现场演奏,最后混音工程师通过调整演奏的单一乐器轨按照原曲伴奏的比例进行混音的伴奏迁移方法相比,本发明可以基于风格向量和归一化乐器信号利用风格迁移模型智能的进行伴奏风格迁移,提高了伴奏风格迁移的效率。
Smart Images

Figure CN117373411B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, device, and readable storage medium for accompaniment style transfer. Background Technology
[0002] Song accompaniment style transfer refers to re-performing or processing a song's accompaniment to make the style of the new accompaniment similar to the original, while maintaining differences. Current methods for accompaniment style transfer are based on manual performance. This method first involves manually listening to the original accompaniment, analyzing and recording the playing style of each instrument, then selecting similar instruments and performing them live in the same style as the original accompaniment. Finally, a mixing engineer mixes the music by adjusting the individual instrument tracks to match the proportions of the original accompaniment. Therefore, current accompaniment style transfer is primarily manual, resulting in low efficiency. A method to improve the efficiency of accompaniment style transfer is needed. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a method, apparatus, device and readable storage medium for accompaniment style transfer, which solves the technical problem of low efficiency in accompaniment style transfer in the prior art.
[0004] To solve the above-mentioned technical problems, the present invention provides a method for accompaniment style transfer, comprising:
[0005] The single instrument track signal in the original instrument signal is normalized to obtain the target normalized instrument signal.
[0006] The style vector of the single instrument track signal is extracted using a style extraction model to obtain the style vector of the target instrument; wherein, the style extraction model is a model for extracting style features of the instrument track signal;
[0007] Based on the original instrument signal, the target normalized instrument signal, and the target instrument style vector, a style transfer model is used to perform style transfer to obtain the target instrument accompaniment signal; wherein, the style transfer model is a model for adapting instrument signals based on style vectors and normalized instrument signals.
[0008] Optionally, the normalization process of the single instrument track signal in the original instrument signal to obtain the target normalized instrument signal includes:
[0009] The original instrument signal is separated using a music separation model to obtain a single instrument track signal with multiple different instrument tracks;
[0010] Each of the individual instrument track signals is normalized to obtain multiple target normalized instrument signals.
[0011] Optionally, before extracting the style vector from the single instrument track signal using the style extraction model to obtain the target instrument style vector, the method further includes:
[0012] Multiple single training instrument signals are normalized to obtain multiple training normalized instrument signals.
[0013] The multiple individual training instrument signals are processed using a style generator to obtain multiple training instrument styles;
[0014] The multiple training normalized instrument signals and the multiple training instrument styles are superimposed to obtain style extraction training samples.
[0015] The machine model is trained based on the style extraction training samples to obtain the style extraction model.
[0016] Optionally, training the machine model based on the style extraction training samples to obtain the style extraction model includes:
[0017] The samples with different training normalized instrument signals superimposed on the same training instrument style in the style extraction training samples are used as positive sample groups.
[0018] The same training normalized instrument signal in the style extraction training samples is superimposed with samples of different training instrument styles to form a negative sample group.
[0019] The style extraction model is obtained by training a machine model based on a triplet loss function using the positive sample group and the negative sample group; wherein the triplet loss function is a function that minimizes the distance between the positive sample groups and maximizes the distance between the negative samples.
[0020] Optionally, before performing style transfer using a style transfer model based on the target normalized instrument signal and the target instrument style vector from the original instrument signal to obtain the target instrument accompaniment signal, the method further includes:
[0021] The style transfer training instrument signal is normalized to obtain the style transfer training normalized instrument signal.
[0022] The style extraction model is used to extract the style vector from the style transfer training instrument signal to obtain the style transfer training instrument signal vector.
[0023] The original style transfer model is trained based on the style transfer training normalized instrument signal and the style transfer training instrument signal vector to obtain the style transfer model.
[0024] Optionally, the step of training the original style transfer model based on the style transfer training normalized instrument signal and the style transfer training instrument signal vector to obtain the style transfer model includes:
[0025] The style transfer model is obtained by training the original style transfer model using the style transfer training normalized instrument signal and the style transfer training instrument signal vector, based on the norm loss minimum absolute value deviation loss function.
[0026] Optionally, the step of using a style transfer model to perform style transfer based on the target normalized instrument signal and the target instrument style vector, derived from the original instrument signal, to obtain the target instrument accompaniment signal includes:
[0027] Determine the number of the target normalized instrument signal and the number of the target instrument style vector;
[0028] When both the target normalized instrument signal and the target instrument style vector are single, the target instrument accompaniment signal is directly generated using the style transfer model.
[0029] When the target normalized instrument signal and the target instrument style vector are not both single, the style transfer model is used to generate multiple adapted instrument track signals, and the multiple adapted instrument track signals are mixed to obtain the target instrument accompaniment signal.
[0030] Optionally, after performing style transfer using a style transfer model based on the target normalized instrument signal and the target instrument style vector according to the original instrument signal to obtain the target instrument accompaniment signal, the method further includes:
[0031] The target instrument accompaniment signal is calibrated based on the musical melody attributes of the original instrument signal to obtain the target calibrated instrument signal.
[0032] Optionally, the calibration of the target instrument accompaniment signal based on the musical melody attributes of the original instrument signal to obtain the target calibrated instrument signal includes:
[0033] Using a loudness calibration instrument signal model, the loudness of the target instrument accompaniment signal is calibrated based on the calibrated loudness of the original instrument signal to obtain the target loudness calibration instrument signal; wherein, the loudness calibration instrument signal model is x. new (t) = ax(t), x new x(t) represents the target loudness calibration instrument signal, and x(t) represents the calibration loudness of the original instrument signal. M represents the number of instrument tracks.
[0034] Optionally, the normalization process of the single instrument track signal in the original instrument signal to obtain the target normalized instrument signal includes:
[0035] The single instrument track signal in the original instrument signal is normalized to obtain the target normalized instrument signal; wherein the single instrument track signal includes at least one of piano track, drum track, bass track and pipa track.
[0036] The present invention also provides an accompaniment style transfer device, comprising:
[0037] The normalized instrument signal generation module is used to normalize the single instrument track signal in the original instrument signal to obtain the target normalized instrument signal.
[0038] The style vector extraction module is used to extract the style vector of the single instrument track signal using a style extraction model to obtain the target instrument style vector; wherein, the style extraction model is a model for extracting the style features of the instrument track signal;
[0039] The style transfer module is used to perform style transfer based on the target normalized instrument signal and the target instrument style vector of the original instrument signal, using a style transfer model to obtain the target instrument accompaniment signal; wherein, the style transfer model is a model for adapting instrument signals based on style vectors and normalized instrument signals.
[0040] The present invention also provides an accompaniment style transfer device, comprising:
[0041] Memory, used to store computer programs;
[0042] A processor is used to implement the steps of the above-described accompaniment style transfer method when executing the computer program.
[0043] The present invention also provides a readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described accompaniment style transfer method.
[0044] As can be seen, this invention normalizes the single instrument track signal in the original instrument signal to obtain the target normalized instrument signal; it then uses a style extraction model to extract the style vector from the single instrument track signal to obtain the target instrument style vector; wherein, the style extraction model is a model for extracting the style features of the instrument track signal; based on the target normalized instrument signal and the target instrument style vector of the original instrument signal, a style transfer model is used to perform style transfer to obtain the target instrument accompaniment signal; wherein, the style transfer model is a model for adapting the instrument signal based on the style vector and the normalized instrument signal. Compared with the current accompaniment transfer method, which involves manually listening to the original accompaniment, analyzing and recording the playing style of each instrument, then selecting similar instruments, performing a live performance according to the original accompaniment style, and finally having a mixing engineer adjust the proportion of the single instrument track played to match the original accompaniment, this invention can intelligently perform accompaniment style transfer based on the style vector and the normalized instrument signal using a style transfer model, thus improving the efficiency of accompaniment style transfer.
[0045] In addition, the present invention also provides an accompaniment style transfer device, apparatus and readable storage medium, which also have the above-mentioned beneficial effects. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0047] Figure 1 A flowchart of an accompaniment style transfer method provided in an embodiment of the present invention;
[0048] Figure 2 A flowchart illustrating a style extraction model generation method provided in an embodiment of the present invention;
[0049] Figure 3 A flowchart illustrating a style extraction model sample generation method provided in an embodiment of the present invention;
[0050] Figure 4 A flowchart of a style transfer model generation method provided in an embodiment of the present invention;
[0051] Figure 5 A flowchart illustrating a style transfer model generation method provided in an embodiment of the present invention;
[0052] Figure 6 A flowchart of another style transfer model generation method provided in an embodiment of the present invention;
[0053] Figure 7 A flowchart illustrating a style transfer model generation method provided in an embodiment of the present invention;
[0054] Figure 8 This is a schematic diagram of the structure of an accompaniment style transfer device provided in an embodiment of the present invention;
[0055] Figure 9 This is a schematic diagram of the structure of an accompaniment style transfer device provided in an embodiment of the present invention. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0057] Please refer to Figure 1 , Figure 1 A flowchart illustrating a method for transferring accompaniment style according to an embodiment of the present invention. The method may include:
[0058] S100 normalizes the single instrument track signal in the original instrument signal to obtain the target normalized instrument signal.
[0059] This embodiment does not limit the specific original instrument signal. For example, the original instrument signal in this embodiment can be a non-single instrument track signal, that is, the original instrument signal includes multiple single instrument track signals. It is understood that when the original instrument signal is not a single instrument track signal, there are multiple single instrument track signals; or the single instrument track signal in this embodiment can also be a single instrument track signal. This embodiment does not limit the specific single instrument track signal. It is understood that when the original instrument track signal is a single instrument track signal, there is one single instrument track signal. For example, the single instrument track signal can be a piano; or the single instrument track signal can also be a drum track; or the single instrument track signal in this embodiment can also be a bass track.
[0060] It should be noted that the above-mentioned normalization processing of the single instrument track signal in the original instrument signal to obtain the target normalized instrument signal may include: normalizing the single instrument track signal in the original instrument signal to obtain the target normalized instrument signal; wherein, the single instrument track signal includes at least one of piano track, drum track, bass track, and pipa track. In this embodiment, the single instrument track signal is limited to multiple or one type, meaning that the original instrument signal can be a multi-instrument accompaniment signal or a single-instrument accompaniment signal, which improves the comprehensiveness of style transfer in the embodiments of the present invention.
[0061] S101, The style vector of a single instrument track signal is extracted using a style extraction model to obtain the style vector of the target instrument; wherein, the style extraction model is a model for extracting the style features of the instrument track signal.
[0062] The style extraction model in this embodiment is a pre-trained machine learning model capable of extracting the style vector of a single instrument track signal; that is, a model that extracts the style features of the instrument track signal. This embodiment does not limit the specific style extraction model, as long as it can extract the style vector of a single instrument track signal. For example, the style extraction model in this embodiment could be a ResUnet model (a deep learning model based on residual connections); or it could be an LSTM (Long Short-Term Memory) model; or it could be a TCN (Temporal Convolutional Network).
[0063] S102, based on the target normalized instrument signal and the target instrument style vector of the original instrument signal, style transfer is performed using a style transfer model to obtain the target instrument accompaniment signal; wherein, the style transfer model is a model for adapting the instrument signal based on the style vector and the normalized instrument signal.
[0064] The style transfer model in this embodiment is pre-trained and can adapt instrument signals based on style vectors and normalized instrument signals. Specifically, the style transfer model in this embodiment can adapt (transfer) instrument signal A to obtain instrument signal B, where instrument signal B can be any signal different from the original instrument signal. The input parameters of the style transfer model in this embodiment are the style vector and the normalized instrument signal. This embodiment can use the style transfer model to superimpose or concatenate the normalized instrument signal and the instrument signal vector corresponding to the initial instrument signal to obtain a new instrument signal.
[0065] It should be noted that the above-mentioned target normalized instrument signal and target instrument style vector based on the original instrument signal, using a style transfer model to perform style transfer to obtain the target instrument accompaniment signal, may include: determining the number of target normalized instrument signals and target instrument style vectors; when both the target normalized instrument signal and the target instrument style vector are single, directly generating the target instrument accompaniment signal using the style transfer model; when not both the target normalized instrument signal and the target instrument style vector are single, generating multiple adapted instrument track signals using the style transfer model, and mixing the multiple adapted instrument track signals to obtain the target instrument accompaniment signal. In this embodiment of the invention, when there is one target normalized instrument signal and one target instrument style vector, the target instrument accompaniment signal can be directly obtained using the style transfer model. This target accompaniment signal can be any instrument signal different from the original instrument accompaniment signal. When there are multiple target normalized instrument signals and multiple target instrument style vectors, determining each target normalized instrument signal and target instrument style corresponding to each single instrument signal, obtaining multiple target instrument accompaniment signals, and then mixing the multiple target instrument accompaniment signals to obtain the final target instrument accompaniment signal. This embodiment can ensure that a single target normalized instrument signal and target instrument style vector do not need to be mixed, depending on the number of target normalized instrument signals and target instrument style vectors, thereby improving the accuracy and efficiency of style transfer.
[0066] It should be noted that after obtaining the target instrument accompaniment signal by performing style transfer using a style transfer model based on the target normalized instrument signal and the target instrument style vector from the original instrument signal, the process may further include: calibrating the target instrument accompaniment signal based on the musical melody attributes of the original instrument signal to obtain a target calibrated instrument signal. This embodiment of the invention can calibrate the target instrument accompaniment signal based on the musical melody attributes of the original instrument signal to obtain a target calibrated instrument signal. The musical melody attributes in this embodiment can be notes, pitch, volume, loudness, etc. This embodiment improves the consistency of the musical melody attributes between the target instrument accompaniment signal and the original instrument accompaniment signal by calibrating the target instrument accompaniment signal according to the musical melody attributes, thereby enhancing the style transfer effect.
[0067] It should be further explained that the above-mentioned calibration of the target instrument accompaniment signal based on the musical melody attributes of the original instrument signal to obtain the target calibrated instrument signal may include:
[0068] Using a loudness calibration instrument signal model, the loudness of the target instrument accompaniment information is calibrated based on the calibrated loudness of the original instrument signal to obtain the target loudness calibration instrument signal; wherein, the loudness calibration instrument signal model is x. new (t) = ax(t), x newx(t) represents the target loudness calibration instrument signal, and x(t) represents the calibration loudness of the original instrument signal. M represents the number of instrument tracks. In this embodiment of the invention, to ensure that the proportions of different instruments in the adapted accompaniment remain consistent with the original accompaniment, the loudness of the target instrument accompaniment signal is calibrated.
[0069] The accompaniment style transfer provided in this invention embodiment can include: normalizing the single instrument track signal in the original instrument signal to obtain a target normalized instrument signal; extracting the style vector of the single instrument track signal using a style extraction model to obtain a target instrument style vector; wherein, the style extraction model is a model for extracting the style features of the instrument track signal; and performing style transfer using a style transfer model based on the target normalized instrument signal and the target instrument style vector of the original instrument signal to obtain the target instrument accompaniment signal; wherein, the style transfer model is a model for adapting the instrument signal based on the style vector and the normalized instrument signal. It can be seen that, compared with the current accompaniment style transfer method which first involves manually listening to the original accompaniment, analyzing and recording the playing style of each instrument, then selecting similar instruments, performing a live performance according to the original accompaniment style, and finally having a mixing engineer adjust the proportion of the single instrument track played to match the original accompaniment, this invention can intelligently perform style transfer on the accompaniment instrument signal based on the style vector and the normalized instrument signal using a pre-trained style transfer model, thus improving the efficiency of accompaniment style transfer. Furthermore, this embodiment limits the single instrument track signal to be either multiple or a single type, meaning the original instrument signal can be a multi-instrument accompaniment signal or a single-instrument accompaniment signal, thus improving the comprehensiveness of style transfer in this embodiment. Moreover, this embodiment ensures that a single target normalized instrument signal and target instrument style vector do not require mixing, based on the different numbers of target normalized instrument signals and target instrument style vectors, improving the accuracy and efficiency of style transfer. Furthermore, this embodiment improves the consistency of musical melody attributes between the target instrument accompaniment signal and the original instrument accompaniment signal by calibrating the target instrument accompaniment signal according to musical melody attributes, thereby enhancing the style transfer effect. Finally, this embodiment calibrates the loudness of the target instrument accompaniment signal to ensure that the proportions of different instruments in the adapted accompaniment remain consistent with the original accompaniment.
[0070] For a clearer understanding of this invention, please refer to the following details. Figure 2 , Figure 2 A flowchart of a style extraction model generation method provided in an embodiment of the present invention may specifically include:
[0071] S200 performs normalization processing on multiple single training instrument signals to obtain multiple normalized training instrument signals.
[0072] The structural framework diagram corresponding to this embodiment is as follows: Figure 3 As shown, Figure 3 The flowchart of a style extraction model sample generation method provided in this embodiment of the invention describes a method for style extraction model sample generation. The method involves performing style normalization on each instrument signal and then superimposing style to generate positive and negative sample groups, thereby training the original style transfer model based on the positive and negative sample groups.
[0073] This embodiment does not limit the specific method of obtaining multiple individual training instrument signals. For example, this embodiment can use a music separation model to separate the initial instrument signal to obtain multiple individual training instrument signals; or this embodiment can directly acquire multiple individual training instrument signals. This embodiment can normalize the style of the signals by performing operations such as EQ (EQ is an abbreviation for equalizer. Its basic function is to adjust the timbre by increasing or decreasing the gain of one or more frequency bands of the sound), DRC (Dynamic Compression), and panning.
[0074] S201 uses a style generator to process multiple single training instrument signals to obtain multiple training instrument styles.
[0075] This embodiment can utilize a style generator to process multiple single training instrument signals to obtain multiple training instrument styles. The style generator in this embodiment can be the pymixconsole tool (a headless multitrack mixing console in Python), which generates different types of styles by configuring different parameters, including operations such as Gain (tuner), Equalizer, Reverb, and Compressor.
[0076] S202, multiple training normalized instrument signals and multiple training instrument styles are superimposed to obtain style extraction training samples.
[0077] In order to enable the model to learn the style vector of the instrument signal better, this embodiment adopts triple loss. That is, different signals superimposed with samples of the same style can be used as positive sample groups, and the same signal superimposed with samples of different styles can be used as negative sample groups. By minimizing the distance between positive sample groups and maximizing the distance between negative sample groups, the vector extracted by the model is highly correlated with the style.
[0078] S203, the machine model is trained based on style extraction training samples to obtain the style extraction model.
[0079] This embodiment does not limit the specific machine model. The machine model in this embodiment can be a Convolutional Neural Network (CNN) model; or the machine learning model in this embodiment can also be a Recurrent Neural Network (RNN) model; or the model in this embodiment can also be a BERT model (Bidirectional Encoder Representation from Transformers, a pre-trained language representation model), etc. It should be noted that when training the model in this embodiment, the input can be a signal of fixed length, such as 1 second; the output can be a vector of fixed length, such as a 2048-bit one-dimensional vector; the model can use a classic model for a similar task, such as ResUnet (a convolutional neural network model).
[0080] It should be noted that the above-mentioned training of the machine model based on style extraction training samples to obtain the style extraction model can include: taking samples from the style extraction training samples that are different training normalized instrument signals superimposed with the same training instrument style as positive sample groups; taking samples from the style extraction training samples that are the same training normalized instrument signals superimposed with different training instrument styles as negative sample groups; and using the positive sample groups and the negative sample groups to train the machine model based on the triple loss function to obtain the style extraction model; wherein, the triple loss function is a function that minimizes the distance between positive sample groups and maximizes the distance between negative samples. It can be understood that, in order to better enable the model to learn the style vector of the instrument signal, a triple loss function is used, that is, taking samples of different signals superimposed with the same style as positive sample groups, and taking samples of the same signal superimposed with different styles as negative sample groups, by minimizing the distance between positive sample groups and maximizing the distance between negative samples, the vector extracted by the model is highly correlated with the style.
[0081] This invention involves normalizing multiple single training instrument signals to obtain multiple normalized training instrument signals; processing these signals using a style generator to obtain multiple training instrument styles; superimposing the normalized training instrument signals and the training instrument styles to obtain style extraction training samples; and training a machine model based on these training samples to obtain a style extraction model. As can be seen, this invention can train a style extraction model based on normalized instrument signals and style vectors to obtain a style extraction model that can intelligently extract style vectors from instrument signals, allowing for direct subsequent style vector extraction. Furthermore, training the style extraction model based on a triplet loss function minimizes the distance between positive sample groups and maximizes the distance between negative samples, thereby distinguishing between highly similar but dissimilar samples and improving the accuracy of subsequent style extraction.
[0082] For a clearer understanding of this invention, please refer to the following details. Figure 4 , Figure 4 A flowchart of a style transfer model generation method provided in an embodiment of the present invention may specifically include:
[0083] S400 normalizes the style transfer training instrument signal to obtain the style transfer training normalized instrument signal.
[0084] The flowchart of the style transfer model generation method corresponding to this embodiment is shown below. Figure 5 , Figure 5 This is a flowchart illustrating a style transfer model generation method provided in an embodiment of the present invention. Figure 5 The solid lines in the diagram represent the training process of the style transfer model, while the dashed lines represent the usage process. Both the training and usage processes include normalizing the style of the instrument signal and extracting the style vector of the instrument signal, which is then input into the style transfer model for training or usage to obtain a new instrument signal. During the training process, this new instrument signal can also be used for further training.
[0085] S401, use the style extraction model to extract the style vector of the style transfer training instrument signal, and obtain the style transfer training instrument signal vector.
[0086] S402, based on the style transfer training normalized instrument signal and the style transfer training instrument signal vector, the original style transfer model is trained to obtain the style transfer model.
[0087] This embodiment does not limit the specific type of the original style transfer model. For example, the original style transfer model in this embodiment may be ResUnet; or the original style transfer model in this embodiment may be TCN.
[0088] It should be noted that the above-mentioned training of the original style transfer model based on the style transfer training normalized instrument signal and the style transfer training instrument signal vector to obtain the style transfer model can include: using the style transfer training normalized instrument signal and the style transfer training instrument signal vector to train the original style transfer model based on the norm loss minimum absolute value deviation loss function to obtain the style transfer model. In this embodiment, when training the original style transfer model, the l1-loss (norm loss minimum absolute value deviation loss function) can be used to train the style transfer model.
[0089] This invention provides an embodiment that trains the original style transfer model based on style transfer training normalized instrument signals and style transfer training instrument signal vectors to obtain a style transfer model. Compared to the current method of manually adapting instruments, the style transfer model provided by this invention can perform intelligent style transfer on the original instrument signals, improving the efficiency of style transfer.
[0090] For a clearer understanding of this invention, please refer to the following details. Figure 6 , Figure 6 A flowchart of another accompaniment style transfer method provided in an embodiment of the present invention may specifically include:
[0091] The S600 uses a music separation model to separate the original instrument signals, resulting in single instrument track signals with different instrument tracks from multiple instrument tracks.
[0092] In this embodiment, when the original instrument signal is not a single instrument track signal, a music separation model is used to separate the original instrument signal, resulting in single instrument track signals with multiple different instrument tracks. This embodiment is not limited to a specific music separation model. For example, the music separation model can be a band-split recurrent neural network (BSRNN); or the music separation model can be other neural network models.
[0093] S601 performs normalization processing on each individual instrument track signal to obtain multiple target normalized instrument signals.
[0094] S602, the style extraction model is used to extract the style vector of each individual instrument track signal to obtain multiple target instrument style vectors; wherein, the style extraction model is a model for extracting the style features of the instrument track signal.
[0095] S603, based on multiple target normalized instrument signals and multiple target instrument style vectors of the original instrument signal, uses a style transfer model to perform style transfer to obtain the target instrument accompaniment signal; wherein, the style transfer model is a model for adapting instrument signals based on style vectors and normalized instrument signals.
[0096] This invention can promptly separate the original instrument signals using a music separation model, obtaining multiple individual instrument track signals. By replacing manual performance with music separation and style transfer, the efficiency of adaptation is improved while reducing the cost of adaptation.
[0097] For a clearer understanding of this invention, please refer to the following details. Figure 7 , Figure 7 A flowchart illustrating a style transfer model generation method provided in this embodiment of the invention may specifically include:
[0098] The S700 uses a music separation model to separate the accompaniment of the original song, resulting in multiple individual instrument track signals.
[0099] In this embodiment, the single track signal can be a piano track, guitar track, bass track, etc. This embodiment can use a music separation model (such as BSRNN) to separate the song accompaniment into single instrument track signals, such as a piano track, drum track, bass track, etc.
[0100] S701 uses a style extraction model to extract the style of all single instrument track signals, resulting in multiple target style vectors; where the style extraction model is a model for extracting instrument style vectors.
[0101] The training process of the style extraction model in this embodiment is as follows:
[0102] (1) Normalize the single instrument track signal to obtain normalized signal data;
[0103] (2) Extract the style type from a single instrument track signal to obtain style samples of different types;
[0104] (3) Based on the triplet loss function, the machine model is trained by positive and negative sample groups composed of signal normalized data and different styles to obtain the style extraction model.
[0105] S702 normalizes multiple single instrument track signals to obtain normalized instrument signals.
[0106] S703 generates target instrument track signals based on multiple normalized instrument signals and multiple target style vectors using a style transfer model; wherein, the style transfer model is a model obtained based on training normalized signals and training style vectors.
[0107] In this embodiment, the target instrument track signal is the mixed instrument track signal.
[0108] S704: Obtain the original instrument calibration loudness of the original song accompaniment, and use the original instrument calibration loudness to calibrate the loudness of the target instrument track signal to obtain the target accompaniment; wherein, the proportions of different instruments in the target accompaniment are consistent with those in the original accompaniment.
[0109] The following is a description of an accompaniment style transfer device provided by an embodiment of the present invention. The accompaniment style transfer device described below and the accompaniment style transfer method described above can be referred to and correspond to each other.
[0110] Please refer to the details. Figure 8 , Figure 8 A schematic diagram of a accompaniment style transfer device provided in an embodiment of the present invention may include:
[0111] The normalized instrument signal generation module 100 is used to normalize the single instrument track signal in the original instrument signal to obtain the target normalized instrument signal.
[0112] The style vector extraction module 200 is used to extract the style vector of the single instrument track signal using a style extraction model to obtain the target instrument style vector; wherein, the style extraction model is a model for extracting the style features of the instrument track signal;
[0113] The style transfer module 300 is used to perform style transfer based on the target normalized instrument signal and the target instrument style vector of the original instrument signal, using a style transfer model to obtain the target instrument accompaniment signal; wherein, the style transfer model is a model for adapting instrument signals based on style vectors and normalized instrument signals.
[0114] Furthermore, based on the above embodiments, the normalized musical instrument signal generation module 100 may include:
[0115] The separation unit is used to separate the original instrument signal using a music separation model to obtain the single instrument track signal with multiple different instrument tracks;
[0116] The normalization unit is used to normalize each of the individual instrument track signals to obtain multiple target normalized instrument signals.
[0117] Furthermore, based on any of the above embodiments, the above-mentioned accompaniment style transfer device may further include:
[0118] The training data normalization processing module is used to normalize multiple single training instrument signals to obtain multiple training normalized instrument signals.
[0119] The training instrument style generation unit is used to process the multiple single training instrument signals using a style generator to obtain multiple training instrument styles.
[0120] The style extraction training sample generation unit is used to superimpose the multiple training normalized instrument signals and the multiple training instrument styles to obtain style extraction training samples.
[0121] The style extraction model training unit is used to train the machine model based on the style extraction training samples to obtain the style extraction model.
[0122] Furthermore, based on the above embodiments, the style extraction model training unit may include:
[0123] A positive sample group generation subunit is used to superimpose samples of the same training instrument style on different training normalized instrument signals in the style extraction training samples as a positive sample group.
[0124] The negative sample group generation sub-unit is used to superimpose samples of different training instrument styles on the same training normalized instrument signal in the style extraction training samples as a negative sample group.
[0125] The model training subunit based on the triplet loss function is used to train the machine model based on the triplet loss function using the positive sample group and the negative sample group to obtain the style extraction model; wherein, the triplet loss function is a function that minimizes the distance of the positive sample group and maximizes the distance of the negative sample group.
[0126] Furthermore, based on any of the above embodiments, the above-mentioned accompaniment style transfer device may further include:
[0127] The style transfer training normalized instrument signal generation unit is used to normalize the style transfer training instrument signal to obtain the style transfer training normalized instrument signal.
[0128] The style transfer training instrument signal vector generation unit is used to extract the style vector from the style transfer training instrument signal using the style extraction model, and obtain the style transfer training instrument signal vector.
[0129] The style transfer model generation unit is used to train the original style transfer model based on the style transfer training normalized instrument signal and the style transfer training instrument signal vector to obtain the style transfer model.
[0130] Furthermore, based on the above embodiments, the style transfer model generation unit may include:
[0131] A style transfer model generation subunit is used to train the original style transfer model based on the norm loss minimum absolute value deviation loss function using the style transfer training normalized instrument signal and the style transfer training instrument signal vector, so as to obtain the style transfer model.
[0132] Furthermore, based on any of the above embodiments, the style transfer module 300 may include:
[0133] The number determination unit is used to determine the number of the target normalized instrument signal and the target instrument style vector;
[0134] A direct generation unit is used to directly generate the target instrument accompaniment signal using the style transfer model when both the target normalized instrument signal and the target instrument style vector are single.
[0135] The mixing unit is used to generate multiple adapted instrument track signals using the style transfer model when the target normalized instrument signal and the target instrument style vector are not both single, and to mix the multiple adapted instrument track signals to obtain the target instrument accompaniment signal.
[0136] Furthermore, based on any of the above embodiments, the above-mentioned accompaniment style transfer device may further include:
[0137] The calibration module is used to calibrate the target instrument accompaniment signal based on the musical melody attributes of the original instrument signal to obtain the target calibrated instrument signal.
[0138] Furthermore, based on the above embodiments, the calibration module may include:
[0139] A loudness calibration unit is used to perform loudness calibration on the target instrument accompaniment signal based on the calibrated loudness of the original instrument signal using a loudness calibration instrument signal model, thereby obtaining a target loudness calibration instrument signal; wherein, the loudness calibration instrument signal model is x new (t) = ax(t), x new x(t) represents the target loudness calibration instrument signal, and x(t) represents the calibration loudness of the original instrument signal.
[0140] M represents the number of instrument tracks.
[0141] Furthermore, based on any of the above embodiments, the normalized musical instrument signal generation module 100 may include:
[0142] A single normalized instrument signal is generated to normalize the single instrument track signal in the original instrument signal to obtain the target normalized instrument signal; wherein the single instrument track signal includes at least one of piano track, drum track, bass track and pipa track.
[0143] It should be noted that the order of the modules and units in the aforementioned accompaniment style transfer device can be changed without affecting the logic.
[0144] An accompaniment style transfer device provided in this embodiment of the invention may include: a normalized instrument signal generation module 100, used to normalize a single instrument track signal in the original instrument signal to obtain a target normalized instrument signal; a style vector extraction module 200, used to extract a style vector from the single instrument track signal using a style extraction model to obtain a target instrument style vector; wherein the style extraction model is a model for extracting style features of the instrument track signal; and a style transfer module 300, used to perform style transfer based on the target normalized instrument signal and the target instrument style vector of the original instrument signal, using a style transfer model to obtain a target instrument accompaniment signal; wherein the style transfer model is a model for adapting instrument signals based on style vectors and normalized instrument signals. As can be seen, compared with the current method of accompaniment style transfer, which first involves manually listening to the original accompaniment, analyzing and recording the playing style of each instrument, then selecting similar instruments, performing them live according to the original accompaniment, and finally having the mixing engineer adjust the single instrument track to match the proportion of the original accompaniment, this invention can intelligently perform style transfer on the accompaniment instrument signal based on style vectors and normalized instrument signals using a pre-trained style transfer model, thus improving the efficiency of accompaniment style transfer. Furthermore, this embodiment limits the single instrument track signal to be either multiple or a single type, meaning the original instrument signal can be a multi-instrument accompaniment signal or a single-instrument accompaniment signal, improving the comprehensiveness of style transfer in this embodiment. Moreover, this embodiment ensures that a single target normalized instrument signal and target instrument style vector do not require mixing, improving the accuracy and efficiency of style transfer, by calibrating the target instrument accompaniment signal according to musical melody attributes. Furthermore, this embodiment improves the consistency of musical melody attributes between the target instrument accompaniment signal and the original instrument accompaniment signal, enhancing the style transfer effect, by calibrating the target instrument accompaniment signal according to musical melody attributes. Furthermore, this embodiment calibrates the loudness of the target instrument accompaniment signal to maintain consistency between the proportions of different instruments in the adapted accompaniment and the original accompaniment. Finally, this embodiment can train a style extraction model based on the normalized instrument signal and style vector to obtain a style extraction model that can intelligently extract style vectors from instrument signals, allowing for direct subsequent extraction of style vectors using this model. Furthermore, by training the style extraction model based on the triplet loss function, the distance between positive samples is minimized and the distance between negative samples is maximized, thereby distinguishing between extremely similar samples of different classes and improving the accuracy of subsequent style extraction.
[0145] The following is an introduction to an accompaniment style transfer device provided by an embodiment of the present invention. The accompaniment style transfer device described below can be referred to in correspondence with the accompaniment style transfer method described above.
[0146] Please refer to Figure 9 , Figure 9 A schematic diagram of a accompaniment style transfer device provided in an embodiment of the present invention may include:
[0147] Memory 10 is used to store computer programs;
[0148] Processor 20 is used to execute computer programs to implement the above-described accompaniment style transfer method.
[0149] The memory 10, processor 20, and communication interface 30 all communicate with each other through the communication bus 40.
[0150] In this embodiment of the invention, the memory 10 is used to store one or more programs. The programs may include program code, which includes computer operation instructions. In this embodiment of the invention, the memory 10 may store programs for implementing the following functions:
[0151] The single instrument track signal in the original instrument signal is normalized to obtain the target normalized instrument signal.
[0152] A style extraction model is used to extract the style vector of a single instrument track signal to obtain the style vector of the target instrument; where the style extraction model is a model for extracting the style features of the instrument track signal.
[0153] Based on the original instrument signal, the target normalized instrument signal, and the target instrument style vector, a style transfer model is used to perform style transfer to obtain the target instrument accompaniment signal; wherein, the style transfer model is a model for adapting the instrument signal based on the style vector and the normalized instrument signal.
[0154] In one possible implementation, the memory 10 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; and the data storage area may store data created during use.
[0155] Furthermore, memory 10 may include read-only memory and random access memory, providing instructions and data to the processor. A portion of the memory may also include NVRAM. The memory stores operating systems and operating instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operating instructions may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic tasks and handling hardware-based tasks.
[0156] Processor 20 can be a central processing unit (CPU), an application-specific integrated circuit, a digital signal processor, a field-programmable gate array, or other programmable logic device. Processor 20 can be a microprocessor or any conventional processor. Processor 20 can call programs stored in memory 10.
[0157] The communication interface 30 can be an interface for the communication module, used to connect with other devices or systems.
[0158] Of course, it should be noted that, Figure 9 The structure shown does not constitute a limitation on the accompaniment style transfer device in the embodiments of the present invention. In practical applications, the accompaniment style transfer device may include more than Figure 9 More or fewer components as shown, or combinations of certain components.
[0159] The following describes the readable storage medium provided in the embodiments of the present invention. The computer-readable storage medium described below can be referred to in correspondence with the accompaniment style transfer method described above.
[0160] The present invention also provides a readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described accompaniment style transfer method.
[0161] The readable storage medium may include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0162] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0163] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0164] Finally, it should be noted that in this document, relationships such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0165] The above provides a detailed description of the accompaniment style transfer method, apparatus, device, and readable storage medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for transferring accompaniment style, characterized in that, include: The single instrument track signal in the original instrument signal is normalized to obtain the target normalized instrument signal. The style vector of the single instrument track signal is extracted using a style extraction model to obtain the style vector of the target instrument; wherein, the style extraction model is a model for extracting style features of the instrument track signal; Based on the original instrument signal, the target normalized instrument signal, and the target instrument style vector, a style transfer model is used to perform style transfer to obtain the target instrument accompaniment signal; wherein, the style transfer model is a model for adapting instrument signals based on style vectors and normalized instrument signals; Specifically, based on the original instrument signal, the target normalized instrument signal, and the target instrument style vector, a style transfer model is used to perform style transfer to obtain the target instrument accompaniment signal, including: Determine the number of the target normalized instrument signal and the number of the target instrument style vector; When both the target normalized instrument signal and the target instrument style vector are single, the target instrument accompaniment signal is directly generated using the style transfer model. When the target normalized instrument signal and the target instrument style vector are not both single, the style transfer model is used to generate multiple adapted instrument track signals, and the multiple adapted instrument track signals are mixed to obtain the target instrument accompaniment signal. Before performing style transfer using a style transfer model based on the target normalized instrument signal and the target instrument style vector, to obtain the target instrument accompaniment signal, the process further includes: The style transfer training instrument signal is normalized to obtain the style transfer training normalized instrument signal. The style extraction model is used to extract style vectors from the style transfer training instrument signals. Extract the style transfer training instrument signal vector; The original style transfer model is trained based on the style transfer training normalized instrument signal and the style transfer training instrument signal vector to obtain the style transfer model.
2. The accompaniment style transfer method according to claim 1, characterized in that, The process of normalizing the single instrument track signal in the original instrument signal to obtain the target normalized instrument signal includes: The original instrument signal is separated using a music separation model to obtain a single instrument track signal with multiple different instrument tracks; Each of the individual instrument track signals is normalized to obtain multiple target normalized instrument signals.
3. The accompaniment style transfer method according to claim 1, characterized in that, Before extracting the style vector from the single instrument track signal using the style extraction model to obtain the target instrument style vector, the method further includes: Multiple single training instrument signals are normalized to obtain multiple training normalized instrument signals. The multiple individual training instrument signals are processed using a style generator to obtain multiple training instrument styles; The multiple training normalized instrument signals and the multiple training instrument styles are superimposed to obtain style extraction training samples. The machine model is trained based on the style extraction training samples to obtain the style extraction model.
4. The accompaniment style transfer method according to claim 3, characterized in that, The step of training the machine model based on the style extraction training samples to obtain the style extraction model includes: The samples with different training normalized instrument signals superimposed on the same training instrument style in the style extraction training samples are used as positive sample groups. The same training normalized instrument signal in the style extraction training samples is superimposed with samples of different training instrument styles to form a negative sample group. The style extraction model is obtained by training a machine model based on a triplet loss function using the positive sample group and the negative sample group; wherein the triplet loss function is a function that minimizes the distance between the positive sample groups and maximizes the distance between the negative samples.
5. The accompaniment style transfer method according to claim 1, characterized in that, The process of training the original style transfer model based on the style transfer training normalized instrument signal and the style transfer training instrument signal vector to obtain the style transfer model includes: The style transfer model is obtained by training the original style transfer model using the style transfer training normalized instrument signal and the style transfer training instrument signal vector, based on the norm loss minimum absolute value deviation loss function.
6. The accompaniment style transfer method according to claim 1, characterized in that, After obtaining the target instrument accompaniment signal by performing style transfer using a style transfer model based on the target normalized instrument signal and the target instrument style vector, the process further includes: The target instrument accompaniment signal is calibrated based on the musical melody attributes of the original instrument signal to obtain the target calibrated instrument signal.
7. The accompaniment style transfer method according to claim 6, characterized in that, The calibration of the target instrument accompaniment signal based on the musical melody attributes of the original instrument signal to obtain the target calibrated instrument signal includes: Using a loudness calibration instrument signal model, the loudness of the target instrument accompaniment information is calibrated based on the calibrated loudness of the original instrument signal to obtain the target loudness calibration instrument signal; wherein, the loudness calibration instrument signal model is... , The instrument signal is calibrated for the target loudness. To calibrate the loudness of the original instrument signal. M represents the number of instrument tracks.
8. The accompaniment style transfer method according to claim 1, characterized in that, The process of normalizing the single instrument track signal in the original instrument signal to obtain the target normalized instrument signal includes: The single instrument track signal in the original instrument signal is normalized to obtain the target normalized instrument signal; wherein the single instrument track signal includes at least one of piano track, drum track, bass track and pipa track.
9. A device for transferring accompaniment style, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the accompaniment style transfer method as described in any one of claims 1 to 8 when executing the computer program.
10. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the accompaniment style transfer method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Music style migration method and device, model training method and device and storage medium
CN112216257A
Music style migration method and device, computer equipment and storage medium
CN116469359A