Music generation method and device, storage medium and electronic equipment
Through the method of determining and highlighting target music components in the music generation model, the problem of poor component isolation in AI-generated music is solved, and the clear expression of music components and the improvement of playback effect is achieved.
Patent Information
- Application Number
- CN202510339669.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
AI Technical Summary
In the existing AI-generating music, the isolation between vocals or instruments and other parts is poor, resulting in the performance of various components in the music being mixed, unable to highlight the focus of the music and affecting the listening experience.
By introducing encoder and decoder into the music generation model, the target music components to be prominent in the original audio are determined, and enhanced feature representations are generated through feature extraction networks and feature fusion, ensuring that the target music components are prominent in the generated target audio.
It is realized that when generating music, at least one music component is clearly expressed, which improves the playback effect of the music and the user's listening experience.
Smart Images

Figure CN120183364A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a music generation method, apparatus, storage medium, and electronic device. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, especially the progress of deep learning technology, it has made it possible to generate artificial intelligence music (AI-Generated Music). AI-generated music uses a deep learning model to learn the structure, style, and characteristics of music from a large amount of music data and generate creative and artistic music works. However, in the music generated in this way, the isolation between the human voice or musical instruments and other parts is poor, and the expressions of various components in the music are mixed with each other, making it impossible to highlight the key points of the music. For example, it is difficult to clearly hear the lyrics or the specific performance of musical instruments, which affects the listening experience. Summary of the Invention
[0003] In view of this, the purpose of this application is to provide a music generation method, apparatus, storage medium, and electronic device, which can select target music components in the original audio for highlighting, so as to ensure that at least one music component in the generated target audio is clearly expressed and improve the playback effect of the generated music.
[0004] An embodiment of this application provides a music generation method, which is implemented based on a music generation model. The method includes:
[0005] Determine at least one target music component to be highlighted in the original audio; wherein, the at least one target music component forms the original audio according to the art form of expression;
[0006] The music generation model includes an encoder and a decoder. Based on the original audio and the encoder, determine at least one enhanced feature representation of the original audio; wherein, the enhanced feature representation refers to the feature representation in which the target music component is enhanced compared to the original audio;
[0007] Input the at least one enhanced feature representation into the decoder to generate a target audio with the target music component highlighted.
[0008] Furthermore, the music generation model further includes at least one feature extraction network corresponding to the at least one target music component one by one;
[0009] The determining of the enhanced feature representation of the original audio based on the original audio and the encoder includes:
[0010] Input the original audio into the encoder to obtain the overall feature representation of the original audio;
[0011] Based on the original audio and at least one feature extraction network, obtain at least one target feature representation corresponding to the target music component;
[0012] Perform feature fusion on the overall feature representation and the at least one target feature representation to obtain the enhanced feature representation.
[0013] Further, training the music generation model includes:
[0014] Input the sample audio into the encoder to obtain the sample overall feature representation of the sample audio;
[0015] Input the sample audio into at least one feature extraction network respectively to obtain at least one sample partial feature representation corresponding to each target music component in the original audio;
[0016] Perform feature processing on the sample overall feature representation and the at least one sample partial feature representation to obtain the sample enhanced feature representation;
[0017] Input the sample enhanced feature representation into the decoder to obtain the enhanced audio;
[0018] Determine a first loss function according to the sample enhanced feature representation, the enhanced audio and the sample audio, and iteratively train the music generation model based on the first loss function until the preset training completion condition is satisfied, then obtain the trained music generation model.
[0019] Further, the determining of the first loss function according to the sample enhanced feature representation, the enhanced audio and the sample audio includes:
[0020] Determine a fusion loss function according to the audio parameters of the sample enhanced feature representation and the sample audio; wherein, the fusion loss function is used to measure the performance effect of the sample enhanced feature representation on the target music component;
[0021] Determine a reconstruction loss function according to the enhanced audio and the sample audio; wherein, the reconstruction loss function is used to measure the difference between the audio data reconstructed by the decoder and the sample audio;
[0022] Determine a generative adversarial loss function according to the enhanced audio and the sample audio; wherein, the music generation model and the discriminator are trained against each other during the training process;
[0023] Based on the fusion loss function, the reconstruction loss function and the generative adversarial loss function, determine the first loss function.
[0024] Further, training the music generation model further includes:
[0025] Input the sample audio into the encoder, and perform vector quantization processing on the output result to obtain the overall sample feature representation of the sample audio;
[0026] Input the sample audio into at least one feature extraction network respectively to obtain at least one sample partial feature representation corresponding to each target music component in the original audio;
[0027] Determine a second loss function according to at least one sample partial feature representation and the overall sample feature representation, and perform iterative training on the music generation model based on the second loss function until the preset training completion condition is met, and obtain the trained music generation model.
[0028] Further, inputting the sample audio into at least one feature extraction network respectively to obtain at least one sample partial feature representation corresponding to each target music component in the original audio includes:
[0029] When the target music component is a vocal type component, input the sample audio into a text conversion network corresponding to the voice type component to obtain a sample text feature representation corresponding to each vocal type component in the sample audio;
[0030] When the target music component is an instrument type component, input the sample audio into a music feature extraction network corresponding to the instrument type component to obtain a sample audio feature representation corresponding to each instrument type component in the sample audio.
[0031] Further, when the target music component includes at least two vocal type components and / or at least two instrument type components, inputting the sample audio into at least one feature extraction network respectively to obtain at least one sample partial feature representation corresponding to each target music component in the original audio includes:
[0032] Obtain independent audio tracks corresponding to each vocal type component and / or each instrument type component;
[0033] Input the independent audio track corresponding to each target music component into the corresponding type of feature extraction network to obtain the sample partial feature representation corresponding to each target music component in the sample audio.
[0034] An embodiment of the present application also provides a music generation device. The music generation device is implemented based on a music generation model. The music generation model includes an encoder and a decoder. The music generation device includes:
[0035] A music component determination module, configured to determine at least one target music component to be highlighted in the original audio; wherein, at least one target music component forms the original audio according to an artistic expression form;
[0036] A feature determination module, configured to determine at least one enhanced feature representation of the original audio based on the original audio and an encoder; wherein, the enhanced feature representation refers to a feature representation in which the target music component is enhanced compared to the original audio.
[0037] A generation module, configured to input the at least one enhanced feature representation into a decoder to generate a target audio with the target music component highlighted.
[0038] An embodiment of the present application further provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the music generation method as described above are executed.
[0039] An embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the music generation method as described above are executed.
[0040] A music generation method, device, storage medium, and electronic device provided by an embodiment of the present application are implemented based on a music generation model. The music generation model includes an encoder and a decoder. Determine at least one target music component to be highlighted in the original audio; based on the original audio and the encoder, determine at least one enhanced feature representation of the original audio, where the enhanced feature representation refers to a feature representation in which the target music component is enhanced compared to the original audio; in this way, during the music generation process, the target music component in the target audio generated based on the enhanced feature representation is highlighted, so as to ensure that the various music components in the target audio are not easily mixed with each other, at least one music component can be clearly expressed, and the playback effect of the generated music is improved.
[0041] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specific preferred embodiments are given in conjunction with the accompanying drawings and described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 Shows a flowchart of a music generation method provided by an embodiment of the present application;
[0044] Figure 2 FIG. 1 shows one of the schematic structural diagrams of a music generation model provided by an embodiment of the present application;
[0045] Figure 3 FIG. 2 shows one of the schematic diagrams of the training process of a music generation model provided by an embodiment of the present application;
[0046] Figure 4 FIG. 3 shows another schematic diagram of the training process of a music generation model provided by an embodiment of the present application;
[0047] Figure 5 FIG. 4 shows the schematic structural diagram of a music generation device provided by an embodiment of the present application;
[0048] Figure 6 FIG. 5 shows the schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application to be protected, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, every other embodiment obtained by those skilled in the art without creative efforts belongs to the scope of protection of the present application.
[0050] Through research, it has been found that in recent years, with the rapid development of artificial intelligence technology, especially the progress of deep learning technology, it has become possible to generate music using artificial intelligence (AI-Generated Music). AI-generated music uses a deep learning model to learn the structure, style, and features of music from a large amount of music data and generate creative and artistic music works. However, in the music generated in this way, the isolation between vocals or musical instruments and other parts is poor, and the performance of each component within the music is prone to be mixed with each other, making it impossible to highlight the key points of the music. For example, it is difficult to clearly hear the lyrics or the specific performance of musical instruments, which affects the listening experience.
[0051] Based on this, the embodiments of the present application provide a music generation method to ensure that at least one music component in the generated target audio is clearly expressed and improve the playback effect of the generated music.
[0052] Please refer to Figure 1 , Figure 1The flowchart of a music generation method provided by an embodiment of the present application. This music generation method is implemented based on a music generation model, as Figure 1 shown in
[0053] S101. Determine at least one target music component to be highlighted in the original audio.
[0054] Among them, the at least one target music component forms the original audio according to the art form of expression; the music component may include a vocal type component and an instrument type component. In different music art forms of expression, the original audio may be formed by the cooperation of multiple vocal type components (pure vocal chorus), or may be formed by the cooperation of multiple instrument type components (pure instrument ensemble), or may be formed by the cooperation of at least one vocal type component and at least one instrument type component (song singing).
[0055] More specifically, the vocal type components in different art forms of expression may include different roles in a duet, different voices in a chorus, as well as the lead singer, harmony, rap, and recitation distinguished according to the role, etc.; the instrument type components in different art forms of expression may include melodic instruments, harmonic instruments, rhythmic instruments, and decorative sound effects; specific instruments may include guitars, pianos, basses, etc.
[0056] Here, the original audio may be audio data in an actual performance obtained by means of copying, recording, etc., or may be virtual audio generated by other models.
[0057] The manner of determining at least one target music component to be highlighted in the original audio may include manual setting, automatic analysis according to music generation requirements, and automatic determination by analyzing the original audio through music processing software. Corresponding to the above examples, the target music component may be a vocal, the sound of a certain instrument (such as a guitar sound), a certain singer in a chorus, and a certain voice in a chorus, etc.
[0058] S102. Based on the original audio and an encoder, determine at least one enhanced feature representation of the original audio.
[0059] Among them, the music generation model includes an encoder and a decoder, and the music model is a pre-trained model. The enhanced feature representation refers to the feature representation in which the target music component is enhanced compared with the original audio, and the enhanced feature representation corresponds one-to-one with the target music component. The enhanced feature representation contains more feature information related to the target music component, so that the target music component is enhanced in the enhanced feature representation compared with that in the original audio.
[0060] In a possible implementation manner, an enhanced feature representation can be obtained by co-training a music generation model based on the overall feature representation of the sample audio and the partial feature representation corresponding to the target music component in the sample audio during the training process. The feature information corresponding to the target music component in the enhanced feature representation obtained by the encoder trained based on this method is more prominent.
[0061] Specific enhancement methods include feature fusion and parameter learning. For feature fusion, it means obtaining an enhanced feature representation by fusing the overall feature representation of the original audio and the partial feature representation corresponding to the target music component, so that the feature information corresponding to the target music component in the enhanced feature representation is more prominent; for parameter learning, it means that in addition to extracting the overall feature representation of the audio, the encoder learns the knowledge of extracting the partial feature representation corresponding to the target music component during the training process, so that the feature information corresponding to the target music component in the enhanced feature representation output by the trained encoder is more prominent.
[0062] S103. Input the at least one enhanced feature representation into a decoder to generate a target audio with the target music component being prominent.
[0063] In this step, the decoder in the music generation model decodes the enhanced feature representation, and the decoded result is the target audio.
[0064] Because the feature information corresponding to the target music component in the enhanced feature representation is more prominent, the target music component in the generated target audio is also correspondingly prominent. In this way, different music components in the original audio can be highlighted according to needs, ensuring that the various music components in the target audio are not easily mixed with each other, at least one music component is clearly expressed, improving the playback effect of the generated music and enhancing the user's listening experience.
[0065] Next, the model structure and training process of the music generation model in the embodiments of the present application will be specifically introduced.
[0066] In the first possible implementation manner, please refer to Figure 2 , Figure 2 which is one of the schematic structural diagrams of a music generation model provided by the embodiments of the present application; the music generation model includes: an encoder, a decoder, and at least one feature extraction network corresponding to at least one target music component one by one. Exemplarily, Figure 2 the music generation model in
[0067] includes one feature extraction network corresponding to one target music component.
[0068] Step a11. Input the original audio into the encoder to obtain the overall feature representation of the original audio.
[0069] In this step, the encoder extracts the feature information of the whole original audio to obtain the overall feature representation of the original audio.
[0070] Step a12: Based on the original audio and at least one feature extraction network, obtain at least one target feature representation corresponding to the target music component.
[0071] In this step, the original audio can be directly input into the feature extraction network corresponding to each target music component respectively to extract the feature information of the target music component in the original audio, and obtain the partial feature representation corresponding to each target music component output by each feature extraction network. Or, the original audio can also be processed first to extract the track of the target music component, and then only the track of the target music component is input into the feature extraction network corresponding to the target music component to extract the feature information of the target music component in the track, so as to improve the accuracy of the partial feature representation corresponding to the target music component through the pre-separation step.
[0072] Exemplarily, when the target music component is a vocal type component, the encoder can be used to extract the music features of the whole original audio file, and the HuBERT (Hidden-Unit BERT, a self-supervised learning method for speech representation learning) model can be used as the feature extraction network to extract the text features of the vocals.
[0073] It should be noted that the target music component can be one or more; when there are multiple target components, the original audio can be respectively input into the feature extraction network corresponding to each target music component to obtain the target feature representation corresponding to each target music component output by each feature extraction network.
[0074] Step a13: Perform feature fusion on the overall feature representation and the at least one target feature representation to obtain the enhanced feature representation.
[0075] Here, the feature fusion methods include direct addition or concatenation. Exemplarily, the overall feature representation A and the partial feature representation B corresponding to the vocal type component are both 10*15 matrices; direct addition means directly adding the corresponding bits of the A matrix and the B matrix; concatenation means concatenating the A matrix and the B matrix into a 15*20 matrix.
[0076] After fusion, it is also necessary to perform a vector quantization (VQ) operation on the fusion result to obtain a discrete codebook as an enhanced feature representation. Vector quantization maps the continuous model output to a discrete codebook, achieving data compression and feature discretization, and thus efficiently modeling complex data through the "prototype vectors" of the codebook. In this way, the subsequent decoder, as a neural network (such as a convolutional network, Transformer, or fully connected network), can reconstruct the discrete codebook into a signal in the original data domain, such as the target audio waveform in the embodiments of this application.
[0077] Please refer to Figure 3 , Figure 3 which is one of the schematic diagrams of the training process of a music generation model provided by the embodiments of this application; as Figure 3 shown in
[0078] Step a21: Input the sample audio into the encoder to obtain the sample overall feature representation of the sample audio.
[0079] Step a22: Input the sample audio into at least one feature extraction network respectively to obtain at least one sample partial feature representation corresponding to each target music component in the original audio.
[0080] Step a23: Perform feature processing on the sample overall feature representation and at least one sample partial feature representation to obtain a sample enhanced feature representation. Among them, feature processing includes feature fusion of the sample overall feature representation and at least one sample partial feature representation and vector quantization of the fusion result hs.
[0081] Step a24: Input the sample enhanced feature representation into the decoder to obtain enhanced audio.
[0082] Step a25: Determine a first loss function according to the sample enhanced feature representation, the enhanced audio, and the sample audio, and iteratively train the music generation model based on the first loss function until the preset training completion condition is met, and then obtain the trained music generation model.
[0083] Among them, the preset training completion condition includes that the loss value of the first loss function is less than a preset threshold or the number of iterative training times reaches a preset number, etc.
[0084] Here, the first loss function can characterize the quantization difference between the sample enhanced feature representation and the sample audio, as well as the quantization difference between the enhanced audio and the sample audio, so as to guide the parameter optimization of the music generation model in iterative training, making the generated sample enhanced feature representation and enhanced audio be able to more clearly highlight the target music components.
[0085] More specifically, the first loss function can be determined in the following manner:
[0086] Step a31: Determine a fusion loss function according to the sample enhanced feature representation and the audio parameters of the sample audio.
[0087] Among them, the fusion loss function is used to measure the performance effect of the sample enhanced feature representation on the target music component. Here, the fusion loss function can be a CTC loss function (Connectionist temporal classification loss). The CTC loss function can be used to determine the loss between two unaligned results. In the embodiments of the present application, the CTC loss is used to calculate the loss between the sample enhanced feature representation obtained after fusion and the audio parameters of the sample audio. This loss can reflect whether the fusion result can completely and clearly reflect the target music component. The specific form of the fusion loss function can refer to the prior art, and the present application does not make any restrictions here.
[0088] Exemplarily, when the target music component is a vocal type component, the audio parameters of the sample audio refer to the lyric text of the sample audio. The CTC loss is used to calculate the loss between the fusion result hs and the lyric text text. This loss can reflect whether the fusion result can completely and clearly reflect the vocal part.
[0089] Step a32: Determine a reconstruction loss function according to the enhanced audio and the sample audio.
[0090] Among them, the reconstruction loss function is used to measure the difference between the audio data reconstructed by the decoder and the sample audio. The reconstruction loss function measures the difference between the data reconstructed by the decoder and the original input data, aiming to ensure that after the input data is compressed by the encoder, the decoder can restore the original data as much as possible. The specific form of the reconstruction loss function can refer to the prior art, and the present application does not make any restrictions here.
[0091] Step a33: Determine a generative adversarial loss function according to the enhanced audio and the sample audio.
[0092] Among them, during the training process, the music generation model serves as a generator and forms a GAN network with a discriminator, and the two compete with each other and are alternately trained. The training objective of the music generation model is to minimize the probability that the generated data is recognized by the discriminator; the training objective of the discriminator is to maximize the ability to distinguish between real data and generated data. The specific form of the generative adversarial loss function can refer to the prior art, and the present application does not make any restrictions here.
[0093] Step a34: Determine the first loss function based on the fusion loss function, the reconstruction loss function, and the generative adversarial loss function.
[0094] In this step, the first loss function can be determined by weighted summation, that is, based on the fusion loss function, the reconstruction loss function, the generative adversarial loss function, and the weights corresponding to each loss function, determine the first loss function. The weights corresponding to different loss functions are pre-set parameters, and the magnitude of the weights is related to the information contained in the generated audio. If the weight corresponding to the CTC loss function increases while the weights of the other loss functions remain unchanged, the information related to the target music component (such as semantic information) will be more prominent, and relatively speaking, the music quality will also be affected to a certain extent. On the contrary, if the weight of the CTC loss function remains unchanged while the weights of the reconstruction loss function and the GAN loss function increase, the expression of semantic information will be affected to a certain extent, but the music quality will increase.
[0095] In the second possible implementation manner, the music generation model includes: an encoder and a decoder.
[0096] Then, in specific implementation, step S102 may include: inputting the original audio into the encoder to obtain the enhanced feature representation of the original audio.
[0097] In this step, the encoder extracts the feature information of the original audio, and then the enhanced feature representation of the original audio can be obtained.
[0098] Please refer to Figure 4 , Figure 4 which is the second schematic diagram of the training process of a music generation model provided by an embodiment of the present application; as Figure 4 shown in
[0099] Step b1: Input the sample audio into the encoder, and perform vector quantization processing on the output result to obtain the sample overall feature representation of the sample audio.
[0100] In this step, the sample audio can be input into the encoder for feature information extraction, and vector quantization (VQ) operation is performed on the output result of the encoder to obtain the sample overall feature representation hs of the sample audio.
[0101] Step b2: Input the sample audio into at least one feature extraction network respectively to obtain at least one sample partial feature representation corresponding to each target music component in the original audio.
[0102] In this step, the sample audio can be directly input into at least one feature extraction network respectively, and feature information extraction is performed on the target music components corresponding to the sample audio, so as to obtain at least one sample partial feature representation corresponding to each target music component output by each feature extraction network.
[0103] Step b3: Determine a second loss function according to at least one sample partial feature representation and the sample overall feature representation, and iteratively train the music generation model based on the second loss function until the trained music generation model is obtained when a preset training completion condition is satisfied.
[0104] Among them, the preset training completion condition includes that the loss value of the loss function is less than a preset threshold or the number of iterative training times reaches a preset number, etc.
[0105] To facilitate the calculation of the second loss function, it is first necessary to align at least one sample partial feature representation and the sample overall feature representation. In one example, a neural network model can be used for feature representation alignment. Specifically, the sample overall feature representation hs is input into the neural network model, and the neural network model can process the sample overall feature representation according to the network parameters of the pre-set feature extraction network, so that the processed sample overall feature representation hs' can be aligned with the sample partial feature representation output by the feature extraction network.
[0106] Here, the second loss function can characterize the quantization difference between the sample partial feature representation and the sample overall feature representation, so as to guide the parameter optimization of the music generation model in iterative training, so that the encoder learns the knowledge of at least one feature extraction network, and the generated sample overall feature representation can more clearly highlight the target music components.
[0107] In step b3, at least one sample partial feature representation and the sample overall feature are input into the loss calculation module to obtain a second loss function for measuring the difference between the two. By adjusting the encoder parameters during training through the second loss function, the encoder can learn the knowledge of extracting the partial feature representations corresponding to the target music components, so as to directly train the encoder into a model that can clearly express the effects of different target music components and be directly used in subsequent generation applications.
[0108] Similarly, the loss function can also include a reconstruction loss function for measuring the difference between the audio data reconstructed by the decoder and the sample audio, and a generative adversarial loss function of the generator when the music generation model acts as a generator and a discriminator to compete with each other, and the final second loss function is determined by a weighted fusion of the three.
[0109] Next, the working process of the music generation model in the embodiments of the present application will be described in conjunction with specific examples.
[0110] In one example, when the target music component is a vocal type component, the sample audio can be input into a text conversion network corresponding to the speech type component to obtain the sample text feature representation corresponding to each vocal type component in the sample audio.
[0111] Here, when the target music component is a vocal type component, the text conversion network can include TTS models such as the whisper model and the HuBERT model. The text conversion network can extract the text features corresponding to the speech contained in the sample audio.
[0112] In another example, when the target music component is an instrument type component, the sample audio is input into a music feature extraction network corresponding to the instrument type component to obtain the sample audio feature representation corresponding to each instrument type component in the sample audio.
[0113] Here, when the target music component is an instrument type component, the music feature extraction network can be the MERT model. The MERT model can extract the music feature representations of different instruments. MERT can also extract the mixed features of songs and texts, not limited to the music features of background music.
[0114] In another example, considering the complexity of songs in practice, the sample audio can include at least two vocal type components and / or at least two instrument type components. For example, multi-person chorus, multi-instrument ensemble, multi-person chorus with multi-instrument accompaniment, etc.
[0115] At this time, the independent audio track corresponding to each vocal type component and / or each instrument type component can be obtained first. For example, the sample audio can be automatically separated into audio tracks by audio separation tools such as Demucs, Spleeter, and OpenUnmix to obtain the independent audio track corresponding to each target music component. For a chorus, the audio tracks can be distinguished according to the voices, and for an audio containing multiple instruments, it can be divided into multiple audio tracks such as guitar, piano, and bass according to the instrument types.
[0116] After that, the independent audio track corresponding to each target music component is input into the corresponding type of feature extraction network to obtain the sample partial feature representation corresponding to each target music component in the sample audio. It should be noted that when there are multiple audio tracks, each audio track needs to be separately extracted to calculate the loss function and trained to ensure the isolation between each audio track.
[0117] A music generation method provided by an embodiment of the present application is implemented based on a music generation model, and the music generation model includes an encoder and a decoder. The method includes: determining at least one target music component to be highlighted in the original audio; based on the original audio and the encoder, determining at least one enhanced feature representation of the original audio, where the enhanced feature representation refers to the feature representation in which the target music component is enhanced compared to the original audio; in this way, during the music generation process, the feature corresponding to the target music component in the enhanced feature representation determined by the encoder is enhanced compared to other music components, and thus the target music component in the target audio generated based on the enhanced feature representation is highlighted, so as to ensure that the various music components in the target audio are not easily mixed with each other, at least one music component can be clearly expressed, and the playback effect of the generated music is improved.
[0118] Among them, the target music component can not only be a human voice, but also an instrument, as well as different voices and / or different roles in a chorus, etc. Therefore, the target audio generated by the embodiment of the present application can flexibly highlight different music components according to requirements, can be applied to more music generation scenarios, and improves the user's listening experience.
[0119] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a music generation device provided by an embodiment of the present application. The music generation device is implemented based on a music generation model, and the music generation model includes an encoder and a decoder. As Figure 5 shown in, the music generation device 500 includes:
[0120] A music component determination module 510, configured to determine at least one target music component to be highlighted in the original audio; among them, the at least one target music component forms the original audio according to the art performance form;
[0121] A feature determination module 520, configured to determine at least one enhanced feature representation of the original audio based on the original audio and the encoder; where the enhanced feature representation refers to the feature representation in which the target music component is enhanced compared to the original audio;
[0122] A generation module 530, configured to input the at least one enhanced feature representation into the decoder to generate a target audio with the target music component highlighted.
[0123] Furthermore, the music generation model further includes at least one feature extraction network corresponding to at least one target music component one by one; when the feature determination module 520 is used to determine the enhanced feature representation of the original audio based on the original audio and the encoder, the feature determination module 520 is configured to:
[0124] Input the original audio into the encoder to obtain an overall feature representation of the original audio;
[0125] Based on the original audio and at least one feature extraction network, obtain at least one target feature representation corresponding to the target music component;
[0126] Perform feature fusion on the overall feature representation and the at least one target feature representation to obtain the enhanced feature representation.
[0127] Further, the music generation device 500 further includes: a training module (not shown in the figure); when training the music generation model, the training module is configured to:
[0128] Input the sample audio into the encoder to obtain a sample overall feature representation of the sample audio;
[0129] Input the sample audio into at least one feature extraction network respectively to obtain at least one sample partial feature representation corresponding to each target music component in the original audio;
[0130] Perform feature processing on the sample overall feature representation and the at least one sample partial feature representation to obtain a sample enhanced feature representation;
[0131] Input the sample enhanced feature representation into the decoder to obtain enhanced audio;
[0132] Determine a first loss function according to the sample enhanced feature representation, the enhanced audio and the sample audio, and perform iterative training on the music generation model based on the first loss function until the preset training completion condition is satisfied, and obtain the trained music generation model.
[0133] Further, when the training module is configured to determine the first loss function according to the sample enhanced feature representation, the enhanced audio and the sample audio, the training module is configured to:
[0134] Determine a fusion loss function according to the audio parameters of the sample enhanced feature representation and the sample audio; wherein, the fusion loss function is used to measure the performance effect of the sample enhanced feature representation on the target music component;
[0135] Determine a reconstruction loss function according to the enhanced audio and the sample audio; wherein, the reconstruction loss function is used to measure the difference between the audio data reconstructed by the decoder and the sample audio;
[0136] Determine a generative adversarial loss function according to the enhanced audio and the sample audio; wherein, the music generation model and the discriminator are trained against each other during the training process;
[0137] Determine the first loss function based on the fusion loss function, the reconstruction loss function, and the generative adversarial loss function.
[0138] Further, when the training module trains the music generation model, the training module is further configured to:
[0139] Input the sample audio into the encoder, and perform vector quantization processing on the output result to obtain the overall feature representation of the sample audio.
[0140] Input the sample audio into at least one feature extraction network respectively to obtain at least one sample partial feature representation corresponding to each target music component in the original audio.
[0141] Determine a second loss function according to at least one sample partial feature representation and the overall feature representation of the sample audio, and iteratively train the music generation model based on the second loss function until the trained music generation model is obtained when a preset training completion condition is satisfied.
[0142] Further, when the training module is configured to input the sample audio into at least one feature extraction network respectively to obtain at least one sample partial feature representation corresponding to each target music component in the original audio, the training module is configured to:
[0143] When the target music component is a vocal type component, input the sample audio into the text conversion network corresponding to the speech type component to obtain the sample text feature representation corresponding to each vocal type component in the sample audio.
[0144] When the target music component is an instrument type component, input the sample audio into the music feature extraction network corresponding to the instrument type component to obtain the sample audio feature representation corresponding to each instrument type component in the sample audio.
[0145] Further, when the target music component includes at least two vocal type components and / or at least two instrument type components, when the training module is configured to input the sample audio into at least one feature extraction network respectively to obtain at least one sample partial feature representation corresponding to each target music component in the original audio, the training module is configured to:
[0146] Obtain the independent audio track corresponding to each vocal type component and / or each instrument type component.
[0147] Input the independent audio track corresponding to each target music component into the corresponding type of feature extraction network to obtain the sample partial feature representation corresponding to each target music component in the sample audio.
[0148] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As shown in Figure 6 , the electronic device 600 includes a processor 610, a memory 620, and a bus 630.
[0149] The memory 620 stores machine-readable instructions executable by the processor 610. When the electronic device 600 runs, the processor 610 communicates with the memory 620 through the bus 630. When the machine-readable instructions are executed by the processor 610, the steps of the music generation method in the method embodiment as described above can be executed. For the specific implementation manner, reference can be made to the method embodiment, which will not be elaborated herein. Figure 1 shown.
[0150] An embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the music generation method in the method embodiment as described above can be executed. For the specific implementation manner, reference can be made to the method embodiment, which will not be elaborated herein. Figure 1 shown.
[0151] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein.
[0152] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0153] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0154] In addition, in each embodiment of the present application, each functional unit may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit.
[0155] If the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0156] Finally, it should be noted that: the above-mentioned embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting them. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A music generation method, which is implemented based on a music generation model, characterized in that: The method comprises: Determining at least one target music component to be highlighted in the original audio; wherein the at least one target music component forms the original audio according to an artistic expression form; The music generation model includes an encoder and a decoder, and determines at least one enhanced feature representation of the original audio based on the original audio and the encoder; wherein the enhanced feature representation refers to a feature representation in which the target music component is enhanced compared to the original audio; The at least one enhanced feature representation is input into a decoder to generate a target audio with the target music component highlighted.
2. The method according to claim 1, characterized in that The music generation model also includes at least one feature extraction network corresponding one-to-one to at least one target music component; The step of determining an enhanced feature representation of the original audio based on the original audio and the encoder comprises: Inputting the original audio into the encoder to obtain an overall feature representation of the original audio; Based on the original audio and at least one feature extraction network, obtaining at least one target feature representation corresponding to the target music component; Feature fusion is performed on the overall feature representation and the at least one target feature representation to obtain the enhanced feature representation.
3. The method according to claim 2, characterized in that Training the music generation model includes: Inputting the sample audio into the encoder to obtain a sample overall feature representation of the sample audio; Inputting the sample audio into at least one feature extraction network respectively to obtain at least one sample partial feature representation corresponding to each target music component in the original audio; Performing feature processing on the overall feature representation of the sample and at least one partial feature representation of the sample to obtain an enhanced feature representation of the sample; Inputting the sample enhanced feature representation into the decoder to obtain enhanced audio; A first loss function is determined according to the sample enhanced feature representation, the enhanced audio and the sample audio, and the music generation model is iteratively trained based on the first loss function until a preset training completion condition is met, thereby obtaining the trained music generation model.
4. The method according to claim 3, characterized in that The determining of a first loss function according to the sample enhanced feature representation, the enhanced audio and the sample audio comprises: Determine a fusion loss function according to the sample enhanced feature representation and the audio parameters of the sample audio; wherein the fusion loss function is used to measure the performance effect of the sample enhanced feature representation on the target music component; Determine a reconstruction loss function according to the enhanced audio and the sample audio; wherein the reconstruction loss function is used to measure the difference between the audio data reconstructed by the decoder and the sample audio; Determine a generative adversarial loss function based on the enhanced audio and the sample audio; wherein, during the training process, the music generation model and the discriminator are trained against each other; The first loss function is determined based on the fusion loss function, the reconstruction loss function and the generative adversarial loss function.
5. The method according to claim 4, characterized in that Training the music generation model further includes: Inputting the sample audio into the encoder, and performing vector quantization processing on the output result to obtain a sample overall feature representation of the sample audio; Inputting the sample audio into at least one feature extraction network respectively to obtain at least one sample partial feature representation corresponding to each target music component in the original audio; A second loss function is determined based on at least one partial feature representation of the sample and the overall feature representation of the sample, and the music generation model is iteratively trained based on the second loss function until a preset training completion condition is met, thereby obtaining the trained music generation model.
6. The method according to claim 3 or 5, characterized in that: Inputting the sample audio into at least one feature extraction network respectively to obtain at least one sample partial feature representation corresponding to each target music component in the original audio, including: When the target music component is a vocal type component, the sample audio is input into a text conversion network corresponding to the voice type component to obtain a sample text feature representation corresponding to each vocal type component in the sample audio; When the target music component is a musical instrument type component, the sample audio is input into a music feature extraction network corresponding to the musical instrument type component to obtain a sample audio feature representation corresponding to each musical instrument type component in the sample audio.
7. The method according to claim 6, characterized in that When the target music component includes at least two vocal type components and / or at least two instrument type components, the sample audio is respectively input into at least one feature extraction network to obtain at least one sample partial feature representation corresponding to each target music component in the original audio, including: Obtaining independent audio tracks corresponding to each vocal type component and / or each instrument type component; The independent audio track corresponding to each target music component is input into a feature extraction network of a corresponding type to obtain a sample partial feature representation corresponding to each target music component in the sample audio.
8. A music generation device, which is implemented based on a music generation model, characterized in that: The music generation model includes an encoder and a decoder, and the music generation device includes: A music component determination module, used to determine at least one target music component to be highlighted in the original audio; wherein the at least one target music component forms the original audio according to an artistic expression form; A feature determination module, configured to determine at least one enhanced feature representation of the original audio based on the original audio and the encoder; wherein the enhanced feature representation refers to a feature representation in which the target music component is enhanced compared to the original audio; A generating module is used to input the at least one enhanced feature representation into a decoder to generate a target audio with the target music component highlighted.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the music generation method according to any one of claims 1 to 7 are executed.
10. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and the machine-readable instructions are executed by the processor to execute the steps of the music generation method as described in any one of claims 1 to 7.