Singing synthesis method, device and medium based on denoising diffusion probability model

By introducing the denoising diffusion probability model and generative adversarial network into the singing synthesis method, the problem of insufficient naturalness in singing audio synthesis is solved, and efficient and high-quality singing audio generation is achieved.

CN116564270BActive Publication Date: 2025-09-16PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310595288.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-24
Publication Date
2025-09-16
Estimated Expiration
2043-05-24

AI Technical Summary

Technical Problem

Existing singing synthesis methods suffer from over-smoothing problems in the generative models trained on the reconstruction objective, resulting in low naturalness of the synthesized singing audio.

Method used

The mel spectrum features are denoised using a denoising diffusion probability model, and denoised using a generative adversarial network. High-quality singing audio is generated through a target generator.

Benefits of technology

The naturalness of singing audio is improved, the computational cost and processing time are reduced, and the quality of synthesized singing audio is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116564270B_ABST
    Figure CN116564270B_ABST
Patent Text Reader

Abstract

The singing synthesis method, device, and medium based on a denoising diffusion probability model proposed in the embodiments of the present application relate to the field of artificial intelligence technology. The method comprises: obtaining an initial Mel-spectrogram feature of a preset musical score; inputting the initial Mel-spectrogram feature into a preset denoising diffusion probability model for noise processing to obtain a priori noise Mel-spectrogram feature; encoding the noise time step of the priori noise Mel-spectrogram feature to obtain a noise time step feature; inputting the priori noise Mel-spectrogram feature, the noise time step feature, and the initial Mel-spectrogram feature into a preset target generator for denoising to obtain a target denoised Mel-spectrogram feature; performing audio synthesis on the target denoised Mel-spectrogram feature to obtain target synthesized audio data. The embodiments of the present application can improve the naturalness of synthesized singing audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a singing synthesis method, singing synthesis device, electronic device and storage medium based on a denoising diffusion probability model. Background Art

[0002] Singing Voice Synthesis (SVS) refers to the use of computers to synthesize human-like singing voices based on a given musical score, aiming to generate natural and expressive singing voices. The process of singing voice synthesis generally includes: estimating the intermediate acoustic features of the singing voice from the musical score through an acoustic model; then, synthesizing the final audio waveform based on the intermediate acoustic features. Among them, the acoustic model determines the musical naturalness and expressiveness of the synthesized singing voice. The acoustic model can be a long short-term memory neural network, and reconstruction loss is generally used to train the acoustic model to estimate the acoustic features. However, the generative model trained on a simple reconstruction target often has the problem of over-smoothing and low naturalness. Therefore, how to provide a singing synthesis method that can improve the naturalness of synthesized singing audio has become a technical problem that needs to be solved urgently. Summary of the Invention

[0003] The main purpose of the embodiments of the present application is to propose a singing synthesis method, singing synthesis device, electronic device and storage medium based on a denoising diffusion probability model, which can improve the naturalness of synthesized singing audio.

[0004] To achieve the above objectives, a first aspect of an embodiment of the present application provides a singing synthesis method based on a denoising diffusion probability model, the method comprising:

[0005] Get the initial Mel spectrum features of the preset music score;

[0006] Inputting the initial Mel spectrum feature into a preset denoising diffusion probability model for noise processing to obtain a priori noise Mel spectrum feature;

[0007] Encoding the noise-added time step of the prior noise Mel-spectrum feature to obtain a noise-added time step feature;

[0008] Inputting the prior noise Mel spectrum feature, the noise-added time step feature and the initial Mel spectrum feature into a preset target generator for denoising to obtain a target denoised Mel spectrum feature;

[0009] Perform audio synthesis on the target denoised Mel-spectrogram feature to obtain target synthesized audio data.

[0010] In some embodiments, before inputting the prior noise Mel-spectrogram feature, the noise-added time step feature, and the initial Mel-spectrogram feature into a preset target generator for denoising to obtain a target denoised Mel-spectrogram feature, the method further includes:

[0011] Training the target generator specifically includes:

[0012] Inputting the prior noise Mel spectrum feature, the noise-added time step feature and the initial Mel spectrum feature into a preset initial generator for denoising to obtain an intermediate denoised Mel spectrum feature;

[0013] Inputting the intermediate denoised Mel-spectrum feature into the denoising diffusion probability model for noise addition processing to obtain a synthetic noise Mel-spectrum feature; wherein the noise addition time step of the synthetic noise Mel-spectrum feature is the same as the noise addition time step of the prior noise Mel-spectrum feature;

[0014] Inputting the synthetic noise Mel spectrum feature, the prior noise Mel spectrum feature and the noise-added time step feature into a preset discriminator for authenticity identification to obtain an identification label;

[0015] If the identification label is a negative label, a loss is calculated based on the synthetic noise Mel spectrum feature and the prior noise Mel spectrum feature to obtain loss data; the negative label is used to indicate that the synthetic noise Mel spectrum feature is identified as false;

[0016] The parameters of the initial generator are adjusted according to the loss data to obtain the target generator.

[0017] In some embodiments, inputting the prior noise Mel spectrum feature, the noise-added time step feature, and the initial Mel spectrum feature into a preset target generator for denoising to obtain a target denoised Mel spectrum feature includes:

[0018] Performing feature fusion on the prior noise Mel spectrum feature, the noise-added time step feature, and the initial Mel spectrum feature to obtain an initial fused noise Mel spectrum feature;

[0019] Performing dimensionality reduction processing on the initial fused noise Mel spectrum feature to obtain a target fused noise Mel spectrum feature;

[0020] Inputting the target fused noise Mel spectrum feature into a preset first convolution layer for convolution processing to obtain a first denoised Mel spectrum feature;

[0021] Inputting the first denoised Mel-spectrum feature into a preset first normalization layer for normalization processing to obtain a second denoised Mel-spectrum feature;

[0022] Inputting the second denoised Mel spectrum feature into a preset activation function layer for activation processing to obtain an initial denoised Mel spectrum feature;

[0023] The initial denoised Mel-spectrum feature is subjected to dimensionality increase processing to obtain the target denoised Mel-spectrum feature.

[0024] In some embodiments, performing dimensionality-increasing processing on the initial denoised Mel-spectrum feature to obtain the target denoised Mel-spectrum feature includes:

[0025] Inputting the target denoised Mel spectrum feature into a preset convolution unit for convolution processing to obtain an initial upsampled Mel spectrum feature;

[0026] Inputting the initial up-sampled Mel spectrum feature into a preset pixel shuffling unit for up-sampling processing to obtain an intermediate up-sampled Mel spectrum feature;

[0027] The intermediate up-sampled Mel spectrum feature is input into a preset activation function unit for activation processing to obtain the target denoised Mel spectrum feature.

[0028] In some embodiments, obtaining the initial Mel-spectrogram features of the preset music score includes:

[0029] Get the note data and lyrics data of the preset music score;

[0030] Inputting the note data and the lyrics data into a preset lyrics and music encoder for feature encoding to obtain an initial lyrics and music feature sequence; the initial lyrics and music feature sequence includes at least one initial lyrics and music sub-feature;

[0031] Performing duration prediction on the initial lyrics and music features using a preset duration predictor to obtain a predicted duration;

[0032] Aligning the predicted duration with the current duration of the initial lyrics and music features to obtain alignment information;

[0033] Performing feature repetition processing on the initial lyrics and music features according to the alignment information to obtain target lyrics and music features;

[0034] Merging the target lyrics and music sub-features to obtain a target lyrics and music feature sequence;

[0035] The target lyrics and music feature sequence is input into a preset spectrum encoder for spectrum feature encoding to obtain the initial Mel spectrum feature.

[0036] In some embodiments, the spectrum encoder includes a multi-head self-attention layer, a second normalization layer, and a second convolution layer. Inputting the target lyrics and music feature sequence into a preset spectrum encoder for spectrum feature encoding to obtain the initial Mel spectrum feature includes:

[0037] Extracting features of the target lyrics and music feature sequence through the multi-head self-attention layer to obtain a first Mel spectrum feature;

[0038] Normalizing the first mel-spectrogram feature and the target lyrics and music feature sequence through the second normalization layer to obtain a second mel-spectrogram feature;

[0039] The second mel-spectrogram feature is convolved by the second convolutional layer to obtain the initial mel-spectrogram feature.

[0040] In some embodiments, the lyric-music encoder includes a word embedding unit, a multi-head attention unit, a normalization unit, and a feedforward network unit. Inputting the note data and the lyrics data into a preset lyric-music encoder for feature encoding to obtain an initial lyric-music feature sequence includes:

[0041] Performing word embedding processing on the note data and the lyrics data by the word embedding unit to obtain a first lyric feature sequence;

[0042] Performing feature attention processing on the first lyric and music feature sequence through the multi-head attention unit to obtain a second lyric and music feature sequence;

[0043] Normalizing the second lyric-music feature sequence and the first lyric-music feature sequence by the normalization unit to obtain a third lyric-music feature sequence;

[0044] The feedforward network unit is used to extract features from the third lyrics and music feature sequence to obtain the initial lyrics and music feature sequence.

[0045] To achieve the above-mentioned objectives, a second aspect of an embodiment of the present application provides a singing synthesis device based on a denoising diffusion probability model, the device comprising:

[0046] A feature acquisition module is used to obtain the initial Mel spectrum features of the preset music score;

[0047] A priori noise adding module is used to input the initial Mel spectrum feature into a preset denoising diffusion probability model for noise adding processing to obtain a priori noise Mel spectrum feature;

[0048] A time step encoding module is used to encode the noise-added time step of the prior noise Mel spectrum feature to obtain a noise-added time step feature;

[0049] a denoising processing module, configured to input the prior noise Mel spectrum feature, the noise-added time step feature, and the initial Mel spectrum feature into a preset target generator for denoising, thereby obtaining a target denoised Mel spectrum feature;

[0050] The audio synthesis module is used to perform audio synthesis on the target denoised Mel spectrum feature to obtain target synthesized audio data.

[0051] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.

[0052] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes a storage medium, which is a computer-readable storage medium and stores a computer program. When the computer program is executed by a processor, the method described in the first aspect is implemented.

[0053] The present application proposes a singing synthesis method, singing synthesis method device, electronic device and storage medium based on a denoising diffusion probability model. The method uses a denoising diffusion probability model to perform noise processing on the Mel spectrum features to obtain diversified prior noise Mel spectrum features, so as to achieve the diversification of Mel spectrum features. For denoising processing, the prior noise Mel spectrum features, the denoising time step features and the initial Mel spectrum features are input into a preset target generator for denoising processing, which can obtain the target denoised Mel spectrum features, and then the target synthesized audio data is obtained through audio synthesis. In summary, the embodiments of the present application can help improve the naturalness of synthesized singing audio through the combined effect of the denoising diffusion probability model and the target generator. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 1 is a schematic diagram of a system architecture for executing a singing synthesis method based on a denoising diffusion probability model provided by an embodiment of the present application;

[0055] Figure 2 is a flow chart of a singing synthesis method based on a denoising diffusion probability model provided in an embodiment of the present application;

[0056] Figure 3 yes Figure 2 Flowchart of step S101 in FIG.

[0057] Figure 4 yes Figure 3 Flowchart of step S202 in FIG.

[0058] Figure 5 yes Figure 3 Flowchart of step S207 in FIG.

[0059] Figure 6 is a flowchart of a singing synthesis method based on a denoising diffusion probability model provided by another embodiment of the present application;

[0060] Figure 7 yes Figure 2 Flowchart of step S104 in FIG.

[0061] Figure 8 This is a block diagram of the module structure of a singing synthesis device based on a denoising diffusion probability model provided by an embodiment of the present application;

[0062] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0064] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0066] First, let’s analyze some of the terms used in this application:

[0067] Artificial intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It also encompasses the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0068] Natural language processing (NLP): NLP uses computers to process, understand, and apply human languages ​​(such as Chinese and English). A branch of artificial intelligence, NLP is an interdisciplinary field between computer science and linguistics, often referred to as computational linguistics. Natural language processing encompasses grammatical analysis, semantic analysis, and discourse comprehension. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information and image processing, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It encompasses data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and linguistics research related to language computing.

[0069] Singing Voice Synthesis (SVS): Singing is a process that synthesizes vocals based on lyrics and musical notation. While Text-to-Speech (TTS) allows machines to "speak," SVS allows machines to sing. Compared to TTS, SVS requires more input information, such as pitch and tempo information from the musical notation.

[0070] Denoising Diffusion Probabilistic Model (DDPM): is a generative model that has shown promising performance in multiple areas, such as image generation, neural speech coding, speech enhancement, and speech or singing audio synthesis.

[0071] In related technologies, the use of neural networks has driven the development of SVS. Generally, the most prominent neural network-driven SVS systems divide the generation pipeline into two stages. First, an acoustic model estimates the musical score to obtain intermediate acoustic features of the singing voice. Then, a speech coding model is used to generate the final synthesized audio from these acoustic features. The acoustic model determines the musical naturalness and expressiveness of the synthesized singing voice. Many deep learning technologies, such as long short-term memory neural networks (LSTMs) and speech synthesis models (WaveNet), have been deployed as backbone modules for building acoustic models. These systems typically use a reconstruction loss to train acoustic models to estimate acoustic features. However, generative models trained on simple reconstruction objectives often suffer from oversmoothing, resulting in estimates that are close to the mean or median of the target distribution, lacking the characteristics of human singing and resulting in a low degree of naturalness. Therefore, developing a singing synthesis method that improves the naturalness of synthesized audio has become a pressing technical challenge.

[0072] Considering the high performance of DDPM in singing audio synthesis, the embodiment of the present application integrates DDPM into the singing synthesis method, and proposes a singing synthesis method based on a diffusion denoising probability model, aiming to use DDPM to generate diverse Mel spectrum features, thereby obtaining high-quality synthesized audio. Among them, DDPM obtains diverse sample features by adding noise to the Mel spectrum features, but denoising is required in the direction process, and the current denoising process consumes a lot of computing cost and takes a long time. Therefore, the embodiment of the present application also introduces Generative Adversarial Networks (GAN), which implements denoising through conditional generative adversarial networks, improves processing efficiency, and reduces computing cost.

[0073] The singing synthesis method based on the denoising diffusion probability model provided in the embodiment of the present application is applied to the server or terminal, and can also be software running on the server. Figure 1The singing synthesis method based on the denoising diffusion probability model of the embodiment of the present application can be executed solely by the target server 400, or by the first terminal 100, the second terminal 200, or the third terminal 300, or can be executed jointly by the first terminal 100, the second terminal 200, the third terminal 300, and the target server 400. The server side can be configured as an independent physical server, or as a server cluster or distributed system consisting of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the singing synthesis method based on the denoising diffusion probability model, but is not limited to the above forms.

[0074] The present application can be used in many general or special computer system environments or configurations. For example: server computers, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0075] The embodiments of the present application provide a singing synthesis method, a singing synthesis device, an electronic device and a storage medium based on a denoising diffusion probability model, which are specifically illustrated by the following embodiments. First, the singing synthesis method based on a denoising diffusion probability model in the embodiments of the present application is described.

[0076] Figure 2 This is an optional flowchart of the singing synthesis method based on the denoising diffusion probability model provided in an embodiment of the present application, which may include but is not limited to steps S101 to S105.

[0077] Step S101, obtaining initial Mel-spectrogram features of a preset music score;

[0078] Step S102: inputting the initial Mel-spectrum frequency spectrum features into a preset denoising diffusion probability model for noise processing to obtain a priori noise Mel-spectrum frequency spectrum features;

[0079] Step S103, encoding the noise-added time step of the prior noise Mel-spectrum feature to obtain a noise-added time step feature;

[0080] Step S104: inputting the prior noise Mel spectrum feature, the noise time step feature, and the initial Mel spectrum feature into a preset target generator for denoising to obtain a target denoised Mel spectrum feature;

[0081] Step S105 , performing audio synthesis on the target denoised Mel-spectrogram feature to obtain target synthesized audio data.

[0082] In steps S101 to S105 shown in the embodiment of the present application, the Mel spectrum features are subjected to noise processing through a denoising diffusion probability model to obtain diversified prior noise Mel spectrum features, so as to achieve the diversification of Mel spectrum features, which is conducive to obtaining more natural synthesized singing audio from the Mel spectrum features. For denoising, the embodiment of the present application adopts a method of inputting the prior noise Mel spectrum features, the noise time step features and the initial Mel spectrum features into a preset target generator for denoising. Specifically, the noise to be removed can be predicted based on the noise time step and the initial Mel spectrum features, and then the prior noise Mel spectrum features are denoised. This can reduce the complex reasoning in the denoising process and also improve the accuracy of the denoising process. In summary, the embodiment of the present application can help improve the naturalness of the synthesized singing audio and improve the processing efficiency by performing data enhancement on the Mel spectrum features.

[0083] In step S101 of some embodiments, the initial Mel-spectrogram features of a preset musical score are obtained. Conventional singing synthesis methods require specifying a musical score, extracting the acoustic features of the musical score, and then performing speech synthesis based on the acoustic features to obtain synthesized audio. Therefore, the preset musical score in the embodiments of the present application refers to the specified musical score to be converted into synthesized singing audio. Compared to text-to-speech, singing synthesis requires pitch information in addition to lyrics, so the initial Mel-spectrogram features in the embodiments of the present application are used to represent the lyrics features and pitch features of the musical score.

[0084] Specifically, when the singing synthesis method based on the denoising diffusion probability model is applied to the first terminal 100, the initial mel-spectrogram features can be obtained via Bluetooth transmission, wired transmission, or downloading. When the singing synthesis method based on the denoising diffusion probability model is applied to the target server 400, the initial mel-spectrogram features can be uploaded to the target server 400 by the first terminal 100 or downloaded by the target server 400 from another server.

[0085] In addition to directly obtaining the initial Mel-spectrogram features from the first terminal 100 or the target server 400, the lyrics data and note data of the preset music score can also be obtained from the first terminal 100 or the target server 400, and then the spectral features of the lyrics data and note data are extracted to obtain the initial Mel-spectrogram features. Figure 3In one embodiment, step S101 specifically includes but is not limited to steps S201 to S207:

[0086] Step S201, obtaining note data and lyrics data of a preset music score;

[0087] Step S202: Input the note data and lyrics data into a preset lyrics and music encoder for feature encoding to obtain an initial lyrics and music feature sequence; the initial lyrics and music feature sequence includes at least one initial lyrics and music sub-feature;

[0088] Step S203, performing duration prediction on the initial lyrics and music features using a preset duration predictor to obtain a predicted duration;

[0089] Step S204, aligning the predicted duration with the current duration of the initial lyrics and music features to obtain alignment information;

[0090] Step S205, performing feature repetition processing on the initial lyrics and music features according to the alignment information to obtain a target lyrics and music feature sequence;

[0091] Step S206, merging the target lyrics and music sub-features to obtain a target lyrics and music feature sequence;

[0092] Step S207: Input the target lyrics and music feature sequence into a preset spectrum encoder for spectrum feature encoding to obtain initial Mel spectrum features.

[0093] In step S201 of some embodiments, the lyric data refers to the target lyrics corresponding to the song, and the note data refers to the target melody corresponding to the song. After feature conversion of the lyric data, text features can be obtained. These text features are used to represent the characteristics of the lyric information in the song, for example, they can be a sequence of phoneme features after the lyric data is converted. After feature conversion of the note data, melody features can be obtained. These melody features are used to represent the characteristics of the melody information in the song, such as notes, pitch, rhythm, legato, and tenuto.

[0094] In step S202 of some embodiments, the note data and lyric data are input into a preset lyric-music encoder for feature encoding to obtain an initial lyric-music feature sequence; the initial lyric-music feature sequence includes at least one initial lyric-music sub-feature. Specifically, the lyric-music encoder is configured to perform feature conversion on the note data and lyric data to obtain an initial lyric-music feature sequence, each of which includes at least one initial lyric-music sub-feature. Each initial lyric-music sub-feature includes a text feature converted from the lyric data and a melody feature converted from the note data.

[0095] In one embodiment, the word-music encoder includes a word embedding unit, a multi-head attention unit, a normalization unit, and a feed-forward network unit. Figure 4Step S202 specifically includes but is not limited to steps S301 to S304:

[0096] Step S301, performing word embedding processing on the note data and the lyrics data through a word embedding unit to obtain a first lyric feature sequence;

[0097] Step S302: performing feature attention processing on the first lyric and music feature sequence through a multi-head attention unit to obtain a second lyric and music feature sequence;

[0098] Step S303: normalizing the second lyric and music feature sequence and the first lyric and music feature sequence by a normalization unit to obtain a third lyric and music feature sequence;

[0099] Step S304: extract features from the third lyrics and music feature sequence through a feedforward network unit to obtain an initial lyrics and music feature sequence.

[0100] In step S301 of some embodiments, the word embedding unit is used to convert the note data and the lyric data into feature vectors to obtain a first word-music feature sequence. For the lyric data, specifically, the lyric data is segmented using word segmentation technology, the segmented words are converted into word vectors, and then the word vectors are spliced ​​according to the order of the words in the lyric data to obtain a text feature vector. For the note data, the main step is to extract the pitch features in the note to obtain a melody feature vector. The text feature vector and the melody feature vector are superimposed to obtain a first word-music feature sequence. Among them, the conversion of words into word vectors can adopt a one-hot vector representation method or a Word2Vec training algorithm.

[0101] In one example, after obtaining the first lyric and music feature sequence, the first lyric and music feature sequence can be used as the initial lyric and music feature sequence to directly perform spectral feature extraction, thereby obtaining synthesized singing audio. However, in actual applications, it has been found that the synthesized singing audio has poor naturalness performance, and the spectral feature extraction process is very time-consuming. Based on this, the embodiment of the present application proposes to continue to perform feature processing such as steps S302 to S304 on the obtained first lyric and music feature sequence after the word embedding processing in step S301, in order to improve the naturalness of the final synthesized singing audio.

[0102] In step S302 of some embodiments, a multi-headed attention unit is used to pay attention to the features of the first word-music feature sequence, which is equivalent to applying different weights to the features in the first word-music feature sequence, so as to obtain a second word-music feature sequence. It should be noted that the attention model refers to a model that applies different weights to different information that needs to be considered to solve a problem in a specific scenario, applies higher weights to information that is of great help to the problem, and applies lower weights to information that is of little help to the problem, so as to better use this information to solve the problem. Therefore, the embodiment of the present application uses a multi-headed self-attention unit to perform feature attention processing on the first word-music feature sequence, and the obtained second word-music feature sequence helps to obtain more natural singing synthesis audio data.

[0103] In one example, after obtaining the second lyric and music feature sequence, the second lyric and music feature sequence can be used as the initial lyric and music feature sequence to directly perform spectral feature extraction, thereby obtaining synthesized singing audio. However, in this embodiment of the application, in order to further reduce the computational complexity and improve the naturalness of the synthesized audio, the normalization processing in step S303 and the feature extraction in step S304 are also required.

[0104] In step S303 of some embodiments, a normalization unit is used to perform normalization processing, which generally includes batch normalization and layer normalization. By normalizing the second word-music feature sequence by the normalization unit to obtain a third word-music feature sequence, convergence speed and performance can be improved.

[0105] In step S304 of some embodiments, a feedforward neural network (FNN), also known as a feedforward network, is a type of artificial neural network. The feedforward neural network adopts a unidirectional multi-layer structure. Each layer contains a number of neurons. In this neural network, each neuron can receive signals from neurons in the previous layer and generate outputs to the next layer. The 0th layer is called the input layer, the last layer is called the output layer, and the other intermediate layers are called hidden layers (or hidden layers, hidden layers). The hidden layer can be one layer. It can also be multiple layers. In this embodiment, the third word and music feature sequence is extracted through a feedforward network unit, and the obtained initial word and music feature sequence is more conducive to generating synthetic audio data with higher naturalness.

[0106] After obtaining the initial lyric and music feature sequence, spectral feature extraction is required. However, considering the inconsistency in sequence length between the initial lyric and music feature sequence and the mel-spectrogram feature to be synthesized, sequence length adjustment is required. Therefore, in some embodiments, in steps S203 to S206, a duration predictor is first used to obtain a predicted duration. Then, alignment information is obtained based on the predicted duration and the current duration. This alignment information is used to indicate the feature repetition multiple. Feature repetition processing is then performed based on the alignment information to obtain the target lyric and music feature sequence.

[0107] In step S207 of some embodiments, after obtaining the target word-music feature sequence, the target word-music feature sequence is input into a preset spectrum encoder for spectrum feature encoding to obtain initial Mel-spectrum features.

[0108] In one embodiment, the spectrum encoder includes a multi-head self-attention layer, a second normalization layer, and a second convolutional layer. Figure 5 Step S207 specifically includes but is not limited to the following steps S401 to S403:

[0109] Step S401, extracting features from the target lyrics and music feature sequence through a multi-head self-attention layer to obtain the first mel-spectrogram feature;

[0110] Step S402: normalizing the first mel-spectrogram feature and the target lyric-music feature sequence through a second normalization layer to obtain a second mel-spectrogram feature;

[0111] Step S403: performing convolution processing on the second mel-spectrogram feature through a second convolutional layer to obtain an initial mel-spectrogram feature.

[0112] In step S401 of some embodiments, a multi-headed self-attention layer is used to pay attention to the features of the target word-music feature sequence, which is equivalent to applying different weights to the features in the target word-music feature sequence, so as to obtain the first mel-spectrogram feature.

[0113] In step S402 of some embodiments, the first mel-spectrogram feature and the target lyric-music feature sequence are input into a second normalization layer for normalization processing to obtain a second mel-spectrogram feature. The second normalization layer of the embodiment of the present application has similar functions to the normalization unit described above, both of which provide normalization processing, and thus will not be described in detail.

[0114] In step S403 of some embodiments, the second mel-spectrogram feature is input into a second convolutional layer for convolution processing to obtain an initial mel-spectrogram feature.

[0115] In some examples, the initial Mel-spectrogram features can be directly synthesized for audio to obtain synthesized audio data. However, in practice, it is found that the synthesized audio data performs poorly in terms of naturalness, and no matter how the network parameters are adjusted, the model performance cannot be effectively improved. Therefore, after obtaining the initial Mel-spectrogram features, the embodiment of the present application introduces a denoising diffusion probability model to perform noise processing in order to obtain diversified prior noise Mel-spectrogram features. In order to control the computational complexity of the denoising process, a generative adversarial network is introduced. The specific processing flow is described as follows.

[0116] In step S102 of some embodiments, the initial Mel-spectrogram feature is input into a preset denoising diffusion probability model for noise addition to obtain a priori noise Mel-spectrogram feature.

[0117] The denoising diffusion probability model (DDPM) constructs the diffused data and corresponding noise at each time step through a deterministic diffusion process. In an example, the initial Mel spectrum feature X0 is input into the DDPM for noise processing, and the prior noise Mel spectrum features (X1, X2, X3, X4, X5, X6, X7, X8, X9, X10, X11, X12, X13, X14, X15, X16, X17, X18, X29, X30, X31, X42, X53, X64) corresponding to the noise time step are obtained. t-1 、X t ). The subscripts 1, 2, t-1, and t are used to represent the noise-adding time step of the prior noise Mel-spectrogram feature, for example, X t Represents the prior noise Mel spectrum feature at the noise addition time step t.

[0118] DDPM's noise addition process essentially provides training data for the Mel-spectrogram feature reconstruction process. The target generator is then trained to learn the relationship between noise and data. After training, the target generator is generated. Subsequently, a prior noisy Mel-spectrogram feature containing noise is fed into the target generator, which then estimates a denoised Mel-spectrogram feature.

[0119] In one embodiment, referring to Figure 6 The training target generator specifically includes but is not limited to the following steps S501 to S505:

[0120] Step S501: inputting the prior noise Mel spectrum feature, the noise time step feature and the initial Mel spectrum feature into a preset initial generator for denoising to obtain an intermediate denoised Mel spectrum feature;

[0121] Step S502: Input the intermediate denoised Mel-spectrum feature into the denoising diffusion probability model for noise addition to obtain a synthetic noisy Mel-spectrum feature; wherein the noise addition time step of the synthetic noisy Mel-spectrum feature is the same as the noise addition time step of the priori noisy Mel-spectrum feature;

[0122] Step S503: Input the synthetic noise Mel spectrum feature, the prior noise Mel spectrum feature, and the noise time step feature into a preset discriminator for authenticity identification to obtain an identification label;

[0123] Step S504: If the identification label is a negative label, a loss is calculated based on the synthetic noise Mel-spectrum feature and the prior noise Mel-spectrum feature to obtain loss data; the negative label is used to indicate that the synthetic noise Mel-spectrum feature is identified as false;

[0124] Step S505: Adjust the parameters of the initial generator according to the loss data to obtain the target generator.

[0125] Steps S501 to S505 shown in the embodiment of this application involve a total of two noisy mel-spectrogram features, one of which is a priori noisy mel-spectrogram feature and the other is a synthetic noisy mel-spectrogram feature. Because they are subjected to different noise addition processes, the noise addition consistency must be ensured during the training process of this embodiment of the application. In addition to inputting the synthetic noisy mel-spectrogram feature into the discriminator, the priori noisy mel-spectrogram feature and the noisy time step feature must also be input into the discriminator to ensure the discriminator's recognition accuracy.

[0126] It should also be noted that after the discriminator performs authenticity verification, it outputs an identification label. Identification labels include positive and negative labels. A positive label indicates that the synthetic noise Mel-spectrogram feature is identified as authentic, while a negative label indicates that the synthetic noise Mel-spectrogram feature is identified as fake. If the identification label is positive, it indicates that the target generator has been trained and can be used for denoising. If the identification label is negative, it indicates that the intermediate denoised Mel-spectrogram features obtained by the initial generator are inaccurate, and the parameters of the initial generator need to be adjusted until the loss data stabilizes, resulting in the target generator.

[0127] The loss data in step S504 can be obtained by calculating the Euclidean distance or cosine similarity between the synthetic noise Mel-spectrum feature and the prior noise Mel-spectrum feature.

[0128] In step S103 of some embodiments, the noise addition time step of the prior noise Mel-spectrogram feature is encoded to obtain a noise addition time step feature. Referring to step S102, since the noise addition process involves multiple noise addition time steps, it is necessary to determine the noise addition time step corresponding to the prior noise Mel-spectrogram feature being processed. Encoding the noise addition time step to obtain the noise addition time step feature primarily provides auxiliary information to the target generator.

[0129] In step S104 of some embodiments, the prior noise Mel-spectrogram feature, the noise-added time step feature, and the initial Mel-spectrogram feature are input into a preset target generator for denoising to obtain a target denoised Mel-spectrogram feature. Specifically, the target generator uses the difference between the prior noise Mel-spectrogram feature and the initial Mel-spectrogram feature, as well as the noise-added time step, to determine the noise to be removed, and then denoises the prior noise Mel-spectrogram feature to obtain a target denoised Mel-spectrogram feature.

[0130] In one embodiment, referring to Figure 7 Step S104 specifically includes but is not limited to the following steps S601 to S606:

[0131] Step S601, performing feature fusion on the prior noise Mel spectrum feature, the noise time step feature and the initial Mel spectrum feature to obtain the initial fused noise Mel spectrum feature;

[0132] Step S602, performing dimensionality reduction processing on the initial fused noise Mel-spectrum feature to obtain the target fused noise Mel-spectrum feature;

[0133] Step S603: input the target fused noise Mel spectrum feature into a preset first convolution layer for convolution processing to obtain a first denoised Mel spectrum feature;

[0134] Step S604: input the first denoised Mel-spectrum feature into a preset first normalization layer for normalization processing to obtain a second denoised Mel-spectrum feature;

[0135] Step S605: input the second denoised Mel-spectrum feature into a preset activation function layer for activation processing to obtain an initial denoised Mel-spectrum feature;

[0136] Step S606 , performing dimensionality increase processing on the initial denoised Mel-spectrum feature to obtain a target denoised Mel-spectrum feature.

[0137] Specifically, the target generator includes a splicing layer, a dimensionality reduction processing layer, a first convolutional layer, a first normalization layer, an activation function layer, and a dimensionality increase processing layer. The feature fusion of step S601 is performed through the splicing layer to obtain the initial fused noise Mel spectrum feature. Considering that the spectral features do not need to be too high in dimension and to improve processing efficiency, the dimensionality reduction processing layer is used to reduce the feature dimension. The dimensionality reduction processing layer includes a downsampling unit (Downsample) and a matrix reconstruction unit (Reshape).

[0138] After step S601, convolution processing in step S602, normalization processing in step S603, and activation processing in step S604 are performed to further obtain a more specific and clean initial denoised Mel-spectrogram feature from the target fused noisy Mel-spectrogram feature. Finally, the initial denoised Mel-spectrogram feature is increased in dimension by a dimensionality increase layer to maintain the dimensionality of the target denoised Mel-spectrogram feature and the initial fused noisy Mel-spectrogram feature.

[0139] In one embodiment, the dimensionality increase processing layer includes a convolution unit, a pixel shuffling unit, and an activation function unit. Step S606 specifically includes: inputting the target denoised Mel spectrum feature into the convolution unit for convolution processing to obtain an initial upsampled Mel spectrum feature; inputting the initial upsampled Mel spectrum feature into the pixel shuffling unit for upsampling processing to obtain an intermediate upsampled Mel spectrum feature; inputting the intermediate upsampled Mel spectrum feature into the activation function unit for activation processing to obtain the target denoised Mel spectrum feature.

[0140] It should be noted that the convolution unit or convolution layer in the embodiments of the present application may be composed of multiple convolution kernels, and the parameters of each convolution kernel are optimized through the backpropagation algorithm. The purpose of the convolution operation is to extract different features of the input. The low-level convolution kernels may only extract low-level features such as edges, lines, and corners. A network with more layers can iteratively extract more complex features from low-level features.

[0141] The pixel shuffler increases the spatial dimensionality of the output by mapping information from the channel dimension at a given spatial location into spatial blocks in the output, where each channel contributes to a consistent spatial location relative to its neighboring channels during upsampling.

[0142] The activation function layer or activation function unit in the embodiment of the present application is composed of activation functions, and the activation functions include linear rectification function GeLu, Softmax, hyperbolic tangent activation function tanh, etc.

[0143] In step S105 of some embodiments, audio synthesis is performed on the target denoised Mel-spectrogram features to obtain target synthesized audio data. Specifically, audio synthesis is performed on the target denoised Mel-spectrogram features using an audio synthesizer to obtain the target synthesized audio data. The audio synthesizer can be a Parallel WaveGan speech encoder or a HiFi-GAN speech encoder, and the specific configuration can be determined based on actual needs.

[0144] In summary, the present application combines the Diffusion Denoising Probabilistic Model (DDPM) and Generative Adversarial Network (GAN) to construct the backbone of the acoustic model, thereby improving the higher-level expressiveness of synthesized speech and thus enhancing the naturalness of the synthesized singing audio. This not only overcomes the problem of GAN-generated speech being monotonous and lacking diversity, but also overcomes the computationally intensive and slow denoising process of the Diffusion Denoising Probabilistic Model. The present application embodiment can quickly generate highly natural audio data.

[0145] See also Figure 8 The embodiment of the present application also provides a singing synthesis device based on a denoising diffusion probability model, which can implement the above-mentioned singing synthesis method based on a denoising diffusion probability model. Figure 8 This is a block diagram of the module structure of a singing synthesis device based on a denoising diffusion probability model provided in an embodiment of the present application. The device includes: a feature acquisition module 701, a priori denoising module 702, a time step encoding module 703, a denoising processing module 704, and an audio synthesis module 705. The feature acquisition module 701 is used to obtain initial Mel-spectrum features of a preset musical score; the priori denoising module 702 is used to input the initial Mel-spectrum features into a preset denoising diffusion probability model for denoising to obtain priori noise Mel-spectrum features; the time step encoding module 703 is used to encode the denoised time steps of the priori noise Mel-spectrum features to obtain denoised time step features; the denoising processing module 704 is used to input the priori noise Mel-spectrum features, the denoised time step features, and the initial Mel-spectrum features into a preset target generator for denoising to obtain target denoised Mel-spectrum features; and the audio synthesis module 705 is used to perform audio synthesis on the target denoised Mel-spectrum features to obtain target synthesized audio data.

[0146] It should be noted that the specific implementation of the singing synthesis device based on the denoising diffusion probability model is basically the same as the specific embodiment of the singing synthesis method based on the denoising diffusion probability model mentioned above, and will not be repeated here.

[0147] The present application also provides an electronic device comprising: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for communicating between the processor and the memory. When the program is executed by the processor, the aforementioned singing synthesis method based on a denoising diffusion probability model is implemented. The electronic device can be any smart terminal, including a tablet computer and an in-vehicle computer.

[0148] See also Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:

[0149] The processor 801 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0150] The memory 802 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called by the processor 801 to execute the singing synthesis method based on the denoising diffusion probability model in the embodiments of this application.

[0151] Input / output interface 803, used to implement information input and output;

[0152] Communication interface 804, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0153] Bus 805 , which transmits information between various components of the device (e.g., processor 801 , memory 802 , input / output interface 803 , and communication interface 804 );

[0154] The processor 801 , the memory 802 , the input / output interface 803 and the communication interface 804 are connected to each other in communication within the device via a bus 805 .

[0155] An embodiment of the present application also provides a storage medium, which is a computer-readable storage medium used for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned singing synthesis method based on the denoising diffusion probability model.

[0156] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0157] The singing synthesis method based on the denoising diffusion probability model, the singing synthesis device based on the denoising diffusion probability model, the electronic device and the storage medium provided in the embodiment of the present application perform noise processing on the Mel spectrum features through the denoising diffusion probability model to obtain diversified prior noise Mel spectrum features, so as to achieve the diversification of the Mel spectrum features, which is conducive to obtaining more natural synthesized singing audio from the Mel spectrum features. For denoising processing, the embodiment of the present application adopts the method of inputting the prior noise Mel spectrum features, the noise time step features and the initial Mel spectrum features into a preset target generator for denoising processing. Specifically, the noise to be removed can be predicted based on the noise time step and the initial Mel spectrum features, and then the prior noise Mel spectrum features are denoised. This can reduce the complex reasoning in the denoising process and also improve the accuracy of the denoising processing. In summary, the embodiment of the present application can help improve the naturalness of the synthesized singing audio and improve the processing efficiency by performing data enhancement on the Mel spectrum features.

[0158] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0159] It will be understood by those skilled in the art that Figure 2-7 The technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or a combination of certain steps, or different steps.

[0160] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0161] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0162] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0163] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0164] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0165] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0166] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0167] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for enabling an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0168] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A singing synthesis method based on a denoising diffusion probability model, characterized in that: The method comprises: Get the initial Mel spectrum features of the preset music score; Inputting the initial Mel spectrum feature into a preset denoising diffusion probability model for noise processing to obtain a priori noise Mel spectrum feature; Encoding the noise-added time step of the prior noise Mel-spectrum feature to obtain a noise-added time step feature; Inputting the prior noise Mel spectrum feature, the noise-added time step feature and the initial Mel spectrum feature into a preset target generator for denoising to obtain a target denoised Mel spectrum feature; Perform audio synthesis on the target denoised Mel-spectrogram feature to obtain target synthesized audio data.

2. The method according to claim 1, characterized in that Before inputting the prior noise Mel spectrum feature, the noise-added time step feature, and the initial Mel spectrum feature into a preset target generator for denoising to obtain a target denoised Mel spectrum feature, the method further includes: Training the target generator specifically includes: Inputting the prior noise Mel spectrum feature, the noise-added time step feature and the initial Mel spectrum feature into a preset initial generator for denoising to obtain an intermediate denoised Mel spectrum feature; Inputting the intermediate denoised Mel-spectrum feature into the denoising diffusion probability model for noise addition processing to obtain a synthetic noise Mel-spectrum feature; wherein the noise addition time step of the synthetic noise Mel-spectrum feature is the same as the noise addition time step of the prior noise Mel-spectrum feature; Inputting the synthetic noise Mel spectrum feature, the prior noise Mel spectrum feature and the noise-added time step feature into a preset discriminator for authenticity identification to obtain an identification label; If the identification label is a negative label, a loss is calculated based on the synthetic noise Mel spectrum feature and the prior noise Mel spectrum feature to obtain loss data; the negative label is used to indicate that the synthetic noise Mel spectrum feature is identified as false; The parameters of the initial generator are adjusted according to the loss data to obtain the target generator.

3. The method according to claim 1, characterized in that The step of inputting the prior noise Mel spectrum feature, the noise-added time step feature, and the initial Mel spectrum feature into a preset target generator for denoising to obtain a target denoised Mel spectrum feature includes: Performing feature fusion on the prior noise Mel spectrum feature, the noise-added time step feature, and the initial Mel spectrum feature to obtain an initial fused noise Mel spectrum feature; Performing dimensionality reduction processing on the initial fused noise Mel spectrum feature to obtain a target fused noise Mel spectrum feature; Inputting the target fused noise Mel spectrum feature into a preset first convolution layer for convolution processing to obtain a first denoised Mel spectrum feature; Inputting the first denoised Mel-spectrum feature into a preset first normalization layer for normalization processing to obtain a second denoised Mel-spectrum feature; Inputting the second denoised Mel spectrum feature into a preset activation function layer for activation processing to obtain an initial denoised Mel spectrum feature; The initial denoised Mel-spectrum feature is subjected to dimensionality increase processing to obtain the target denoised Mel-spectrum feature.

4. The method according to claim 3, characterized in that The step of performing dimensionality-increasing processing on the initial denoised Mel-spectrum feature to obtain the target denoised Mel-spectrum feature includes: Inputting the target denoised Mel spectrum feature into a preset convolution unit for convolution processing to obtain an initial upsampled Mel spectrum feature; Inputting the initial up-sampled Mel spectrum feature into a preset pixel shuffling unit for up-sampling processing to obtain an intermediate up-sampled Mel spectrum feature; The intermediate up-sampled Mel spectrum feature is input into a preset activation function unit for activation processing to obtain the target denoised Mel spectrum feature.

5. The method according to any one of claims 1 to 4, characterized in that The step of obtaining the initial Mel-spectrogram features of the preset music score includes: Get the note data and lyrics data of the preset music score; Inputting the note data and the lyrics data into a preset lyrics and music encoder for feature encoding to obtain an initial lyrics and music feature sequence; the initial lyrics and music feature sequence includes at least one initial lyrics and music sub-feature; Performing duration prediction on the initial lyrics and music features using a preset duration predictor to obtain a predicted duration; Aligning the predicted duration with the current duration of the initial lyrics and music features to obtain alignment information; Performing feature repetition processing on the initial lyrics and music features according to the alignment information to obtain target lyrics and music features; Merging the target lyrics and music sub-features to obtain a target lyrics and music feature sequence; The target lyrics and music feature sequence is input into a preset spectrum encoder for spectrum feature encoding to obtain the initial Mel spectrum feature.

6. The method according to claim 5, characterized in that The spectrum encoder includes a multi-head self-attention layer, a second normalization layer and a second convolution layer. The target lyrics and music feature sequence is input into a preset spectrum encoder for spectrum feature encoding to obtain the initial Mel spectrum feature, including: Extracting features of the target lyrics and music feature sequence through the multi-head self-attention layer to obtain a first Mel spectrum feature; Normalizing the first mel-spectrogram feature and the target lyrics and music feature sequence through the second normalization layer to obtain a second mel-spectrogram feature; The second mel-spectrogram feature is convolved by the second convolutional layer to obtain the initial mel-spectrogram feature.

7. The method according to claim 5, characterized in that The lyric-music encoder includes a word embedding unit, a multi-head attention unit, a normalization unit, and a feedforward network unit. The note data and the lyrics data are input into a preset lyric-music encoder for feature encoding to obtain an initial lyric-music feature sequence, including: Performing word embedding processing on the note data and the lyrics data by the word embedding unit to obtain a first lyric feature sequence; Performing feature attention processing on the first lyric and music feature sequence through the multi-head attention unit to obtain a second lyric and music feature sequence; Normalizing the second lyric-music feature sequence and the first lyric-music feature sequence by the normalization unit to obtain a third lyric-music feature sequence; The feedforward network unit is used to extract features from the third lyrics and music feature sequence to obtain the initial lyrics and music feature sequence.

8. A singing synthesis device based on a denoising diffusion probability model, characterized in that: The device comprises: A feature acquisition module is used to obtain the initial Mel spectrum features of the preset music score; A priori noise adding module is used to input the initial Mel spectrum feature into a preset denoising diffusion probability model for noise adding processing to obtain a priori noise Mel spectrum feature; A time step encoding module, configured to encode the noise-added time step of the prior noise Mel-spectrum feature to obtain a noise-added time step feature; a denoising processing module, configured to input the prior noise Mel spectrum feature, the noise-added time step feature, and the initial Mel spectrum feature into a preset target generator for denoising, thereby obtaining a target denoised Mel spectrum feature; The audio synthesis module is used to perform audio synthesis on the target denoised Mel spectrum feature to obtain target synthesized audio data.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Acoustic model post-processing method based on probability diffusion model, server and readable memory

    CN114512114A

  • Speech synthesis method and device, electronic equipment and storage medium

    CN115641834A