A singing voice beautification method and system based on streaming matching

Through the vocal beautification method based on stream matching, the timbre characteristics and phoneme posterior probability map of the singing voice are extracted, and a multi-dimensional vocal expression sequence and vocal Mel score are generated, which solves the problem of insufficient vocal expression in the existing technology, and achieves higher quality and natural singing generation.

CN119479686BActive Publication Date: 2025-07-01JINAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510007602.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-07-01
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The existing singing beautification technology focuses on pitch correction, and fails to fully model and optimize the expressiveness of the singing, resulting in the generated singing lacking the expressiveness and emotional transmission of professional singers, and the beautified singing appears stiff and unnatural in emotional expression.

Method used

A method of singing vocal beautification based on stream matching is proposed. By obtaining singing vocal data and music score data, timbre characteristics and phoneme posterior probability map are extracted, multi-dimensional singing expression sequences are generated, and the voice-based mel score is converted through the vocoder to obtain the beautified singing voice.

Benefits of technology

It significantly improves the expressiveness and naturalness of the singing, making the generated singing quality higher, the listening feeling is smoother and expressive, and can effectively model and optimize the multi-dimensional expressive parameters of the singing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479686B_ABST
    Figure CN119479686B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and discloses a singing voice beautification method and system based on streaming matching. The method includes obtaining singing voice data and musical score data; extracting timbre features and phoneme posterior probability maps from the singing voice data; generating a multi-dimensional singing voice expressiveness sequence according to the musical score data and the phoneme posterior probability maps; generating a voice mel spectrogram according to the multi-dimensional singing voice expressiveness sequence, the phoneme posterior probability maps and the timbre features; and inputting the voice mel spectrogram into a vocoder for conversion processing to obtain the beautified singing voice. The present invention can optimize the output singing voice in terms of intonation, timbre and expressiveness, can significantly improve the expressiveness and naturalness of the singing voice, make the generated singing voice of higher quality, smoother to listen to and more expressive.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to a singing voice beautification method and system based on streaming matching. Background Art

[0002] With the rapid development of deep learning technology, singing voice beautification technology has gradually been applied to the improvement of audio quality, which can correct the deficiencies of the input singing voice and make it closer to the singing style of professional singers. Such technology provides a new solution for the automatic correction and beautification of singing voices, and has achieved certain results especially in pitch correction, which can effectively solve problems such as pitch deviation, thereby improving audio quality and making the singing voice sound more accurate and pleasant.

[0003] However, most of the existing singing voice beautification technologies only focus on the correction of a single dimension of pitch, and fail to comprehensively model and optimize the expressiveness of singing voices. Expressiveness includes multi-dimensional parameters such as pitch, singing style, tension, and energy, which jointly determine the natural fluency and emotional expression of singing voices. Due to ignoring these dimensions, the generated singing voice, although accurate in pitch, lacks the expressiveness and emotional transmission in the singing of professional singers. In addition, in the generation process of the existing technology, the dynamic relationship between the expressiveness parameters and the time steps of the singing voice content lacks consistency, resulting in the beautified singing voice being rigid and unnatural in emotional expression, with low overall expressiveness, and it is difficult to meet the requirements for high-quality singing voices in actual application scenarios. Summary of the Invention

[0004] In order to improve the defects of low singing voice expressiveness and quality existing in the existing singing voice beautification technology, the present invention proposes the following technical solutions:

[0005] In a first aspect, the present invention proposes a singing voice beautification method based on streaming matching, including:

[0006] Obtain singing voice data and musical score data;

[0007] Extract timbre features and phoneme posterior probability maps from the singing voice data;

[0008] Generate a multi-dimensional singing voice expressiveness sequence according to the musical score data and the phoneme posterior probability map;

[0009] Generate a voice Mel spectrogram according to the multi-dimensional singing voice expressiveness sequence, the phoneme posterior probability map, and the timbre features;

[0010] Input the voice Mel spectrogram into a vocoder for conversion processing to obtain the beautified singing voice.

[0011] As a preferred technical solution, the multi-dimensional singing voice expressiveness sequence includes a pitch frame-level sequence, a singing style frame-level sequence, an energy frame-level sequence, and a tension frame-level sequence.

[0012] As a preferred technical solution, a multi-dimensional singing expressiveness sequence is generated based on the music score data and the phoneme posterior probability map, including:

[0013] Construct and train a streaming matching network;

[0014] Perform dimensionality expansion processing on the phoneme posterior probability map;

[0015] According to the phoneme posterior probability map after dimensionality expansion and the music score data, use the trained streaming matching network for feature mapping to generate a pitch frame-level sequence;

[0016] According to the music score data and the pitch frame-level sequence, use the trained streaming matching network for feature mapping to generate a singing style frame-level sequence;

[0017] According to the music score data, the pitch frame-level sequence, and the phoneme posterior probability map, use the trained streaming matching network for feature mapping to generate an energy frame-level sequence;

[0018] According to the music score data, the pitch frame-level sequence, the energy frame-level sequence prediction, and the phoneme posterior probability map, use the trained streaming matching network for feature mapping to generate a target tension frame-level sequence.

[0019] As a preferred technical solution, construct a streaming matching network and use the streaming matching network for feature mapping, including:

[0020] Model the mapping relationship between the input distribution and the target distribution, and its expression is as follows:

[0021]

[0022] Where, represents the distribution of the j th expressiveness parameter at the time step t , represents the target distribution of the j th expressiveness parameter, represents the input distribution of the j th expressiveness parameter;

[0023] Through integral solution, map the input distribution to the target distribution , and use as the final singing expressiveness sequence output , and its expression is as follows:

[0024] ,

[0025] Where:

[0026]

[0027]

[0028]

[0029] In the formula, is the vector velocity obtained by the streaming matching network in the learning j expressiveness parameter stage, is the state distribution at the time step t of, represents the model parameters of the streaming matching network in the learning j expressiveness parameter, is the input score data, is the phoneme posterior probability map, is the phoneme posterior probability map after dimension expansion.

[0030] As a preferred technical solution, when using the streaming matching network for feature mapping, the method further includes predicting the vector velocity of the streaming matching network, including:

[0031] Obtain the phoneme posterior probability map, time step, and expressiveness parameter in real time;

[0032] Extract score features including rhythm, style, and pitch range from the score data;

[0033] Perform position encoding on the time step through the sine position encoding module to generate position encoding information;

[0034] Use the score features, phoneme posterior probability map, and position encoding information as input features and input them into the dilated convolution module, where the dilated convolution module consists of multiple dilated convolution blocks, and each dilated convolution block performs the following operations:

[0035] Perform dilated convolution operations on the input features;

[0036] Perform activation processing on the input features through the tanh activation function and the sigmoid activation function;

[0037] Add the output of the dilated convolution block to the input features through residual connection;

[0038] Accumulate the output results of all dilated convolution blocks, and perform dimension integration on the accumulated results to generate the final vector velocity.

[0039] As a preferred technical solution, training the streaming matching network includes:

[0040] Construct a training set using score data and singing voice data;

[0041] Extract the music score feature parameters including notes, singing voices, and note durations from the music score data, and extract the expressiveness parameters including energy, pitch, and tension from the singing voice data;

[0042] Using the music score feature parameters, expressiveness parameters, phoneme posterior probability map, and timbre features as training labels, and using the training set data, train the streaming matching network by minimizing the objective function Train the streaming matching network.

[0043] As a preferred technical solution, generate a voice Mel spectrogram based on the multi-dimensional singing voice expressiveness sequence, phoneme posterior probability map, and timbre features, including:

[0044] Take the phoneme posterior probability map as the query matrix Q, take the multi-dimensional singing voice expressiveness sequence as the key matrix K and value matrix V, fuse the features of the multi-dimensional singing voice expressiveness sequence and the phoneme posterior probability map through the multi-head attention mechanism to obtain the multi-attention fusion feature, and its expression is as follows:

[0045]

[0046] =softmax ( * / )* ,

[0047] wherein, represents the concatenation operation, represents the attention head j of the attention calculation, is the output weight matrix in the multi-head attention mechanism, represents the dimension of;

[0048] Generate a voice Mel spectrogram based on the multi-attention fusion feature, phoneme posterior probability map, and timbre features.

[0049] As a preferred technical solution, generate a voice Mel spectrogram based on the multi-attention fusion feature, phoneme posterior probability map, and timbre features, including:

[0050] After integrating the dimensions of the multi-attention fusion feature and the phoneme posterior probability map, add them to the timbre feature to obtain the fused input feature matrix;

[0051] Successively pass the input feature matrix through four DITBlocks for feature adjustment processing to obtain a high-dimensional feature matrix, where each DITBlock calculates the scaling parameter through a layer perceptron , With the bias parameter , and perform two-layer AdaLN operations on the input feature matrix according to the following formula for feature adjustment processing:

[0052]

[0053]

[0054] In the formula, x is the input feature matrix;

[0055] Compress the high-dimensional feature matrix through a one-dimensional convolution operation to generate a voice Mel spectrogram.

[0056] As a preferred technical solution, each DITBlock calculates the scaling parameter and the bias parameter through a layer perceptron according to the introduced conditional matrix; among them, the conditional matrix of the first DITBlock is the time-step information matrix, which is used to align the input feature matrix with the time-step information, and the conditional matrices of the latter three DITBlocks are the timbre feature matrices, which are used to perform adaptive adjustment of the timbre features of the input feature matrix;

[0057] Among them, the acquisition steps of the time-step information matrix include:

[0058] Extract the time-step information from the phoneme posterior probability map to construct the time-step information matrix;

[0059] Extract the timbre feature parameters from the timbre features to construct the timbre feature matrix.

[0060] In the second aspect, the present invention also proposes a singing beautification system based on streaming matching, which is applied to the singing beautification method based on streaming matching described in any one of the schemes in the first aspect, including:

[0061] An acquisition module, which is used to acquire singing data and score data;

[0062] An extraction module, which is used to extract timbre features and phoneme posterior probability maps from the singing data;

[0063] A first generation module, which is used to generate a multi-dimensional singing expressiveness sequence according to the score data and the phoneme posterior probability map;

[0064] A second generation module, which is used to generate a voice Mel spectrogram according to the multi-dimensional singing expressiveness sequence, the phoneme posterior probability map and the timbre features;

[0065] A conversion module, which is used to input the voice Mel spectrogram into a vocoder for conversion processing to obtain the beautified singing voice.

[0066] The beneficial effects of the present invention at least include:

[0067] By extracting timbre features and phoneme posterior probability maps, the present invention can accurately capture the basic timbre information and speech content features of the singing voice; by generating a multi-dimensional singing voice expressiveness sequence from the musical score data and the phoneme posterior probability map, the expressiveness parameters of the singing voice are effectively modeled; combining these features to generate a speech Mel spectrogram to ensure the coordination and unity of pitch, timbre, and expressiveness parameters; finally, the vocoder performs conversion processing on the speech Mel spectrogram, so that the output singing voice is optimized in terms of intonation, timbre, and expressiveness, can significantly improve the expressiveness and naturalness of the singing voice, make the generated singing voice of higher quality, smoother to listen to, and more expressive. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 It is a schematic flowchart of a singing voice beautification method based on streaming matching provided by an embodiment of the present invention.

[0069] Figure 2 It is a schematic diagram of generating a multi-dimensional singing voice expressiveness sequence according to musical score data and a phoneme posterior probability map provided by an embodiment of the present invention.

[0070] Figure 3 It is a schematic diagram of the improved WaveNet architecture in an embodiment of the present invention.

[0071] Figure 4 It is a schematic diagram of generating a beautified singing voice according to a multi-dimensional singing voice expressiveness sequence, a phoneme posterior probability map, and timbre features provided by an embodiment of the present invention.

[0072] Figure 5 It is a schematic diagram of a streaming matching decoder based on the DIT architecture provided by an embodiment of the present invention.

[0073] Figure 6 It is a pitch map of the input singing voice provided by an embodiment of the present invention.

[0074] Figure 7 It is a pitch map of the output beautified singing voice provided by an embodiment of the present invention.

[0075] Figure 8 It is an architecture diagram of a singing voice beautification system based on streaming matching provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0076] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred technical solutions. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred technical solutions are only for explaining the present invention, rather than limiting the protection scope of the present invention.

[0077] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0078] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0079] Embodiment 1

[0080] This embodiment proposes a singing beautification method based on streaming matching, as Figure 1 shown. Figure 1 is a schematic flowchart of a singing beautification method based on streaming matching provided by an embodiment of the present invention. The method includes the following steps:

[0081] S1: Obtain singing data and score data;

[0082] S2: Extract timbre features and phoneme posterior probability maps from the singing data;

[0083] S3: Generate a multi-dimensional singing expressiveness sequence according to the score data and the phoneme posterior probability map;

[0084] S4: Generate a speech Mel spectrogram according to the multi-dimensional singing expressiveness sequence, the phoneme posterior probability map, and the timbre features;

[0085] S5: Input the speech Mel spectrogram into a vocoder for conversion processing to obtain the beautified singing voice.

[0086] As an exemplary illustration, in the specific implementation process, taking a segment of a user-recorded singing voice as the input and combining it with the corresponding song score data, first, feature extraction is performed on the input singing voice data to obtain timbre features (such as timbre stability and clarity) and a phoneme posterior probability graph (PPG, a feature graph representing the singing voice content and time step information); then, using the input score data and the phoneme posterior probability graph, a multi-dimensional singing voice expressiveness sequence is generated, including multiple expressiveness dimensions such as pitch, timbre, tension, and energy, ensuring the modeling of the rich expressiveness of the singing voice; next, the generated multi-dimensional singing voice expressiveness sequence, the phoneme posterior probability graph, and the timbre features are fused and processed to generate the corresponding speech Mel spectrogram; finally, the generated speech Mel spectrogram is input into a vocoder (such as the NSF-HIFIGAN vocoder), and combined with the pitch parameter sequence, conversion processing is performed to obtain the beautified singing voice.

[0087] It can be understood that by extracting the timbre features and the phoneme posterior probability graph, the basic timbre information and the speech content features of the singing voice can be accurately captured; by generating a multi-dimensional singing voice expressiveness sequence from the score data and the phoneme posterior probability graph, the expressiveness parameters of the singing voice are effectively modeled; combining these features to generate the speech Mel spectrogram ensures the coordination and unity of pitch, timbre, and expressiveness parameters; finally, the vocoder performs conversion processing on the speech Mel spectrogram, optimizing the output singing voice in terms of pitch accuracy, timbre, and expressiveness, significantly enhancing the expressiveness and naturalness of the singing voice, making the generated singing voice of higher quality, with a smoother listening experience and more expressiveness.

[0088] Embodiment 2

[0089] This embodiment makes improvements on the basis of the singing voice beautification method based on streaming matching proposed in Embodiment 1.

[0090] In this embodiment, the multi-dimensional singing voice expressiveness sequence includes a pitch frame-level sequence, a singing style frame-level sequence, an energy frame-level sequence, and a tension frame-level sequence.

[0091] In this embodiment, according to the score data and the phoneme posterior probability graph, a multi-dimensional singing voice expressiveness sequence is generated, including:

[0092] Construct and train a streaming matching network;

[0093] Perform dimension expansion processing on the phoneme posterior probability graph;

[0094] According to the dimension-expanded phoneme posterior probability graph and the score data, use the trained streaming matching network for feature mapping to generate a pitch frame-level sequence;

[0095] According to the score data and the pitch frame-level sequence, use the trained streaming matching network for feature mapping to generate a singing style frame-level sequence;

[0096] According to the music score data, the pitch frame-level sequence, and the phoneme posterior probability map, use the trained streaming matching network for feature mapping to generate an energy frame-level sequence;

[0097] According to the music score data, the pitch frame-level sequence, the predicted energy frame-level sequence, and the phoneme posterior probability map, use the trained streaming matching network for feature mapping to generate a target tension frame-level sequence.

[0098] In this embodiment, a streaming matching network is constructed, and using the streaming matching network for feature mapping includes:

[0099] Model the mapping relationship between the input distribution and the target distribution, and its expression is as follows:

[0100]

[0101] Among them, represents the distribution of the j th expressiveness parameter at the time step t , represents the target distribution of the j th expressiveness parameter, represents the input distribution of the j th expressiveness parameter;

[0102] Through integral solution, map the input distribution to the target distribution , and take as the output of the final singing expressiveness sequence , and its expression is as follows:

[0103] ,

[0104] Among them:

[0105]

[0106]

[0107]

[0108] In the formula, is the vector velocity obtained by the streaming matching network in the stage of learning the j th expressiveness parameter, is the state distribution at the time step t , represents the model parameters of the streaming matching network in the stage of learning the j th expressiveness parameter, is the input music score data, is the posterior probability map of phonemes, is the posterior probability map of phonemes after dimensionality expansion, , and are all combined multi-dimensional matrices concatenated on their input dimensions.

[0109] As an exemplary illustration, in the specific implementation process, as Figure 2 shown, Figure 2 is the schematic diagram of generating a multi-dimensional singing expressiveness sequence according to the score data and the posterior probability map of phonemes provided by the embodiment of the present invention. The input singing data and score data are used as the initial data sources to extract the timbre characteristics and the posterior probability map of phonemes from the singing data. Through the dimensionality expansion operation, is obtained. The score data is processed by a score encoder to generate score latent features representing information such as melody, style, and pitch range. A streaming matching network is constructed, and corresponding predictors (pitch predictor, singing style predictor, energy predictor, and tension predictor) are designed based on different expressiveness parameters (pitch, singing style, energy, tension). The pitch predictor, singing style predictor, energy predictor, and tension predictor are used to construct

[0110] Among them, the pitch predictor learns the pitch frame-level sequence according to the posterior probability map of phonemes after dimensionality expansion and the score data; the singing style predictor learns the singing style frame-level sequence according to the score data and the pitch frame-level sequence; the energy predictor learns the energy frame-level sequence according to the score data, the pitch frame-level sequence, and the posterior probability map of phonemes; the tension predictor learns the target tension frame-level sequence according to the score data, the pitch frame-level sequence, the predicted energy frame-level sequence, and the posterior probability map of phonemes. The expressiveness parameters are matched with the input data in terms of time steps to ensure a synchronous relationship in the time sequence dimension. In the streaming matching network, through integral solution, the input distribution is mapped to each target distribution to generate the corresponding multi-dimensional expressiveness parameter sequence.

[0111] In this embodiment, when using the streaming matching network for feature mapping, the method further includes predicting the vector velocity of the streaming matching network, including:

[0112] Obtaining the posterior probability map of phonemes, time steps, and expressiveness parameters in real time;

[0113] Extracting score features including rhythm, style, and pitch range from the score data;

[0114] Encoding the time steps through a sine position encoding module to generate position encoding information;

[0115] The music score features, phoneme posterior probability maps, and positional encoding information are used as input features and input into the dilated convolution module, where the dilated convolution module consists of multiple dilated convolution blocks, and each dilated convolution block performs the following operations:

[0116] Perform a dilated convolution operation on the input features;

[0117] Activate the input features through the tanh activation function and the sigmoid activation function;

[0118] Add the output of the dilated convolution block to the input features through a residual connection;

[0119] Accumulate the output results of all dilated convolution blocks, and integrate the accumulated results in terms of dimensions to generate the final vector velocity.

[0120] As an exemplary illustration, this embodiment uses a modified WaveNet architecture to predict the vector velocity, as Figure 3 shown, Figure 3 is the schematic diagram of the modified WaveNet architecture in the embodiment of the present invention. In the specific implementation process, the phoneme posterior probability map is processed by the first convolutional network (Conv1d) to generate a high-dimensional PPG feature matrix [B, 256, 16000]. The input time step t is processed by the sine positional encoding module (SinPosEmb) and the multi-layer linear network (Linear) to map the time step features to time encoding information of [B, 256, 1]. The input music score features and expressiveness parameters input the preprocessed PPG features, time step features, and music score expressiveness features into the dilated convolution module through a stacking operation (stack). The dilated convolution module consists of 20 layers of dilated convolution (dilated_conv) and gated activation units (gatedactivation). Each layer of dilated convolution realizes multi-scale time feature extraction and matching through different dilation rates, and obtains a feature matrix of [B, 512, 16000]. The features output by each layer of dilated convolution are accumulated through an addition operation to achieve layer-by-layer feature fusion, forming a music score expressiveness feature matrix of [B, 320, 16000].

[0121] Among them, the gated activation unit , uses the tanh function and sigmoid function to multiply. l is the length of the number of input channels. The first half of the input channels is used as the tanh input, and the second half is used as the sigmoid input. tanh The function is the hyperbolic tangent function, tanh and sigmoid both belong to saturation functions. Through the tanh function and sigmoid function, the problem of gradient disappearance can be solved.

[0122] The final output of the dilated convolution module is processed by a second convolutional network (Conv1d) to generate a feature matrix of [B, 256, 16000]. Multiple feature matrices are added together (Sum) to achieve the final fusion of features and dimension integration. Finally, the fused features are further processed through the output channels to generate the vector velocity.

[0123] In this embodiment, training the streaming matching network includes:

[0124] Constructing a training set using music score data and singing voice data;

[0125] Extracting music score feature parameters including notes, singing styles, and note durations from the music score data, and extracting expressive force parameters including energy, pitch, and tension from the singing voice data;

[0126] Using the music score feature parameters, expressive force parameters, phoneme posterior probability map, and timbre features as training labels, and using the training set data to train the streaming matching network by minimizing the objective function Training the streaming matching network.

[0127] As an exemplary illustration, for the extraction of energy, the root mean square of the speech data is directly calculated frame by frame. For example, if the speech signal is a discrete signal x[n] and the length of each frame after framing is N, the formula for the root mean square energy of the k-th frame is: . Where is the root mean square energy of the k-th frame, is the speech signal sample value of the k -th frame, N is the number of samples in each frame, n is the sample index.

[0128] For the extraction of pitch, a pitch extraction neural network is used to directly extract it from the audio by inputting the speech.

[0129] For the extraction of tension, it can be obtained according to the formula , where r is a ratio, and its calculation formula is . In the formula, is the total harmonic, is the half harmonic, RMS is the root mean square of the data.

[0130] In this embodiment, according to the multi-dimensional singing voice expressiveness sequence, phoneme posterior probability map, and timbre features, generating the speech Mel spectrogram includes:

[0131] Taking the phoneme posterior probability map as the query matrix Q, the multi-dimensional singing expressiveness sequence as the key matrix K and value matrix V, the features of the multi-dimensional singing expressiveness sequence and the phoneme posterior probability map are fused through the multi-head attention mechanism to obtain the multi-attention fusion features, and its expression is as follows:

[0132]

[0133] =softmax ( * / )* ,

[0134] where, represents the concatenation operation, represents the attention head j of the attention calculation, is the output weight matrix in the multi-head attention mechanism, represents the dimension of, represents the square root of the dimension of to prevent the formula value from being too large;

[0135] According to the multi-attention fusion features, the phoneme posterior probability map and the timbre features, a speech Mel spectrogram is generated.

[0136] In this embodiment, generating a speech Mel spectrogram according to the multi-attention fusion features, the phoneme posterior probability map and the timbre features includes:

[0137] After integrating the dimensions of the multi-attention fusion features and the phoneme posterior probability map, adding them to the timbre features to obtain a fused input feature matrix;

[0138] The input feature matrix is sequentially processed through four DITBlocks for feature adjustment to obtain a high-dimensional feature matrix, where each DITBlock calculates the scaling parameter , and the bias parameter , and performs two-layer AdaLN operations on the input feature matrix according to the following formula for feature adjustment processing:

[0139]

[0140]

[0141] In the formula, x is the input feature matrix;

[0142] The high-dimensional feature matrix is dimensionally compressed through a one-dimensional convolution operation to generate a voice Mel spectrogram.

[0143] In this embodiment, each DITBlock calculates the scaling parameter and the bias parameter through a layer perceptron according to the introduced conditional matrix; among them, the conditional matrix of the first DITBlock is the time step information matrix, which is used to align the input feature matrix with the time step information, and the conditional matrices of the latter three DITBlocks are the timbre feature matrices, which are used to adaptively adjust the timbre features of the input feature matrix;

[0144] Among them, the steps for obtaining the time step information matrix include:

[0145] Extract the time step information from the phoneme posterior probability map to construct the time step information matrix;

[0146] Extract the timbre feature parameters from the timbre features to construct the timbre feature matrix.

[0147] As an exemplary illustration, as Figure 4 shown, Figure 4 is the schematic diagram for generating the beautified singing voice according to the multi-dimensional singing expressiveness sequence, phoneme posterior probability map and timbre features provided by the embodiment of the present invention. In the specific implementation process, the phoneme posterior probability map and the multi-dimensional expressiveness sequence are input into the expressiveness multi-head attention mechanism. There are a total of four heads in the multi-head attention mechanism, namely the singing style head, the energy head, the tension head and the pitch head. In the process of generating the voice Mel spectrogram, a streaming matching decoder based on the DIT architecture is used, as Figure 5 shown, Figure 5 is the schematic diagram of the streaming matching decoder based on the DIT architecture provided by the embodiment of the present invention. The timbre features ([B, 160, 1]) are processed through a Linear layer to generate an intermediate feature matrix with a dimension of [B, 256, 1]. The phoneme posterior probability map and the multi-attention fusion features are concatenated into an input matrix ([B, 160, 16000]), and features are extracted through the Conv1d layer, and the dimension of the output matrix is [B, 256, 16000]. The time step information ([B, 1]) is encoded through SinPosEmb (sine position encoding), and the dimension of the output matrix is [B, 256, 1].

[0148] After dimensional integration, the input matrix and the timbre representation are added and input into four DITBlocks. The structure of each DITBlock is the same. First, six parameters are generated by the conditional matrix through a multi-layer perceptron MLP. , , for AdaLN operations in the DIT network, for the AdaLN calculation of For the AdaLN calculation of This DITBlock can embed the information of the conditional matrix into the input matrix. In the DIT streaming matching, the conditional matrix of the first DITBlock is the time step t , used to embed the time step information. The conditional matrices of the next three consecutive DITBlocks are timbre characterizations, used to embed timbre information and enhance timbre similarity.

[0149] The vocoder selects the NSF-HIFIGAN vocoder that supports pitch input. Compared with ordinary vocoders, it has more information in the pitch dimension and is more suitable for singing. The NSF-HIFIGAN vocoder synthesizes highly expressive singing speech containing pitch, singing style, energy, and tension according to the input Mel spectrogram and combines it with the pitch parameter sequence.

[0150] As an exemplary illustration, as Figure 6 shown, Figure 6This is the pitch map of the input singing voice provided by the embodiments of the present invention. In the specific implementation process, first, the singing voice input and the music score input are received, and the PPG, timbre characterization, and note sequence therein are extracted. Taking a music score sung by a user as an example, the notes are C4, D4, E4, G4, Am4, F4, C4, and G4, and the duration of each note is 0.7 seconds. This melody starts from C4, gradually rises to the high point G4, and then slightly descends. The overall melody line is gentle and warm. The content of the singing voice input by the user is "How deep is my love for you". Without changing the content of the singing voice, by analyzing the PPG and the note sequence, it is input into the singing voice expressiveness generator to generate a multi-dimensional expressiveness sequence including pitch, singing style, energy, and tension. In the specific operation, dynamic prediction of the pitch is performed. At the beginning of the singing voice, an upward glide is predicted to express the emotion of telling a story; at the climax part of the melody (G4), a vibrato is predicted to show the high and tense emotional expression; at the end of the song, a downward glide is predicted to create a warm and gentle emotional ending effect. At the same time, the singing style is carefully modeled. In the low-pitch segment, it is more inclined to chest voice pronunciation to show a thick feeling; in the high-pitch segment, it is more inclined to head voice pronunciation to show the singer's difficult singing skills. In addition, the energy output is dynamically adjusted according to different parts of the melody. The energy is enhanced in the high-pitch part to reflect the strength, while the energy is appropriately reduced in the low-pitch part to appear gentle. For the tension parameter, the system increases the tension in the part with deep emotion to further express the deeper emotional expression. Then, the time relationship between the expressiveness parameters and the PPG is modeled using a streaming matching network and an attention mechanism to ensure the consistency of the expressiveness parameters and the singing voice content in the time dimension. This modeling method enables the system to automatically adjust the expressiveness, making the generated singing voice not only accurate in pitch but also more natural and fluent in emotional expression.

[0151] As Figure 6 shown, in the pitch map of the input singing voice, the melody line is relatively flat and lacks skillful performance. As Figure 7 shown, Figure 7 This is the pitch map of the output beautified singing voice provided by the embodiments of the present invention. From Figure 7 it can be seen that the pitch curve is significantly more layered. By adding singing skills such as vibrato and glide, the singing voice becomes more vivid and expressive. Using the high-quality open-source dataset OpenCpop as the training and test set, the average duration of each music score is about 3 minutes. The test results show that the beautified singing voice generated by the system performs excellently in terms of pitch accuracy, and the out-of-tune parts do not exceed 5%. Moreover, these out-of-tune parts are mainly the turns generated by the system, rather than prediction errors.

[0152] In summary, without changing the content of the singing voice, the present invention dynamically predicts expressiveness parameters such as pitch, singing style, energy, and tension, and combines a streaming matching network to model the temporal relationship, generating a beautified singing voice with accurate pitch and rich expressiveness. This method effectively solves the limitation of traditional singing voice beautification techniques that only target pitch correction and lack expressiveness, providing users with a singing voice output with more emotional expression and professional singing skills.

[0153] Embodiment 3

[0154] As Figure 8 shown, this embodiment proposes a singing voice beautification system based on streaming matching, which is applied to the singing voice beautification method based on streaming matching as described in the above embodiment, and includes: an acquisition module 100, an extraction module 200, a first generation module 300, a second generation module 400, and a conversion module 500.

[0155] Among them, the acquisition module 100 is used to acquire singing voice data and score data; the extraction module 200 is used to extract timbre features and phoneme posterior probability maps from the singing voice data; the first generation module 300 is used to generate a multi-dimensional singing voice expressiveness sequence according to the score data and the phoneme posterior probability map; the second generation module 400 is used to generate a voice Mel spectrogram according to the multi-dimensional singing voice expressiveness sequence, the phoneme posterior probability map, and the timbre features; the conversion module 500 is used to input the voice Mel spectrogram into a vocoder for conversion processing to obtain a beautified singing voice.

[0156] It should be noted that the foregoing explanation of the embodiment of the singing voice beautification method based on streaming matching also applies to the singing voice beautification system based on streaming matching in this embodiment, and will not be elaborated here.

[0157] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or N embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0158] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0159] Any process or method description shown in the flowchart or described otherwise herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations where functions may be executed in a substantially simultaneous manner or in an order opposite to that shown or discussed, according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0160] It should be understood that various parts of the present invention may be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods may be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art may be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays, field programmable gate arrays, etc.

[0161] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0162] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not limitations on the embodiments of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the embodiments here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A singing voice beautification method based on streaming matching, characterized in that: include: Acquire singing data and music score data; Extracting timbre features and phoneme posterior probability maps from singing voice data; According to the score data and the phoneme posterior probability map, a multi-dimensional singing expression sequence including a pitch frame-level sequence, a singing voice frame-level sequence, an energy frame-level sequence and a tension frame-level sequence is generated, including: Build and train a streaming matching network; Perform dimension expansion processing on the phoneme posterior probability map; Based on the dimensionally expanded phoneme posterior probability map and music score data, the trained streaming matching network is used to perform feature mapping to generate a pitch frame-level sequence. According to the score data and pitch frame-level sequence, the trained streaming matching network is used to perform feature mapping to generate the singing frame-level sequence; According to the score data, pitch frame-level sequence and phoneme posterior probability map, the trained streaming matching network is used to perform feature mapping to generate energy frame-level sequence; Based on the score data, pitch frame-level sequence, energy frame-level sequence prediction and phoneme posterior probability map, the trained streaming matching network is used for feature mapping to generate tension frame-level sequence; Generate speech mel spectrogram based on multi-dimensional singing expressiveness sequence, phoneme posterior probability map and timbre characteristics; The speech Mel spectrum is input into the vocoder for conversion processing to obtain the beautified singing voice.

2. The method for beautifying singing voice based on streaming matching according to claim 1, characterized in that: Construct a streaming matching network and use it for feature mapping, including: The mapping relationship between modeling input distribution and target distribution is expressed as follows: in, Indicates j The expressiveness parameter is in the time step t The distribution of time, Indicates j The target distribution of expressive parameters, Indicates the j The input distribution of the expressive parameters; By integral solution, the input distribution Mapping to target distribution ,Will Output as the final vocal expressive sequence , whose expression is as follows: , in: In the formula, For the streaming matching network in learning j The vector velocity obtained in the expressive parameter stage, For the time step t The state distribution of Indicates that the streaming matching network is learning j Model parameters for expressiveness parameters, is the input score data, is the phoneme posterior probability map, It is the phoneme posterior probability map after dimension expansion.

3. The method for beautifying singing voice based on streaming matching according to claim 2, characterized in that: When the streaming matching network is used for feature mapping, the method also includes predicting a vector velocity of the streaming matching network, including: Get phoneme posterior probability maps, time steps, and expressiveness parameters in real time; Extracting music score features including rhythm, style and range from music score data; The time step is position-encoded by a sinusoidal position encoding module to generate position encoding information; The music score features, phoneme posterior probability map and position encoding information are input as input features into the dilated convolution module, where the dilated convolution module consists of multiple dilated convolution blocks, each of which performs the following operations: Perform dilated convolution operation on the input features; The input features are activated through the tanh activation function and the sigmoid activation function; The output of the dilated convolutional block is added to the input features through a residual connection; The output results of all dilated convolution blocks are accumulated and dimensionally integrated to generate the final vector velocity.

4. The method for beautifying singing voice based on streaming matching according to claim 2, characterized in that: Training a streaming matching network, including: Use music score data and singing data to build a training set; Extracting music score feature parameters including notes, singing style and note duration from the music score data, and extracting expressive parameters including energy, pitch and tension from the singing voice data; The music score feature parameters, expressiveness parameters, phoneme posterior probability maps and timbre features are used as training labels, and the training set data is used to minimize the objective function. Train a streaming matching network.

5. The method for beautifying singing voice based on streaming matching according to claim 2, characterized in that: Generate speech mel spectrogram based on multi-dimensional singing expressiveness sequence, phoneme posterior probability map and timbre characteristics, including: The phoneme posterior probability map is used as the query matrix Q, and the multi-dimensional singing voice expressiveness sequence is used as the key-value matrix K and V. The features of the multi-dimensional singing voice expressiveness sequence and the phoneme posterior probability map are fused through the multi-head attention mechanism to obtain the multi-attention fusion feature, which is expressed as follows: =softmax ( * / )* , in, Represents a splicing operation, Represents the attention head j Attention calculation, is the output weight matrix in the multi-head attention mechanism, express Dimensions; Generate speech mel-spectrogram based on multi-attention fusion features, phoneme posterior probability map and timbre features.

6. The method for beautifying singing voice based on streaming matching according to claim 5, characterized in that: Generate speech mel spectrum based on multi-attention fusion features, phoneme posterior probability map and timbre features, including: After dimensional integration of the multi-attention fusion features and the phoneme posterior probability map, they are added to the timbre features to obtain the fused input feature matrix; The input feature matrix is ​​processed by four DITBlocks in turn to obtain a high-dimensional feature matrix, where each DITBlock calculates the scaling parameters through a layer perceptron. , With bias parameters , and perform two layers of AdaLN operations on the input feature matrix to perform feature adjustment according to the following formula: In the formula, x is the input feature matrix; The high-dimensional feature matrix is ​​compressed through a one-dimensional convolution operation to generate a speech mel spectrum.

7. The method for beautifying singing voice based on streaming matching according to claim 6, characterized in that: Each DITBlock calculates the scaling parameters and bias parameters through the layer perceptron according to the introduced conditional matrix; the conditional matrix of the first DITBlock is the time step information matrix, which is used to align the input feature matrix with the time step information, and the conditional matrices of the last three DITBlocks are the timbre feature matrix, which is used to adaptively adjust the timbre features of the input feature matrix; Among them, the steps of obtaining the time step information matrix include: Extract time step information from the phoneme posterior probability map and construct a time step information matrix; The timbre feature parameters are extracted from the timbre features and a timbre feature matrix is ​​constructed.

8. A singing voice beautification system based on streaming matching, applied to the singing voice beautification method based on streaming matching as claimed in any one of claims 1 to 7, characterized in that it comprises: An acquisition module is used to acquire singing data and music score data; An extraction module, used for extracting timbre features and phoneme posterior probability maps from singing data; The first generation module is used to generate a multi-dimensional singing voice expressiveness sequence according to the music score data and the phoneme posterior probability map; The second generation module is used to generate speech mel spectrum according to the multi-dimensional singing expressiveness sequence, the phoneme posterior probability map and the timbre characteristics; The conversion module is used to input the speech Mel spectrum into the vocoder for conversion processing to obtain the beautified singing voice.

Citation Information

Patent Citations

  • Speech synthesis method and related device and equipment

    CN114220414A

  • Singing synthesis method, computer equipment and storage medium

    CN117238273A