Huangmuo singing cloning method and system, storage medium and electronic equipment

The synthesis of the Huangmei Opera singing meer spectrum through the cloning model has solved the problem that singing voice cloning in the existing technology is difficult to accurately capture the timbre details, and the personalized customization of the high-fidelity Huangmei Opera singing voice timbre is realized, which improves the naturalness and emotional expression of singing synthesis.

CN120340445APending Publication Date: 2025-07-18HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510661840.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing voice cloning technology is difficult to accurately capture and reproduce the unique tone details of each singer, resulting in a lack of personality and recognition of the generated voice, especially in the Huangmei Opera singing style, the synthesis results are insufficient in fluency, rhythm and emotional expression.

Method used

The cloned model is used to synthesize the target singer's Huangmei Opera singing mel spectrum. Through the speaker encoder, Mel style encoder, music score encoder, duration predictor, length regulator, pitch diffuser and decoder, the details such as tremolo, drag cavity and tone are accurately captured to generate high-fidelity Huangmei Opera singing audio.

Benefits of technology

It realizes accurate capture and reproduction of the tone of each singer, and improves the personalized customization ability of synthesized singing, especially in the tone performance of singers outside the domain, meeting the requirements of Huangmei Opera for delicate tone and emotional communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340445A_ABST
    Figure CN120340445A_ABST
Patent Text Reader

Abstract

The invention provides a Huangmuo singing cloning method and system, a storage medium and electronic equipment, and relates to the technical field of singing cloning. And synthesizing the Mel spectrum of the Huangmuo singing cavity of the target singer through a clone model, wherein the clone model comprises a speaker encoder, a Mel style encoder, a music score encoder, a time length predictor, a length regulator, a pitch diffuser and a decoder. According to the method, on the basis of the unique singing cavity characteristics of the Huangmuo, the advanced singing synthesis and cloning technology is utilized, details such as tremolo, drag cavities and timbres are accurately captured through deep learning, high-fidelity reproduction of the singing cavity and the timbres of the Huangmuo is achieved, and meanwhile a personalized customization scheme is provided for professional deduction and enthusiasts. Therefore, inheritance of traditional art is facilitated, and a new path is provided for digital culture creativity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of singing voice cloning technology, and particularly relates to a Huangmei Opera singing voice cloning method, system, storage medium and electronic device. Background Art

[0002] With the wide application of singing voice cloning technology in the music creation and entertainment industries, these technologies have gradually matured, providing new ideas for audio processing and intelligent creation.

[0003] However, the existing synthesized audio has insufficient timbre similarity, and it is often difficult to accurately capture and reproduce the unique timbre details of each singer, resulting in the generated singing voice lacking personality and recognition. Summary of the Invention

[0004] (1) Technical Problems to be Solved

[0005] Aiming at the deficiencies of the existing technology, the present invention provides a Huangmei Opera singing voice cloning method, system, storage medium and electronic device, which solves the technical problem that the existing singing voice cloning technology is difficult to accurately capture and reproduce the unique timbre details of each singer.

[0006] (2) Technical Solutions

[0007] To achieve the above object, the present invention is realized through the following technical solutions:

[0008] In the first aspect, the present invention provides a Huangmei Opera singing voice cloning method, which synthesizes the Mel spectrogram of the Huangmei Opera singing voice of a target singer through a cloning model. The cloning model includes a speaker encoder, a Mel style encoder, a score encoder, a duration predictor, a length regulator, a pitch diffuser and a decoder; the Huangmei Opera singing voice cloning method includes:

[0009] Obtain a Huangmei Opera singing sample and corresponding score data, as well as the original audio of the target singer, and perform preprocessing on the Mel spectrogram of the original audio of the target singer to obtain the original Mel spectrogram of the target singer;

[0010] Model the singing style of the original Mel spectrogram of the target singer through the Mel style encoder to obtain a one-dimensional style vector; process the score data and the style vector through the score encoder to obtain a set of feature representations related to speech generation; process the original speech signal of the target singer through the speaker encoder to obtain the voiceprint feature of the target singer;

[0011] Concatenate the feature representations related to speech generation and the voiceprint feature of the target singer. The concatenated features are processed by a duration predictor, a length regulator and a pitch diffuser. The processed features are concatenated with the style vector and subjected to positional encoding to obtain acoustic features;

[0012] Decode the acoustic features through a decoder to generate the Mel spectrogram of the Huangmei Opera singing voice of the target singer.

[0013] Preferably, the speaker encoder adopts a generalized end-to-end model.

[0014] Preferably, the Mel style encoder includes a spectrum processing module, a temporal modeling module, and a global style modeling module;

[0015] Among them,

[0016] The spectrum processing module is used to send the input original Mel spectrogram into a fully connected layer, and through two layers of fully connected layers, convert each frame of the Mel spectrogram into a hidden sequence;

[0017] The temporal modeling module is used to connect the hidden sequence to a subsequent gated convolutional neural network to capture the temporal information in the original Mel spectrogram;

[0018] The global style modeling module is used to utilize the multi-head attention mechanism combined with residual connections for the output result of the temporal modeling module to extract global style information, and perform modeling at the Mel spectrogram frame level; and perform temporal average processing on the output of the multi-head attention to obtain a one-dimensional style vector.

[0019] Preferably, the score encoder includes a Branchformer module, and the Branchformer module includes a global feature extraction branch based on multi-head attention and a local feature extraction branch based on a convolutional gated multi-layer perceptron.

[0020] Preferably, the decoder adopts a flow matching decoder.

[0021] Preferably, the loss function L CFM (θ) in the training process of the flow matching decoder includes:

[0022]

[0023] Among them, t represents the time step, q(x1) represents the target distribution, and p t (x|x1) represents the conditional probability path; v t (x; θ) represents the vector field output by the neural network controlled by the parameter θ; u t (x|x1) represents the target vector field; x1 represents the Mel spectrogram feature frame; x represents the Mel spectrogram.

[0024] Preferably, the Huangmei Opera singing voice cloning method further includes: processing the Mel spectrogram of the Huangmei Opera singing voice of the target singer through a neural vocoder to obtain the Huangmei Opera singing voice audio of the target singer.

[0025] Second aspect, the present invention provides a Huangmei Opera singing voice cloning system, which synthesizes the Mel spectrogram of the Huangmei Opera singing voice of the target singer through a cloning model. The cloning model includes a speaker encoder, a Mel style encoder, a music score encoder, a duration predictor, a length regulator, a pitch diffuser and a decoder; the Huangmei Opera singing voice cloning system includes:

[0026] A data acquisition and processing module, configured to acquire Huangmei Opera singing samples and corresponding music score data, as well as the original audio of the target singer, and preprocess the original audio of the target singer to obtain the original Mel spectrogram of the target singer;

[0027] An encoding module, configured to model the singing style of the original Mel spectrogram of the target singer through a Mel style encoder to obtain a one-dimensional style vector; process the music score data and the style vector through a music score encoder to obtain a set of feature representations related to speech generation; process the original speech signal of the target singer through a speaker encoder to obtain the voiceprint feature of the target singer;

[0028] An acoustic feature acquisition module, configured to splice the feature representations related to speech generation and the voiceprint feature of the target singer. The spliced features are processed by a duration predictor, a length regulator and a pitch diffuser, and the processed features are spliced with the style vector and subjected to positional encoding to obtain acoustic features;

[0029] A decoding module, configured to decode the acoustic features through a decoder to generate the Mel spectrogram of the Huangmei Opera singing voice of the target singer.

[0030] Third aspect, the present invention provides a computer-readable storage medium, which stores a computer program for Huangmei Opera singing voice cloning. Wherein, the computer program enables a computer to execute the Huangmei Opera singing voice cloning method as described above.

[0031] Fourth aspect, the present invention provides an electronic device, including:

[0032] One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors. The programs include those for executing the Huangmei Opera singing voice cloning method as described above.

[0033] (III) Beneficial effects

[0034] The present invention provides a Huangmei Opera singing voice cloning method, system, storage medium and electronic device. Compared with the prior art, the following beneficial effects are achieved:

[0035] Compared with the existing technologies, the Huangmei Opera singing voice cloning method provided by the present invention can accurately capture and reproduce the unique timbre details of each singer, shows better performance for the timbres of singers not seen outside the domain, and can provide personalized customization solutions for Huangmei Opera enthusiasts. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0037] Figure 1 It is the overall architecture diagram of the cloning model in the embodiment of the present invention;

[0038] Figure 2 It is the calculation process of the similarity matrix in the speaker encoder;

[0039] Figure 3 It is the structural schematic diagram of the Mel-style encoder;

[0040] Figure 4 It is the structural schematic diagram of the Branchformer module in the music score encoder;

[0041] Figure 5 It is the structural schematic diagram of the flow prediction network;

[0042] Figure 6 It is the schematic diagram of the flow change process. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are clearly and completely described below. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0044] By providing a Huangmei Opera singing voice cloning method, system, storage medium, and electronic device in the embodiments of the present application, the technical problem that the existing singing voice cloning technology is difficult to accurately capture and reproduce the unique timbre details of each singer is solved, and stronger generalization ability and higher timbre reduction degree are achieved.

[0045] The overall idea of the technical solutions in the embodiments of the present application to solve the above technical problems is as follows:

[0046] With the wide application of singing voice cloning technology in music creation and the entertainment industry, it provides new ideas for audio processing and intelligent creation. However, the synthesized audio of existing singing voice cloning technology lacks sufficient similarity in timbre, often making it difficult to accurately capture and reproduce the unique timbre details of each singer, resulting in the generated singing voices lacking personality and recognition. At the same time, when applying existing cloning models to Huangmei Opera singing, the synthesized results have obvious deficiencies in fluency, rhythm, and emotional expression, with a subjective listening feeling being rather rigid and difficult to meet the strict requirements of Huangmei Opera for delicate timbre and emotional conveyance.

[0047] To solve the above problems, embodiments of the present invention are based on the unique singing characteristics of Huangmei Opera, utilize advanced singing voice synthesis and cloning technology, accurately capture details such as vibrato, sustained notes, and timbre through deep learning, achieve high-fidelity reproduction of Huangmei Opera singing and timbre, and at the same time provide personalized customization solutions for professional performances and enthusiasts. This not only helps the inheritance of traditional art but also provides a new path for digital cultural creativity.

[0048] To better understand the above technical solution, the above technical solution will be described in detail below in combination with the accompanying drawings of the specification and specific implementation manners.

[0049] Embodiments of the present invention provide a method for cloning Huangmei Opera singing. This method synthesizes the Mel spectrogram of the Huangmei Opera singing of a target singer through a cloning model, as Figure 1 shown. The cloning model includes a speaker encoder, a Mel style encoder, a score encoder, a duration predictor, a length regulator, a pitch diffuser, and a decoder. The specific data processing process is as follows:

[0050] S1. Obtain Huangmei Opera singing samples and corresponding score data, as well as the original audio of the target singer, and perform preprocessing on the original audio of the target singer to obtain the original Mel spectrogram of the target singer;

[0051] S2. Model the singing style of the original Mel spectrogram of the target singer through the Mel style encoder to obtain a one-dimensional style vector; process the score data and the style vector through the score encoder to obtain a set of feature representations related to speech generation; process the original speech signal of the target singer through the speaker encoder to obtain the voiceprint feature of the target singer;

[0052] S3. Concatenate the feature representations related to speech generation and the voiceprint feature of the target singer. The concatenated features are processed by the duration predictor, the length regulator, and the pitch diffuser. The processed features are concatenated with the style vector and subjected to positional encoding to obtain acoustic features;

[0053] S4. Decode the acoustic features through the decoder to generate the Mel spectrogram of the Huangmei Opera singing of the target singer.

[0054] Compared with the prior art, the Huangmei Opera singing voice cloning method provided by the embodiments of the present invention can accurately capture and reproduce the unique timbre details of each singer, and also shows better performance for the timbres of singers not seen outside the domain, and can provide personalized customization solutions for Huangmei Opera enthusiasts.

[0055] It should be noted that although the embodiments of the present invention are a Huangmei Opera singing voice cloning method, in fact, it can also be applied to the cloning of singing voices of various operas such as Peking Opera, Cantonese Opera, and Yu Opera, as long as the singing samples and music scores are adaptively adjusted. Of course, it can also be applied to the cloning of pop song singing voices.

[0056] The following is a detailed description of each step:

[0057] In step S1, obtain the Huangmei Opera singing samples and the corresponding music score data, as well as the original audio of the target singer, and perform preprocessing of the Mel spectrogram on the original audio of the target singer to obtain the original Mel spectrogram of the target singer. Specifically, it includes:

[0058] It should be noted that the Huangmei Opera singing samples are Huangmei Opera singing samples marked with the dragging tune marks.

[0059] The preprocessing process of generating the original Mel spectrogram of the target singer specifically includes pre-emphasis, framing, windowing, Fourier transform, etc.

[0060] In step S2, based on the Mel style encoder, model the singing style of the original Mel spectrogram of the target singer to obtain a one-dimensional style vector; based on the music score encoder, process the music score data and the style vector to obtain a set of feature representations closely related to speech generation; based on the speaker encoder, process the original speech signal of the target singer to obtain the voiceprint feature of the target singer. Specifically, it includes:

[0061] It should be noted that the above speaker encoder is pre-trained separately to extract the identity information of the speaker. The Mel style encoder, the music score encoder, and the subsequent decoder are trained together.

[0062] In the embodiments of the present invention, in order to extract stable speaker features, the speaker encoder adopts a Generalized End to End (GE2E) model, which uses a generalized end-to-end loss to optimize the parameters and can directly learn stable and highly discriminative speaker feature vectors from the audio.

[0063] During the training process of the speaker encoder, a data batch needs to be constructed. Assuming the number of speakers is N and M voice segments are taken from each speaker, an N×M sentence can be formed. Use the LSTM neural network to extract multiple sentence embeddings e from these sentences ij , where the subscript i represents the i-th speaker and the subscript j represents the j-th voice segment of this speaker. In this way, the corresponding speaker embedding c can be calculated according to the sentence embedding e ij , and the calculation formula is: i

[0064]

[0065] Subsequently, calculate the cosine similarity between each speaker embedding c k and all sentence embeddings e ij , a similarity matrix can be obtained, and the calculation formula for each value in the similarity matrix is:

[0066] S i,j,k =w·cos(e ij ,c k )+b

[0067] where w and b are trainable parameters used to adjust the similarity scaling and translation.

[0068] The process of calculating the similarity matrix is as shown in the appendix Figure 2 . It can be clearly seen from the figure that the audio data undergoes feature extraction, the LSTM model calculates the embedding vector, and then the similarity matrix is obtained. If the sentence and the speaker can be matched, that is, the colored part in the figure, the calculated cosine similarity is larger at this time. If the sentence and the speaker do not match, that is, the gray part in the figure, the calculated cosine similarity will be smaller.

[0069] The purpose of the Mel-style encoder is to extract a style vector s that represents the style features from the reference speech V. This vector contains style attributes such as the speaker identity and prosody of the speech. The structure of the Mel-style encoder is as shown in the appendix Figure 3 and is composed of the following three major modules.

[0070] Spectrum processing module: Send the original Mel spectrum of the target singer as input to the fully connected layer. Through two layers of fully connected layers, each frame of the Mel spectrogram is converted into a hidden sequence;

[0071] Time modeling module: Connect the hidden sequence to the subsequent gated convolutional neural network, which can capture the temporal information in the original Mel spectrum of the target singer and enhance the modeling ability for prosody features.

[0072] ​Global Style Modeling Module: The output of the time modeling module is used with a multi-head attention mechanism combined with residual connections to extract global style information, and the modeling is performed at the Mel spectrogram frame level, which can ensure that the encoder can effectively extract style information even in short speech samples. Finally, temporal averaging is performed on the output of the multi-head attention to obtain a one-dimensional style vector s.

[0073] The score encoder encodes score information including notes, durations, phonemes, etc. To better enhance the model's ability to capture and modulate style features, a Style-Adaptive Layer Normalization (SALN) is introduced in the Branchformer module of the score encoder to normalize the outputs of the two branches respectively. The structure of the Branchformer module is as Figure 4 shown. Specifically, in the two branches of the global feature extraction branch based on multi-head attention and the local feature extraction branch based on Convolutional Gated Multi-Layer Perceptron (CGMLP), style-adaptive normalization (SALN) is performed using the features extracted by the Mel style encoder. For a given feature vector h = (h1, h2,..., h H ), the normalized vector y = (y1, y2,..., y H ) can be derived. The specific formula for normalization is as follows:

[0074]

[0075] SALN(h, s) = g(s)·y + b(s)

[0076] where H is the vector dimension, g(s) and b(s) are the adaptive parameters obtained by passing the style vector s through a fully connected layer, μ represents the mean, and σ represents the variance.

[0077] Through SALN, style information can be effectively injected into different branches, adaptively adjusting the feature distribution, which can better learn the style characteristics of the singing voice and improve the style consistency and naturalness of the synthesized singing voice.

[0078] In step S3, the feature representation closely related to speech generation and the voiceprint feature of the target singer are concatenated. The concatenated features are processed by a duration predictor, a length regulator, and a pitch diffuser. The features obtained after processing are concatenated with the style vector and passed through positional encoding to obtain acoustic features. The specific implementation process is as follows:

[0079] The features closely related to speech generation are concatenated with the voiceprint features of the target singer. The concatenated features are used by a Duration Predictor to predict the duration of phonemes to match the real pronunciation. The prediction conditions are expanded by a Length Regulator, and then the expanded information is input into a Pitch Diffusion Prediction to predict the pitch, outputting a feature. The output feature is concatenated with a style vector, and the concatenated feature undergoes positional encoding to obtain an acoustic feature containing information such as phonemes, duration, and pitch.

[0080] In step S4, the acoustic feature is decoded by a decoder to generate the Mel spectrogram of the Huangmei Opera singing of the target singer. Specifically, it includes:

[0081] It should be noted that in the embodiment of the present invention, a Conditional flowmatching Decoder (CFM decoder) is adopted, and its training process is mainly as follows:

[0082] Assume that x represents the Mel spectrogram sample in the data space R d and follows an unknown data distribution q(x). The probability density path is defined as a time-dependent probability density function p t :[0,1]×R d →R d >0. Its goal is to construct an evolution path from a simple prior distribution such as the Gaussian distribution p0(x) = N(x; 0, I) to the target distribution p1(x), so as to achieve efficient sampling of the target distribution. Here, p1(x) is defined as an approximation of q(x), i.e., p1(x)≈q(x), and its true trajectory is approximated by adding a small amount of white noise perturbation. The continuous normalizing flow CNF defines the evolution process through a vector field v t :[0,1]×R d →R d and is specifically described by an Ordinary Differential Equation (ODE), i.e., the formula.

[0083]

[0084] where φ t is called the flow mapping, representing the evolution result of the Mel spectrogram x at time t. The vector field v t controls the direction and speed of the evolution and needs to be learned through training. The corresponding flow prediction network is shown in Appendix Figure 5 .

[0085] By solving the initial value problem of this ODE, the flow mapping φt Gradually map the initial distribution \(p_0(x)\) to the target distribution \(p_1(x)\), where \(p\) t (x) is the sample marginal distribution under the action of \(\varphi\) t .

[0086] Suppose there is an ideal target vector field \(u\) t generating a probability path from \(p_0\) to \(p_1\approx q\), then the loss function of flow matching is defined as follows.

[0087]

[0088] where \(t\sim U(0, 1)\) is a uniformly sampled time step, and \(v\) t (x; \(\theta\)) is the vector field output by a neural network controlled by the parameter \(\theta\). However, since it is difficult to directly obtain the true distribution \(p\) in practice t and its corresponding vector field \(u\) t , by using the conditional flow matching loss function, according to the known conditions, the difficulty of directly solving the unconditional vector field can be avoided, and it is easier to obtain the true probability distribution. The conditional flow matching loss function is as follows:

[0089]

[0090] where \(t\) is the time step, \(q(x_1)\) represents the target distribution (the distribution that needs to be obtained finally, the real Mel spectrogram of Huangmei Opera singing), and \(p\) t (x) represents the probability path, that is, the distribution sequence from the initial noise distribution \(p_0\) to the final target distribution \(p_1\). \(p\) t (x|x_1) represents the conditional probability path, that is, the true distribution during the flow evolution process. \(\theta\) is the neural network parameter and is a control parameter.

[0091] By optimizing the loss value, the model is trained and optimized. This optimization objective is essentially to learn a neural network vector field \(v\) that approximates the true vector field \(u\) in the conditional probability space t . Compared with directly fitting the evolution path of the target distribution, this method can utilize the known conditions to achieve efficient training and ensure that the gradient information is consistent with the theoretical path, thus ensuring the correctness of model parameter updates. t

[0092] To accelerate the training speed, optimal transport is used to keep the minimum distance for each step of the evolution process. The specific operation is that the network only approximates the target along a "straight line" at each step, and the result can reduce spectral distortion. Rewrite the above conditional flow matching loss function to obtain the following loss function:

[0093]

[0094] where the flow mapping is defined as: \(\varphi\)t (x) = (1 - (1 - σ min )t)x0 + tx1, u t (φ t (x)|x1) is the target vector field, where x0 ∼ N(0, I) is the noise sample, x1 is the Mel-spectrum feature frame, c is the conditional mean predicted by the musical score information, d represents the embedding information of the speaker, and σ min is a very small noise parameter. The conditional flow matching of optimal transport has linear time-invariant characteristics and only depends on x0 and x1, which can simplify the training and improve the efficiency. As attached Figure 6 denotes the process of x t change. Therefore, through conditional flow matching, Mel-spectrums with rich details can be efficiently generated, and the generation of content containing timbre speaker information can also be flexibly controlled through conditions.

[0095] At time t, the original Mel-spectrum x of the target singer t is used as one of the input conditions, and the acoustic features are decoded by the trained flow matching decoder to generate the Mel-spectrum of the Huangmei Opera singing voice of the target singer.

[0096] In the specific implementation process, the generated Mel-spectrum goes through post-processing modules such as a neural vocoder, and finally a high-quality Huangmei Opera singing voice audio is synthesized, thus realizing the effective cloning of the original audio timbre.

[0097] The experiment of the embodiment of the present invention uses the Ubuntu20.04 operating system, and the host configuration is Intel(R) Core(TM) i7-12700F@2.10GHz, RAM32.0GB, RTX3090 GPU, and is implemented using the Pycharm software, where the Python version is 3.8.

[0098] The cloning of singing voice timbre and style is essentially a research on the intra-domain and extra-domain generalization of singing voice synthesis. To ensure that the model can fully learn the timbre features and clone the singing voice audio with similar timbre, the training data needs to contain enough speakers (singers) with different timbres. Therefore, the high-quality and low-quality Huangmei Opera singing voice datasets constructed in Chapter 3 are used in the experiment. To evaluate the cloning performance of the model in intra-domain (visible singers) and extra-domain (invisible singers) scenarios, all the audio of 10 singers is reserved as the extra-domain test set (invisible singers) from the 6-hour Huangmei Opera dataset, and at the same time, the audio of 10 singers is randomly selected from the training set as the intra-domain test set (visible singers). Similarly, the model is trained until convergence.

[0099] The Mean Opinion Score (MOS) generally evaluates the comprehensive performance of synthetic audio in terms of naturalness, sound quality, and artistic expressiveness. For voice cloning, an additional Similarity Mean Opinion Score (SMOS) is required. To more accurately evaluate the quality of synthetic audio and analyze the performance of the system, this paper uses some objective evaluation metrics, including Mel Cepstral Distortion (MCD) and Fundamental Frequency Frame Error (FFE). In the performance evaluation of singing voice cloning technology, Speaker Encoder Cosine Similarity (SECS) is used to evaluate the similarity between the synthetic singing voice and the original singing voice timbre.

[0100] 1. Comparative Experiment Analysis

[0101] The comparative experiments are shown in Tables 1 and 2, including in-domain and out-of-domain comparisons.

[0102] Table 1 Comparative Experiment Results of Huangmei Opera Aria Cloning (In-domain)

[0103]

[0104] Table 2 Comparative Experiment Results of Huangmei Opera Aria Cloning (Out-of-domain)

[0105]

[0106] It can be seen from both subjective and objective evaluation metrics that the StyleFlowSinger model in the embodiments of the present invention demonstrates superior performance. The synthesized Huangmei Opera has higher naturalness and similarity, and can achieve more accurate timbre cloning of the singers seen in the in-domain training set. Under in-domain conditions, the MOS score is 3.93 and the SMOS score is 3.97. Under out-of-domain conditions, the MOS score is 3.85 and the SMOS score is 3.62.

[0107] 2. Ablation Experiment Analysis

[0108] The ablation experiment analyzes the effectiveness of the key modules in the StyleFlowSinger model proposed in the embodiments of the present invention. Specifically, it analyzes the effectiveness of the speaker encoder, Mel style encoder, and flow matching decoder introduced into the model. The test set used in the ablation experiment is also tested according to unseen singers out-of-domain and seen singing voices in-domain, and the results are shown in Table 3.

[0109] Table 3 Analysis Results of Ablation Experiment

[0110]

[0111] Among them, "w / o SpkEnc" in the table means not using a speaker encoder in the model. "w / o SLAN" means that the Mel-style vectors in the model do not use style adaptive normalization to be incorporated into the encoder, but are directly added to the output result of the encoder. "w / o MelStyleEncoder" means not using a Mel-style encoder in the model. "w / o CFMDecoder" means not using a decoder composed of a flow matching decoder Branchformer block in the model.

[0112] It can be seen from the results shown in Table 3 that the decoding conditions incorporating timbre information can better guide the decoder to generate Mel spectrograms with the target timbre. Generally speaking, each key module plays a complementary role in timbre information modeling and Mel spectrogram reconstruction. The experimental results show that StyleFlowSinger can better achieve high-fidelity audio synthesis and accurate timbre cloning in combining the multi-dimensional timbre information provided by the speaker encoder and the Mel-style encoder, and using the flow matching decoder to effectively guide the generation process. The experimental data not only demonstrates the effectiveness of each module in objective metrics, but also shows good performance in subjective scoring, verifying the rationality and effectiveness of the overall model design.

[0113] An embodiment of the present invention provides a Huangmei Opera singing voice cloning system. The system synthesizes the Mel spectrogram of the Huangmei Opera singing voice of a target singer through a cloning model. The cloning model includes a speaker encoder, a Mel-style encoder, a score encoder, a duration predictor, a length regulator, a pitch diffuser, and a decoder. The system includes:

[0114] A data acquisition and processing module, configured to acquire Huangmei Opera singing samples and corresponding score data, as well as the original audio of the target singer, and perform preprocessing on the Mel spectrogram of the original audio of the target singer to obtain the original Mel spectrogram of the target singer;

[0115] An encoding module, configured to model the singing style of the original Mel spectrogram of the target singer through a Mel-style encoder to obtain a one-dimensional style vector; process the score data and the style vector through a score encoder to obtain a set of feature representations closely related to speech generation; process the original speech signal of the target singer through a speaker encoder to obtain the voiceprint feature of the target singer;

[0116] An acoustic feature acquisition module, configured to splice the feature representations closely related to speech generation and the voiceprint feature of the target singer. The spliced features are processed by a duration predictor, a length regulator, and a pitch diffuser, and the processed features are spliced with the style vector and subjected to positional encoding to obtain acoustic features;

[0117] A decoding module, configured to decode acoustic features through a flow-matching decoder to generate a Mel spectrogram of the Huangmei Opera singing voice of the target singer.

[0118] It can be understood that the Huangmei Opera singing voice cloning system provided by the embodiments of the present invention corresponds to the above-mentioned Huangmei Opera singing voice cloning method. The explanations, examples, beneficial effects, etc. of the relevant content can refer to the corresponding content in the Huangmei Opera singing voice cloning method, which will not be elaborated here.

[0119] The embodiments of the present invention also provide a computer-readable storage medium, which stores a computer program for Huangmei Opera singing voice cloning. Wherein, the computer program enables a computer to execute the Huangmei Opera singing voice cloning method as described above.

[0120] The embodiments of the present invention also provide an electronic device, including: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors. The programs include those for executing the Huangmei Opera singing voice cloning method as described above.

[0121] In summary, compared with the prior art, the following beneficial effects are achieved:

[0122] 1. Compared with the prior art, the Huangmei Opera singing voice cloning method provided by the present invention can accurately capture and reproduce the unique timbre details of each singer, and also shows better performance for the timbres of singers not seen outside the domain, and can provide personalized customization solutions for Huangmei Opera lovers.

[0123] 2. Conditional flow matching models the complete distribution of real spectrograms and adopts a uniform step-by-step sampling and denoising reconstruction mechanism, avoiding the blurred averaging effect brought by MSE. The reconstruction process will retain details and textures, and will not ignore information such as sharp edges and high frequencies of the spectrogram, thereby reducing over-smoothing, distortion, and artifacts, and generating a more natural and high-fidelity Mel spectrogram.

[0124] 3. Flow matching single-step sampling improves the inference speed. Only one forward pass of the neural network is required, which is friendly to the device and more suitable for deployment. It can generate and infer faster without a complex step scheduler.

[0125] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0126] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the same; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for cloning Huangmei Opera singing tunes, characterized in that Synthesize the Mel spectrogram of the Huangmei Opera singing voice of the target singer through a cloning model, where the cloning model includes a speaker encoder, a Mel style encoder, a musical score encoder, a duration predictor, a length regulator, a pitch diffuser, and a decoder; the Huangmei Opera singing voice cloning method includes: Obtain Huangmei Opera singing samples and corresponding musical score data, as well as the original audio of the target singer, and perform preprocessing on the Mel spectrogram of the original audio of the target singer to obtain the original Mel spectrogram of the target singer; Model the singing style of the original Mel spectrogram of the target singer through the Mel style encoder to obtain a one-dimensional style vector; process the musical score data and the style vector through the musical score encoder to obtain a set of feature representations related to speech generation; process the original speech signal of the target singer through the speaker encoder to obtain the voiceprint feature of the target singer; Concatenate the feature representations related to speech generation and the voiceprint feature of the target singer. The concatenated features are processed by the duration predictor, the length regulator, and the pitch diffuser. The features obtained after processing are concatenated with the style vector and subjected to positional encoding to obtain acoustic features; Decode the acoustic features through the decoder to generate the Mel spectrogram of the Huangmei Opera singing voice of the target singer.

2. The Huangmei Opera singing voice cloning method according to claim 1, characterized in that The speaker encoder uses a generalized end-to-end model.

3. The Huangmei Opera aria cloning method according to claim 1, wherein The Mel style encoder includes a spectral processing module, a temporal modeling module, and a global style modeling module; Among them, The spectral processing module is used to send the input original Mel spectrogram into a fully connected layer. Through two layers of fully connected layers, each frame of the Mel spectrogram is converted into a hidden sequence; The temporal modeling module is used to connect the hidden sequence to a subsequent gated convolutional neural network to capture the temporal information in the original Mel spectrogram; The global style modeling module is used to extract global style information from the output result of the temporal modeling module using a multi-head attention mechanism combined with residual connections, and performs modeling at the Mel spectrogram frame level; and performs temporal averaging on the output of the multi-head attention to obtain a one-dimensional style vector.

4. The Huangmei Opera singing voice cloning method according to claim 1, characterized in that, The musical score encoder includes a Branchformer module, and the Branchformer module includes a global feature extraction branch based on multi-head attention and a local feature extraction branch based on a convolutional gated multi-layer perceptron.

5. The Huangmei Opera aria cloning method according to claim 1, characterized in that The decoder uses a flow matching decoder.

6. The Huangmei Opera singing voice cloning method according to claim 5, characterized in that, The loss function L CFM (θ) of the flow matching decoder during the training process includes: where t represents the time step, q(x1) represents the target distribution, and p t (x|x1) represents the conditional probability path; v t (x; θ) represents the vector field output by the neural network controlled by the parameter θ; u t (x|x1) represents the target vector field; x1 represents the Mel spectrogram feature frame; x represents the Mel spectrogram.

7. The Huangmei Opera aria cloning method according to any one of claims 1 to 6, characterized in that, The Huangmei Opera singing voice cloning method further includes: processing the Mel spectrogram of the Huangmei Opera singing voice of the target singer through a neural vocoder to obtain the Huangmei Opera singing voice audio of the target singer.

8. A Huangmei Opera singing voice cloning system, characterized in that, Synthesize the Mel spectrogram of the Huangmei Opera singing voice of the target singer through a cloning model, where the cloning model includes a speaker encoder, a Mel style encoder, a musical score encoder, a duration predictor, a length regulator, a pitch diffuser, and a decoder; the Huangmei Opera singing voice cloning system includes: A data acquisition and processing module, configured to obtain Huangmei Opera singing samples and corresponding musical score data, as well as the original audio of the target singer, and perform preprocessing on the Mel spectrogram of the original audio of the target singer to obtain the original Mel spectrogram of the target singer; An encoding module, configured to model the singing style of the original Mel spectrogram of the target singer through a Mel-style encoder to obtain a one-dimensional style vector; process the sheet music data and the style vector through a sheet music encoder to obtain a set of feature representations related to speech generation; process the original speech signal of the target singer through a speaker encoder to obtain the voiceprint feature of the target singer; An acoustic feature acquisition module, configured to splice the feature representations related to speech generation and the voiceprint feature of the target singer. The spliced features are processed by a duration predictor, a length regulator, and a pitch diffuser. The processed features are spliced with the style vector and subjected to positional encoding to obtain acoustic features; A decoding module, configured to decode the acoustic features through a decoder to generate the Mel spectrogram of the Huangmei Opera singing of the target singer.

9. A computer-readable storage medium, characterized in that, It stores a computer program for Huangmei Opera singing cloning, wherein the computer program causes a computer to execute the Huangmei Opera singing cloning method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, Including: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors. The programs include those for executing the Huangmei Opera singing cloning method according to any one of claims 1 to 7.