Dialect speech synthesis method and device, electronic equipment and storage medium
By obtaining the text characteristics of dialect pronunciation and generating the optimized Mel spectrum, combined with the speech synthesis model, the problem of low-resource dialect pronunciation quality is solved, and high-quality dialect pronunciation synthesis under low-resource conditions is achieved.
Patent Information
- Application Number
- CN202510231627.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-03
AI Technical Summary
When the existing speech synthesis model deals with low-resource dialect speech, the generation quality has dropped significantly and cannot meet the actual application needs.
By obtaining the text features of the target dialect pronunciation, a Mel spectrum is generated and inputting it into a pre-trained optimization model to obtain the optimized Mel spectrum. Then, the optimized Mel spectrum is input into the speech synthesis model to generate the speech waveform of the target dialect speech.
In the case of scarcity of low-resource dialect corpus, the quality of dialect pronunciation synthesis is improved and the constraints on model performance by corpus scarcity is improved.
Smart Images

Figure CN120089126A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech synthesis technology, and particularly to a method, device, electronic device and storage medium for dialect speech synthesis. Background Art
[0002] Current mainstream speech synthesis models generally rely on a large amount of high-quality labeled corpus and standard Mandarin speech data. In the face of low-resource dialect speech or scarce data scenarios, the generation quality of existing models significantly decreases and cannot meet the actual application requirements. The research on low-resource dialect speech synthesis not only has important significance for protecting language diversity and cultural inheritance, but also has broad application prospects in meeting regional intelligent speech interaction needs. Traditional speech synthesis methods face the following main challenges when dealing with low-resource dialect speech: there is a lack of high-quality text for low-resource dialects, the performance of the generated speech quality is poor, and the difficulty of generating high-quality dialect speech is high. Summary of the Invention
[0003] Embodiments of this application disclose a method, device, electronic device and storage medium for dialect speech synthesis, which can improve the quality of dialect speech synthesis in the case of scarce low-resource dialect corpus.
[0004] In a first aspect of embodiments of this application, a method for dialect speech synthesis is disclosed. The method includes:
[0005] Obtain text features corresponding to the target dialect speech, where the text features include phoneme feature vectors and / or character feature vectors corresponding to the text and position encoding vectors of the text sequence;
[0006] Obtain the Mel spectrogram of the target dialect speech according to the text features, and input the Mel spectrogram into a pre-trained optimization model to obtain the optimized Mel spectrogram. The optimization model is trained from a first initial model according to a first spectrogram corresponding to a dialect speech sample and a second spectrogram, and the resolution of the second spectrogram is higher than that of the first spectrogram;
[0007] Input the optimized Mel spectrogram into a pre-trained speech synthesis model to obtain the speech waveform of the target dialect speech. The speech synthesis model is trained from a second initial model according to the speech waveform of the speech sample and the optimized Mel spectrogram corresponding to the speech sample. The speech sample includes the dialect speech sample and the Mandarin speech sample corresponding to the dialect speech.
[0008] As an optional implementation manner, in the first aspect of this embodiment, the first initial model includes an encoding module. Before inputting the Mel spectrogram into the pre-trained optimization model to obtain the optimized Mel spectrogram, the method further includes:
[0009] Train the first initial model according to a preset first loss function, a first spectrum corresponding to a dialect voice sample, and a second spectrum to obtain the trained optimized model. The preset first loss function includes a first multi-scale perceptual loss function and a reconstruction loss function. Among them, the first multi-scale perceptual loss function is determined according to the number of network layers of the encoding module of the first initial model and a multi-scale feature extraction function, and the reconstruction loss function is determined according to the difference value between the output value after the first spectrum is input into the first initial model and the second spectrum.
[0010] As an optional implementation manner, in the first aspect of this embodiment, the first initial model includes a skip connection module and a decoding module. The encoding module includes at least one encoding unit, and the encoding unit is configured to perform a downsampling operation on the first spectrum to obtain an output feature map; the skip connection module is configured to connect the feature maps output by different encoding units in the encoding module to the decoding unit corresponding to the encoding unit in the decoding module; the decoding module includes at least one of the decoding units, and the decoding unit is configured to perform an upsampling operation on the output feature map of the previous-level decoding unit and fuse the feature map in the skip connection module; or, perform an upsampling operation on the output feature map of the encoding module and fuse the feature map in the skip connection module to obtain the second spectrum.
[0011] As an optional implementation manner, in the first aspect of this embodiment, the second initial model includes a generator and a discriminator. Before inputting the optimized Mel spectrum into a pre-trained speech synthesis model to obtain the speech waveform of the target dialect voice, the method further includes:
[0012] Train the second initial model according to a preset second loss function, the speech waveform of the speech sample, and the optimized Mel spectrum corresponding to the speech sample to obtain the trained speech synthesis model. The preset second loss function includes a second multi-scale perceptual loss function and a dynamic balance loss function. The second multi-scale perceptual loss function is determined according to the feature distance between the speech waveform generated by the generator and the speech waveform of the speech sample, and the dynamic balance loss function is determined according to the discrimination result of the discriminator on the speech waveform generated by the generator and the discrimination result on the speech waveform of the speech sample.
[0013] As an optional implementation manner, in the first aspect of this embodiment, the method further includes:
[0014] In the first stage of the second initial model training, the parameters of the generator are adjusted through the second multi-scale perception loss function;
[0015] In the second stage of the second initial model training, the parameters of the generator and the parameters of the discriminator are adjusted through the dynamic balance loss function;
[0016] In the third stage of the second initial model training, the parameters of the generator are adjusted through at least one loss function in the second loss function.
[0017] As an optional implementation manner, in the first aspect of this embodiment, the method further includes:
[0018] Dynamically adjusting the loss weight of the dialect speech sample and the loss weight of the Mandarin speech sample corresponding to the dialect speech according to the model performance index during the training of the second initial model, where the model performance index includes the error value between the speech waveform generated by the generator and the speech waveform of the speech sample and the discrimination accuracy rate of the discriminator for the speech waveform generated by the generator.
[0019] As an optional implementation manner, in the first aspect of this embodiment, the dynamically adjusting the loss weight of the dialect speech sample and the loss weight of the Mandarin speech sample corresponding to the dialect speech according to the model performance index during the training of the second initial model includes:
[0020] When the error value is greater than or equal to a preset first threshold, or when the discrimination accuracy rate is greater than or equal to a preset second threshold, increase the loss weight of the dialect speech sample and decrease the loss weight of the Mandarin speech sample corresponding to the dialect speech.
[0021] The second aspect of the embodiments of the present application discloses a dialect speech synthesis device, and the device includes:
[0022] A text feature acquisition module, configured to acquire text features corresponding to a target dialect speech, where the text features include a phoneme feature vector and / or a character feature vector corresponding to the text and a position encoding vector of the text sequence;
[0023] A spectrum optimization module, configured to obtain the Mel spectrum of the target dialect speech according to the text features, and input the Mel spectrum into a pre-trained optimization model to obtain the optimized Mel spectrum, where the optimization model is trained by using a first spectrum corresponding to a dialect speech sample and a second spectrum, and the resolution of the second spectrum is higher than the resolution of the first spectrum;
[0024] A speech synthesis module, configured to input the optimized Mel spectrogram into a pre-trained speech synthesis model to obtain a speech waveform of the target dialect speech. The speech synthesis model is trained from a second initial model based on the speech waveforms of speech samples and the corresponding optimized Mel spectrograms of the speech samples. The speech samples include the dialect speech samples and the Mandarin speech samples corresponding to the dialect speech.
[0025] A third aspect of the embodiments of the present application discloses an electronic device, including a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor implements the method as described above.
[0026] A fourth aspect of the embodiments of the present application discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method as described above is implemented.
[0027] Compared with the related art, the embodiments of the present application at least include the following beneficial effects:
[0028] Obtain the text features corresponding to the target dialect speech, obtain the Mel spectrogram of the target dialect speech according to the text features, input the Mel spectrogram into a pre-trained optimization model to obtain an optimized Mel spectrogram; input the optimized Mel spectrogram into a pre-trained speech synthesis model to obtain a speech waveform of the target dialect speech. By using the optimization model to improve the quality of the Mel spectrogram of the dialect speech, and using the speech synthesis model to obtain the speech waveform of the target dialect speech, compared with the traditional single model, it can improve the restriction of scarce corpus on the model performance, so as to improve the quality of dialect speech synthesis in the case of scarce low-resource dialect corpus. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0030] Figure 1 It is a flowchart of a dialect speech synthesis method in an embodiment;
[0031] Figure 2 It is a schematic structural diagram of a first initial model in an embodiment;
[0032] Figure 3 It is a flowchart of a training strategy method of a second initial model in an embodiment;
[0033] Figure 4It is a general schematic diagram of a dialect speech synthesis method in an embodiment;
[0034] Figure 5 It is a block diagram of a dialect speech synthesis device in an embodiment;
[0035] Figure 6 It is a structural block diagram of an electronic device in an embodiment. Specific embodiments
[0036] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0037] In order to facilitate a clear description of the technical solutions in the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. For example, the first instruction and the second instruction are used to distinguish different user instructions, and their order is not limited. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.
[0038] It should be noted that in the present application, words such as "exemplarily" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplarily" or "for example" is intended to present relevant concepts in a specific manner.
[0039] In addition, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, and c can represent: a, or b, or c, or a and b, or a and c, or b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0040] In addition, the terms "including" and "having" and any variations thereof in the embodiments of the present application and the accompanying drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0041] Embodiments of the present application disclose a dialect speech synthesis method, device, electronic device, and storage medium, which can improve the quality of dialect speech synthesis in the case of scarce low-resource dialect corpus.
[0042] The following will be described in detail with reference to the accompanying drawings.
[0043] Please refer to Figure 1 , Figure 1 which is a flowchart of the dialect speech synthesis method in an embodiment. As Figure 1 shown, the method may include the following steps:
[0044] Step 101, obtain text features corresponding to the target dialect speech, where the text features include a phoneme feature vector and / or a character feature vector corresponding to the text and a position encoding vector of the text sequence.
[0045] The speech synthesis method provided by the embodiments of the present application can be applied to a terminal device, which may include, but is not limited to, a mobile phone, a tablet computer, a desktop computer, a laptop computer, a portable terminal, a vehicle-mounted terminal, a personal computer (Personal Computer, PC), etc. The computer system in the terminal device may be Windows, Linux, IOS, or Unix, which is not specifically limited herein.
[0046] In the embodiments of the present application, the target dialect speech is the speech content of a specific dialect that the user expects to perform speech synthesis. Optionally, the terminal device may receive the text corresponding to the target dialect speech input by the user and extract the text features through a text encoder. The text features may be a combination of a phoneme feature vector corresponding to the text and a position encoding vector of the text sequence, or a combination of a character feature vector corresponding to the text and a position encoding vector of the text sequence, or a combination of a phoneme feature vector, a character feature vector, and a position encoding vector of the text sequence corresponding to the text, etc. The embodiments of the present invention do not make a limitation.
[0047] A phoneme is the smallest unit in speech and is the smallest speech unit divided according to the natural properties of speech. The phoneme feature vector corresponding to the text is a vector representation used to characterize the acoustic and speech features of each phoneme in the text. Dividing the text corresponding to the target dialect speech into a phoneme sequence, an acoustic model or a deep learning method can be used to map the phoneme into a feature vector. The text encoder can be an autoencoder architecture of deep learning, such as a Transformer model, a Long-Short Term Memory (LSTM) model, etc. Inputting the phoneme into the neural network model, the phoneme is converted into a feature vector through the encoding layer.
[0048] In some embodiments, each character in the text corresponding to the target dialect speech can also be mapped into a feature vector through a text encoder. For example, using deep learning-based character embedding technology, the character passes through the embedding layer of the neural network and is mapped into a vector with a fixed dimension. The position encoding vector of the text sequence is a vector representation that adds position information to each position in the text sequence, which can enhance the time dimension information of the output sequence features of the text encoder. As an optional implementation manner, position encoding methods such as trigonometric function-based position encoding, learning-based position encoding, and relative position encoding can be used to generate the position encoding vector of the text sequence.
[0049] Step 102: Obtain the Mel spectrogram of the target dialect speech according to the text features, and input the Mel spectrogram into a pre-trained optimization model to obtain an optimized Mel spectrogram. The optimization model is obtained by training a first initial model according to the first spectrogram and the second spectrogram corresponding to the dialect speech sample, and the resolution of the second spectrogram is higher than that of the first spectrogram.
[0050] In the embodiments of the present application, the terminal device can input the obtained text features into a sequence-to-sequence spectrogram generation model to output a Mel spectrogram. Optionally, the spectrogram generation model can be a Tacotron model, a Transformer model, a FastSpeech model, etc., which is not limited in the embodiments of the present invention. The Mel spectrogram of the generated target dialect speech includes the fundamental frequency, amplitude, and energy distribution of the target dialect speech, which are the key features for generating the target dialect speech.
[0051] In some embodiments, the first spectrogram is a Mel spectrogram with a relatively low resolution extracted from the dialect speech sample, and the second spectrogram is also a Mel spectrogram extracted from the dialect speech sample, but the resolution of the second spectrogram is higher than that of the first spectrogram. The second spectrogram can more accurately reflect the subtle changes of the dialect speech in frequency and time and contains more detailed information of the dialect speech.
[0052] Use the first spectrum and the second spectrum as training data to train the first initial model. During the training process, the first initial model can learn how to map from the low-resolution first spectrum to the high-resolution second spectrum. For example, the first initial model can continuously adjust its own weight parameters so that when the first spectrum is input, the output is as close as possible to the second spectrum. The first initial model after training is the optimized model that can perform optimization operations such as improving the resolution of the Mel spectrum.
[0053] In some embodiments, the first initial model includes an encoding module. Before inputting the Mel spectrum into the pre-trained optimized model to obtain the optimized Mel spectrum, the dialect speech synthesis method further includes:
[0054] Train the first initial model according to a preset first loss function, the first spectrum corresponding to the dialect speech sample, and the second spectrum to obtain the trained optimized model. The preset first loss function includes a first multi-scale perceptual loss function and a reconstruction loss function. Among them, the first multi-scale perceptual loss function is determined according to the number of network layers of the encoding module of the first initial model and the multi-scale feature extraction function, and the reconstruction loss function is determined according to the difference value between the output value after the first spectrum is input into the first initial model and the second spectrum.
[0055] As an alternative implementation, the training process of the first initial model is as follows:
[0056] First, use the first spectrum as input data and input it into the first initial model. The first initial model processes the first spectrum to obtain an output spectrum.
[0057] Secondly, use the first multi-scale perceptual loss function to calculate the perceptual differences between the spectrum output by the first initial model and the second spectrum at multiple scales. At the same time, use the reconstruction loss function to calculate the difference between the output spectrum and the second spectrum, such as calculating the mean square error or the mean absolute error. The total training loss can be a combination of the first multi-scale perceptual loss function and the reconstruction loss function, for example, it can be a weighted sum of the two to balance the training requirements in different aspects.
[0058] Finally, according to the calculated total loss, update the parameters of the first initial model through the backpropagation algorithm, and continuously adjust the parameters of the first initial model in multiple iterations so that when the first spectrum is input, the output spectrum of the first initial model is closer to the second spectrum.
[0059] After the above training process, the first initial model after training completion is the optimized model in this application. Input the Mel spectrogram of the target dialect speech to be optimized into the optimized model. The optimized model will perform optimization processing on the input Mel spectrogram according to the learned patterns and features, and output the optimized Mel spectrogram. The resolution of the optimized Mel spectrogram is higher than that of the unoptimized Mel spectrogram.
[0060] The encoding module of the first initial model is used to encode the input spectral data (such as Mel spectrogram), convert it into a representation form more suitable for subsequent processing and optimization. The encoding module can include multiple convolutional layers, pooling layers, etc., and can convert the input spectral information into an abstract feature representation. The first multi-scale perceptual loss function is used to measure the perceptual difference between the spectrograms generated by the first initial model at different scales and the target spectrogram. The purpose is to capture the perceptual characteristics of the spectrogram at different time and frequency resolutions and improve the speech naturalness of the optimized spectrogram. The expression of the first multi-scale perceptual loss function is as follows:
[0061]
[0062] where L perceptual represents the multi-scale perceptual loss value, M represents the number of layers of the feature extraction network used by the encoding module, m is an index variable used to traverse the M-layer feature extraction network, and its value range is from 1 to M. represents the output spectrogram obtained by processing the first spectrogram through the first initial model, S represents the second spectrogram. φ m is the feature extraction function, which is used to extract the features of the m-th feature layer from the input spectrogram (which can be or S), and this feature can be the output of a certain convolutional layer in the encoding module.
[0063] represents calculating the distance between the two feature vectors extracted by the feature extraction function, that is, the sum of the absolute values of the differences of the corresponding elements of the two feature vectors. The multi-scale perceptual loss function is used to calculate the perceptual differences between the generated spectrogram of the first initial model and the second spectrogram on multiple feature layers and average these differences. Exemplarily, its calculation process is as follows:
[0064] For each feature layer m, use the feature extraction function to extract features from the generated spectrogram of the first initial model and the second spectrogram S to obtain Then calculate the distance between these two feature vectors This represents the feature difference between the generated spectrogram and the second spectrogram S on the m-th feature layer. Finally, add up the feature differences of all M feature layers and divide by the number of feature layers M to obtain the average multi-scale perceptual loss value L perceptual。
[0065] By minimizing the first multi-scale perceptual loss function, the first initial model can not only focus on the pixel-level differences in the spectrum during training, but also on the differences in the feature representations at different levels of the model, so that the generated spectrum is closer to the second spectrum at multiple scales that are more in line with human perception, improving the quality and naturalness of the generated spectrum.
[0066] The reconstruction loss function is used to measure the difference between the generated spectrum of the first initial model and the second spectrum. The expression of the reconstruction loss function is as follows:
[0067]
[0068] where L recon represents the reconstruction loss value, N is the number of data points, that is, the total number of frames in the spectrum data, and i is an index variable used to iterate through these N data points, with a value range from 1 to N. represents the value of the reconstructed spectrum obtained by the model at the i-th data point, that is, the value of the spectrum output by the model after inputting the first spectrum into the first initial model, and S i represents the value of the second spectrum at the i-th data point.
[0069] Exemplarily, the calculation process of the reconstruction loss function is as follows:
[0070] For each data point i (from 1 to N) in the spectrum, calculate the absolute value of the difference between the reconstructed spectrum and the second spectrum S i . Then sum up the absolute values of the differences for all N data points, and finally divide by the number of data points N to obtain the average reconstruction loss L recon .
[0071] By minimizing the reconstruction loss function, the generated spectrum of the first initial model can be made close to the target spectrum (the second spectrum) to improve the quality of the generated spectrum.
[0072] The total loss of the first initial model is the sum of the reconstruction loss and the multi-scale perceptual loss, and its expression is:
[0073] L total = λ 1 L recon + λ 2 L perceptual (3)
[0074] where λ 1 and λ 2 are preset weight coefficients, L total represents the total loss, L recon represents the reconstruction loss, L perceptualRepresents the first multi-scale perception loss.
[0075] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of the first initial model in an embodiment. In the embodiments of the present application, the first initial model after training is the optimized model. As Figure 2 shown, the first initial model includes an encoding module, a skip connection module, and a decoding module. The encoding module includes at least one encoding unit, and the encoding unit is used to perform downsampling on the first spectrum to obtain an output feature map; the skip connection module is used to connect the feature maps output by different encoding units in the encoding module to the decoding unit corresponding to the encoding unit in the decoding module; the decoding module includes at least one decoding unit, and the decoding unit is used to perform upsampling on the output feature map of the previous-level decoding unit and fuse the feature maps in the skip connection module; or, to perform upsampling on the output feature map of the encoding module and fuse the feature maps in the skip connection module to obtain the second spectrum.
[0076] As an alternative embodiment, Figure 2 the encoding module in may include 3 encoding units, and each encoding unit may include two convolutional layers and a downsampling layer. The convolutional layer slides a convolutional kernel on the input spectrum (such as the first spectrum) and performs a convolutional operation with the spectrum data to extract the feature information in the first spectrum and generate a feature map. In an encoder with multiple convolutional layers, each convolutional layer can learn different levels of features. The convolutional layer closer to the input layer can learn the local features of the input spectrum. As the number of layers deepens, the subsequent convolutional layers can learn more complex and semantic features based on the features extracted by the previous convolutional layers.
[0077] In some embodiments, the downsampling layer may be a pooling layer or a causal convolutional layer, etc. In the embodiments of the present invention, the causal convolutional layer is taken as an example for illustration, and the causal convolution formula is as follows:
[0078]
[0079] where y(t) represents the output feature at time step t, x(t-k) represents the input feature at time step t-k, the value range of k is from 0 to K-1, and K is the size of the convolutional kernel. For example, when k = 0, x(t) is the input feature at the current time step; when k = 1, x(t-1) is the input feature at the previous time step, and so on. ω k is the weight of the convolutional kernel, and for each time step delay k, there is a corresponding weight ω k , and these weights determine the contribution degree of the input features at different time steps to the current output.
[0080] The downsampling layer uses causal convolution, which can reduce the data dimension while retaining as much temporal information as possible in the input spectral data, ensuring that the model does not disrupt the temporal order and dependencies of the spectral data during the feature learning process.
[0081] As Figure 2 shown, in some embodiments, Figure 2 the decoding module may include 3 decoding units. The number of decoding units is the same as the number of encoding units. Each decoding unit includes an upsampling layer and a convolutional layer. In the embodiments of the present application, the number of convolutional layers in the encoding module and the decoding module is not limited. The first encoding unit in the encoding module corresponds to the last decoding unit in the decoding module, the second encoding unit in the encoding module corresponds to the second last decoding unit in the decoding module, and so on.
[0082] The skip connection module is used to connect the feature maps output by different encoding units in the encoding module to the corresponding decoding units in the decoding module, that is, to connect the original spectral information and the spectral features obtained by downsampling to the decoding module to achieve the generation of high-quality Mel spectrograms. In some embodiments, the first decoding unit in the decoding module is used to perform an upsampling operation on the output feature map of the encoding module and fuse the feature map in the skip connection module, and the remaining decoding units are used to perform an upsampling operation on the output feature map of the previous decoding unit and fuse the feature map in the skip connection module.
[0083] In the embodiments of the present application, the upsampling layer in the decoding unit is described by taking the transposed convolutional layer as an example. It can be understood that the upsampling layer can also adopt other methods, such as dilated convolution, interpolation method, etc. The embodiments of the present application do not make specific limitations. The upsampling operation of the decoding module enlarges the feature map through transposed convolution. The transposed convolution formula is as follows:
[0084] Y l =W l TConv *X l +F l (5)
[0085] where W l TConv is the transposed convolution kernel of the l-th layer, X l is the feature map input to the l-th layer, * represents the convolution operation, that is, using the transposed convolution kernel to perform a convolution operation on the input feature map, F l is the feature map of the corresponding layer in the encoding module output through the skip connection module, and Y l represents the output of the l-th transposed convolutional layer.
[0086] The decoding unit gradually enlarges the feature map size through transposed convolution. The final output Y of the decoding moduleoutput It can be expressed as follows:
[0087]
[0088] Wherein, represents the transposed convolutional kernel of the last layer in the decoding module, and X final represents the input feature map of the last layer in the decoding module.
[0089] In the embodiment of the present application, the resolution of the Mel spectrogram corresponding to the target dialect speech is improved by optimizing the model, the optimization of the Mel spectrogram quality is realized, and the quality of dialect speech synthesis can be improved in the case of scarce low-resource dialect corpus.
[0090] Step 103, input the optimized Mel spectrogram into a pre-trained speech synthesis model to obtain the speech waveform of the target dialect speech. The speech synthesis model is trained from a second initial model according to the speech waveforms of speech samples and the optimized Mel spectrograms corresponding to the speech samples. The speech samples include dialect speech samples and Mandarin speech samples corresponding to the dialect speech.
[0091] In the embodiment of the present application, the speech waveform of the speech sample is used as the target output, and the optimized Mel spectrogram corresponding to the speech sample is used as the input to train the second initial model. During the training process, the second initial model can learn the mapping relationship from the Mel spectrogram to the speech waveform. The trained second initial model is the speech synthesis model.
[0092] In some embodiments, the second initial model includes a generator and a discriminator. Before inputting the optimized Mel spectrogram into the pre-trained speech synthesis model to obtain the speech waveform of the target dialect speech, the dialect speech synthesis method further includes:
[0093] Train the second initial model according to a preset second loss function, the speech waveform of the speech sample, and the optimized Mel spectrogram corresponding to the speech sample to obtain the trained speech synthesis model. The preset second loss function includes a second multi-scale perceptual loss function and a dynamic balance loss function. The second multi-scale perceptual loss function is determined according to the feature distance between the speech waveform generated by the generator and the speech waveform of the speech sample, and the dynamic balance loss function is determined according to the discrimination result of the discriminator on the speech waveform generated by the generator and the discrimination result on the speech waveform of the speech sample.
[0094] As an optional implementation manner, the structure of the second initial model may be a HiFi-GAN model, including a generator and a discriminator. The generator is used to convert the input Mel spectrogram into a speech waveform, and the discriminator is used to distinguish the speech waveform generated by the generator from the real speech waveform. The training process of the second initial model is as follows:
[0095] The optimized Mel spectrogram corresponding to the voice sample is input into the second initial model. The generator of the model will generate a corresponding voice waveform, and then the generated voice waveform is compared with the voice waveform of the voice sample, and the difference value between them is calculated using the second loss function. According to this difference value, the parameters of the second initial model are updated through the backpropagation algorithm, enabling the second initial model to gradually adjust the internal weights of the generator and discriminator to reduce the value of the second loss function.
[0096] In some embodiments, to ensure that the voice waveform generated by the generator is consistent with the voice waveform of the voice sample in both the frequency domain and the time domain, the second multi-scale perceptual loss function is used. This loss function extracts features through a pre-trained acoustic model and calculates the feature distance between the generated voice waveform and the voice waveform of the voice sample, enabling the model to pay more attention to the key feature regions of the generated voice. The expression of the second multi-scale perceptual loss function is as follows:
[0097]
[0098] where ψ i is the feature extraction result of the i-th pre-trained acoustic model, is the voice waveform of the generated voice, y is the voice waveform of the voice sample, and N is the number of network layers of the pre-trained acoustic model. represents calculating the feature distance between the generated voice waveform and the voice waveform of the voice sample.
[0099] By using the pre-trained acoustic model to extract features and calculate distances from different levels and perspectives, the second initial model pays more attention to the key feature regions of the generated voice, thereby improving the quality of the generated voice and enabling it to better approximate the voice sample at multiple scales.
[0100] In some embodiments, to avoid the mode collapse problem caused by the unbalanced confrontation between the generator and the discriminator, a dynamic balance loss function is introduced during the training process of the second initial model. By adjusting the confrontation intensity between the generator and the discriminator, the stable convergence of the model is ensured. The loss functions of the generator and the discriminator are respectively:
[0101]
[0102] where L G represents the loss of the generator, L D represents the loss of the discriminator, represents the discrimination result of the discriminator on the voice waveform generated by the generator, that is, the discriminator believes that The probability of being real data, D(y) represents the discrimination result of the discriminator on the speech waveform of the speech sample, that is, the probability that the discriminator believes y is real data, and β is the weight coefficient used to balance the importance of the two parts of the loss, L aux is the auxiliary loss, which can be defined according to specific requirements. E[log(D(y)] represents the expectation of the discrimination result of the discriminator on the speech waveform of the speech sample, represents the expectation of the discrimination result of the discriminator on the speech waveform generated by the generator.
[0103] In the embodiments of the present application, a multi-stage training strategy is adopted to optimize the second initial model. Please refer to Figure 3 , Figure 3 is a flowchart of the training strategy method of the second initial model in an embodiment. As Figure 3 shown, the training strategy method may include the following steps:
[0104] Step 301, in the first stage of training the second initial model, adjust the parameters of the generator through the second multi-scale perceptual loss function.
[0105] In the first stage of training the second initial model, fix the parameters of the discriminator and only optimize the parameters of the generator. The purpose is to ensure that the generator can accurately generate speech waveforms from the mel spectrogram. As described above, the second multi-scale perceptual loss function extracts features through a pre-trained acoustic model and calculates the feature distance between the speech waveform of the generated speech and the speech waveform of the sample speech. The generator will continuously adjust its internal parameters according to the loss value of the second multi-scale perceptual loss to reduce the difference in multi-scale features between the speech waveform of the generated speech and the speech waveform of the sample speech.
[0106] Step 302, in the second stage of training the second initial model, adjust the parameters of the generator and the parameters of the discriminator through the dynamic balance loss function.
[0107] The dynamic balance loss function will adjust the balance between the generator and the discriminator according to the discrimination result of the discriminator on the speech waveform generated by the generator and the discrimination result on the speech waveform of the speech sample. When the discriminator can accurately distinguish between the generated speech waveform and the sample speech waveform, the dynamic balance loss function will prompt the generator to further optimize the generated speech waveform to make it more difficult to be recognized by the discriminator; at the same time, it will also adjust the parameters of the discriminator to enable it to better adapt to the changes of the generator and continue to maintain the discrimination ability for the generated speech waveform and the real speech waveform. This dynamic interaction can avoid the situation where one of the generator or the discriminator is too powerful, resulting in unstable training or poor model performance. By jointly optimizing the generator and the discriminator in the second stage of training the second initial model and enhancing the adaptability of the generator to the feedback of the discriminator, the quality of dialect speech synthesis can be further improved.
[0108] Step 303, in the third stage of the second initial model training, adjust the parameters of the generator through at least one loss function in the second loss function.
[0109] The third stage of the second initial model training is the fine-tuning stage. In this stage, at least one loss function in the second loss function is combined to optimize the generator as a whole to enhance the high-fidelity characteristics of the generator, that is, to make the generated speech waveform as close as possible to the sample speech waveform in all aspects. Optionally, in the third stage, the second multi-scale perceptual loss function can be used to adjust the parameters of the generator, or the dynamic balance loss function can be used to adjust the parameters of the generator, or a combination of the second multi-scale perceptual loss function and the dynamic balance loss function can be used to adjust the parameters of the generator.
[0110] In the embodiments of the present application, the multi-stage training strategy method trains and optimizes the generator and the discriminator in a targeted manner by using different loss functions in different stages, can gradually improve the performance of the second initial model, significantly reduce the dependence of the model on data, and at the same time improve the naturalness and authenticity of the generated speech.
[0111] In some embodiments, during the training of the second initial model, the dialect speech synthesis method further includes:
[0112] Dynamically adjust the loss weight of the dialect speech sample and the loss weight of the corresponding Mandarin speech sample according to the model performance metrics during the training of the second initial model. The model performance metrics include the error value between the speech waveform generated by the generator and the speech waveform of the speech sample and the discrimination accuracy of the discriminator for the speech waveform generated by the generator.
[0113] In some embodiments, the error value between the speech waveform generated by the generator and the speech waveform of the speech sample is used to measure the difference between the speech waveform generated by the generator and the sample waveform. Exemplarily, various methods can be used to calculate the error value, such as mean square error (MSE), mean absolute error (MAE), or more complex perceptual loss functions, etc. The error value reflects the performance of the generator in generating speech waveforms. A smaller error value indicates that the generated speech waveform is closer to the sample speech waveform, indicating better performance of the generator in acoustic characteristics.
[0114] The discrimination accuracy of the discriminator for the speech waveform generated by the generator represents the correct degree of judgment of the discriminator. A higher discrimination accuracy means that the discriminator can better distinguish the two, indicating that the speech waveform generated by the generator is significantly different from the sample speech waveform, and the generated speech waveform is easily recognized as fake by the discriminator, while a lower discrimination accuracy indicates that the speech waveform generated by the generator is less different from the sample speech waveform, making it difficult for the discriminator to distinguish.
[0115] In some embodiments, dynamically adjusting the loss weight of the dialect speech samples and the loss weight of the Mandarin speech samples corresponding to the dialect speech according to the model performance metrics during the training process of the second initial model includes:
[0116] In the case where the error value is greater than or equal to a preset first threshold, or in the case where the discrimination accuracy is greater than or equal to a preset second threshold, increase the loss weight of the dialect speech samples and decrease the loss weight of the Mandarin speech samples corresponding to the dialect speech.
[0117] As an alternative implementation, the loss weights of the dialect speech samples and the loss weights of the Mandarin speech samples corresponding to the dialect speech can be dynamically adjusted through a mixed loss function. The expression of the mixed loss function is as follows:
[0118] L total =λ·L dialect +(1 - λ)·L mandarin (10)
[0119] Where L dialect is the loss of the dialect speech samples, L mandarin is the loss of the Mandarin speech samples corresponding to the dialect speech, λ is the dynamic weight, the value of λ is greater than 0.8, and during the training process of the second initial model, λ can be dynamically adjusted according to the model performance metrics to enhance the weight of the dialect samples during the training process, increase the loss weight of the dialect speech samples so that the Mandarin speech samples provide auxiliary distribution information for the dialect speech synthesis model, and the dialect speech samples dominate the generation direction of the model under the weight priority.
[0120] The preset first threshold is a preset standard value. When the error value is greater than or equal to the preset first threshold, it indicates that the difference between the speech waveform generated by the generator and the sample speech waveform is relatively large. In order to improve the quality of dialect speech generation, it is necessary to increase the loss weight of the dialect speech samples and decrease the loss weight of the Mandarin speech samples corresponding to the dialect speech.
[0121] The preset second threshold is also a preset standard value. The preset second threshold can be the same as or different from the preset first threshold. When the discrimination accuracy is greater than or equal to the preset second threshold, it means that the discriminator can better distinguish the two, indicating that the difference between the speech waveform generated by the generator and the sample speech waveform is relatively large. It is necessary to increase the loss weight of the dialect speech samples and decrease the loss weight of the Mandarin speech samples corresponding to the dialect speech.
[0122] In the embodiments of the present application, dynamically adjusting weights according to model performance indicators can highlight the dominant position of dialect speech samples in training, enabling the model to be more effectively optimized towards generating high-quality dialect speech, thereby improving the quality of the generated dialect speech.
[0123] In the embodiments of the present application, the quality of the Mel spectrogram of dialect speech is improved by optimizing the model, and the speech waveform of the target dialect speech is obtained through the speech synthesis model. Compared with the traditional single model, it can improve the restriction of scarce corpus on the model performance, thereby improving the quality of dialect speech synthesis in the case of scarce low-resource dialect corpus.
[0124] Please refer to Figure 4 , Figure 4 which is a general schematic diagram of the dialect speech synthesis method in an embodiment. As Figure 4 shown, the terminal device receives the text corresponding to the target dialect speech input by the user, and extracts the text features corresponding to the target dialect speech from the text. In the spectrogram generation stage 401, the text features are converted into the Mel spectrogram corresponding to the target dialect speech through the spectrogram generation model. In the spectrogram optimization stage 402, the low-resolution Mel spectrogram is converted into a high-resolution Mel spectrogram through the optimization model to improve the quality of the Mel spectrogram corresponding to the target dialect speech. In the speech synthesis stage 403, the optimized Mel spectrogram is converted into the speech waveform corresponding to the target dialect speech through the speech synthesis model. The specific implementation manners of the above method can refer to the foregoing embodiments and will not be elaborated herein.
[0125] In the embodiments of the present application, the quality of the Mel spectrogram of dialect speech is improved by optimizing the model, and the speech waveform of the target dialect speech is obtained through the speech synthesis model. Compared with the traditional single model, it can improve the restriction of scarce corpus on the model performance, thereby improving the quality of dialect speech synthesis in the case of scarce low-resource dialect corpus.
[0126] Please refer to Figure 5 , Figure 5 which is a block diagram of the dialect speech synthesis device in an embodiment. As Figure 5 shown, the dialect speech synthesis device 500 includes: a text feature acquisition module 510, a spectrogram optimization module 520, and a speech synthesis module 530.
[0127] The text feature acquisition module 510 is used to acquire the text features corresponding to the target dialect speech, where the text features include the phoneme feature vector and / or character feature vector corresponding to the text and the position encoding vector of the text sequence;
[0128] A spectrum optimization module 520 is configured to obtain the Mel spectrum of the target dialect voice according to the text feature, input the Mel spectrum into a pre-trained optimization model to obtain the optimized Mel spectrum. The optimization model is obtained by training a first initial model according to the first spectrum and the second spectrum corresponding to the dialect voice sample, and the resolution of the second spectrum is higher than that of the first spectrum;
[0129] A voice synthesis module 530 is configured to input the optimized Mel spectrum into a pre-trained voice synthesis model to obtain the voice waveform of the target dialect voice. The voice synthesis model is obtained by training a second initial model according to the voice waveform of the voice sample and the optimized Mel spectrum corresponding to the voice sample. The voice sample includes the dialect voice sample and the Mandarin voice sample corresponding to the dialect voice.
[0130] In some embodiments, the first initial model includes an encoding module. Before inputting the Mel spectrum into the pre-trained optimization model to obtain the optimized Mel spectrum, the spectrum optimization module 520 is specifically configured to:
[0131] Train the first initial model according to a pre-set first loss function, the first spectrum and the second spectrum corresponding to the dialect voice sample to obtain the trained optimization model. The pre-set first loss function includes a first multi-scale perceptual loss function and a reconstruction loss function. The first multi-scale perceptual loss function is determined according to the number of network layers of the encoding module of the first initial model and the multi-scale feature extraction function, and the reconstruction loss function is determined according to the difference value between the output value after the first spectrum is input into the first initial model and the second spectrum.
[0132] In some embodiments, the second initial model includes a generator and a discriminator. Before inputting the optimized Mel spectrum into the pre-trained voice synthesis model to obtain the voice waveform of the target dialect voice, the voice synthesis module 530 is specifically configured to:
[0133] Train the second initial model according to a pre-set second loss function, the voice waveform of the voice sample and the optimized Mel spectrum corresponding to the voice sample to obtain the trained voice synthesis model. The pre-set second loss function includes a second multi-scale perceptual loss function and a dynamic balance loss function. The second multi-scale perceptual loss function is determined according to the feature distance between the voice waveform generated by the generator and the voice waveform of the voice sample, and the dynamic balance loss function is determined according to the discrimination result of the discriminator on the voice waveform generated by the generator and the discrimination result on the voice waveform of the voice sample.
[0134] In some embodiments, the speech synthesis module 530 is specifically configured to:
[0135] In the first stage of training the second initial model, adjust the parameters of the generator through the second multi-scale perceptual loss function;
[0136] In the second stage of training the second initial model, adjust the parameters of the generator and the parameters of the discriminator through the dynamic balance loss function;
[0137] In the third stage of training the second initial model, adjust the parameters of the generator through at least one loss function in the second loss function.
[0138] In some embodiments, the speech synthesis module 530 is further configured to:
[0139] Dynamically adjust the loss weight of the dialect speech sample and the loss weight of the Mandarin speech sample corresponding to the dialect speech according to the model performance index during the training of the second initial model, where the model performance index includes the error value between the speech waveform generated by the generator and the speech waveform of the speech sample and the discrimination accuracy of the discriminator for the speech waveform generated by the generator;
[0140] In the case where the error value is greater than or equal to a preset first threshold, or in the case where the discrimination accuracy is greater than or equal to a preset second threshold, increase the loss weight of the dialect speech sample and decrease the loss weight of the Mandarin speech sample corresponding to the dialect speech.
[0141] The description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0142] Please refer to Figure 6 , Figure 6 which is a block diagram of an electronic device in an embodiment. As Figure 6 shown, the electronic device 600 may include: a processor 610, a memory 620, and a bus 630.
[0143] Among them, the processor 610 calls the executable program code stored in the memory 620 to execute any dialect speech synthesis method disclosed in the embodiments of the present application. Those skilled in the art can understand that Figure 6 the structure of the electronic device shown in
[0144] The processor 610 is used to execute obtaining text features corresponding to the target dialect voice, obtaining the Mel spectrogram of the target dialect voice according to the text features, inputting the Mel spectrogram into a pre-trained optimization model to obtain the optimized Mel spectrogram, and inputting the optimized Mel spectrogram into a pre-trained speech synthesis model to obtain the speech waveform of the target dialect voice. For specific details, please refer to the detailed description in the method examples and will not be elaborated here.
[0145] In the embodiments of the present application, the processor 610 may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0146] The memory 620 can be used to store software programs and modules. The processor 610 executes various functional applications and data processing of the electronic device by running the software programs and modules stored in the memory 620. The memory 620 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the electronic device. In addition, the memory 620 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0147] In the embodiments of the present application, the processor 610 and the memory 620 are connected through a bus 630. The bus 630 is shown in Figure 6 in thick lines. The connection manners between other components are only for illustrative purposes and are not limited thereto. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 6 only one thick line is shown in, but it does not mean that there is only one bus or one type of bus.
[0148] The embodiments of the present application disclose a computer-readable storage medium that stores a computer program, wherein the computer program, when executed by a processor, implements the methods described in the above embodiments.
[0149] The embodiments of the present application disclose a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program, when executed by a processor, implements the methods described in the above embodiments.
[0150] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disc, a ROM, etc.
[0151] Any reference to memory, storage, database, or other media as used herein may include non-volatile and / or volatile memory. Suitable non-volatile memory may include ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which is used as an external cache. By way of illustration and not limitation, RAM can be of various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus DRAM (RDRAM), and direct Rambus DRAM (DRDRAM).
[0152] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. Those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present application. It should be noted that "a plurality of" in the present application includes "two or more".
[0153] In various embodiments of the present application, it should be understood that the magnitude of the serial numbers of the above processes does not necessarily mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0154] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0155] In addition, each functional unit in the embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0156] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0157] The above has introduced in detail a dialect voice synthesis method, device, electronic device and storage medium disclosed in the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A dialect speech synthesis method, characterized in that: The method comprises: Acquire text features corresponding to the target dialect speech, wherein the text features include a phoneme feature vector and / or a character feature vector corresponding to the text and a position encoding vector of the text sequence; Obtaining a Mel spectrum of the target dialect speech according to the text features, and inputting the Mel spectrum into a pre-trained optimization model to obtain an optimized Mel spectrum, wherein the optimization model is obtained by training a first initial model according to a first spectrum and a second spectrum corresponding to the dialect speech sample, wherein a resolution of the second spectrum is higher than a resolution of the first spectrum; The optimized Mel spectrum is input into a pre-trained speech synthesis model to obtain a speech waveform of the target dialect speech, wherein the speech synthesis model is obtained by training a second initial model based on the speech waveform of a speech sample and the optimized Mel spectrum corresponding to the speech sample, wherein the speech sample includes the dialect speech sample and a Mandarin speech sample corresponding to the dialect speech.
2. The method according to claim 1, characterized in that The first initial model includes a coding module. Before inputting the Mel spectrum into a pre-trained optimization model to obtain the optimized Mel spectrum, the method further includes: The first initial model is trained according to a preset first loss function, a first spectrum corresponding to the dialect speech sample, and a second spectrum to obtain the trained optimization model, wherein the preset first loss function includes a first multi-scale perceptual loss function and a reconstruction loss function, wherein the first multi-scale perceptual loss function is determined according to the number of network layers of the encoding module of the first initial model and a multi-scale feature extraction function, and the reconstruction loss function is determined according to the difference value between the output value after the first spectrum is input into the first initial model and the second spectrum.
3. The method according to claim 2, characterized in that The first initial model includes a skip connection module and a decoding module, the encoding module includes at least one encoding unit, and the encoding unit is used to perform a downsampling operation on the first spectrum to obtain an output feature map; the skip connection module is used to connect the feature maps output by different encoding units in the encoding module to the decoding units corresponding to the encoding units in the decoding module; the decoding module includes at least one decoding unit, and the decoding unit is used to perform an upsampling operation on the output feature map of the previous level decoding unit and fuse the feature map in the skip connection module; or, to perform an upsampling operation on the output feature map of the encoding module and fuse the feature map in the skip connection module to obtain the second spectrum.
4. The method according to claim 1, characterized in that: The second initial model includes a generator and a discriminator. Before inputting the optimized Mel spectrum into a pre-trained speech synthesis model to obtain the speech waveform of the target dialect speech, the method further includes: The second initial model is trained according to a preset second loss function, the speech waveform of the speech sample, and the optimized Mel spectrum corresponding to the speech sample to obtain the trained speech synthesis model, wherein the preset second loss function includes a second multi-scale perceptual loss function and a dynamic balance loss function, wherein the second multi-scale perceptual loss function is determined according to the characteristic distance between the speech waveform generated by the generator and the speech waveform of the speech sample, and the dynamic balance loss function is determined according to the identification result of the discriminator on the speech waveform generated by the generator and the identification result of the speech waveform of the speech sample.
5. The method according to claim 4, characterized in that The method further comprises: In the first stage of the second initial model training, adjusting the parameters of the generator by the second multi-scale perceptual loss function; In the second stage of the second initial model training, adjusting the parameters of the generator and the parameters of the discriminator by means of the dynamic balance loss function; In the third stage of the second initial model training, the parameters of the generator are adjusted by at least one loss function in the second loss functions.
6. The method according to claim 4, characterized in that The method further comprises: The loss weight of the dialect speech sample and the loss weight of the Mandarin speech sample corresponding to the dialect speech are dynamically adjusted according to the model performance index of the second initial model during the training process, and the model performance index includes the error value between the speech waveform generated by the generator and the speech waveform of the speech sample and the identification accuracy rate of the discriminator on the speech waveform generated by the generator.
7. The method according to claim 6, characterized in that The dynamically adjusting the loss weight of the dialect speech sample and the loss weight of the Mandarin speech sample corresponding to the dialect speech according to the model performance index of the second initial model during the training process includes: When the error value is greater than or equal to a preset first threshold, or when the identification accuracy is greater than or equal to a preset second threshold, the loss weight of the dialect speech sample is increased, and the loss weight of the Mandarin speech sample corresponding to the dialect speech is reduced.
8. A dialect speech synthesis device, characterized in that: The device comprises: A text feature acquisition module, used to acquire text features corresponding to the target dialect speech, wherein the text features include a phoneme feature vector and / or a character feature vector corresponding to the text and a position coding vector of the text sequence; a spectrum optimization module, configured to obtain a Mel spectrum of the target dialect speech according to the text features, and input the Mel spectrum into a pre-trained optimization model to obtain the optimized Mel spectrum, wherein the optimization model is obtained by training a first initial model according to a first spectrum and a second spectrum corresponding to the dialect speech sample, wherein the resolution of the second spectrum is higher than that of the first spectrum; A speech synthesis module is used to input the optimized Mel spectrum into a pre-trained speech synthesis model to obtain a speech waveform of the target dialect speech, wherein the speech synthesis model is obtained by training a second initial model based on the speech waveform of a speech sample and the optimized Mel spectrum corresponding to the speech sample, wherein the speech sample includes the dialect speech sample and a Mandarin speech sample corresponding to the dialect speech.
9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Semi-supervised dialect emotion speech synthesis system based on hybrid experts
CN120299449A
Method, device and equipment for training translation model, medium and product
CN121093982A
Dialect emotion speech synthesis method
CN121922102A