Speech synthesis method and device, computer device and storage medium

By synthesizing textual latent vectors, prosodic latent vectors, and user-encoded vectors, and utilizing an attention mechanism to generate target acoustic features, the problem of poor performance in existing speech synthesis technologies is solved, achieving more natural speech synthesis.

CN115359780BActive Publication Date: 2025-12-26PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210897499.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-28
Publication Date
2025-12-26
Estimated Expiration
2042-07-28

AI Technical Summary

Technical Problem

Existing speech synthesis technology is ineffective, failing to effectively combine text content, prosodic style, and user voice timbre, resulting in unnatural synthesized speech.

Method used

By acquiring textual latent vectors, prosodic latent vectors, and user-encoded vectors, attention mechanisms are used to synthesize target acoustic features, and speech synthesis is performed based on these features, ensuring that the target audio file is related to the text content, prosodic style, and user voice timbre.

Benefits of technology

It improves the naturalness of speech synthesis, making the generated audio files more consistent with the expected text content, rhythmic style and user voice timbre, thus enhancing the speech synthesis effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359780B_ABST
    Figure CN115359780B_ABST
Patent Text Reader

Abstract

The application discloses a speech synthesis method and device, a computer and a storage medium, and relates to the technical field of speech synthesis. The method comprises the following steps: processing a text sequence to obtain a text hidden vector; extracting prosodic features from a prosodic reference audio to obtain a prosodic hidden vector; obtaining a user coding vector corresponding to a user identifier; synthesizing the text hidden vector, the prosodic hidden vector and the user coding vector to obtain target acoustic features; and performing speech synthesis based on the target acoustic features to obtain a target audio file corresponding to the text sequence. The method makes the obtained target audio file not only related to the text content corresponding to the text sequence, but also related to the prosodic style in the prosodic reference audio and the user voice timbre corresponding to the user identifier, which helps to guarantee the speech synthesis effect of the target audio file and improve the naturalness of the synthesized speech.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech synthesis, and in particular to a speech synthesis method and device, computer equipment and a storage medium. BACKGROUND

[0002] With the development of computer technology and digital signal processing technology, speech synthesis technology has begun to develop, and current TTS technology has been widely used in information exchange and broadcasting, etc. Tacotron is an end-to-end TTS generation model. The so-called "end-to-end" is to directly synthesize speech from character text, breaking down the barriers between various traditional components. The text can be directly synthesized into acoustic features through a model, and then the acoustic features are generated into audio files through a vocoder. Even the text can be input into the model to directly generate an audio file and skip the intermediate vocoder link. The existing speech synthesis is generally a simple synthesis of text content, resulting in poor speech synthesis effect. SUMMARY

[0003] The embodiments of the present application provide a speech synthesis method and device, computer equipment and a storage medium to solve the problem of poor speech synthesis effect.

[0004] A speech synthesis method comprises:

[0005] processing a text sequence to obtain a text hidden vector;

[0006] extracting prosodic features from a prosodic reference audio to obtain a prosodic hidden vector;

[0007] obtaining a user code vector corresponding to a user identifier;

[0008] synthesizing the text hidden vector, the prosodic hidden vector and the user code vector using an attention mechanism to obtain target acoustic features;

[0009] performing speech synthesis based on the target acoustic features to obtain a target audio file corresponding to the text sequence.

[0010] A speech synthesis device comprises:

[0011] a text hidden vector obtaining module configured to process a text sequence to obtain a text hidden vector;

[0012] a prosodic hidden vector obtaining module configured to extract prosodic features from a prosodic reference audio to obtain a prosodic hidden vector;

[0013] a user code vector obtaining module configured to obtain a user code vector corresponding to a user identifier;

[0014] an acoustic feature acquisition module configured to synthesize the text hidden vector, the prosody hidden vector and the user encoding vector by using an attention mechanism, and acquire a target acoustic feature;

[0015] a target audio file acquisition module configured to perform speech synthesis based on the target acoustic feature, and acquire a target audio file corresponding to the text sequence.

[0016] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the speech synthesis method when executing the computer program.

[0017] A computer readable storage medium stores a computer program, and the computer program is executable on a processor to implement the speech synthesis method.

[0018] The speech synthesis method, device, computer device and storage medium described above determine a text hidden vector based on a text sequence, so that the text hidden vector contains text content and can be used for subsequent encoding and synthesis processing, thereby ensuring the feasibility of speech synthesis. A prosody hidden vector is determined based on a prosody reference audio, so that the prosody hidden vector can learn the prosody style in the prosody reference audio. A user encoding vector corresponding to a user identifier is acquired, so that the user encoding vector can learn the user voice timbre. The text hidden vector, the prosody hidden vector and the user encoding vector are synthesized to form a target acoustic feature, and speech synthesis is performed based on the target acoustic feature, so that the target audio file acquired not only relates to the text content corresponding to the text sequence, but also relates to the prosody style in the prosody reference audio and the user voice timbre corresponding to the user identifier, which helps to ensure the speech synthesis effect of the target audio file and improve the naturalness of the synthesized speech. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0020] Figure 1 is an application environment diagram of the speech synthesis method in an embodiment of the present application;

[0021] Figure 2 is a flowchart of the speech synthesis method in an embodiment of the present application;

[0022] Figure 3 is another flowchart of the speech synthesis method in an embodiment of the present application;

[0023] Figure 4is another flow chart of the speech synthesis method in an embodiment of the present application;

[0024] Figure 5 is another flow chart of the speech synthesis method in an embodiment of the present application;

[0025] Figure 6 is another flow chart of the speech synthesis method in an embodiment of the present application;

[0026] Figure 7 is another flow chart of the speech synthesis method in an embodiment of the present application;

[0027] Figure 8 is another flow chart of the speech synthesis method in an embodiment of the present application;

[0028] Figure 9 is a schematic diagram of a speech synthesis device in an embodiment of the present application;

[0029] Figure 10 is a schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0030] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0031] The speech synthesis method provided by the embodiments of the present application can be applied in an application environment as shown in Figure 1 . Specifically, the speech synthesis method is applied in a speech synthesis system, which includes a client and a server as shown in Figure 1 . The client and the server communicate through a network to realize multi-user speech synthesis. The client, also known as the user end, is a program that provides local services for clients corresponding to the server. The client can be installed on, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers.

[0032] In an embodiment, as shown in Figure 2 , a speech synthesis method is provided. Taking the server in Figure 1 as an example, the method includes the following steps:

[0033] S201: processing a text sequence to obtain a text hidden vector;

[0034] S202: rhythm feature extraction is performed on the rhythm reference audio to obtain a rhythm hidden vector;

[0035] S203: a user coding vector corresponding to the user identifier is obtained;

[0036] S204: the text hidden vector, the rhythm hidden vector, and the user coding vector are synthesized to obtain target acoustic features;

[0037] S205: speech synthesis is performed based on the target acoustic features to obtain a target audio file corresponding to the text sequence.

[0038] The text sequence is a sequence formed by text content that needs to be synthesized by speech. The text hidden vector is a vector formed by vector conversion of the text sequence, specifically a hidden vector formed by encoding the text sequence using an encoder network. The hidden vector is an intermediate vector output by a specific network.

[0039] As an example, in step S201, the server obtains a text sequence that needs to be synthesized by speech, analyzes and vector converts the text sequence, and obtains a corresponding text hidden vector. The text hidden vector can be understood as an intermediate vector related to the text content. For example, the server can analyze the text sequence to obtain a phoneme sequence, encode the phoneme sequence to obtain specific format data acceptable to the network, and then process the specific format data through the encoder network to obtain a hidden layer feature representation, i.e., the text hidden vector. Understandably, the text hidden vector is a vector representation of the text sequence that needs to be synthesized by speech, providing a text basis for subsequent speech synthesis.

[0040] The rhythm reference audio is pre-set to provide audio of a rhythm style as a reference object. Rhythm refers to prosody and rhythm, which can be understood as the format of level and stress in speech and the rules of rhyming, or as information such as speaking pauses or speed. The rhythm hidden vector is a hidden vector formed by rhythm style extraction and encoding of the rhythm reference audio.

[0041] As an example, in step S202, the server can extract the prosody style of the prosody reference audio to obtain a prosody hidden vector, using the prosody reference audio set by default or selected by the user. The prosody hidden vector refers to an intermediate vector learned from the prosody reference audio and related to the prosody style. In this example, the server can use the prosody reference audio set by default, and the prosody hidden vector extracted from the prosody reference audio is a pre-extracted hidden vector, which can be a prosody hidden vector extracted by manual labeling or a prosody hidden vector extracted intelligently by a prosodic encoder. For example, in the manual labeling mode, different prosody reference audios are labeled so that the same prosody style has the same value. This method is based on human ear perception and is subjective. For example, in the prosodic encoder mode, different prosody reference audios are extracted for feature extraction, and the spectral features are extracted as prosody style information, and then the prosody style information is encoded to obtain the prosody hidden vector. Understandably, the prosody hidden vector is a hidden vector learned from the prosody reference audio and related to the prosody, which provides a prosody basis for subsequent speech synthesis and ensures that the final synthesized target audio file learns the prosody style in the prosody reference audio.

[0042] wherein the user identifier is an identifier for distinguishing different users. The user encoding vector is a vector for reflecting the speech of different users. As an example, the user encoding vector can be a simple digital number ID, or an acoustic feature, which can be set autonomously according to user needs.

[0043] As an example, in step S203, the server can obtain a user encoding vector corresponding to at least one user identifier. The user encoding vector can be understood as a vector formed by encoding the audio data formed by the user corresponding to each user identifier. In this example, the user encoding vector corresponding to the user identifier is determined so that the user encoding vector can reflect the voice timbre, and the synthesis effect of the target audio file formed by subsequent speech synthesis is ensured, so that the user identifier corresponds to the user, and the naturalness of speech synthesis is improved.

[0044] As an example, in step S204, after obtaining the text hidden vector, the prosody hidden vector, and the user encoding vector, the server can fuse the text hidden vector, the prosody hidden vector, and the user encoding vector to obtain a target vector, and then input the target vector to the decoder for decoding to obtain a target acoustic feature. The target acoustic feature is related to the text content, the prosody style, and the voice timbre of the user, which helps to ensure the synthesis effect of speech synthesis.

[0045] In this example, the server can first align the text hidden vector and the prosody hidden vector in time length to ensure consistency in time length, then perform first splicing and fusion on the text hidden vector and the prosody hidden vector after time length alignment to obtain a fused hidden vector, then perform second splicing and fusion on the fused hidden vector and the user encoding vector to obtain a fused target vector, and finally input the fused target vector into the decoder to obtain the target acoustic feature.

[0046] As an example, in step S205, after obtaining the target acoustic feature, the server can use a pre-set vocoder to encode and synthesize the target acoustic feature of the fusion of the text content, the prosody style and the user's voice timbre, etc., to obtain a target audio file corresponding to the text sequence, so that the target audio file is not only related to the text content corresponding to the text sequence, but also related to the prosody style in the prosody reference audio and the user voice timbre corresponding to the user identifier, which helps to ensure the voice synthesis effect of the target audio file and improve the naturalness of the synthesized voice.

[0047] In the voice synthesis method provided in this embodiment, the text hidden vector is determined based on the text sequence, so that it not only contains the text content but also can be used for subsequent encoding and synthesis processing, thereby ensuring the feasibility of voice synthesis; the prosody hidden vector is determined based on the prosody reference audio, so that it can learn the prosody style in the prosody reference audio; the user encoding vector corresponding to the user identifier is obtained, so that it can learn the user voice timbre; and the target acoustic feature formed by synthesizing the text hidden vector, the prosody hidden vector and the user encoding vector is used for voice synthesis, so that the obtained target audio file is not only related to the text content corresponding to the text sequence, but also related to the prosody style in the prosody reference audio and the user voice timbre corresponding to the user identifier, which helps to ensure the voice synthesis effect of the target audio file and improve the naturalness of the synthesized voice.

[0048] In an embodiment, as shown in Figure 3 Step S201, i.e., processing the text sequence to obtain the text hidden vector, includes:

[0049] S301: analyzing the text sequence to obtain a phoneme sequence;

[0050] S302: performing spatial vector conversion on the phoneme sequence to obtain a phoneme feature vector;

[0051] S303: using a phoneme encoder to encode the phoneme feature vector to obtain the text hidden vector.

[0052] The phoneme sequence is a sequence related to phonemes, and specifically a sequence obtained by performing character-phoneme conversion on the text sequence. A phoneme (phone) is the smallest unit of speech divided according to the natural properties of speech, and is analyzed according to the pronunciation action in a syllable. One action constitutes one phoneme.

[0053] As an example, in step S301, the server can parse the text sequence by using a pre-set model for text-to-phoneme conversion, including but not limited to a G2P (Grapheme-to-Phoneme) model, to obtain a phoneme sequence. In this example, the server can use an RNN and LSTM-based G2P model to convert the text sequence into a phoneme sequence to obtain the phoneme sequence.

[0054] wherein the phoneme feature vector is to represent each phoneme vector by a corresponding spatial vector.

[0055] As an example, in step S302, the server performs spatial embedding coding on the phoneme sequence, i.e., represents each phoneme in the phoneme sequence by a corresponding spatial vector to obtain a phoneme feature vector. In this example, the obtained phoneme feature vector can be a continuous variable or a discrete variable. The continuous variable is represented as a related floating-point data format on data, and if it is a discrete variable representation, it can also be represented as a set integer vector or a floating-point data vector.

[0056] wherein the phoneme encoder is a functional module for implementing phoneme coding. As an example, the phoneme encoder can be constructed by using an LSTM layer or a transformer layer.

[0057] As an example, in step S303, the server can encode the phoneme feature vector by using the phoneme encoder, and determine the output result of the phoneme encoder as the hidden layer feature of the text sequence, i.e., the text hidden vector. In this example, the server can use a transformer layer-based phoneme encoder, which is provided with four layers of feedforward transformer layers. By using the four layers of feedforward transformer layers, the attention mechanism in the transformer layer is used to enhance the learning of the time sequence attention in the text sequence, and to improve the recognition accuracy of the text hidden vector obtained by the phoneme encoder.

[0058] In the speech synthesis method provided by the embodiment, the text sequence is first parsed to obtain a phoneme sequence formed by the smallest speech unit (i.e., a phoneme), thereby improving the basis for speech synthesis; then the phoneme sequence is subjected to spatial embedding coding to obtain a phoneme feature vector, so as to convert the phoneme sequence into a phoneme feature vector that can be subjected to coding calculation, thereby ensuring the coding feasibility; finally, the phoneme encoder is used to code the phoneme feature vector to obtain a text hidden vector output by the text sequence coding, thereby providing a text basis for subsequent speech synthesis. Understandably, when the phoneme encoder uses a four-layer feedforward transformer layer encoder, the attention mechanism in the transformer layer is used to enhance the learning of the timing attention in the text sequence, thereby improving the recognition accuracy of the text hidden vector obtained by the phoneme encoder coding.

[0059] In an embodiment, as shown in Figure 4 Step S202, the prosody reference audio is subjected to prosody feature extraction to obtain a prosody hidden vector, including:

[0060] S401: The prosody reference audio is subjected to prosody feature extraction to obtain a prosody style code;

[0061] S402: The prosody encoder is used to code the prosody style code and the prosody reference audio to obtain a prosody feature vector;

[0062] S403: The duration control module is used to perform duration alignment processing on the prosody feature vector to obtain a prosody hidden vector.

[0063] The prosody reference audio is audio that is pre-set to provide a prosody style as a reference object. The prosody style code is a code obtained by coding the prosody feature extracted from the prosody reference audio.

[0064] As an example, in step S401, the server can use the prosody reference audio set by default or selected by the user to extract the prosody style information related to the speech pronunciation manner of the prosody reference audio, code the extracted prosody style information, and obtain a prosody style code that can be subjected to subsequent model processing. The prosody style information is a prosody expression feature irrelevant to the text content. The prosody style code is a coding result of the prosody style information. In this example, the server uses an encoder-decoder mode, and specifically uses but is not limited to a Mel-GAN vocoder to extract the prosody feature of the prosody reference audio and obtain the prosody style code.

[0065] The Mel-GAN vocoder is a pre-trained vocoder, and a training process of the Mel-GAN vocoder includes the following steps: obtaining training text and training audio corresponding to the training text; performing spectrum extraction on the training audio to obtain a real spectrum of the training audio; generating audio according to the training text to obtain generated audio; performing spectrum extraction on the generated audio to obtain a generated spectrum corresponding to the generated audio; using a loss function to calculate a model loss value between the real spectrum and the generated spectrum; and if the model loss value is less than a preset value, determining that the Mel-GAN vocoder is trained.

[0066] The prosody encoder is an encoder for implementing prosody coding. As an example, the prosody encoder can adopt a two-layer two-dimensional convolutional network, which is more accurate in encoding output of a prosody feature vector than a one-layer two-dimensional convolutional network, and has higher processing efficiency than a multi-layer two-dimensional convolutional network.

[0067] As an example, in step S402, the server can use the prosody encoder to perform encoding processing on the input prosody style code and prosody reference audio. Specifically, the prosody reference audio can be first converted into a prosody spectrum, and then processed by a two-dimensional convolutional network to output spectrum feature information. The spectrum feature information and the prosody style code can be spliced or fused in other ways to output a prosody feature vector represented in a two-dimensional matrix form.

[0068] The duration control module is a module for implementing duration alignment processing.

[0069] As an example, in step S403, after obtaining the prosody feature vector represented in a two-dimensional matrix form, the server uses the duration control module to perform duration alignment on the prosody feature vector to obtain a prosody latent vector after duration alignment. In this example, the prosody feature vector is a two-dimensional matrix reflecting the prosodic phoneme-time correspondence relationship. When the duration control module is used to perform duration alignment on the prosody feature vector, the column vector can be expanded to obtain an expanded two-dimensional matrix as the prosody latent vector. Generally, the duration of audio is the size of the frame sequence, for example, 1 second of audio is divided into frames according to a frame block size of 20 ms to obtain a two-dimensional matrix with a size of 500, where the row vector represents the duration. In the prosody feature vector in a two-dimensional matrix form after coding based on the prosody style code and the prosody reference audio, a plurality of audio frames are correspondingly represented, and the prosody feature vector needs to be copied and expanded according to the prediction of the duration control module to output the prosody latent vector.

[0070] In the speech synthesis method provided by the embodiment, the prosody reference audio is subjected to prosody feature extraction according to the default prosody reference audio or the prosody reference audio selected by the user, and prosody style coding is obtained, so as to improve the basis of speech synthesis; then the prosody encoder is used to code the prosody style coding and the prosody reference audio, and prosody feature vectors are obtained, so as to ensure the feasibility of coding; finally, the time length control module is used to perform time length alignment processing on the prosody feature vectors, and prosody hidden vectors are obtained, so as to provide a text basis for subsequent speech synthesis. Understandably, the prosody encoder adopts a two-layer two-dimensional convolution network, and compared with a one-layer two-dimensional convolution network, the prosody feature vectors output by the two-layer two-dimensional convolution network are more accurate, and compared with a multi-layer two-dimensional convolution network, the efficiency of speech synthesis is improved.

[0071] In an embodiment, as shown in Figure 5 In step S203, the user coding vector corresponding to the user identifier is obtained, including:

[0072] S501: querying the identifier coding table based on the user identifier to determine whether the user identifier is a preset identifier in the identifier coding table;

[0073] S502: if the user identifier is the preset identifier, determining the preset coding vector corresponding to the preset identifier as the user coding vector corresponding to the user identifier;

[0074] S503: if the user identifier is not the preset identifier, obtaining the user audio data corresponding to the user identifier, and determining the user coding vector corresponding to the user identifier based on the user audio data.

[0075] The identifier coding table is an information table formed according to different preset identifiers in the training set and the preset coding vectors corresponding to the preset identifiers. The preset identifier is an identifier for uniquely identifying a user, which is set in advance and stored in the identifier coding table.

[0076] As an example, in step S501, after obtaining the user identifier, the server can query the preset identifier coding table based on the user identifier to determine whether the user identifier is a preset identifier in the identifier coding table.

[0077] The preset coding vector is a coding vector formed by feature extraction on preset audio data corresponding to the preset identifier, and the preset coding vector reflects the speaking habit of the user corresponding to the preset identifier.

[0078] As an example, in step S502, the server stores the preset encoding vector formed by encoding the preset audio data corresponding to the user identifier in the identification encoding table in association with the preset identifier when the user identifier is the preset identifier in the identification encoding table. Therefore, the preset encoding vector corresponding to the preset identifier can be directly determined as the user encoding vector corresponding to the user identifier, and the efficiency of obtaining the user encoding vector can be improved.

[0079] The audio data corresponding to the user identifier is audio data formed by the user speaking corresponding to the user identifier.

[0080] As an example, in step S503, the server needs to obtain the user audio data corresponding to the user identifier in real time when the user identifier is not the preset identifier in the identification encoding table. Then, the user encoding vector corresponding to the user identifier is determined by encoding the user audio data, and the real-time performance of obtaining the user encoding vector is ensured.

[0081] In the example, the server is provided with a user encoding module. The user encoding module is a coding module for distinguishing the speech synthesis of different users. The user encoding module can be implemented in multiple ways. The simplest way is to use a digital number ID as a user identifier to distinguish different users, and to obtain a user encoding vector by using an embedding layer. This way of obtaining a user encoding vector has the advantages of convenience and simplicity. The user encoding module can also use acoustic features such as X-vector and d-vector to distinguish different user voices and obtain a user encoding vector. The user encoding vector can reflect the voice tone, and the synthesis effect of the target audio file formed by subsequent speech synthesis can be ensured. The user identifier corresponds to the user, and the naturalness of speech synthesis is improved.

[0082] In the speech synthesis method provided in the embodiment, whether the user identifier is a preset identifier in the identification encoding table is determined. When the user identifier is the preset identifier, the preset encoding vector corresponding to the preset identifier is directly determined as the user encoding vector corresponding to the user identifier, and the efficiency of obtaining the user encoding vector is improved. When the user identifier is not the preset identifier, the user audio data is encoded to obtain the user encoding vector, and the real-time performance of obtaining the user encoding vector is ensured. It can be understood that the user encoding vector corresponding to the user identifier is determined, the user encoding vector can reflect the voice tone, the synthesis effect of the target audio file formed by subsequent speech synthesis can be ensured, the user identifier corresponds to the user, and the naturalness of speech synthesis is improved.

[0083] In an embodiment, as shown in Figure 6As shown, step S503, determining the user code vector corresponding to the user identifier based on the user audio data, comprises:

[0084] S601: performing feature extraction on the user audio data to obtain a first spectral feature;

[0085] S602: segmenting the first spectral feature to obtain N second spectral features;

[0086] S603: sequentially outputting the N second spectral features to the convolutional neural network for processing to obtain a first hidden vector corresponding to the N second spectral features;

[0087] S604: performing mean and variance calculation on the first hidden vector corresponding to the N second spectral features to determine a hidden vector mean and a hidden vector variance corresponding to the N second spectral features;

[0088] S605: concatenating the hidden vector mean and the hidden vector variance corresponding to the N second spectral features to obtain a second hidden vector;

[0089] S606: inputting the second hidden vector to the convolutional neural network for processing to obtain a user code vector corresponding to the user identifier.

[0090] Wherein, the first spectral feature is a spectral feature extracted from the user audio data.

[0091] As an example, in step S601, the server obtains the user audio data corresponding to the user identifier in real time, performs feature extraction on the user audio data, and specifically extracts the spectral feature corresponding to the user audio data as the first spectral feature, providing a basis for subsequent code vector generation.

[0092] Wherein, the second spectral feature is a spectral feature obtained by segmenting the first spectral feature.

[0093] As an example, in step S602, the server segments the extracted first spectral feature according to a pre-set spectral feature segmentation strategy, and takes each segmented spectral feature as a second spectral feature. The spectral feature segmentation strategy is a strategy pre-set for segmenting spectral features, which can be segmenting according to a fixed time length or segmenting according to a user-defined segmentation standard.

[0094] The first convolutional neural network is a convolutional neural network used for processing the second spectral feature. In this example, the second convolutional neural network is a DNN network, and the DNN network (Deep Neural Networks) is a deep neural network. The neural network layers in the DNN can be divided into three categories: input layer, hidden layer, and output layer. The hidden layer in the middle can be divided into multiple layers. The first hidden vector is the output value of the second spectral feature after the convolutional neural network.

[0095] As an example, in step S602, the server inputs each segmented second spectral feature into the first convolutional neural network in turn. Specifically, a convolutional neural network composed of 9 fully connected layers is input. The output value of the convolutional neural network is determined as the first hidden vector corresponding to the second spectral feature. The first hidden vector here can be understood as an intermediate variable output by the convolutional neural network after processing the second spectral feature.

[0096] As an example, in step S603, after obtaining the first hidden vectors corresponding to the N second spectral features, the server can use the mean calculation formula and the variance calculation formula to calculate the first hidden vectors corresponding to the N second spectral features, respectively, to determine the hidden vector mean and the hidden vector variance corresponding to the N second spectral features.

[0097] As an example, in step S604, after obtaining the hidden vector mean and the hidden vector variance corresponding to the N second spectral features, the server can use a specific splicing strategy or according to a pre-set splicing order to splice the hidden vector mean and the hidden vector variance corresponding to the N second spectral features, and determine the splicing result as the second hidden vector.

[0098] The second convolutional neural network is a convolutional neural network used for processing the second hidden vector. The user encoding vector is a voiceprint recognition vector X-vector, which can accept an input of any length and convert it into a fixed-length feature expression.

[0099] As an example, in step S606, the server inputs the second hidden vector into the second convolutional neural network for processing. Specifically, the second hidden vector is input into a 4-layer second convolutional neural network to obtain a voiceprint recognition vector X-vector. The voiceprint recognition vector X-vector is used as the user encoding vector corresponding to the user identifier. In this example, the user encoding vector can be directly generated without passing through the fully connected layer.

[0100] In the speech synthesis method provided by the embodiment, the user audio data is first subjected to feature extraction to obtain a first spectrum feature to provide a basis for subsequent encoding vector synthesis; then the second spectrum feature is segmented so that each second spectrum feature has less information quantity, facilitating subsequent processing and ensuring the feasibility of encoding generation; then the second spectrum feature is input into the first convolutional neural network, and the first hidden vector output by the first convolutional neural network is subjected to mean and variance calculation and splicing to obtain a second hidden vector to provide a basis for subsequent encoding vector synthesis; finally, the second hidden vector is input into the second convolutional neural network to obtain a user encoding vector. The generated user encoding vector is a voiceprint recognition vector X-vector, which can accept an input of any length and convert it into a fixed-length feature expression. In the convolutional neural network training, a data enhancement strategy including noise and reverberation is introduced to make the model stronger against noise and reverberation interference.

[0101] In an embodiment, as shown in Figure 7 Step S204, i.e., synthesizing the text hidden vector, the prosody hidden vector and the user encoding vector to obtain the target acoustic feature, includes:

[0102] S701: processing the text hidden vector and the prosody hidden vector by using an attention mechanism to obtain a fusion hidden vector;

[0103] S702: synthesizing the fusion hidden vector and the user encoding vector to obtain the target acoustic feature.

[0104] The attention mechanism used is cross attention, which can be understood as calculating a weight based on similarity and then performing weighted averaging. The fusion hidden vector is a vector obtained by synthesizing the text hidden vector and the prosody hidden vector through the attention mechanism.

[0105] As an example, in step S701, the server uses the attention mechanism, specifically the cross attention mechanism, to process the text hidden vector and the prosody hidden vector. The text hidden vector can be used as the query of the attention mechanism, and the prosody hidden vector can be used as the key of the attention mechanism to calculate the attention score. The calculated attention score is determined as the fusion hidden vector.

[0106] As an example, in step S702, the server inputs the fusion hidden vector and the user encoding vector into the decoder for decoding. The output result of the decoder is determined as the target acoustic feature.

[0107] In the speech synthesis method provided by the embodiment, the attention mechanism is used to synthesize the text hidden vector and the prosody hidden vector, to obtain a fusion hidden vector, to prepare for obtaining a target acoustic feature. Finally, the fusion vector and a user encoding vector are input into a decoder for decoding to obtain the target acoustic feature, so that the obtained target acoustic feature is not only related to the text hidden vector and the prosody hidden vector, but also related to the user encoding vector, which helps to guarantee the speech synthesis effect of the target audio file and improve the naturalness of the synthesized speech.

[0108] In an embodiment, as shown in FIG. 7, S701, i.e., using the attention mechanism to process the text hidden vector and the prosody hidden vector to obtain a fusion hidden vector, includes the following steps: Figure 8

[0109] S801: using the attention mechanism to calculate the similarity of the text hidden vector and the prosody hidden vector to obtain a vector similarity;

[0110] S802: using a softmax layer to normalize the vector similarity to obtain a prosody weight value;

[0111] S803: performing weighted processing on the text hidden vector and the prosody weight value to obtain a fusion hidden vector.

[0112] The vector similarity is obtained by calculating the vector similarity of the text hidden vector and the prosody hidden vector. The softmax layer is an activation function that can normalize a numerical vector into a probability distribution vector.

[0113] As an example, in step S801, the server uses the attention mechanism to calculate the similarity of the text hidden vector and the prosody hidden vector to obtain a vector similarity F(query, key), where query is the text hidden vector, key is the prosody hidden vector, and F is the similarity calculation of the text hidden vector and the prosody hidden vector, to obtain the vector similarity. The similarity calculation can be completed by dot product or implemented by a fully connected layer.

[0114] As an example, in step S802, after calculating the vector similarity F(query, key) corresponding to the text hidden vector and the prosody hidden vector, the server can use the softmax layer to normalize the vector similarity F(query, key) to a value between 0 and 1, to determine the prosody weight value, i.e., Softmax(F(query, key)).

[0115] ​As an example, in step S803, the server further performs a weighted summation of the text latent vector α and the prosodic weight value Softmax(F(query,key)), i.e., ∑α*Softmax(F(query,key)), to obtain the fused latent vector. In this example, the text latent vector and the prosodic latent vector need to be continuously optimized through training on the training set.

[0116] In the speech synthesis method provided in this embodiment, the server uses an attention mechanism to calculate the similarity between the text latent vector and the prosodic latent vector, which demonstrates the feasibility of the solution. Then, a softmax layer is used to normalize the vector similarity to obtain the prosodic weight value. Finally, the prosodic weight value and the text latent vector are weighted and summed to obtain the fused latent vector, which lays the foundation for synthesis with user encoding and prepares for enhancing prosodic style control in multi-user speech synthesis.

[0117] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0118] In one embodiment, a speech synthesis device is provided, which corresponds one-to-one with the speech synthesis methods described in the above embodiments. For example... Figure 9 As shown, the speech synthesis device includes a text latent vector acquisition module 901, a prosodic latent vector acquisition module 902, a user-encoded vector acquisition module 903, a target acoustic feature acquisition module 904, and a target audio file acquisition module 905. Detailed descriptions of each functional module are as follows:

[0119] The text latent vector acquisition module 901 is used to process text sequences and acquire text latent vectors;

[0120] The prosodic latent vector acquisition module 902 is used to extract prosodic features from the prosodic reference audio and obtain the prosodic latent vector.

[0121] User encoding vector acquisition module 903 is used to acquire the user encoding vector corresponding to the user identifier;

[0122] The target acoustic feature acquisition module 904 is used to synthesize the text latent vector, prosodic latent vector and user-encoded vector using an attention mechanism to acquire the target acoustic features;

[0123] The target audio file acquisition module 905 is used to perform speech synthesis based on the target acoustic features and acquire the target audio file corresponding to the text sequence.

[0124] In one embodiment, the text latent vector acquisition module 901 includes:

[0125] The phoneme sequence acquisition unit is configured to parse the text sequence and acquire a phoneme sequence.

[0126] The phoneme feature vector acquisition unit is configured to perform spatial vector conversion on the phoneme sequence and acquire a phoneme feature vector.

[0127] The text hidden vector acquisition unit is configured to encode the phoneme feature vector and acquire a text hidden vector.

[0128] In an embodiment, the prosody hidden vector acquisition module 902 comprises:

[0129] The prosody style code acquisition unit is configured to perform prosody feature extraction on the prosody reference audio and acquire a prosody style code.

[0130] The prosody feature vector acquisition unit is configured to encode the prosody style code and the prosody reference audio and acquire a prosody feature vector.

[0131] The prosody hidden vector acquisition unit is configured to perform duration alignment processing on the prosody feature vector and acquire a prosody hidden vector.

[0132] In an embodiment, the user code vector acquisition module 903 comprises:

[0133] The preset identifier judgment unit is configured to query an identifier code table based on the user identifier and judge whether the user identifier is a preset identifier in the identifier code table.

[0134] The first user code vector determination unit is configured to, if the user identifier is the preset identifier, determine a preset code vector corresponding to the preset identifier as a user code vector corresponding to the user identifier.

[0135] The second user code vector determination unit is configured to, if the user identifier is not the preset identifier, acquire user audio data corresponding to the user identifier, and determine a user code vector corresponding to the user identifier based on the user audio data.

[0136] In an embodiment, the second user code vector determination unit comprises:

[0137] The first spectral feature acquisition subunit is configured to perform feature extraction on the user audio data and acquire a first spectral feature.

[0138] The second spectral feature acquisition subunit is configured to segment the first spectral feature and acquire N second spectral features.

[0139] The first hidden vector acquisition subunit is configured to sequentially output the N second spectral features to a first convolutional neural network for processing and acquire first hidden vectors corresponding to the N second spectral features.

[0140] a mean-variance determination subunit, configured to perform mean and variance calculation on the first hidden vectors corresponding to the N second spectral features, to determine hidden vector mean and hidden vector variance corresponding to the N second spectral features;

[0141] a second hidden vector acquisition subunit, configured to splice the hidden vector mean and the hidden vector variance corresponding to the N second spectral features, to acquire a second hidden vector;

[0142] a user encoding vector acquisition subunit, configured to input the second hidden vector into the second convolutional neural network for processing, to acquire a user encoding vector corresponding to the user identifier.

[0143] In an embodiment, the target acoustic feature acquisition module 904 comprises:

[0144] a fusion hidden vector acquisition unit, configured to process the text hidden vector and the prosody hidden vector by using an attention mechanism, to acquire a fusion hidden vector;

[0145] a target acoustic feature acquisition unit, configured to synthesize the fusion hidden vector and the user encoding vector, to acquire a target acoustic feature.

[0146] In an embodiment, the fusion hidden vector acquisition unit comprises:

[0147] a vector similarity acquisition subunit, configured to perform similarity calculation on the text hidden vector and the prosody hidden vector by using an attention mechanism, to acquire a vector similarity;

[0148] a prosody weight value acquisition subunit, configured to perform normalization processing on the vector similarity, to acquire a prosody weight value;

[0149] a fusion hidden vector acquisition subunit, configured to perform weighted processing based on the text hidden vector and the prosody weight value, to acquire a fusion hidden vector.

[0150] The specific limitations on the speech synthesis device can be referred to the limitations on the speech synthesis method in the foregoing, which will not be repeated here. Each module in the speech synthesis device described above can be realized by software, hardware and combinations thereof, in whole or in part. Each module described above can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to each module.

[0151] In an embodiment, a computer device is provided, which can be a server, and an internal structure diagram of the computer device can be as shown in Figure 9The computer device includes a processor, a memory, a network interface and a database connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store data used or generated in the process of executing the speech synthesis method. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is configured to be executed by the processor to implement a speech synthesis method.

[0152] In an embodiment, a computer device is provided, including a memory, a processor and a computer program stored in the memory and executable on the processor, the processor being configured to implement the speech synthesis method in the above embodiments when executing the computer program, for example Figure 2 S201-S205, or Figures 3 to 8 For brevity, details are not repeated here. Alternatively, the processor is configured to implement the functions of each module / unit in the embodiment of the speech synthesis device when executing the computer program, for example Figure 9 For brevity, details are not repeated here. Alternatively, the processor is configured to implement the functions of each module / unit in the embodiment of the speech synthesis device when executing the computer program, for example

[0153] In an embodiment, a computer readable storage medium is provided, the computer readable storage medium storing a computer program, the computer program being configured to implement the speech synthesis method in the above embodiments when executed by a processor, for example Figure 2 S201-S205, or Figures 3 to 8 For brevity, details are not repeated here. Alternatively, the processor is configured to implement the functions of each module / unit in the embodiment of the speech synthesis device when executing the computer program, for example Figure 9 For brevity, details are not repeated here. Alternatively, the processor is configured to implement the functions of each module / unit in the embodiment of the speech synthesis device when executing the computer program, for example

[0154] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0155] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0156] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, but not limit it. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features. The modification or replacement does not make the essence of the corresponding technical solution deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A speech synthesis method characterized by, The method comprises the following steps: processing a text sequence to obtain a text hidden vector; extracting prosody features from a prosody reference audio to obtain prosody style information related to the prosody of the prosody reference audio, encoding the extracted prosody style information to obtain prosody style encoding; the prosody reference audio is an audio that is pre-set to provide a prosody style as a reference object; the prosody style information is a prosody expression feature that is independent of the text content; the prosody style encoding is the encoding result of the prosody style information; performing spectral conversion on the prosody reference audio to obtain prosody reference spectrum, processing the prosody reference spectrum using a two-dimensional convolution network to output spectral feature information, and splicing or fusing the spectral feature information and the prosody style encoding to output a prosody feature vector in the form of a two-dimensional matrix; the prosody feature vector is a two-dimensional matrix reflecting the prosody phoneme-time correspondence; using a time length control module to perform time length alignment on the prosody feature vector, and expanding the column vector to obtain an expanded two-dimensional matrix as a prosody hidden vector; obtaining a user code vector corresponding to a user identifier; synthesizing the text hidden vector, the prosody hidden vector, and the user code vector using an attention mechanism to obtain a target acoustic feature; performing speech synthesis based on the target acoustic feature to obtain a target audio file corresponding to the text sequence.

2. The speech synthesis method of claim 1, wherein, The method comprises the following steps: analyzing the text sequence to obtain a phoneme sequence; performing spatial vector conversion on the phoneme sequence to obtain a phoneme feature vector; encoding the phoneme feature vector to obtain a text hidden vector.

3. The speech synthesis method of claim 1, wherein, The method comprises the following steps: querying an identifier code table based on the user identifier to determine whether the user identifier is a preset identifier in the identifier code table; if the user identifier is the preset identifier, determining a preset code vector corresponding to the preset identifier as the user code vector corresponding to the user identifier; if the user identifier is not the preset identifier, obtaining user audio data corresponding to the user identifier, and determining the user code vector corresponding to the user identifier based on the user audio data.

4. The speech synthesis method of claim 3, wherein, The method comprises the following steps: extracting features from the user audio data to obtain a first spectral feature; segmenting the first spectral feature to obtain N second spectral features; outputting the N second spectral features to a first convolutional neural network in sequence for processing to obtain first hidden vectors corresponding to the N second spectral features; performing mean and variance calculation on the first hidden vectors corresponding to the N second spectral features to determine hidden vector mean and hidden vector variance corresponding to the N second spectral features; splicing the hidden vector mean and the hidden vector variance corresponding to the N second spectral features to obtain a second hidden vector; inputting the second hidden vector into a second convolutional neural network for processing to obtain the user code vector corresponding to the user identifier.

5. The speech synthesis method of claim 1, wherein, The attention mechanism is used to synthesize the text hidden vector, the prosody hidden vector and the user code vector to obtain a target acoustic feature. The attention mechanism is used to process the text hidden vector and the prosody hidden vector to obtain a fusion hidden vector. The fusion hidden vector and the user code vector are synthesized to obtain a target acoustic feature.

6. The speech synthesis method of claim 5, wherein, The attention mechanism is used to process the text hidden vector and the prosody hidden vector to obtain a fusion hidden vector, including: The attention mechanism is used to calculate the similarity of the text hidden vector and the prosody hidden vector to obtain a vector similarity. The vector similarity is normalized to obtain a prosody weight value. The text hidden vector and the prosody weight value are weighted to obtain a fusion hidden vector.

7. A speech synthesis apparatus characterized by comprising: It includes: A text hidden vector acquisition module is configured to process a text sequence to obtain a text hidden vector. A prosody hidden vector acquisition module is configured to extract prosody features from a prosody reference audio, extract prosody style information related to the prosody reference audio from the prosody reference audio, encode the extracted prosody style information, and obtain prosody style encoding. The prosody reference audio is an audio that is set in advance to provide a prosody style as a reference object. The prosody style information is a prosody expression feature that is independent of the text content. The prosody style encoding is an encoding result of the prosody style information. The prosody reference audio is converted into a spectrum to obtain a prosody reference spectrum. A two-dimensional convolution network is used to process the prosody reference spectrum to output spectral feature information. The spectral feature information and the prosody style encoding are spliced or fused to output a prosody feature vector represented in a two-dimensional matrix form. The prosody feature vector is a two-dimensional matrix reflecting the prosodic phoneme-time correspondence. A time length control module is used to align the time length of the prosody feature vector, expand the column vector, and obtain an expanded two-dimensional matrix as a prosody hidden vector. A user code vector acquisition module is configured to obtain a user code vector corresponding to a user identifier. A target acoustic feature acquisition module is configured to use an attention mechanism to synthesize the text hidden vector, the prosody hidden vector and the user code vector to obtain a target acoustic feature. A target audio file acquisition module is configured to perform speech synthesis based on the target acoustic feature to obtain a target audio file corresponding to the text sequence.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the speech synthesis method of any one of claims 1-6.

9. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 8. The computer program is executed by the processor to implement the speech synthesis method of any one of claims 1-6.

Citation Information

Patent Citations

  • Speaker recognition method, device and equipment and storage medium

    CN111508505A

  • Voice synthesis method and device thereof, electronic equipment, storage medium and program product

    CN114005428A