Text-guided speech synthesis method and device, computer equipment and storage medium

By performing style labeling and scene noise injection on the speech data sets of different languages, and training the acoustic model, the existing speech synthesis model has solved the problems of single language, limited style description and low efficiency and quality, and achieved efficient and high-quality speech synthesis in multilingual and multi-scene.

CN120015011AActive Publication Date: 2025-05-16PING AN TECH (SHENZHEN) CO LTD

Patent Information

Application Number
CN202510192011.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-16
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The existing pronunciation synthesis model has problems such as single language, limited style description, and poor efficiency and quality of pronunciation synthesis.

Method used

By obtaining speech data sets and text data sets of different languages, performing style label annotation and scene noise injection, inputting them into the pre-constructed acoustic model, training style encoder, reference encoder, text encoder, acoustic structure and vocoder, to obtain the final speech synthesis model.

Benefits of technology

Multilingual text-guided speech synthesis is realized, which improves the generalization ability and robustness of the model in a noisy environment, improves the naturalness, fidelity and multi-scene applicability of speech synthesis, and improves the efficiency and real-timeness of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015011A_ABST
    Figure CN120015011A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to a text-guided speech synthesis method, which comprises the following steps: carrying out style label labeling and scene noise injection on a speech data set to obtain a reference speech set; inputting the reference voice set and the text data set into an acoustic model; encoding the style label through a style encoder to obtain style encoding features; encoding the reference voice through a reference encoder to obtain reference voice encoding features; encoding the text through a text encoder to obtain text encoding features; inputting all the coding features into an acoustic structure to obtain voice acoustic features; and inputting the voice acoustic features into a vocoder to synthesize a waveform, obtaining a predicted synthesized voice, and training the predicted synthesized voice to obtain a voice synthesis model. The invention further provides a text-guided speech synthesis device, computer equipment and a storage medium. In addition, the invention also relates to a block chain technology, and the to-be-converted text can be stored in a block chain. The speech synthesis efficiency and quality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a text-guided speech synthesis method, apparatus, computer equipment and storage medium. Background Art

[0002] With the continuous breakthroughs in artificial intelligence technology, speech synthesis large models have ushered in unprecedented development in recent years. As an important means of human-computer interaction, speech synthesis technology has penetrated into multiple fields such as intelligent customer service, voice assistants, education and training, greatly improving user experience and efficiency. Especially in scenarios such as financial services and medical services, speech synthesis technology is used in intelligent voice customer service to answer financial knowledge, promote medical knowledge, etc. Its natural and smooth voice and personalized expression ability bring customers more humanized financial, medical and other scenario services. Especially in the financial field, speech synthesis technology is also widely used in scenarios such as telephone banking and smart investment advisors, which enhances customer trust and satisfaction by synthesizing realistic voices.

[0003] In recent years, with people's increasing attention to privacy protection and the increasing demand for diversified and flexible speech synthesis audio, an innovative "text-guided speech synthesis method" has emerged. This method achieves precise control and personalized customization of speech synthesis by introducing two key prmopts (i.e. input parameters), namely content prompt (content prompt, which refers to the text to be synthesized) and style prompt (style prompt, which refers to the text describing the speech style). Content prompts are the specific text content that users want to synthesize, while style prompts are used to describe the desired speech style, such as speaking speed, pitch, emotion, etc.

[0004] However, although text-guided speech synthesis methods have made significant progress in the field of speech synthesis, existing models still have a series of defects and shortcomings. First, the language monotony problem is a major limitation of current models. Existing large speech synthesis models can usually only process content cues and style cues in a single language, and cannot achieve the fusion of multilingual content and style conversion, which limits the wide application of models in the context of globalization. Secondly, the limited description of style is also a major shortcoming of current models. When using style cues to describe speech style, the model is mainly limited to several aspects such as speech speed, pitch, emotion, signal-to-noise ratio, gender, accent, etc., and lacks the ability to simulate speech style in different noise scenarios. Therefore, it is difficult to synthesize realistic audio in complex and diverse real-life scenarios (such as making phone calls on the roadside, eating in restaurants, etc.). Finally, the limitations of the acoustic model also restrict the performance of large speech synthesis models. At present, the acoustic model in the model is mainly based on the diffusion model, which has certain shortcomings in synthesis speed, audio quality and adaptability to specific scenarios, and it is difficult to meet users' needs for efficient and high-quality speech synthesis. Summary of the invention

[0005] The purpose of the embodiments of the present application is to propose a text-guided speech synthesis method, apparatus, computer device and storage medium to solve the technical problems of the existing speech synthesis with a single language, limited style description, and poor efficiency and quality of speech synthesis.

[0006] In order to solve the above technical problems, the present application embodiment provides a text-guided speech synthesis method, which adopts the following technical solution:

[0007] Acquire speech data sets and text data sets in different languages, wherein the speech data in the speech data sets and the text data in the text data sets are paired data;

[0008] Performing style labeling and scene noise injection on the speech data set to obtain a reference speech set;

[0009] Inputting the reference speech set and the text data set into a pre-built acoustic model, wherein the acoustic model includes a style encoder, a reference encoder, a text encoder, an acoustic structure, and a vocoder;

[0010] Encoding the style labels of the text data set and the reference speech set by the style encoder to obtain style coding features;

[0011] Encoding the reference speech of the reference speech set by the reference encoder to obtain reference speech coding features;

[0012] Encoding the text dataset by the text encoder to obtain text encoding features;

[0013] Inputting the style coding feature, the reference speech coding feature and the text coding feature into the acoustic structure to obtain speech acoustic features;

[0014] Performing waveform synthesis on the acoustic features of the speech by the vocoder to obtain predicted synthesized speech;

[0015] According to a preset loss function, the loss is calculated according to the reference speech and the predicted synthesized speech, model parameters are adjusted based on the loss, and iterative training is continued until an iteration stop condition is met to obtain a final speech synthesis model;

[0016] The text to be converted is obtained and input into the speech synthesis model to obtain the target synthesized speech.

[0017] Furthermore, the step of performing style tagging and scene noise injection on the speech data set to obtain a reference speech set includes:

[0018] Determining a style label for each piece of speech data in the speech data set according to a preset style dimension, and annotating the speech data based on the style label to obtain an annotated speech set;

[0019] A scene noise data set is obtained, and a data fusion algorithm is used to fuse the scene noise data set with the annotated speech set to obtain a reference speech set containing noise.

[0020] Furthermore, the step of determining the style label of each speech data in the speech data set according to the preset style dimension includes:

[0021] Analyzing the speech data using a speech signal processing algorithm to obtain corresponding speech speed, pitch and signal-to-noise ratio;

[0022] Determine a speech rate label and a tone label of the speech data according to the speech rate and the tone;

[0023] Comparing the signal-to-noise ratio with a preset signal-to-noise ratio threshold to obtain a signal-to-noise ratio label;

[0024] Input the speech data set into a trained emotion recognition model for emotion classification to obtain corresponding emotion labels;

[0025] Using the trained gender recognition model, determine the gender of the speaker in the speech data set and obtain a corresponding gender label;

[0026] Inputting the speech data set into a trained accent recognition model, determining the accent type of the speech data, and obtaining an accent label based on the accent type;

[0027] The speaking rate label, the pitch label, the signal-to-noise ratio label, the emotion label, the gender label, and the accent label are aggregated to obtain a style label.

[0028] Furthermore, the style encoder includes a BERT embedding layer, a spatial expansion layer and a style encoding layer, and the step of encoding the style tags of the text dataset and the reference speech set by the style encoder to obtain style encoding features includes:

[0029] Performing vector conversion on the style labels of the text dataset and the reference speech set through the BERT embedding layer to obtain a style embedding vector;

[0030] Inputting the style embedding vector into the spatial extension layer, and obtaining a style extension vector by concatenating the style embedding vector and the introduced style hint vector;

[0031] The style features in the style extension vector are extracted through the self-attention mechanism of the style encoding layer to obtain style encoding features.

[0032] Furthermore, the text encoder includes a text embedding layer, a Transformer encoding layer and a pooling layer, and the step of encoding the text dataset by the text encoder to obtain text encoding features includes:

[0033] Performing vector embedding on the text data in the text data set through the text embedding layer to obtain a text embedding vector;

[0034] Inputting the text embedding vector into the Transformer encoding layer, extracting semantic features through a multi-layer self-attention mechanism, and obtaining semantic encoding features;

[0035] The semantic coding features are pooled by the pooling layer to obtain text coding features.

[0036] Furthermore, the acoustic structure includes a reversible transform layer, a stream decoder and a residual layer, and the step of inputting the style coding feature, the reference speech coding feature and the text coding feature into the acoustic structure to obtain the speech acoustic feature includes:

[0037] Performing feature transformation on the style coding feature, the reference speech coding feature and the text coding feature through the reversible transformation layer to obtain a transformation coding feature;

[0038] Decoding the transform coded features by the stream decoder to obtain an acoustic feature sequence;

[0039] The acoustic feature sequence is input into the residual layer for residual connection to obtain speech acoustic features.

[0040] Furthermore, the step of calculating the loss according to the reference speech and the predicted synthesized speech according to a preset loss function includes:

[0041] Calculating the adversarial loss and L1 loss between the reference speech and the predicted synthesized speech;

[0042] Calculating the similarity between the style coding feature and the reference speech coding feature to obtain a feature correlation loss;

[0043] The adversarial loss, the L1 loss and the feature-related loss are weightedly summed to obtain the final loss.

[0044] In order to solve the above technical problems, the present application also provides a text-guided speech synthesis device, which adopts the following technical solution:

[0045] An acquisition module, used to acquire speech data sets and text data sets in different languages, wherein the speech data in the speech data set and the text data in the text data set are paired data;

[0046] A label injection module, used for labeling the speech data set with style labels and injecting scene noise to obtain a reference speech set;

[0047] An input module, configured to input the reference speech set and the text data set into a pre-built acoustic model, wherein the acoustic model includes a style encoder, a reference encoder, a text encoder, an acoustic structure, and a vocoder;

[0048] A style encoding module, used for encoding the style labels of the text data set and the reference speech set through the style encoder to obtain style encoding features;

[0049] A speech encoding module, used for encoding the reference speech of the reference speech set by the reference encoder to obtain reference speech encoding features;

[0050] A text encoding module, used for encoding the text data set through the text encoder to obtain text encoding features;

[0051] An acoustic feature extraction module, used for inputting the style coding feature, the reference speech coding feature and the text coding feature into the acoustic structure to obtain speech acoustic features;

[0052] A speech prediction module, used for performing waveform synthesis on the speech acoustic features through the vocoder to obtain predicted synthesized speech;

[0053] An iteration module, configured to calculate a loss according to the reference speech and the predicted synthesized speech according to a preset loss function, adjust model parameters based on the loss, and continue iterative training until an iteration stop condition is met, thereby obtaining a final speech synthesis model;

[0054] The speech synthesis module is used to obtain the text to be converted, input it into the speech synthesis model, and obtain the target synthesized speech.

[0055] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:

[0056] The computer device comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the text-guided speech synthesis method as described above when executing the computer-readable instructions.

[0057] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0058] The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the text-guided speech synthesis method as described above.

[0059] Compared with the prior art, this application has the following beneficial effects:

[0060] The present application provides a text-guided speech synthesis method. By acquiring speech datasets and text datasets in different languages, and annotating the speech datasets with style tags and injecting scene noise, multi-language text-guided speech synthesis can be achieved. The model can learn to accurately recognize and process speech signals in a noisy environment, thereby improving the generalization ability and robustness of the model, while simulating real scenes to improve the multi-scenario applicability of speech synthesis. By inputting a reference speech set and a text dataset into a pre-built acoustic model, training the style encoder, reference encoder, text encoder, acoustic structure and vocoder of the acoustic model, and obtaining the final speech synthesis model, the naturalness and realism of speech synthesis can be improved, the adaptability and generalization ability of the model can be enhanced, and the efficiency and real-time performance of speech synthesis can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the scheme in the present application, a brief introduction is given below to the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0062] Figure 1is an exemplary system architecture diagram to which the present application may be applied;

[0063] Figure 2 is a flow chart of an embodiment of a text-guided speech synthesis method according to the present application;

[0064] Figure 3 is a structural schematic diagram of an embodiment of a text-guided speech synthesis device according to the present application;

[0065] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION

[0066] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by technicians in the technical field of the present application; the terms used in the specification of the application herein are only for the purpose of describing specific embodiments and are not intended to limit the present application; the terms "including" and "having" and any variations thereof in the specification and claims of the present application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of the present application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0067] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0068] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0069] like Figure 1 As shown, the system architecture 100 may include a terminal device 101, a network 102 and a server 103. The terminal device 101 may be a laptop 1011, a tablet computer 1012 or a mobile phone 1013. The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links or optical fiber cables.

[0070] The user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0071] The terminal device 101 can be any electronic device with a display screen and supporting web browsing. In addition to a laptop computer 1011, a tablet computer 1012 or a mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV), a laptop computer, a desktop computer, etc.

[0072] The server 103 may be a server that provides various services, such as a background server that provides support for a web page displayed on the terminal device 101 .

[0073] It should be noted that the text-guided speech synthesis method provided in the embodiment of the present application is generally executed by a server / terminal device, and accordingly, the text-guided speech synthesis device is generally arranged in the server / terminal device.

[0074] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0075] Continue to refer Figure 2 , shows a flowchart of an embodiment of a text-guided speech synthesis method according to the present application, comprising the following steps:

[0076] Step S201 , obtaining a speech data set and a text data set in different languages, wherein the speech data in the speech data set and the text data in the text data set are paired data.

[0077] The speech data sets in different languages ​​contain a large amount of speech data in different languages, including but not limited to Chinese, English, Korean, Japanese, etc. Among them, Chinese also includes speech data in multiple dialects.

[0078] Speech datasets can be obtained from public open source speech datasets, such as the Chinese speech open source dataset Aishell3 and the English speech open source dataset LibriTTS. The Aishell3 dataset covers a large number of rich and diverse Chinese speech samples, including Chinese speech data with multi-dimensional features such as different genders, ages, accents, and emotional expressions.

[0079] The original speech data set is preprocessed, including noise removal, silent segment removal, amplitude normalization, voice quality adjustment, etc., to obtain the preprocessed speech data set. The text data set corresponding to the speech data set is obtained through the speech recognition algorithm, and the speech data set is annotated and the text data set is matched with the corresponding acoustic features to establish a mapping relationship between text and speech, providing supervised learning data for subsequent model training.

[0080] The choice of speech dataset depends on the specific scenario and needs. For example, in the financial field, the speech dataset can be speech data containing financial professional terms, financial market data, stock trading information, financial reports, etc., or it can be recordings of conversations between financial customer service and customers; in the medical field, the speech dataset can be speech data containing medical diagnosis, treatment plans, drug names, disease names, etc., or it can be speech data containing medical guidance information such as drug instructions, health examination instructions, disease prevention suggestions, etc., or it can be recordings of conversations between doctors and patients of different ages and genders during the diagnosis and treatment process.

[0081] Step S202 , annotating the speech data set with style labels and injecting scene noise to obtain a reference speech set.

[0082] Specifically, the style label of each speech data in the speech data set is determined according to the preset style dimension, and the speech data is annotated based on the style label to obtain a labeled speech set; a scene noise data set is obtained, and a data fusion algorithm is used to fuse the scene noise data set with the annotated speech set to obtain a reference speech set containing noise.

[0083] The preset style dimensions include speaking speed, pitch, emotion, signal-to-noise ratio, gender, accent and other dimensions. The style label of each speech data in the speech data set is analyzed by the preset speech processing algorithm, and then the corresponding style label is annotated on the speech data using the annotation tool, that is, the style labels include speaking speed label, pitch label, signal-to-noise ratio label, emotion label, gender label and accent label.

[0084] According to the characteristics of the noise environment in the actual application scenario, data with similar noise characteristics are screened from open source datasets such as AudioSet to construct a scene noise dataset; a data fusion algorithm is used to fuse the annotated speech set with the screened scene noise dataset to generate a reference speech set containing noise. The reference speech set is speech data with a sense of real scene.

[0085] Among them, the data fusion algorithm can adopt a weighted average algorithm, a signal superposition method, a Gaussian mixture model (GMM), etc.

[0086] In some optional implementations, the reference speech set containing noise is enriched and expanded through data enhancement technology, including changing the speech speed, volume, pitch, etc., to obtain an enhanced reference speech set.

[0087] Through style label annotation and scene noise injection, the description of special noise scenes is further expanded on the basis of style dimension, thereby greatly enriching the style diversity of speech synthesis. In addition, the model can learn the changing patterns and characteristics of speech under different noise backgrounds, so that when synthesizing speech, it can accurately simulate the speech style in various special noise scenes according to needs, providing a more realistic and practical solution for speech synthesis in fields such as virtual reality and the application of intelligent voice assistants in complex environments, and significantly improving the applicability and expressiveness of speech synthesis models in complex actual scenarios.

[0088] Step S203: input the reference speech set and the text data set into a pre-built acoustic model, wherein the acoustic model includes a style encoder, a reference encoder, a text encoder, an acoustic structure and a vocoder.

[0089] In this embodiment, the pre-built acoustic model is a new acoustic model based on the Rectified Flow architecture, including a style encoder, a reference encoder, a text encoder, an acoustic structure, and a vocoder. The style encoder focuses on encoding Style Prompt (speech style description text). With its powerful semantic understanding and feature extraction capabilities, it can accurately and deeply encode various style information in Style Prompt, laying a solid foundation for the precise shaping of the subsequent synthetic speech style; the reference encoder encodes the reference speech in the reference speech set, providing a key reference basis for model learning; the text encoder is mainly responsible for encoding Context Prompt (text to be synthesized), and can perform efficient feature extraction and semantic understanding of the input text data, converting Context Prompt into Text Embedding with rich semantic information, and providing accurate and comprehensive text information guidance for subsequent speech synthesis; the acoustic structure (Acoustic Model) is built based on the Rectified Flow architecture; the vocoder uses the HifiGAN vocoder. The final waveform is synthesized based on HifiGAN technology.

[0090] In a specific implementation, the acoustic model also includes an input layer, through which the input reference speech set and text data set are preprocessed to obtain input data in a data format suitable for model processing.

[0091] Step S204: encoding the style labels of the text data set and the reference speech set through a style encoder to obtain style encoding features.

[0092] In this embodiment, the style encoder adopts an improved BERT model, including a BERT embedding layer, a spatial expansion layer and a style encoding layer. The style label corresponding to each reference speech in the text data set and the reference speech set is input into the style encoder, and the style description text in the text data set and the style label is processed in turn by the BERT embedding layer, the spatial expansion layer and the style encoding layer to obtain the style encoding features for style prompts.

[0093] Furthermore, the style labels of the text dataset and the reference speech set are converted into vectors through the BERT embedding layer to obtain a style embedding vector; the style embedding vector is input into the space extension layer, and the style extension vector is obtained by concatenating the style embedding vector and the introduced style hint vector; the style features in the style extension vector are extracted through the self-attention mechanism of the style encoding layer to obtain the style encoding features.

[0094] The style tag text is input into the BERT embedding layer for vector conversion. The vector conversion of the embedding layer includes token embedding, segmentation embedding and position embedding of the style tag text.

[0095] Among them, word embedding converts each word into a vector representation of fixed dimension to represent the main semantic information of the word; segment embedding is used to represent the structure of the sentence to distinguish different sentences or paragraphs; position embedding provides information about the position of the word in the sentence, adding temporal information to the attention mechanism. The vectors corresponding to word embedding, segment embedding and position embedding are added together to obtain the style embedding vector.

[0096] The spatial expansion layer introduces a learnable style cue vector pool and spatially expands the style embedding vector to enrich the feature representation. Specifically, the style embedding vector is input into the spatial expansion layer, and the pre-trained cue vector pool is obtained through the spatial expansion layer. The cue vector pool contains S style cue vectors, where S is a positive integer; k pre-trained style cue vectors are obtained from the cue vector pool, where k is a positive integer and less than or equal to S; the style embedding vector is concatenated with the k style cue vectors to obtain a style expansion vector. By introducing the style cue vector, additional information or constraints are introduced on the basis of the original feature space, thereby expanding the feature space and enabling the model to capture more features related to style information.

[0097] The style encoding layer is a stack of multiple connected Transformer encoders. Through a multi-layer self-attention mechanism and a feedforward neural network, it captures the semantic information and contextual relationships in the style extension vector to obtain style encoding features.

[0098] In some optional implementations, the style encoder also includes a linear layer, which is connected after the last layer of Transformer encoder, and performs a linear transformation on the features output by the last layer of Transformer encoder to obtain the final style coding features for output.

[0099] Style encoding through a style encoder can improve the semantic understanding and feature extraction capabilities of style texts and accurately extract style information from style tags.

[0100] Step S205: Encode the reference speech of the reference speech set by using a reference encoder to obtain reference speech coding features.

[0101] In this embodiment, the reference encoder adopts a Transformer encoder structure, which is composed of multiple encoder layers stacked together. The semantic and contextual features of the reference speech are extracted through the multi-head attention mechanism of each encoder layer to obtain the reference speech coding features.

[0102] During the model training process, the reference speech will be enabled by default, and the encoding result of the reference speech by the reference encoder will provide a key reference for model learning. However, during the inference phase, in order to optimize system resource utilization and improve efficiency, the reference speech will be discarded.

[0103] In some optional implementations, the reference encoder further includes a linear layer, which is connected after the last layer of the encoder, and performs a linear transformation on the features output by the last layer of the encoder to obtain the final reference speech coding features for output.

[0104] Step S206, encoding the text data set through a text encoder to obtain text encoding features.

[0105] In this embodiment, the text encoder includes a text embedding layer, a Transformer encoding layer and a pooling layer. The text dataset is input into the text encoder and passes through the text embedding layer, the Transformer encoding layer and the pooling layer.

[0106] Furthermore, the text data in the text dataset is vectorized through the text embedding layer to obtain the text embedding vector; the text embedding vector is input into the Transformer encoding layer, and semantic features are extracted through the multi-layer self-attention mechanism to obtain the semantic encoding features; the semantic encoding features are pooled through the pooling layer to obtain the text encoding features.

[0107] Among them, the text encoder adopts a multi-layer Transformer encoding structure, which includes multiple encoders connected in a stacked manner, initializes the encoder model parameters, determines the number of layers, number of attention heads, hidden layer dimensions and other hyperparameters of the encoder, and obtains the text encoder. That is, the Transformer encoding layer of the text encoder is composed of multiple Transformer encoders connected, and each Transformer encoder includes a multi-head attention layer and a feedforward neural network layer.

[0108] The text dataset is converted to a vector through the text embedding layer, and the obtained text embedding vector is passed to the Transformer encoder. Through the multi-layer self-attention mechanism and feedforward neural network, the semantic information and contextual relationship of the text in the text embedding vector are captured to obtain the encoded intermediate feature representation. In the last layer of the Transformer encoder, the semantic encoding features of the final output of each position are obtained, and the semantic encoding features are pooled through the pooling layer, such as average pooling or maximum pooling, to obtain the text encoding features.

[0109] In some optional implementations, the text encoder further includes a linear layer, which is connected after the pooling layer and outputs the text encoding features of the pooling operation after linear transformation.

[0110] By encoding the input text through a text encoder, efficient feature extraction and semantic understanding of the input text information can be performed, and text encoding features with rich semantic information can be obtained, providing accurate and comprehensive text information guidance for subsequent speech synthesis work.

[0111] Step S207: input the style coding feature, the reference speech coding feature and the text coding feature into the acoustic structure to obtain the speech acoustic feature.

[0112] In this embodiment, the acoustic structure is constructed based on the Rectified Flow architecture, including a reversible transformation layer, a stream decoder and a residual layer. The reversible transformation layer is used to transform the style coding features, the reference speech coding features and the text coding features to obtain the transformed coding features; the transformed coding features are decoded by the stream decoder to obtain an acoustic feature sequence; the acoustic feature sequence is input into the residual layer for residual connection to obtain the speech acoustic features.

[0113] Among them, speech acoustic features include Mel spectrogram, duration information, pitch (F0) and energy. The Mel spectrogram represents a visual representation of the frequency content of the speech signal; the duration information represents the duration of each phoneme or syllable; the pitch (F0) represents the fundamental frequency of the speech signal, which is related to human pitch perception; and the energy represents the intensity or loudness of the speech signal.

[0114] Specifically, the style coding features, reference speech coding features and text coding features are concatenated to obtain concatenated coding features; the concatenated coding features are input into the reversible transformation layer, which includes an affine coupling layer, a permutation layer and a scale transformation layer, and the concatenated coding features are processed in turn by the affine coupling layer, the permutation layer and the scale transformation layer to obtain the transformed coding features. The stream decoder adopts the Transformer decoding structure, which is composed of multiple decoders stacked together. The long-distance dependency between the transformed coding features is captured through the multi-head attention mechanism and the cross-attention mechanism of each encoder to obtain the predicted acoustic feature sequence; the acoustic feature sequence is residually connected through the residual layer, which enhances the model's expression ability and information transmission ability, avoids the gradient disappearance problem, and improves the training stability of the model.

[0115] Among them, the affine coupling layer linearly transforms the features through learnable parameters; the permutation layer rearranges the feature dimensions; and the scale transformation layer scales the features. A gating mechanism is introduced in each reversible transformation layer, and a gating signal is generated through a sigmoid activation function to dynamically adjust the transformation parameters, so that the model can adaptively control the strength of each layer of transformation.

[0116] In some optional implementations of this embodiment, the acoustic structure may be provided with multiple reversible transformation layers, a gating mechanism may be introduced in each reversible transformation layer, a gating signal may be generated through a sigmoid activation function, and the transformation parameters may be dynamically adjusted, so that the model can adaptively control the intensity of each layer of transformation.

[0117] By predicting the acoustic features of speech through acoustic structure, the efficiency and accuracy of the Rectified Flow architecture in processing complex acoustic tasks are fully utilized, effectively overcoming the bottleneck of traditional acoustic models in synthesis speed, making the speech synthesis process smoother and more efficient, and significantly improving the quality and speed of synthesized speech.

[0118] Step S208: synthesize the waveform of the speech acoustic features through a vocoder to obtain predicted synthesized speech.

[0119] In this embodiment, the vocoder adopts a HifiGAN vocoder, and the input speech acoustic features are converted into a high-fidelity speech waveform through a GAN generator network in the vocoder, that is, the predicted synthesized speech is obtained.

[0120] The synthesized speech is post-processed, including denoising and volume normalization, to obtain the final high-fidelity speech output.

[0121] Step S209, according to the preset loss function, the loss is calculated according to the reference speech and the predicted synthesized speech, the model parameters are adjusted based on the loss, and the iterative training is continued until the iteration stop condition is met to obtain the final speech synthesis model.

[0122] Specifically, the adversarial loss and L1 loss between the reference speech and the predicted synthesized speech are calculated; the similarity between the style coding features and the reference speech coding features is calculated to obtain the feature-related loss; the adversarial loss, L1 loss and feature-related loss are weighted and summed to obtain the final loss.

[0123] In this embodiment, the vocoder further includes a discriminator. During the model training phase, the discriminator evaluates the difference between the waveform of the generated predicted synthesized speech and the waveform of the real reference speech to obtain the adversarial loss.

[0124] Among them, the L1 loss calculation formula is as follows:

[0125]

[0126] Where n is the number of reference speech in the reference speech set; y i represents the i-th reference speech in the reference speech set; y' i represents the predicted synthesized speech corresponding to the i-th reference speech.

[0127] In the inference stage, in order to optimize system resource utilization and improve efficiency, the reference speech will be discarded. To achieve this smooth transition, in this embodiment, a calculation strategy is introduced in the training process, that is, calculating the cosine similarity between the style coding features and the reference speech coding features as the feature-related loss, so that even without the reference speech in the inference stage, the synthetic speech that meets the expected style can still be accurately generated.

[0128] By optimizing the model using multiple losses such as adversarial loss, L1 loss, and feature-related loss, the model prediction accuracy can be improved, the model robustness can be enhanced, the model generalization ability can be promoted, and the naturalness and intelligibility of the synthesized speech can be improved.

[0129] In this embodiment, the Adam or SGD optimizer is used to adjust the model parameters of the acoustic model according to the loss, and the adjusted model is iteratively trained until the iteration stop condition is met, that is, the number of iterations reaches a preset number or the loss does not change significantly, and the final speech synthesis model is output.

[0130] Step S210, obtaining the text to be converted, inputting it into the speech synthesis model, and obtaining the target synthesized speech.

[0131] In this embodiment, a speech synthesis model is obtained for inputting text to be converted, and style encoding is performed on the text to be converted through an internal style encoder to obtain style semantic features; the text to be converted is input into a text encoder to extract text semantic features to obtain text semantic features; the style semantic features and text semantic features are input into an acoustic structure to perform acoustic feature prediction to obtain speech acoustic features; and speech waveform synthesis is performed on the speech acoustic features through a vocoder to obtain target synthesized speech.

[0132] It should be emphasized that in order to further ensure the privacy and security of the text to be converted, the above text to be converted can also be stored in a node of a blockchain.

[0133] The blockchain referred to in this application is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.

[0134] The present application can realize text-guided speech synthesis in multiple languages ​​by acquiring speech datasets and text datasets in different languages, and annotating the speech datasets with style tags and injecting scene noise. The application can enable the model to learn to accurately recognize and process speech signals in a noisy environment, thereby improving the generalization ability and robustness of the model, while simulating real scenes to improve the multi-scenario applicability of speech synthesis. By inputting the reference speech set and text dataset into a pre-built acoustic model, the style encoder, reference encoder, text encoder, acoustic structure and vocoder of the acoustic model are trained to obtain the final speech synthesis model, which can improve the naturalness and realism of speech synthesis, enhance the adaptability and generalization ability of the model, and improve the efficiency and real-time performance of speech synthesis.

[0135] In some optional implementations of this embodiment, the step of determining the style label of each speech data in the speech data set according to the preset style dimension includes:

[0136] The speech data is analyzed using a speech signal processing algorithm to obtain the corresponding speech speed, pitch and signal-to-noise ratio;

[0137] Determine a speech rate label and a pitch label of the speech data according to the speech rate and the pitch;

[0138] Comparing the signal-to-noise ratio with a preset signal-to-noise ratio threshold to obtain a signal-to-noise ratio label;

[0139] Input the speech data set into the trained emotion recognition model for emotion classification and obtain the corresponding emotion label;

[0140] Use the trained gender recognition model to determine the gender of the speaker in the speech dataset and obtain the corresponding gender label;

[0141] Input the speech data set into the trained accent recognition model, determine the accent type of the speech data, and obtain the accent label based on the accent type;

[0142] The speaking rate labels, pitch labels, signal-to-noise ratio labels, emotion labels, gender labels, and accent labels are aggregated to obtain the style label.

[0143] In this embodiment, a speech signal processing algorithm is used to analyze the three dimensions of speech rate, signal-to-noise ratio, and pitch. By accurately analyzing and calculating characteristic parameters such as frequency, amplitude, and duration of the speech signal, the speech rate, signal-to-noise ratio, and pitch changes of the speech data can be accurately determined.

[0144] Among them, the speech speed of the voice data is divided into different levels such as slow, normal, and fast; the tone types of the voice data include flat tone, rising tone, falling tone, etc.; the signal-to-noise ratio of the voice data is divided into high, medium, and low levels.

[0145] The speech rate, pitch and signal-to-noise ratio of each voice data are compared with the preset range standard to determine the corresponding speech rate label, pitch label and signal-to-noise ratio label. For example, if the calculated signal-to-noise ratio is equal to the preset signal-to-noise ratio threshold, it is a medium signal-to-noise ratio; if the calculated signal-to-noise ratio is less than the preset signal-to-noise ratio threshold, it is a low signal-to-noise ratio; if the calculated signal-to-noise ratio is higher than the preset signal-to-noise ratio threshold, it is a high signal-to-noise ratio.

[0146] Use the trained emotion recognition model to predict the emotion of the speech data in the speech data set and determine the emotion category of the speech data, such as positive, negative, neutral, surprised, worried, etc. Use the trained gender recognition model to determine the gender of the speaker corresponding to the speech data and obtain the gender of the speech data. The gender recognition model can be trained and optimized using the support vector machine algorithm. Use the trained accent recognition model to determine the accent type of the speech data through speech recognition and accent recognition algorithms.

[0147] Finally, the speech data's speed labels, pitch labels, signal-to-noise ratio labels, emotion labels, gender labels, and accent labels are summarized and annotated to obtain an annotated speech set.

[0148] By determining the style labels of speech data, this application can greatly enrich the style diversity of speech synthesis and improve the naturalness and expressiveness of speech synthesis.

[0149] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0150] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0151] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through computer-readable instructions, and the computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0152] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0153] Further references Figure 3 , as a response to the above Figure 2 The present application provides an embodiment of a text-guided speech synthesis device, and the device embodiment is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0154] like Figure 3As shown, the text-guided speech synthesis device 300 described in this embodiment includes: an acquisition module 301, a label injection module 302, an input module 303, a style encoding module 304, a speech encoding module 305, a text encoding module 306, an acoustic feature extraction module 307, a speech prediction module 308, an iteration module 309 and a speech synthesis module 310. Among them:

[0155] The acquisition module 301 is used for acquiring a speech data set and a text data set in different languages, wherein the speech data in the speech data set and the text data in the text data set are paired data;

[0156] The annotation injection module 302 is used to perform style label annotation and scene noise injection on the speech data set to obtain a reference speech set;

[0157] The input module 303 is used to input the reference speech set and the text data set into a pre-built acoustic model, wherein the acoustic model includes a style encoder, a reference encoder, a text encoder, an acoustic structure and a vocoder;

[0158] The style encoding module 304 is used to encode the style labels of the text data set and the reference speech set through the style encoder to obtain style encoding features;

[0159] The speech encoding module 305 is used to encode the reference speech of the reference speech set through the reference encoder to obtain reference speech coding features;

[0160] The text encoding module 306 is used to encode the text data set through the text encoder to obtain text encoding features;

[0161] The acoustic feature extraction module 307 is used to input the style coding feature, the reference speech coding feature and the text coding feature into the acoustic structure to obtain speech acoustic features;

[0162] The speech prediction module 308 is used to perform waveform synthesis on the speech acoustic features through the vocoder to obtain predicted synthesized speech;

[0163] The iteration module 309 is used to calculate the loss according to the reference speech and the predicted synthesized speech according to a preset loss function, adjust the model parameters based on the loss, and continue iterative training until the iteration stop condition is met to obtain the final speech synthesis model;

[0164] The speech synthesis module 310 is used to obtain the text to be converted, input it into the speech synthesis model, and obtain the target synthesized speech.

[0165] It should be emphasized that in order to further ensure the privacy and security of the text to be converted, the above text to be converted can also be stored in a node of a blockchain.

[0166] Based on the above-mentioned text-guided speech synthesis device 300, by acquiring speech data sets and text data sets in different languages, and annotating the speech data sets with style tags and injecting scene noise, multi-language text-guided speech synthesis can be realized, and the model can be learned to accurately recognize and process speech signals in a noisy environment, thereby improving the generalization ability and robustness of the model, while simulating real scenes to improve the multi-scene applicability of speech synthesis; by inputting the reference speech set and text data set into the pre-built acoustic model, training the style encoder, reference encoder, text encoder, acoustic structure and vocoder of the acoustic model, and obtaining the final speech synthesis model, the naturalness and realism of speech synthesis can be improved, the adaptability and generalization ability of the model can be enhanced, and the efficiency and real-time performance of speech synthesis can be improved.

[0167] In some optional implementations, the annotation injection module 302 includes:

[0168] A labeling submodule, used to determine the style label of each voice data in the voice data set according to a preset style dimension, and label the voice data based on the style label to obtain a labeled voice set;

[0169] The injection submodule is used to obtain a scene noise data set, and adopt a data fusion algorithm to fuse the scene noise data set with the annotated speech set to obtain a reference speech set containing noise.

[0170] Through style label annotation and scene noise injection, the style diversity of speech synthesis is greatly enriched. In addition, when synthesizing speech, the speech style in various special noise scenes can be accurately simulated according to needs, which significantly improves the applicability and expressiveness of the speech synthesis model in complex actual scenarios.

[0171] In some optional implementations of this embodiment, the marking submodule is further used to:

[0172] Analyzing the speech data using a speech signal processing algorithm to obtain corresponding speech speed, pitch and signal-to-noise ratio;

[0173] Determine a speech rate label and a tone label of the speech data according to the speech rate and the tone;

[0174] Comparing the signal-to-noise ratio with a preset signal-to-noise ratio threshold to obtain a signal-to-noise ratio label;

[0175] Input the speech data set into a trained emotion recognition model for emotion classification to obtain corresponding emotion labels;

[0176] Using the trained gender recognition model, determine the gender of the speaker in the speech data set and obtain a corresponding gender label;

[0177] Inputting the speech data set into a trained accent recognition model, determining the accent type of the speech data, and obtaining an accent label based on the accent type;

[0178] The speaking rate label, the pitch label, the signal-to-noise ratio label, the emotion label, the gender label, and the accent label are aggregated to obtain a style label.

[0179] By determining the style labels of speech data, the style diversity of speech synthesis can be greatly enriched, and the naturalness and expressiveness of speech synthesis can be improved.

[0180] In some optional implementations, the style encoder includes a BERT embedding layer, a spatial expansion layer, and a style encoding layer, and the style encoding module 304 includes:

[0181] An embedding submodule, configured to perform vector conversion on the style labels of the text dataset and the reference speech set through the BERT embedding layer to obtain a style embedding vector;

[0182] An extension submodule, used for inputting the style embedding vector into the spatial extension layer, and obtaining a style extension vector by concatenating the style embedding vector and the introduced style hint vector;

[0183] The attention calculation submodule is used to extract the style features in the style extension vector through the self-attention mechanism of the style encoding layer to obtain the style encoding features.

[0184] Style encoding through a style encoder can improve the semantic understanding and feature extraction capabilities of style texts and accurately extract style information from style tags.

[0185] In some optional implementations of this embodiment, the text encoder includes a text embedding layer, a Transformer encoding layer and a pooling layer, and the text encoding module 306 includes:

[0186] A text embedding submodule, used to perform vector embedding on the text data in the text data set through the text embedding layer to obtain a text embedding vector;

[0187] An encoding submodule, used for inputting the text embedding vector into the Transformer encoding layer, extracting semantic features through a multi-layer self-attention mechanism, and obtaining semantic encoding features;

[0188] The pooling submodule is used to perform a pooling operation on the semantic coding features through the pooling layer to obtain text coding features.

[0189] By encoding the input text through a text encoder, efficient feature extraction and semantic understanding of the input text information can be performed, and text encoding features with rich semantic information can be obtained, providing accurate and comprehensive text information guidance for subsequent speech synthesis work.

[0190] In some optional implementations, the acoustic structure includes a reversible transform layer, a stream decoder, and a residual layer, and the acoustic feature extraction module 307 is further used to:

[0191] Performing feature transformation on the style coding feature, the reference speech coding feature and the text coding feature through the reversible transformation layer to obtain a transformation coding feature;

[0192] Decoding the transform coded features by the stream decoder to obtain an acoustic feature sequence;

[0193] The acoustic feature sequence is input into the residual layer for residual connection to obtain speech acoustic features.

[0194] By predicting the acoustic features of speech through acoustic structure, the bottleneck of traditional acoustic models in synthesis speed is effectively overcome, making the speech synthesis process smoother and more efficient, and significantly improving the quality and speed of synthesized speech.

[0195] In some optional implementations of this embodiment, the iteration module 309 is further configured to:

[0196] Calculating the adversarial loss and L1 loss between the reference speech and the predicted synthesized speech;

[0197] Calculating the similarity between the style coding feature and the reference speech coding feature to obtain a feature correlation loss;

[0198] The adversarial loss, the L1 loss and the feature-related loss are weightedly summed to obtain the final loss.

[0199] By optimizing the model using multiple losses such as adversarial loss, L1 loss, and feature-related loss, the model prediction accuracy can be improved, the model robustness can be enhanced, the model generalization ability can be promoted, and the naturalness and intelligibility of the synthesized speech can be improved.

[0200] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0201] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected and communicated through a system bus. It should be noted that the figure only shows a computer device 4 with a memory 41, a processor 42, and a network interface 43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (Application Specific Integrated Circuit, ASIC), programmable gate arrays (Field-Programmable Gate Array, FPGA), digital processors (Digital Signal Processor, DSP), embedded devices, etc.

[0202] The computer device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may interact with a user through a keyboard, a mouse, a remote controller, a touch pad, or a voice control device.

[0203] The memory 41 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as a hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk equipped on the computer device 4, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), etc. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions of a text-guided speech synthesis method, etc. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.

[0204] The processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run computer-readable instructions or process data stored in the memory 41, such as computer-readable instructions for running the text-guided speech synthesis method.

[0205] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0206] By obtaining speech datasets and text datasets in different languages, and annotating the speech datasets with style tags and injecting scene noise, we can achieve text-guided speech synthesis in multiple languages, and enable the model to learn to accurately recognize and process speech signals in a noisy environment, thereby improving the generalization and robustness of the model, while simulating real scenarios to improve the multi-scenario applicability of speech synthesis; by inputting the reference speech set and text dataset into the pre-built acoustic model, training the style encoder, reference encoder, text encoder, acoustic structure and vocoder of the acoustic model, and obtaining the final speech synthesis model, we can improve the naturalness and realism of speech synthesis, enhance the adaptability and generalization ability of the model, and at the same time improve the efficiency and real-time performance of speech synthesis.

[0207] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the text-guided speech synthesis method as described above.

[0208] By obtaining speech datasets and text datasets in different languages, and annotating the speech datasets with style tags and injecting scene noise, we can achieve text-guided speech synthesis in multiple languages, and enable the model to learn to accurately recognize and process speech signals in a noisy environment, thereby improving the generalization and robustness of the model, while simulating real scenarios to improve the multi-scenario applicability of speech synthesis; by inputting the reference speech set and text dataset into the pre-built acoustic model, training the style encoder, reference encoder, text encoder, acoustic structure and vocoder of the acoustic model, and obtaining the final speech synthesis model, we can improve the naturalness and realism of speech synthesis, enhance the adaptability and generalization ability of the model, and at the same time improve the efficiency and real-time performance of speech synthesis.

[0209] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0210] Obviously, the embodiments described above are only some embodiments of the present application, rather than all embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application is described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions recorded in the aforementioned specific implementation methods, or to perform equivalent replacement of some of the technical features therein. Any equivalent structure made using the contents of the specification and drawings of this application, directly or indirectly used in other related technical fields, is similarly within the scope of patent protection of this application.

Claims

1. A text-guided speech synthesis method, characterized in that: The steps include: Acquire speech data sets and text data sets in different languages, wherein the speech data in the speech data sets and the text data in the text data sets are paired data; Performing style labeling and scene noise injection on the speech data set to obtain a reference speech set; Inputting the reference speech set and the text data set into a pre-built acoustic model, wherein the acoustic model includes a style encoder, a reference encoder, a text encoder, an acoustic structure, and a vocoder; Encoding the style labels of the text data set and the reference speech set by the style encoder to obtain style coding features; Encoding the reference speech of the reference speech set by the reference encoder to obtain reference speech coding features; Encoding the text dataset by the text encoder to obtain text encoding features; Inputting the style coding feature, the reference speech coding feature and the text coding feature into the acoustic structure to obtain speech acoustic features; Performing waveform synthesis on the acoustic features of the speech by the vocoder to obtain predicted synthesized speech; According to a preset loss function, the loss is calculated according to the reference speech and the predicted synthesized speech, model parameters are adjusted based on the loss, and iterative training is continued until an iteration stop condition is met to obtain a final speech synthesis model; The text to be converted is obtained and input into the speech synthesis model to obtain the target synthesized speech.

2. The text-guided speech synthesis method according to claim 1, characterized in that: The step of annotating the speech data set with style tags and injecting scene noise to obtain a reference speech set comprises: Determining a style label for each piece of speech data in the speech data set according to a preset style dimension, and annotating the speech data based on the style label to obtain an annotated speech set; A scene noise data set is obtained, and a data fusion algorithm is used to fuse the scene noise data set with the annotated speech set to obtain a reference speech set containing noise.

3. The text-guided speech synthesis method according to claim 2, characterized in that: The step of determining the style label of each piece of speech data in the speech data set according to the preset style dimension comprises: Analyzing the speech data using a speech signal processing algorithm to obtain corresponding speech speed, pitch and signal-to-noise ratio; Determine a speech rate label and a tone label of the speech data according to the speech rate and the tone; Comparing the signal-to-noise ratio with a preset signal-to-noise ratio threshold to obtain a signal-to-noise ratio label; Input the speech data set into a trained emotion recognition model for emotion classification to obtain corresponding emotion labels; Using the trained gender recognition model, determine the gender of the speaker in the speech data set and obtain a corresponding gender label; Inputting the speech data set into a trained accent recognition model, determining the accent type of the speech data, and obtaining an accent label based on the accent type; The speaking rate label, the pitch label, the signal-to-noise ratio label, the emotion label, the gender label, and the accent label are aggregated to obtain a style label.

4. The text-guided speech synthesis method according to claim 1, characterized in that: The style encoder includes a BERT embedding layer, a spatial expansion layer and a style encoding layer. The step of encoding the style tags of the text data set and the reference speech set by the style encoder to obtain style encoding features includes: Performing vector conversion on the style labels of the text dataset and the reference speech set through the BERT embedding layer to obtain a style embedding vector; Inputting the style embedding vector into the spatial extension layer, and obtaining a style extension vector by concatenating the style embedding vector and the introduced style hint vector; The style features in the style extension vector are extracted through the self-attention mechanism of the style encoding layer to obtain style encoding features.

5. The text-guided speech synthesis method according to claim 1, characterized in that: The text encoder includes a text embedding layer, a Transformer encoding layer and a pooling layer. The step of encoding the text dataset by the text encoder to obtain text encoding features includes: Performing vector embedding on the text data in the text data set through the text embedding layer to obtain a text embedding vector; Inputting the text embedding vector into the Transformer encoding layer, extracting semantic features through a multi-layer self-attention mechanism, and obtaining semantic encoding features; The semantic coding features are pooled by the pooling layer to obtain text coding features.

6. The text-guided speech synthesis method according to claim 1, characterized in that: The acoustic structure includes a reversible transform layer, a stream decoder and a residual layer, and the step of inputting the style coding feature, the reference speech coding feature and the text coding feature into the acoustic structure to obtain the speech acoustic feature includes: Performing feature transformation on the style coding feature, the reference speech coding feature and the text coding feature through the reversible transformation layer to obtain a transformation coding feature; Decoding the transform coded features by the stream decoder to obtain an acoustic feature sequence; The acoustic feature sequence is input into the residual layer for residual connection to obtain speech acoustic features.

7. The text-guided speech synthesis method according to claim 1, characterized in that: The step of calculating the loss according to the reference speech and the predicted synthesized speech according to the preset loss function comprises: Calculating the adversarial loss and L1 loss between the reference speech and the predicted synthesized speech; Calculating the similarity between the style coding feature and the reference speech coding feature to obtain a feature correlation loss; The adversarial loss, the L1 loss and the feature-related loss are weightedly summed to obtain the final loss.

8. A text-guided speech synthesis device, characterized in that: include: An acquisition module, used to acquire speech data sets and text data sets in different languages, wherein the speech data in the speech data set and the text data in the text data set are paired data; A label injection module, used for labeling the speech data set with style labels and injecting scene noise to obtain a reference speech set; An input module, configured to input the reference speech set and the text data set into a pre-built acoustic model, wherein the acoustic model includes a style encoder, a reference encoder, a text encoder, an acoustic structure, and a vocoder; A style encoding module, used for encoding the style labels of the text data set and the reference speech set through the style encoder to obtain style encoding features; A speech encoding module, used for encoding the reference speech of the reference speech set by the reference encoder to obtain reference speech encoding features; A text encoding module, used for encoding the text data set through the text encoder to obtain text encoding features; An acoustic feature extraction module, used for inputting the style coding feature, the reference speech coding feature and the text coding feature into the acoustic structure to obtain speech acoustic features; A speech prediction module, used for performing waveform synthesis on the speech acoustic features through the vocoder to obtain predicted synthesized speech; An iteration module, configured to calculate a loss according to the reference speech and the predicted synthesized speech according to a preset loss function, adjust model parameters based on the loss, and continue iterative training until an iteration stop condition is met, thereby obtaining a final speech synthesis model; The speech synthesis module is used to obtain the text to be converted, input it into the speech synthesis model, and obtain the target synthesized speech.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the text-guided speech synthesis method according to any one of claims 1 to 7 when executing the computer-readable instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the text-guided speech synthesis method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Speech clone model training method and device, speech synthesis method and device and related equipment

    CN116403558A

  • Speech synthesis method and device, computer equipment and storage medium

    CN119068863A

  • Speech synthesis method and apparatus, and device and storage medium

    WO2022178941A1

Cited By

  • End-to-end speech synthesis method and device, computer equipment and storage medium

    CN121075306A

  • An end-to-end speech synthesis method, apparatus, computer device, and storage medium

    CN121075306B

  • ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion

    CN121354597A