Speech synthesis device and training method thereof, electronic equipment and storage medium

By dynamically selecting an expert network for speech synthesis, the problems of multi-model deployment and high cost in existing technologies are solved, efficient and fast speech synthesis adaptability is achieved, and the scenario adaptability and efficiency of speech synthesis are improved.

CN120673743APending Publication Date: 2025-09-19MOORE THREADS TECH CO LTD

Patent Information

Application Number
CN202510969977.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing speech synthesis technology, when faced with changing application scenarios and user needs, requires training and deploying multiple models for different scenarios, which increases the deployment and management costs of the models. In addition, it leads to high training and inference costs in fast-response interaction scenarios, affecting the user experience.

Method used

The voiceprint encoding module is used to extract the voiceprint features of the reference audio, and the audio coding generation module dynamically selects the target expert network from multiple expert networks for speech synthesis. The dynamic routing mechanism is used to automatically select the appropriate expert network, reduce the parameters involved in the inference calculation, and improve the efficiency of speech synthesis.

Benefits of technology

It significantly improves the scenario adaptability of speech synthesis, reduces the amount of calculation by about 30%, reduces inference latency, and supports rapid iteration of new application scenarios, improving the efficiency and adaptability of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673743A_ABST
    Figure CN120673743A_ABST
Patent Text Reader

Abstract

The invention relates to a voice synthesis device and a training method thereof, electronic equipment and a storage medium, and the method comprises a voiceprint coding module which receives an input reference audio and extracts voiceprint features of the reference audio; the audio code generation module is used for determining a target expert network from a plurality of expert networks to execute a voice synthesis operation according to the input text to be synthesized and the voiceprint features so as to obtain a generated audio code sequence; wherein the plurality of expert networks correspond to different application scenes; and the audio decoding module is used for converting the audio coding sequence output by the audio coding generation module into an audio signal. According to the embodiment of the invention, through the dynamically selected target expert network, the voice synthesis of the application scene suitable for the input to-be-synthesized text and the voiceprint feature can be executed, and the scene adaptability of the synthesized voice is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a speech synthesis device and a training method thereof, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of artificial intelligence (AI), speech synthesis (STS) has been widely used in various voice interaction and text reading scenarios. Speech synthesis technology, particularly that based on deep learning, has made significant progress in recent years. In particular, STS based on the autoregressive (AR) Large Language Model (LLM) has significantly improved the naturalness of synthesized speech, now capable of producing audio that closely resembles the intonation and rhythm of human speech.

[0003] Despite this, existing technologies still have certain limitations when faced with diverse application scenarios and user needs. For example, in a human-computer interaction system, the speech synthesis system needs to synthesize different styles of speech according to different scenarios. Sometimes natural speech intonation is needed for oral conversations, sometimes a lively and vivid tone is needed for storytelling, and sometimes a serious tone is needed for news broadcasts. To meet these diverse needs, traditional methods usually require training and deploying multiple models for different scenarios. This not only increases the cost of model deployment and management, but also requires manual control of the specific model used during use, making it difficult to adapt to the application needs of complex scenarios.

[0004] Another approach is to increase the model's parameter size to improve speech synthesis, thereby ensuring synthesis quality in different scenarios. However, in related technologies, since all model parameters are involved in training and inference, larger models mean higher training and inference costs. During training, larger models require longer training times or more hardware resources. During inference, larger models may result in higher response latency and require more hardware resources to support the same number of concurrent users. This is particularly noticeable in interactive scenarios that require fast responses, affecting the user experience. Summary of the Invention

[0005] The present disclosure proposes a speech synthesis technology solution.

[0006] According to one aspect of the present disclosure, there is provided a speech synthesis apparatus, comprising:

[0007] A voiceprint encoding module receives an input reference audio and extracts the voiceprint features of the reference audio;

[0008] An audio code generation module, which determines a target expert network from multiple expert networks based on the input text to be synthesized and voiceprint features, performs speech synthesis operations, and generates an audio code sequence; wherein the multiple expert networks correspond to different application scenarios;

[0009] The audio decoding module converts the audio code sequence output by the audio code generation module into an audio signal.

[0010] In a possible implementation, the audio coding generation module includes a routing module and multiple expert networks, wherein:

[0011] The routing module determines, based on the text to be synthesized and the voiceprint features, a target expert network that matches the scenario of the text to be synthesized and the voiceprint features from a plurality of expert networks;

[0012] The target expert network uses the voiceprint feature as a reference to generate an audio coding sequence corresponding to the text to be synthesized.

[0013] In a possible implementation, the audio code generation module includes an embedding feature extraction module and an attention feature extraction module, wherein:

[0014] The embedding feature extraction module extracts embedding features from the text to be synthesized;

[0015] The attention feature extraction module extracts attention features from the text to be synthesized based on a multi-head attention mechanism;

[0016] The routing module fuses the embedded features and the attention features to obtain fused features, and determines a target expert network from multiple expert networks based on the fused features.

[0017] In a possible implementation, the routing module determines an activation score of each expert network based on the fusion feature; normalizes the activation score of each expert network to obtain an activation probability corresponding to each expert network; and determines a target expert network based on the activation probability.

[0018] In a possible implementation, the network structure of the attention feature extraction module and the expert network is a large language model network structure.

[0019] According to another aspect of the present disclosure, a training method for a speech synthesis device is provided, for training the above-mentioned device, the method comprising:

[0020] Determining a text sample and a reference audio sample corresponding to the audio sample, wherein the reference audio sample and the audio sample belong to the same speaker;

[0021] Encoding the audio sample to obtain an audio coding sequence sample;

[0022] Using the reference audio sample and the text sample as inputs of the speech synthesis device, and obtaining a predicted audio code sequence output by an audio code generation module of the speech synthesis device;

[0023] Based on the loss between the predicted audio coding sequence and the audio coding sequence samples, parameters of the speech synthesis device are adjusted.

[0024] In a possible implementation, adjusting parameters of the speech synthesis device based on the loss between the predicted audio coding sequence and the audio coding sequence samples includes:

[0025] When the training sample contains a scene label, the parameters of the routing module in the audio coding generation module are adjusted to maximize the activation probability of the expert network corresponding to the scene label.

[0026] In a possible implementation, adjusting parameters of the speech synthesis device based on the loss between the predicted audio coding sequence and the audio coding sequence samples includes:

[0027] Determine the ratio f of the number of samples assigned to the i-th expert network in the current training batch to the total number of samples T in the current batch i ;

[0028] Determine the average value P of the probability that all samples in the current batch are assigned to the i-th expert network i ;

[0029] Based on f i and P i The product of is used to construct a loss function, and the parameters of the routing module in the audio coding generation module are adjusted with the goal of minimizing the loss function, so that the audio coding generation module evenly distributes samples to each expert.

[0030] According to another aspect of the present disclosure, there is provided a speech synthesis method, comprising:

[0031] receiving an input reference audio, and extracting a voiceprint feature of the reference audio;

[0032] Determine a target expert network from multiple expert networks based on the input text to be synthesized and voiceprint features to perform speech synthesis operations, thereby generating an audio coding sequence; wherein the multiple expert networks correspond to different application scenarios;

[0033] The audio code sequence is converted into an audio signal.

[0034] In one possible implementation, the method of determining a target expert network from a plurality of expert networks to perform speech synthesis operations based on the input text to be synthesized and voiceprint features to obtain a generated audio code sequence includes:

[0035] Based on the text to be synthesized and the voiceprint features, determining a target expert network that matches the scenario of the text to be synthesized and the voiceprint features from a plurality of expert networks;

[0036] The target expert network uses the voiceprint feature as a reference to generate an audio coding sequence corresponding to the text to be synthesized.

[0037] In one possible implementation, the method of determining a target expert network from a plurality of expert networks to perform speech synthesis operations based on the input text to be synthesized and voiceprint features to obtain a generated audio code sequence includes:

[0038] Extracting embedded features from the text to be synthesized;

[0039] Extracting attention features from the text to be synthesized based on a multi-head attention mechanism;

[0040] The embedded features and the attention features are fused to obtain fused features, and a target expert network is determined from multiple expert networks based on the fused features.

[0041] In a possible implementation, determining, from a plurality of expert networks, a target expert network that matches a scenario of the text to be synthesized and the voiceprint feature based on the text to be synthesized and the voiceprint feature includes:

[0042] Based on the fusion features, an activation score of each expert network is determined; the activation score of each expert network is normalized to obtain an activation probability corresponding to each expert network; and a target expert network is determined based on the activation probability.

[0043] In a possible implementation, the extraction of attention features from the text to be synthesized based on the multi-head attention mechanism is implemented by an attention feature extraction module, and the network structure of the attention feature extraction module and the expert network is a large language model network structure.

[0044] According to one aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to call the instructions stored in the memory to execute the above method.

[0045] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above method is implemented.

[0046] In the disclosed embodiment, the voiceprint encoding module receives input reference audio and extracts the voiceprint features of the reference audio; the audio encoding generation module determines a target expert network from multiple expert networks based on the input text to be synthesized and voiceprint features to perform speech synthesis operations and obtain a generated audio coding sequence; wherein the multiple expert networks correspond to different application scenarios; the audio decoding module converts the audio coding sequence output by the audio encoding generation module into an audio signal. Thus, through a dynamic routing mechanism, an expert network is automatically selected to perform speech synthesis operations based on the input text to be synthesized and voiceprint features. Since multiple expert networks correspond to different application scenarios, the dynamically selected target expert network can perform speech synthesis suitable for the application scenario of the input text to be synthesized and voiceprint features, significantly improving the scenario adaptability of the synthesized speech.

[0047] Furthermore, by activating only the target expert network among multiple expert networks, unselected expert networks are not involved in speech synthesis operations and, therefore, do not participate in the inference calculations during speech synthesis. This reduces the number of parameters involved in inference calculations during speech synthesis and improves speech synthesis efficiency. Compared to a single large model, the model's computational workload is reduced by approximately 30%, significantly reducing inference latency. Furthermore, when adding new application scenarios, only the corresponding expert network needs to be trained, without reconstructing the entire model, enabling rapid iteration.

[0048] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, rather than limiting the present disclosure. Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0050] Figure 1 A block diagram of a speech synthesis apparatus according to an embodiment of the present disclosure is shown.

[0051] Figure 2 A block diagram of a speech synthesis apparatus according to an embodiment of the present disclosure is shown.

[0052] Figure 3 A block diagram illustrating a single-layer network structure according to an embodiment of the present disclosure.

[0053] Figure 4 A flowchart of a training method for a speech synthesis device according to an embodiment of the present disclosure is shown.

[0054] Figure 5A diagram showing the training architecture of an audio encoder according to an embodiment of the present disclosure is shown.

[0055] Figure 6 A flowchart of a speech synthesis method according to an embodiment of the present disclosure is shown.

[0056] Figure 7 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0057] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0058] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0059] The term "and / or" herein simply describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent the existence of three situations: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" herein refers to any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C.

[0060] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0061] Figure 1 A block diagram of a speech synthesis device according to an embodiment of the present disclosure is shown. Figure 1 As shown, the device includes:

[0062] The voiceprint encoding module 11 receives the input reference audio and extracts the voiceprint features of the reference audio;

[0063] The audio code generation module 12 determines a target expert network from a plurality of expert networks based on the input text to be synthesized and voiceprint features, performs a speech synthesis operation, and generates an audio code sequence; wherein the plurality of expert networks correspond to different application scenarios;

[0064] The audio decoding module 13 converts the audio code sequence output by the audio code generation module into an audio signal.

[0065] Reference audio is the input audio that serves as the timbre source and guides the timbre characteristics of the synthesized speech. This audio can be a speech clip of any length, containing the voice of the target speaker used as a reference. It is a speech clip recorded by the target speaker (such as a news broadcast or story reading) that provides reference for timbre, intonation, and style. Examples include a recording of a news anchor (such as "Welcome to today's news") or user-provided custom audio (such as "Hello, I'm the smart assistant Xiao A").

[0066] The voiceprint encoding module converts audio signals into features that represent the speaker's timbre. It extracts the target speaker's unique timbre, intonation, and pronunciation habits from the reference audio and encodes them into a discrete code sequence. This discrete code sequence can be used as a voiceprint feature to distinguish the voices of different speakers. A pre-trained Vector Quantization-Variational AutoEncoder (VQVAE) can be used to compress audio into a discrete code sequence. For example, a 10-second speech segment is input to the voiceprint encoding module, which then outputs a 128-dimensional code sequence.

[0067] The trained voiceprint encoding module can obtain the audio data of the target speaker as the input reference audio, and then extract the voiceprint features from the reference audio. The specific training process can be found in the speech synthesis device training method provided in the present disclosure, which will not be described here.

[0068] The voiceprint encoding module can ensure that the synthesized speech is highly consistent with the speaker's timbre of the reference audio. The voiceprint features extracted by the voiceprint encoding module can be input into the audio coding generation module together with the text to be synthesized. Combining the voiceprint features and the text to be synthesized, the expert network adapted to the scenario is dynamically selected to perform the speech synthesis operation.

[0069] The text to be synthesized is the text content that needs to be converted to speech. This is the text that needs to be input into the audio code generation module. It can be text of any length and includes content that needs to be read aloud. For example, the text to be synthesized might be "The weather is nice today, perfect for a walk." and needs to be converted to speech output.

[0070] Expert networks are subnetworks within the audio code generation module. Each expert network corresponds to a speech synthesis strategy for a specific application scenario and is used to synthesize audio for that specific application scenario. Application scenarios are the target usage environments or task types for speech synthesis, including news broadcasts, story readings, customer service conversations, and other scenarios requiring a specific speech style. Multiple expert networks can correspond to, for example, news broadcast expert networks, story reading expert networks, and daily conversation expert networks. Each expert network, through different parameter configurations and training data, learns a speech synthesis strategy suitable for a specific scenario.

[0071] For example, the news broadcast expert network was trained on a large amount of news broadcast audio data, learning how to generate news broadcast voices with a moderate speaking speed and serious intonation. The story reading expert network was trained on story reading audio data, learning how to generate story reading voices with a slower speaking speed and emotional intonation. The specific training process can be found in the speech synthesis device training method provided in this disclosure and will not be detailed here.

[0072] The audio code generation module performs feature fusion, such as concatenation, on the received text to be synthesized and voiceprint features to form input features. The text embedding dimension is 512, and the voiceprint feature dimension is 128, resulting in a concatenated input dimension of 640. The module then selects an expert network appropriate for the scenario based on the input features. Specifically, the routing module's fully connected layer and a normalized softmax function are used to calculate the activation probability of each expert network. The expert network with the highest probability is selected to perform speech synthesis. For ease of description, this selected expert network is referred to as the target expert network.

[0073] The target expert network outputs an audio code sequence, which is a discretized sequence representing the audio content. This can be a series of discrete codewords, representing the audio content corresponding to the text to be synthesized. The audio code sequence is then passed to the audio decoding module by the target expert network to generate an audio waveform.

[0074] For example, the audio code generation module receives the text to be synthesized, "The weather is nice today, perfect for a walk," and a voiceprint feature vector. The module's internal routing mechanism analyzes the text content and features and determines that the text is suitable for a news broadcast scenario. Therefore, the routing mechanism selects the news broadcast expert network as the target expert network. Based on the input text and voiceprint features, the target expert network generates the corresponding audio code sequence [101, 202, 303, ...], representing the news broadcast-style audio corresponding to the text.

[0075] The audio decoding module then converts the audio code sequence output by the audio code generation module into an audio signal, maps the discretized audio codes into a time-domain waveform, reconstructs a high-quality speech signal, and ultimately outputs a playable audio signal. This module can include a pretrained audio decoder, such as a high-fidelity generative adversarial network (Hifi-GAN). For example, the audio decoding module can use a Hifi-GAN decoder to convert the discrete audio code sequence [101, 202, 303, ...] into a 5-second audio file containing the content "The weather is nice today, perfect for a walk."

[0076] In the disclosed embodiment, the voiceprint encoding module receives input reference audio and extracts the voiceprint features of the reference audio; the audio encoding generation module determines a target expert network from multiple expert networks based on the input text to be synthesized and voiceprint features to perform speech synthesis operations and obtain a generated audio coding sequence; wherein the multiple expert networks correspond to different application scenarios; the audio decoding module converts the audio coding sequence output by the audio encoding generation module into an audio signal. Thus, through a dynamic routing mechanism, an expert network is automatically selected to perform speech synthesis operations based on the input text to be synthesized and voiceprint features. Since multiple expert networks correspond to different application scenarios, the dynamically selected target expert network can perform speech synthesis suitable for the application scenario of the input text to be synthesized and voiceprint features, significantly improving the scenario adaptability of the synthesized speech.

[0077] Furthermore, by activating only the target expert network among multiple expert networks, unselected expert networks are not involved in speech synthesis operations and, therefore, do not participate in the inference calculations during speech synthesis. This reduces the number of parameters involved in inference calculations during speech synthesis and improves speech synthesis efficiency. Compared to a single large model, the model's computational workload is reduced by approximately 30%, significantly reducing inference latency. Furthermore, when adding new application scenarios, only the corresponding expert network needs to be trained, without reconstructing the entire model, enabling rapid iteration.

[0078] In one possible implementation, the audio coding generation module includes a routing module and multiple expert networks, wherein: the routing module determines, based on the text to be synthesized and the voiceprint features, a target expert network that matches the scenario of the text to be synthesized and the voiceprint features from multiple expert networks; the target expert network uses the voiceprint features as a reference to generate an audio coding sequence corresponding to the text to be synthesized.

[0079] The routing module is used to select the most appropriate target expert network from multiple expert networks. This module can evaluate the capabilities of each expert network based on the input text to be synthesized and voiceprint features, and select the expert network with the best processing capabilities for the current input.

[0080] The routing module can use attention mechanisms, classifiers, or other deep learning mechanisms to evaluate the degree of match between the input and each expert network. For example, based on the sentiment, style, and context of the input text, combined with the timbre and style of the voiceprint features, it can determine which expert network can generate speech that matches the scenario of the synthesized text and voiceprint features.

[0081] After the routing module selects a target expert network, it uses the input text to be synthesized and voiceprint features to generate a corresponding audio encoding sequence. The expert network can use a deep learning-based Transformer model (a deep learning model based on an attention mechanism) to learn speech generation rules from the input text to be synthesized and, combined with voiceprint features, generate speech similar to that of the target speaker.

[0082] Suppose the input text to be synthesized is "Today is a sunny day." The voiceprint feature is a 128-dimensional vector representing the timbre of a newscaster. Based on this input information, the routing module selects a news broadcast expert network from multiple expert networks. This expert network generates an audio encoding sequence based on the input text and voiceprint features. This sequence may contain a series of codewords describing the rhythm, intonation, and emotion of the speech. The generated audio encoding sequence is then passed to the audio decoding module to convert it into the final audio signal, producing a clear and natural news broadcast-style speech.

[0083] In the disclosed embodiment, the audio code generation module, through the decision-making of the routing module, selects a target expert network that matches the scenario of the text and voiceprint features to be synthesized. Different expert networks correspond to different application scenarios and can be specially optimized and trained for specific types of input. Therefore, by selecting a target expert network to generate the audio code sequence, the expert network can better understand and process the input text and voiceprint features, thereby generating high-quality, natural and fluent speech.

[0084] The shared architecture of routing modules and expert networks allows for more efficient utilization of computing resources. Although the number of expert networks increases the number of model parameters, during inference, only one expert network needs to be activated at a time, rather than all networks. This reduces computational effort and memory usage, enabling the system to run more efficiently across diverse hardware environments.

[0085] In one possible implementation, the audio coding generation module includes an embedding feature extraction module and an attention feature extraction module, wherein: the embedding feature extraction module extracts embedding features from the text to be synthesized; the attention feature extraction module extracts attention features from the text to be synthesized based on a multi-head attention mechanism; the routing module fuses the embedding features and the attention features to obtain fused features, and determines the target expert network from multiple expert networks based on the fused features.

[0086] The embedding feature extraction module extracts semantic embeddings from the text to be synthesized. These embeddings represent the text's vocabulary, grammar, and context, providing a foundational semantic representation for subsequent processing. The embedding feature extraction module uses pre-trained word embedding models (such as Word2Vec, GloVe, or Transformer) to convert each word in the text into a high-dimensional vector. These vectors represent the semantic and grammatical information of the word and can be concatenated or averaged to generate embedding features for the entire text.

[0087] For example, for the text to be synthesized, "Today's weather is sunny and suitable for outdoor activities," the text is converted into a token sequence through word segmentation and embedding mapping, and then mapped into a high-dimensional vector through the embedding layer. Deep semantic features are further extracted through a feed-forward neural network (FFN) or a Transformer encoder.

[0088] The attention feature extraction module uses a multi-head attention mechanism to extract global dependency features from text, capturing long-range contextual relationships and key semantic focal points. Attention features are generated by calculating the attention weights between each word and every other word in the synthesized text. These features highlight key parts of the text, helping the model better understand and process textual information.

[0089] The attention feature extraction module calculates the association weights for different positions in the text using a matrix of query vectors, key vectors, and value vectors. A multi-head mechanism can parallelize the calculation of multiple sets of attention weights, enhancing feature diversity. Attention features contain contextual information about the text (e.g., the strong association between "sunny" and "weather"), guiding the prosody generation of speech synthesis.

[0090] The routing module generates a scene-adapted fused feature by fusing embedded features with attention features. This fused feature incorporates the text's semantic information, grammatical structure, and the relevance and importance between different parts, enabling accurate selection of the target expert network. The routing module can concatenate or weightedly add the embedded and attention features to generate the fused feature. Based on the fused feature, a fully connected layer and a softmax activation function are used to calculate the activation probability of each expert network, ultimately determining the target expert network from among multiple expert networks.

[0091] In the embodiment of the present disclosure, the audio coding generation module ensures semantic accuracy through the embedded features and optimizes rhythmic naturalness through the attention features through the division of labor between the embedded feature extraction and the attention feature extraction; then the embedded features and the attention features are fused through the routing module, thereby improving the scene adaptation capability of the speech synthesis and thus improving the accuracy of the speech synthesis.

[0092] In a possible implementation, the routing module determines an activation score of each expert network based on the fusion feature; normalizes the activation score of each expert network to obtain an activation probability corresponding to each expert network; and determines a target expert network based on the activation probability.

[0093] Based on the fused features, the routing module uses a specific evaluation model or algorithm (such as a weight matrix in a neural network or a support vector machine) to determine the activation score of each expert network. This activation score reflects the degree to which each expert network adapts to the current fused features. For example, if the text to be synthesized is a news report, the activation score corresponding to the news broadcast expert network will be relatively high.

[0094] The audio coding generation module consists of multiple transformer layers, and the activation score r can be expressed as:

[0095] r=W r concat(e,o)

[0096] Among them, W r is an adjustable weight, i is the input of the entire transformer layer (the output of the previous layer), e is the output of the embedding extraction layer after extracting the input i, the embedding extraction layer is composed of multiple layers of FFN, o is the output of the input i after the multi-head attention mechanism, and concat(,) represents the splicing operation.

[0097] Furthermore, the activation scores of each expert network can be normalized to obtain the activation probability corresponding to each expert network, that is, the activation scores can be converted into probability distributions to ensure that the sum of the probabilities of each expert is 1, which is convenient for subsequent selection. The Softmax function can be applied to the activation scores to generate the activation probability p of each expert.i The probability distribution of :

[0098]

[0099] Among them, r i is the vector in r used to calculate the i-th expert, and r j It is used to sum all r, i and j are positive integers, and N is the number of all experts.

[0100] After determining the activation probability, the expert network that best suits the current input can be selected to perform calculations based on the normalized activation probability.

[0101] For example, the input fusion features are semantic and contextual features that represent news content; the activation scores of the features are calculated, and r = [3.2, 1.5, -0.5, 0.8] is obtained, where each value in r represents the activation score of each expert network; then, the activation scores are normalized to obtain the activation probabilities corresponding to each expert network: p = [0.85, 0.10, 0.02, 0.03]; finally, the expert network with a probability of 0.85 (news expert, probability 0.85) is selected to generate an audio coding sequence with a serious tone.

[0102] In a possible implementation, the network structure of the attention feature extraction module and the expert network is a large language model network structure.

[0103] like Figure 2 As shown, Figure 2 The block diagram of the speech synthesis device according to the embodiment of the present disclosure is shown. In this embodiment, the speech synthesis device includes a voiceprint encoding module, an autoregressive large language model (audio encoding generation module), and an audio decoder.

[0104] The audio code generation module uses a multi-layer Transformer network containing only a decoder as the network structure of the large language model. Each Transformer layer consists of a multi-head attention mechanism and a feed-forward network (FFN). In this embodiment, the multi-head attention mechanism remains unchanged, and the FFN in each Transformer layer is replaced with multiple expert networks, where the network structure of each expert network is the same as the original FFN structure.

[0105] like Figure 3 As shown, Figure 3The block diagram of a single-layer network structure according to an embodiment of the present disclosure is shown. Its input can come from the output of the upper network, and the input is extracted through the embedding feature extraction module and the attention feature extraction module, and then the embedded features and attention features are spliced ​​through the routing module. The routing layer outputs the probability distribution p = [0.85, 0.10 ... 0.03] corresponding to each expert network (for example, the expert networks E1, E2 ... EN in the figure), selects the expert network with the largest probability (0.85) for speech synthesis operation, and then outputs it to the next layer of network. The audio coding generation module may include multiple layers such as Figure 3 The network shown.

[0106] In addition, the present disclosure also provides a training method for a speech synthesis device, which is used to train any speech synthesis device provided by the present disclosure. Figure 4 A flow chart showing a method for training a speech synthesis device according to an embodiment of the present disclosure is shown in FIG. Figure 4 As shown, the method includes:

[0107] In step S21, a text sample and a reference audio sample corresponding to the audio sample are determined, wherein the reference audio sample and the audio sample belong to the same speaker;

[0108] A set of training data includes audio samples, text samples, and reference audio samples. The audio samples include the target speaker's speech data, such as a news broadcast, a story reading, or daily conversation. The text samples are text descriptions of the audio sample content, such as "The weather is nice today, perfect for a walk." The reference audio samples are other audio samples from the same speaker as the audio sample, used to extract voiceprint features and guide the speech synthesis device to generate speech with a timbre similar to that of the target speaker.

[0109] In step S22, the audio sample is encoded to obtain an audio coding sequence sample;

[0110] Before training the speech synthesis device, the audio samples are encoded to obtain code sequence samples. This is because traditional large language models, whose input and output are text tokens (i.e., the basic units of text, such as words and characters), cannot directly process or output audio data. However, here, the large language model can output a series of audio code sequences. Therefore, the audio samples can be encoded into audio code sequences and used as sample labels to guide the training of the large language model.

[0111] This process is based on the Vector Quantized Variational Autoencoder (VQVAE) technique. VQVAE is a deep neural network model that can encode audio. The network takes an audio sequence as input, and after passing through the encoder, the audio data is converted into a series of vectors (i.e., audio vectors). These audio vectors are then vector-quantized to produce discrete audio codes (i.e., audio tokens).

[0112] Vector quantization (VQ) is a technique for converting continuous data into discrete data. In VQVAE, this is achieved by mapping audio vectors to their nearest neighbor vectors in a predefined codebook. This codebook contains a series of discrete vectors, each representing a possible audio feature. By selecting the vector in the codebook that is closest to the audio vector, the continuous audio vector is converted into a discrete audio code.

[0113] Figure 5 The training architecture diagram of the audio encoder according to the embodiment of the present disclosure is shown. When training the audio encoder and vector quantization parameters, the following can be constructed: Figure 5 The training architecture shown in the figure is that the audio data passes through the audio encoder and completes the vector quantization to obtain the result {q1,q 2, …,q T After quantization, the quantized result needs to be decoded to produce the output audio. This process is achieved by the VQVAE decoder. The decoder input is the quantized audio codec, and after decoding, it outputs audio data that is as similar as possible to the input audio.

[0114] During training, the parameters of the VQVAE network are continuously adjusted to minimize the error between the output and input audio. Once training is complete, the network can be used to encode any audio. Specifically, any audio is input, and the result after the intermediate VQ layer is used as the discrete encoding of the audio, which is then used as the true value of the audio encoding output by the large language model.

[0115] Similarly, during the inference phase, after the large language model outputs discrete audio codes, other vocoders (such as Hifi-GAN) can be used to restore the codes to audio waveforms, thus completing the conversion process from text tokens to audio tokens.

[0116] In the embodiment of the present disclosure, continuous audio data is converted into discrete audio coding sequences so that a large language model can process them.

[0117] In step S23, the reference audio sample and the text sample are used as inputs of the speech synthesis device to obtain a predicted audio code sequence output by an audio code generation module of the speech synthesis device;

[0118] The reference audio sample can be input into the voiceprint encoding module to extract the voiceprint features; the text sample can be input into the speech synthesis device as the basis for generating speech. The voiceprint encoding module extracts the voiceprint features of the reference audio, and the audio code generation module selects the appropriate target expert network based on the text sample and voiceprint features to generate the predicted audio code sequence.

[0119] In step S24, parameters of the speech synthesis device are adjusted based on the loss between the predicted audio coding sequence and the audio coding sequence samples.

[0120] Here, the loss function is used to calculate the loss value by comparing the difference between the predicted audio code sequence and the actual audio code sequence sample. Then, the backpropagation algorithm is used to adjust the parameters of the speech synthesis device based on the loss value to optimize the performance of the model.

[0121] By continuously adjusting parameters, the difference between the predicted audio coding sequence and the real audio coding sequence samples is minimized, thereby improving the accuracy and naturalness of speech synthesis.

[0122] In an embodiment of the present disclosure, a text sample and a reference audio sample corresponding to an audio sample are determined, wherein the reference audio sample and the audio sample belong to the same speaker; the audio sample is encoded to obtain an audio code sequence sample; the reference audio sample and the text sample are used as inputs of the speech synthesis device to obtain a predicted audio code sequence output by the audio code generation module of the speech synthesis device; and the parameters of the speech synthesis device are adjusted based on the loss between the predicted audio code sequence and the audio code sequence sample. Thus, since the difference between the predicted audio code sequence and the actual audio code sequence sample is directly compared during the training process and the device is adjusted accordingly, the trained device has a higher accuracy in speech synthesis. By introducing a reference audio sample of the same speaker as one of the inputs, the trained device may be able to better capture the speaker's voice characteristics, intonation and other detailed information in the reference audio, thereby guiding the speech synthesis device to generate more natural and realistic speech.

[0123] In a possible implementation, adjusting parameters of the speech synthesis device based on the loss between the predicted audio coding sequence and the audio coding sequence samples includes:

[0124] When the training sample contains a scene label, the parameters of the routing module in the audio coding generation module are adjusted to maximize the activation probability of the expert network corresponding to the scene label.

[0125] In this implementation, training samples include not only audio samples, corresponding text samples, and reference audio samples, but also scene labels. Scene labels are used to indicate the specific application scenario to which the audio sample belongs, such as news broadcast, storytelling, or spoken dialogue.

[0126] The audio code generation module is responsible for generating a predicted audio code sequence based on the input text and voiceprint features. The routing module, on the other hand, is responsible for dynamically selecting the most appropriate expert network to perform speech synthesis based on the input (including text content and voiceprint features). By adjusting the parameters of the routing module, the probability of different expert networks being activated can be changed.

[0127] Then, the routing module parameters can be adjusted to maximize the probability of the expert network corresponding to the scene label being activated under a given input. In other words, when the input contains a specific scene label, the routing module should be more inclined to select the expert network designed specifically for that scene for speech synthesis.

[0128] This is achieved by introducing an additional cross-entropy loss function during training, which penalizes incorrect selection of the expert network for each scenario. By minimizing this cross-entropy loss function, the parameters of the routing module can be gradually adjusted, increasing the probability of activating the correct expert network.

[0129] In the disclosed embodiment, when training samples include scene labels, the parameters of the routing module in the audio code generation module are adjusted to maximize the activation probability of the expert network corresponding to the scene label. Thus, based on the specific scene label, a dedicated expert network is assigned to speech synthesis for each specific scene, thereby generating speech that better meets the requirements of that scene. Because each expert network is trained for a specific scene, it can generate higher-quality speech.

[0130] In a possible implementation, the parameter adjustment of the speech synthesis device based on the loss between the predicted audio coding sequence and the audio coding sequence sample includes: determining the proportion f of the number of samples assigned to the i-th expert network in the total number of samples in the current batch T in the current training batch. i ; Determine the average value P of the probability that all samples in the current batch are assigned to the i-th expert network i ; Based on f i and P iThe product of is used to construct a loss function, and the parameters of the routing module in the audio coding generation module are adjusted with the goal of minimizing the loss function, so that the audio coding generation module evenly distributes samples to each expert.

[0131] When training a speech synthesis device, multiple training batches are used for training. After the training samples of the current batch are input into the speech synthesis device, the speech synthesis device will assign the corresponding expert network to each sample in the current batch. At this time, the proportion f of the number of samples assigned to the i-th expert network to the total number of samples in the current batch T can be counted. i .

[0132] Furthermore, when the speech synthesis device assigns the corresponding expert network to each sample in the current batch, the probability p of each expert network is calculated. i , determine the average value P of the probability that all samples in the current batch are assigned to the i-th expert network i , that is, for each expert i, accumulate the p of all T samples i (x), and then find the average. The calculation process can be seen in the following formula:

[0133]

[0134] Among them, x is the sample number, p i (x) is the probability that sample x is assigned to expert i, and T is the total number of samples in the current batch.

[0135] Then calculate f for each expert i and P i The product of , to quantify the imbalance of expert i. If an expert is frequently activated (f i High) and the model is highly dependent on it (P i high), the product value is large, indicating that the utilization rate of the expert is too high and the utilization rates of other experts are low, resulting in other experts not being well trained.

[0136] Constructed loss function L b It can be expressed by the following formula:

[0137]

[0138] Among them, N is the number of all experts, that is, the loss function L b To get all experts' corresponding f i ·P i Add and multiply by N to amplify the loss value and ensure that the loss function is sensitive to the imbalance of sample distribution.

[0139] Assume N = 2 experts:

[0140] In the case of balanced distribution: f1 = f2 = 0.5, P1 = P2 = 0.5;

[0141] L b =2×(0.5×0.5+0.5×0.5)=2×(0.25+0.25)=2×0.5=1.

[0142] In the case of unbalanced distribution: f1=1, f2=0, P1=1, P2=0;

[0143] L b =2×(1×1+0×0)=2×(1+0)=2.

[0144] Obviously, by minimizing L b will promote balanced distribution, because L b The smaller it is, the more evenly the training samples are distributed among the experts, so that each expert can be well trained.

[0145] In one possible implementation, the speech synthesis method can be executed by electronic devices such as terminal devices and servers. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be executed by a processor calling computer-readable instructions stored in a memory.

[0146] In addition, the present disclosure also provides a speech synthesis method, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any speech synthesis device provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method section and will not be repeated here.

[0147] Figure 6 A flow chart of a speech synthesis method according to an embodiment of the present disclosure is shown as follows: Figure 6 As shown, the method includes:

[0148] In step S31, receiving an input reference audio and extracting a voiceprint feature of the reference audio;

[0149] In step S32, based on the input text to be synthesized and voiceprint features, a target expert network is determined from multiple expert networks to perform speech synthesis operations to obtain a generated audio code sequence; wherein the multiple expert networks correspond to different application scenarios;

[0150] In step S33, the audio code sequence is converted into an audio signal.

[0151] In one possible implementation, the method of determining a target expert network from a plurality of expert networks to perform speech synthesis operations based on the input text to be synthesized and voiceprint features to obtain a generated audio code sequence includes:

[0152] Based on the text to be synthesized and the voiceprint features, determining a target expert network that matches the scenario of the text to be synthesized and the voiceprint features from a plurality of expert networks;

[0153] The target expert network uses the voiceprint feature as a reference to generate an audio coding sequence corresponding to the text to be synthesized.

[0154] In one possible implementation, the method of determining a target expert network from a plurality of expert networks to perform speech synthesis operations based on the input text to be synthesized and voiceprint features to obtain a generated audio code sequence includes:

[0155] Extracting embedded features from the text to be synthesized;

[0156] Extracting attention features from the text to be synthesized based on a multi-head attention mechanism;

[0157] The embedded features and the attention features are fused to obtain fused features, and a target expert network is determined from multiple expert networks based on the fused features.

[0158] In a possible implementation, determining, from a plurality of expert networks, a target expert network that matches a scenario of the text to be synthesized and the voiceprint feature based on the text to be synthesized and the voiceprint feature includes:

[0159] Based on the fusion features, an activation score of each expert network is determined; the activation score of each expert network is normalized to obtain an activation probability corresponding to each expert network; and a target expert network is determined based on the activation probability.

[0160] In a possible implementation, the extraction of attention features from the text to be synthesized based on the multi-head attention mechanism is implemented by an attention feature extraction module, and the network structure of the attention feature extraction module and the expert network is a large language model network structure.

[0161] This method has a specific technical connection with the internal structure of the computer system, and can solve the technical problem of how to improve the hardware computing efficiency or execution effect (including reducing the amount of data storage, reducing the amount of data transmission, increasing the hardware processing speed, etc.), thereby obtaining the technical effect of improving the internal performance of the computer system in accordance with the laws of nature.

[0162] In some embodiments, the specific implementation of the method provided by the embodiments of the present disclosure can refer to the description of the above device embodiments, and for the sake of brevity, it will not be repeated here.

[0163] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.

[0164] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the above method.

[0165] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0166] The electronic device may be provided as a terminal, a server, or other forms of devices.

[0167] Figure 7 1 shows a block diagram of an electronic device according to an embodiment of the present disclosure. For example, the electronic device 1900 can be provided as a server or a terminal device. Figure 7 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.

[0168] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as a Microsoft Server operating system (Windows Server 2003). TM ), a graphical user interface operating system launched by Apple (Mac OS X TM ), a multi-user, multi-process computer operating system (Unix TM ), a free and open source Unix-like operating system (Linux TM), an open-source Unix-like operating system (FreeBSD TM ) or similar.

[0169] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.

[0170] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0171] Computer-readable storage media can be a tangible device that can hold and store the instructions used by the instruction execution device. Computer-readable storage media can be, for example, (but not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove on which instructions are stored, and any suitable combination thereof. Computer-readable storage media used herein is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (for example, light pulses by fiber optic cables), or electrical signals transmitted by wires.

[0172] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0173] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0174] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0175] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0176] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0177] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0178] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0179] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0180] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0181] If the technical solution of this application involves personal information, the product that applies the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves sensitive personal information, the product that applies the technical solution of this application has obtained the individual's separate consent before processing sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information; among which, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0182] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A speech synthesis device, characterized in that: include: A voiceprint encoding module receives an input reference audio and extracts the voiceprint features of the reference audio; An audio code generation module, which determines a target expert network from multiple expert networks based on the input text to be synthesized and voiceprint features, performs speech synthesis operations, and generates an audio code sequence; wherein the multiple expert networks correspond to different application scenarios; The audio decoding module converts the audio code sequence output by the audio code generation module into an audio signal.

2. The device according to claim 1, characterized in that The audio coding generation module includes a routing module and multiple expert networks, wherein: The routing module determines, based on the text to be synthesized and the voiceprint features, a target expert network that matches the scenario of the text to be synthesized and the voiceprint features from a plurality of expert networks; The target expert network uses the voiceprint feature as a reference to generate an audio coding sequence corresponding to the text to be synthesized.

3. The device according to claim 2, characterized in that The audio coding generation module includes an embedding feature extraction module and an attention feature extraction module, wherein: The embedding feature extraction module extracts embedding features from the text to be synthesized; The attention feature extraction module extracts attention features from the text to be synthesized based on a multi-head attention mechanism; The routing module fuses the embedded features and the attention features to obtain fused features, and determines a target expert network from multiple expert networks based on the fused features.

4. The device according to claim 3, characterized in that The routing module determines the activation score of each expert network based on the fusion feature; normalizes the activation score of each expert network to obtain the activation probability corresponding to each expert network; and determines the target expert network based on the activation probability.

5. The device according to claim 3, characterized in that The network structure of the attention feature extraction module and the expert network is a large language model network structure.

6. A training method for a speech synthesis device, characterized in that: For training the apparatus according to any one of claims 1 to 5, the method comprising: Determining a text sample and a reference audio sample corresponding to the audio sample, wherein the reference audio sample and the audio sample belong to the same speaker; Encoding the audio sample to obtain an audio coding sequence sample; Using the reference audio sample and the text sample as inputs of the speech synthesis device, and obtaining a predicted audio code sequence output by an audio code generation module of the speech synthesis device; Based on the loss between the predicted audio coding sequence and the audio coding sequence samples, parameters of the speech synthesis device are adjusted.

7. The method according to claim 6, characterized in that The step of adjusting parameters of the speech synthesis device based on the loss between the predicted audio coding sequence and the audio coding sequence samples includes: When the training sample contains a scene label, the parameters of the routing module in the audio coding generation module are adjusted to maximize the activation probability of the expert network corresponding to the scene label.

8. The method according to claim 6, characterized in that The step of adjusting parameters of the speech synthesis device based on the loss between the predicted audio coding sequence and the audio coding sequence samples includes: Determine the ratio f of the number of samples assigned to the i-th expert network in the current training batch to the total number of samples T in the current batch i ; Determine the average value P of the probability that all samples in the current batch are assigned to the i-th expert network i ; Based on f i and P i The product of is used to construct a loss function, and the parameters of the routing module in the audio coding generation module are adjusted with the goal of minimizing the loss function, so that the audio coding generation module evenly distributes samples to each expert.

9. A speech synthesis method, characterized in that: include: receiving an input reference audio, and extracting a voiceprint feature of the reference audio; Determine a target expert network from multiple expert networks based on the input text to be synthesized and voiceprint features to perform speech synthesis operations, thereby generating an audio coding sequence; wherein the multiple expert networks correspond to different application scenarios; The audio code sequence is converted into an audio signal.

10. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to implement the method according to any one of claims 6 to 9.

11. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 6 to 9 is implemented.

Citation Information

Patent Citations

  • Audio generator and methods for generating an audio signal and training an audio generator

    CA3195578A1

  • Speech recognition method and device, speech recognition model training method and device, medium and equipment

    CN116013257A

  • Speech synthesis method, speech synthesis device, electronic equipment and storage medium

    CN116343747A

  • Speech synthesis method and device, computer equipment and readable storage medium

    CN119516997A

  • Composite style voice generation method and device, equipment and storage medium

    CN119889285A

Cited By

  • Sound duplicating method and related device

    CN122157639A

  • Sound replication method and related apparatus

    CN122157639B