Audio generation methods, devices, storage media and chips

By introducing gating networks and feature extraction networks into the audio synthesis model, the overfitting problem of user-specific speech synthesis models is solved, audio fidelity is improved, computational load is reduced, and speech synthesis efficiency is increased.

CN115240638BActive Publication Date: 2025-11-14BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210887736.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-26
Publication Date
2025-11-14
Estimated Expiration
2042-07-26

AI Technical Summary

Technical Problem

In existing technologies, user-specific speech synthesis models are prone to overfitting due to the small amount of training data, which prevents further improvement in the realism of AI-synthesized speech. In addition, they involve a large amount of computation and have low speech synthesis efficiency.

Method used

A pre-defined audio synthesis model is adopted, including a gating network and multiple feature extraction networks. The target feature extraction network is determined from the feature extraction networks through the gating network to generate target audio data, thereby reducing the amount of computation and improving the realism.

Benefits of technology

It effectively overcomes the overfitting problem, improves the realism of synthesized audio, and significantly reduces the amount of computation required to generate target audio data, thereby improving the efficiency of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240638B_ABST
    Figure CN115240638B_ABST
Patent Text Reader

Abstract

This disclosure relates to an audio generation method, apparatus, storage medium, and chip. The audio generation method inputs target text information into a preset audio synthesis model to obtain audio data with a specified timbre corresponding to the target text information. The preset audio synthesis model includes a gating network and multiple feature extraction networks. The gating network is used to determine a target feature extraction network from the multiple feature extraction networks, and the target feature extraction network is used to determine the target audio data corresponding to the target text information. Thus, by determining the target feature extraction network from the multiple feature extraction networks through the gating network in the preset audio synthesis model, and then determining the target audio data corresponding to the target text information through the target feature extraction network, the overfitting problem that easily occurs in related technologies when training data is limited can be effectively overcome, and the computational load required to generate the target audio data can also be significantly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of electronic equipment technology, and in particular to audio generation methods, apparatus, storage media and chips. Background Technology

[0002] With the rapid development of intelligent voice technology, AI (Artificial Intelligence) synthesized speech sounds more realistic and natural, and is now widely used in various scenarios such as voice assistants, chatbots, audiobooks, and news broadcasting. Customized voice synthesis, based on advanced deep learning technology, requires only a small amount of audio data uploaded by the user to quickly synthesize a voice synthesis model with a unique user tone. Summary of the Invention

[0003] This disclosure provides an audio generation method, apparatus, storage medium, and chip.

[0004] According to a first aspect of the present disclosure, an audio generation method is provided, comprising:

[0005] Obtain target text information;

[0006] The target text information is input into a preset audio synthesis model to obtain the target audio data output by the preset audio synthesis model. The target audio data is the audio data of the specified timbre corresponding to the target text information.

[0007] The preset audio synthesis model includes a gating network and multiple feature extraction networks. Different feature extraction networks are used to extract feature data of different dimensions. The gating network is used to determine the target feature extraction network from the multiple feature extraction networks. The target feature extraction network is used to determine the target audio data corresponding to the target text information.

[0008] Optionally, the step of inputting the target text information into a preset audio synthesis model to obtain the target audio data output by the preset audio synthesis model includes:

[0009] The target text information is input into the gating network to obtain the feature matrix output by the gating network. Different elements in the feature matrix are used to represent the weights of features of different dimensions.

[0010] At least one target feature extraction network is determined from the plurality of feature extraction networks based on the feature matrix;

[0011] The target text information is input into the target feature extraction network to obtain the target feature data output by the target feature extraction network.

[0012] The target audio data is determined based on the target feature data.

[0013] Optionally, determining at least one target feature extraction network from the plurality of feature extraction networks based on the feature matrix includes:

[0014] Based on the feature matrix, determine one or more target feature extraction networks corresponding to a preset number of dimensional features with the largest weights.

[0015] Optionally, the preset audio synthesis model can be trained in the following way:

[0016] Acquire first audio sample data of multiple specified timbres, and first text information corresponding to the first audio sample data;

[0017] Using multiple first audio sample data and the first text information corresponding to each first audio sample data as training data, a preset pre-trained model is trained to obtain the preset audio synthesis model.

[0018] Optionally, the pre-trained model includes a gated network to be determined and multiple feature extraction networks to be determined. The step of training the pre-trained model using multiple first audio sample data and the first text information corresponding to each first audio sample data as training data to obtain the pre-trained audio synthesis model includes:

[0019] The first text information corresponding to each first audio sample data is input into the undetermined gating network to obtain the undetermined feature matrix output by the undetermined gating network. The elements in the undetermined feature matrix are used to characterize the undetermined weights of features of different dimensions.

[0020] The first undetermined loss value is determined based on the undetermined feature matrix using a loss function that maximizes the preset nuclear norm.

[0021] Based on the undetermined feature matrix, determine one or more designated feature extraction networks corresponding to a preset number of dimensions with the largest undetermined weight, and input the first text information into the one or more designated feature extraction networks to obtain the designated feature data output by each designated feature extraction network;

[0022] The specified audio data is determined based on the specified feature data, and the second undetermined loss value of the first loss function is determined based on the specified audio data and the first audio sample data.

[0023] The first target loss value is determined based on the first undetermined loss value and the second undetermined loss value;

[0024] Determine whether the first target loss value is greater than the first preset loss threshold;

[0025] If the first target loss value is greater than the first preset loss threshold, the model parameters of the pre-trained model are adjusted to obtain an updated pre-trained model. The step of inputting the first text information corresponding to each first audio sample data into the undetermined gating network is executed again until the first target loss value is determined to be greater than the first preset loss threshold is determined. Then, if the first target loss value is determined to be greater than the first preset loss threshold, the current pre-trained model is used as the preset audio synthesis model.

[0026] Optionally, the pre-trained model is trained in the following manner:

[0027] Acquire second audio sample data with multiple different timbres, and second text information corresponding to the second audio sample data;

[0028] The preset initial model is trained using the second audio sample data with multiple different timbres and the second text information corresponding to the second audio sample data as training data to obtain the pre-trained model.

[0029] Optionally, the preset initial model includes an initial gating network and multiple initial feature extraction networks. The step of training the preset initial model using the multiple second audio sample data with different timbres and the corresponding second text information as training data to obtain the pre-trained model includes:

[0030] The second text information corresponding to each of the second audio sample data is input into the initial gating network and each of the initial feature extraction networks to obtain the initial feature matrix output by the initial gating network and the initial feature data output by the initial feature extraction network. The elements in the initial feature matrix are used to characterize the initial weights of features of different dimensions.

[0031] The initial feature data output by the multiple initial feature extraction networks are weighted and summed according to the initial weights to obtain the initial audio data;

[0032] The preset initial model is iteratively trained based on the initial audio data and the preset loss function to obtain the pre-trained model.

[0033] Optionally, the preset loss function includes a first loss function and a second loss function, wherein the second loss function is a loss function that maximizes a preset nuclear norm, and the step of iteratively training the preset initial model based on the initial audio data and the preset loss function to obtain the pre-trained model includes:

[0034] Based on the first loss function, a first specified loss value is determined from the initial audio data and the second audio sample data;

[0035] The second specified loss value is determined based on the initial feature matrix and the second loss function;

[0036] A second target loss value is determined based on the first specified loss value and the second specified loss value. If the second target loss value is greater than or equal to the second preset loss threshold, the model parameters of the preset initial model are adjusted to obtain an updated preset initial model. The step of inputting the second text information corresponding to each second audio sample data into the initial gating network and each initial feature extraction network to obtain the initial feature matrix output by the initial gating network and the initial feature data output by the initial feature extraction network is executed again until the second target loss value is determined to be less than the preset loss threshold. Then, the current preset initial model is used as the pre-trained model.

[0037] According to a second aspect of the present disclosure, an audio generation apparatus is provided, comprising:

[0038] The first acquisition module is configured to acquire target text information;

[0039] The first determining module is configured to input the target text information into a preset audio synthesis model to obtain target audio data output by the preset audio synthesis model, wherein the target audio data is audio data of a specified timbre corresponding to the target text information.

[0040] The preset audio synthesis model includes a gating network and multiple feature extraction networks. Different feature extraction networks are used to extract feature data of different dimensions. The gating network is used to determine the target feature extraction network from the multiple feature extraction networks. The target feature extraction network is used to determine the target audio data corresponding to the target text information.

[0041] Optionally, the first determining module is configured to:

[0042] The target text information is input into the gating network to obtain the feature matrix output by the gating network. Different elements in the feature matrix are used to represent the weights of features of different dimensions.

[0043] At least one target feature extraction network is determined from the plurality of feature extraction networks based on the feature matrix;

[0044] The target text information is input into the target feature extraction network to obtain the target feature data output by the target feature extraction network.

[0045] The target audio data is determined based on the target feature data.

[0046] Optionally, the first determining module is configured to:

[0047] Based on the feature matrix, determine one or more target feature extraction networks corresponding to a preset number of dimensional features with the largest weights.

[0048] Optionally, the device further includes:

[0049] The second acquisition module is configured to acquire first audio sample data of multiple specified timbres, and first text information corresponding to the first audio sample data;

[0050] The second determining module is configured to use multiple first audio sample data and the first text information corresponding to each first audio sample data as training data to train a preset pre-trained model to obtain the preset audio synthesis model.

[0051] Optionally, the pre-trained model includes a gated network to be determined and multiple feature extraction networks to be determined, and the second determining module is configured to:

[0052] The first text information corresponding to each first audio sample data is input into the undetermined gating network to obtain the undetermined feature matrix output by the undetermined gating network. The elements in the undetermined feature matrix are used to characterize the undetermined weights of features of different dimensions.

[0053] The first undetermined loss value is determined based on the undetermined feature matrix using a loss function that maximizes the preset nuclear norm.

[0054] Based on the undetermined feature matrix, determine one or more designated feature extraction networks corresponding to a preset number of dimensions with the largest undetermined weight, and input the first text information into the one or more designated feature extraction networks to obtain the designated feature data output by each designated feature extraction network;

[0055] The specified audio data is determined based on the specified feature data, and the second undetermined loss value of the first loss function is determined based on the specified audio data and the first audio sample data.

[0056] The first target loss value is determined based on the first undetermined loss value and the second undetermined loss value;

[0057] Determine whether the first target loss value is greater than the first preset loss threshold;

[0058] If the first target loss value is greater than the first preset loss threshold, the model parameters of the pre-trained model are adjusted to obtain an updated pre-trained model. The step of inputting the first text information corresponding to each first audio sample data into the undetermined gating network is executed again until the first target loss value is determined to be greater than the first preset loss threshold is determined. Then, if the first target loss value is determined to be greater than the first preset loss threshold, the current pre-trained model is used as the preset audio synthesis model.

[0059] Optionally, the device further includes a model training module configured to:

[0060] Acquire second audio sample data with multiple different timbres, and second text information corresponding to the second audio sample data;

[0061] The preset initial model is trained using the second audio sample data with multiple different timbres and the second text information corresponding to the second audio sample data as training data to obtain the pre-trained model.

[0062] Optionally, the preset initial model includes an initial gating network and multiple initial feature extraction networks, and the model training module is configured as follows:

[0063] The second text information corresponding to each second audio sample data is input into the initial gating network and each initial feature extraction network to obtain the initial feature matrix output by the initial gating network and the initial feature data output by the initial feature extraction network. The elements in the initial feature matrix are used to characterize the initial weights of features of different dimensions.

[0064] The initial feature data output by the multiple initial feature extraction networks are weighted and summed according to the initial weights to obtain the initial audio data;

[0065] The preset initial model is iteratively trained based on the initial audio data and the preset loss function to obtain the pre-trained model.

[0066] Optionally, the preset loss function includes a first loss function and a second loss function, wherein the second loss function is a loss function that maximizes a preset nuclear norm, and the model training module is configured as follows:

[0067] Based on the first loss function, a first specified loss value is determined from the initial audio data and the second audio sample data;

[0068] The second specified loss value is determined based on the initial feature matrix and the second loss function;

[0069] A second target loss value is determined based on the first specified loss value and the second specified loss value. If the second target loss value is greater than or equal to the second preset loss threshold, the model parameters of the preset initial model are adjusted to obtain an updated preset initial model. The step of inputting each second audio sample data and the second text information corresponding to the second audio sample data into the initial gating network and each initial feature extraction network to obtain the initial feature matrix output by the initial gating network and the initial feature data output by the initial feature extraction network is executed again until the second target loss value is determined to be less than the preset loss threshold. Then, the current preset initial model is used as the pre-trained model.

[0070] According to a third aspect of the present disclosure, an audio generation apparatus is provided, comprising:

[0071] processor;

[0072] Memory used to store processor-executable instructions;

[0073] The processor is configured as follows:

[0074] The steps to implement the method described in the first aspect above.

[0075] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that stores computer program instructions thereon, which, when executed by a processor, implement the steps of the method described in the first aspect above.

[0076] According to a fifth aspect of the present disclosure, a chip is provided, including a processor and an interface; the processor is configured to read instructions to execute the method described in the first aspect above.

[0077] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:

[0078] This method allows for the input of target text information into a preset audio synthesis model to obtain audio data with a specified timbre corresponding to the target text information. The preset audio synthesis model includes a gating network and multiple feature extraction networks. Different feature extraction networks are used to extract feature data of different dimensions. The gating network is used to determine a target feature extraction network from the multiple feature extraction networks, and the target feature extraction network is used to determine the target audio data corresponding to the target text information. Thus, by using the gating network in the preset audio synthesis model to determine the target feature extraction network from the multiple feature extraction networks, and then using this target feature extraction network to determine the target audio data corresponding to the target text information, the computational load required to generate the target audio data can be significantly reduced while effectively improving the realism of the synthesized audio.

[0079] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0080] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0081] Figure 1 This is a flowchart illustrating an audio generation method according to an exemplary embodiment;

[0082] Figure 2 Based on this disclosure Figure 1 The illustrated embodiment shows a flowchart of an audio generation method;

[0083] Figure 3 This is a flowchart illustrating an audio generation method according to another exemplary embodiment of this disclosure;

[0084] Figure 4 This is a schematic diagram illustrating the structure of a pre-trained model according to an exemplary embodiment of the present disclosure;

[0085] Figure 5 This is a flowchart illustrating a training method for a pre-trained model according to an exemplary embodiment of this disclosure;

[0086] Figure 6 Based on this disclosure Figure 5 The illustrated embodiment shows a flowchart of a training method for a pre-trained model;

[0087] Figure 7 This is a block diagram illustrating an audio generation apparatus according to an exemplary embodiment of the present disclosure;

[0088] Figure 8 Based on this disclosure Figure 7 The illustrated embodiment shows a block diagram of an audio generation apparatus;

[0089] Figure 9 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation

[0090] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0091] It should be noted that all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located, and with the authorization granted by the owner of the relevant device.

[0092] Before detailing the specific implementation methods of this disclosure, the application scenarios of this disclosure are explained below. This disclosure can be applied to human-computer dialogue scenarios, such as achieving multi-turn interaction between the device and the user by synthesizing the user's voice. It can also be applied to voice broadcasting scenarios, such as broadcasting text messages, weather, traffic conditions, and program schedules. It can also be applied to entertainment scenarios, such as allowing users to tell jokes, read novels, sing, etc., using their own voice. Furthermore, it can be applied to educational scenarios, such as teaching by synthesizing the user's own voice.

[0093] In related technologies, a large dataset is typically used to train a complex neural network model via machine learning as a pre-trained base model. Then, a small amount of user speech data is acquired and used as training data to fine-tune the pre-trained base model, resulting in a fine-tuned neural network model. Finally, this fine-tuned neural network model is used as a user-specific voice synthesis model. In the field of deep learning, the size of the training data and the model capacity are crucial to its success. When the training data is large enough, increasing the neural network capacity (number of parameters) can achieve higher prediction accuracy. However, when the training data is small, the complex model structure generates a large number of redundant parameters, leading to overfitting. In other words, most user-specific voice synthesis models in related technologies suffer from overfitting due to insufficient training data, thus stagnating the realism of AI-synthesized speech and preventing further improvement.

[0094] To address the aforementioned technical issues, this disclosure provides an audio generation method, apparatus, storage medium, and chip. The audio generation method inputs target text information into a preset audio synthesis model to obtain audio data with a specified timbre corresponding to the target text information. The preset audio synthesis model includes a gating network and multiple feature extraction networks. Different feature extraction networks are used to extract feature data of different dimensions. The gating network is used to determine a target feature extraction network from the multiple feature extraction networks. This target feature extraction network is then used to determine the target audio data corresponding to the target text information. Thus, by using the gating network in the preset audio synthesis model to determine the target feature extraction network from the multiple feature extraction networks, and then using this target feature extraction network to determine the target audio data corresponding to the target text information, the method effectively overcomes the problems of large model structures, small training data amounts, and overfitting issues in related technologies. This not only improves the realism of the synthesized audio but also significantly reduces the computational load required to generate the target audio data, thereby improving speech synthesis efficiency.

[0095] Figure 1 This is a flowchart illustrating an audio generation method according to an exemplary embodiment, such as... Figure 1 As shown, the audio generation method may include the following steps:

[0096] Step 101: Obtain the target text information.

[0097] The target text information is the text content corresponding to the target audio data to be generated.

[0098] Step 102: Input the target text information into a preset audio synthesis model to obtain the target audio data output by the preset audio synthesis model. The target audio data is the audio data of the specified timbre corresponding to the target text information.

[0099] The preset audio synthesis model includes a gating network and multiple feature extraction networks. Different feature extraction networks are used to extract feature data of different dimensions. The gating network is used to determine the target feature extraction network from the multiple feature extraction networks. The target feature extraction network is used to determine the target audio data corresponding to the target text information.

[0100] For example, in a human-computer dialogue scenario, if a user asks the terminal "How is the weather today?", the terminal will first obtain the reply text used to answer "How is the weather today?" (e.g., Today is sunny, temperature 25℃, light southwest wind). The reply text is then input into the preset voice synthesis model as the target text information, so that the preset voice synthesis model outputs audio data with a specified timbre, such as generating audio data with the user's timbre as "Today is sunny, temperature 25℃, light southwest wind".

[0101] It should be noted that the specified timbre can be the user's own voice timbre, or other specified timbre, such as the timbre of a celebrity, a broadcaster, or a child. The gating network can be any classification network structure in the prior art, used to output the influence weight of each feature extraction network's corresponding dimension on the generation of the target audio data. Then, based on this influence weight, the dimension with the greater influence is selected, and the feature extraction network corresponding to the dimension with the greater influence is used as the target feature extraction network. Thus, the target audio data is obtained based on the gating network and the target feature extraction network. This effectively reduces the actual model structure used and avoids overfitting when the training data volume is small. It should also be pointed out that the feature extraction network can be any network module in the prior art that can be used for feature extraction; this disclosure does not limit it.

[0102] The above technical solution determines the target feature extraction network from multiple feature extraction networks by using a gated network in the preset audio synthesis model. Then, the target feature extraction network determines the target audio data corresponding to the target text information. This effectively overcomes the problems of large model structure, small training data volume, and easy overfitting in related technologies. As a result, it can not only improve the realism of synthesized audio, but also significantly reduce the amount of computation required to generate target audio data and improve speech synthesis efficiency.

[0103] Figure 2 Based on this disclosure Figure 1 The illustrated embodiment presents a flowchart of an audio generation method; as shown in the figure. Figure 2 As shown, step 102 above, which involves inputting the target text information into a preset audio synthesis model to obtain the target audio data output by the preset audio synthesis model, may include:

[0104] Step 1021: Input the target text information into the gating network to obtain the feature matrix output by the gating network. Different elements in the feature matrix are used to represent the weights of features in different dimensions.

[0105] For example, if the feature matrix output by the gated network is D-dimensional, then each element in the feature matrix represents the weight of a feature in one dimension.

[0106] Step 1022: Determine at least one target feature extraction network from the plurality of feature extraction networks based on the feature matrix.

[0107] In this step, one or more target feature extraction networks can be determined based on the feature matrix, corresponding to a preset number of dimensional features with the largest weights.

[0108] For example, the feature extraction network containing the dimension feature with the largest weight in the feature matrix output by the gated network can be used as the target feature extraction network, or the feature extraction network corresponding to the three dimensions with the largest weight can be used as the target feature extraction network. It should be noted that the feature extraction networks corresponding to the three dimensions with the largest weight can be the same or different.

[0109] Step 1023: Input the target text information into the target feature extraction network to obtain the target feature data output by the target feature extraction network.

[0110] Step 1024: Determine the target audio data based on the target feature data.

[0111] In this step, if there are multiple target feature extraction networks, the multiple target feature data can be weighted and summed to obtain the target audio data; if there is only one target feature extraction network, the target feature data can be used as the target audio data.

[0112] The above technical solution, by determining at least one target feature extraction network from the multiple feature extraction networks, and determining the target audio data based on the target feature data output by the target feature extraction network, can effectively reduce the amount of computation required for the audio data generation process of a specified timbre while ensuring the fidelity of the audio data, improve the generation efficiency of audio data of a specified timbre, reduce user waiting time during synthesized speech, and thus effectively improve the user experience.

[0113] Figure 3 This is a flowchart illustrating another exemplary embodiment of the present disclosure of an audio generation method, such as... Figure 3 As shown, the audio generation method includes:

[0114] Step 301: In response to receiving the first instruction, acquire multiple first audio sample data of the specified timbre, and the first text information corresponding to the first audio sample data.

[0115] The first instruction can be a user-triggered instruction for training a preset audio synthesis model. By triggering the first instruction, training of the preset audio synthesis model begins. When the specified timbre is the user's own timbre, the multiple first audio sample data of the specified timbre can be multiple voice information input by the user (e.g., 5 voice sentences, 10 voice sentences, 20 voice sentences, etc.). The first text information corresponding to the first audio sample data can be text data obtained after speech recognition of the user's input voice information, or it can be text information manually entered by the user after inputting the voice information.

[0116] Step 302: Using multiple first audio sample data and the first text information corresponding to each first audio sample data as training data, a preset pre-trained model is trained to obtain the preset audio synthesis model.

[0117] The pre-trained model includes an undetermined gating network and multiple undetermined feature extraction networks.

[0118] For example, Figure 4 This is a schematic diagram illustrating the structure of a pre-trained model according to an exemplary embodiment of this disclosure; as shown below. Figure 4 As shown, multiple feature extraction networks to be determined are represented by N sub-layers (sparsely gated parallel sub-modules), each sub-layer being a feature extraction network to be determined, which is a gating network.

[0119] In this step, the first text information corresponding to each first audio sample data can be input into the undetermined gating network to obtain the undetermined feature matrix output by the undetermined gating network. The elements in the undetermined feature matrix are used to characterize the undetermined weights of features of different dimensions. The pre-trained model is iteratively trained according to the undetermined feature matrix through a loss function that maximizes the preset nuclear norm to obtain the preset audio synthesis model.

[0120] It should be noted that the above-described implementation method of iteratively training the pre-trained model using a loss function that maximizes the predefined kernel norm based on the undetermined feature matrix to obtain the predefined audio synthesis model is as follows: After inputting the first text information corresponding to a first audio sample data into the undetermined gating network, the current undetermined feature matrix is ​​obtained, and calculation is performed based on the undetermined feature matrix. Where X is the feature matrix to be determined, B is the batch size during training, and one or more designated feature extraction networks corresponding to the preset number of features with the largest undetermined weights are determined according to the feature matrix to be determined; the first text information is input into the one or more designated feature extraction networks to obtain the designated feature data output by each designated feature extraction network; the current designated audio data is determined according to the designated feature data, and the second undetermined loss value of the first loss function is determined according to the designated audio data and the first audio sample data; the first target loss value is determined according to the first undetermined loss value and the second undetermined loss value, and it is determined whether the first target loss value is greater than the first preset loss threshold; if the first target loss value is greater than the first preset loss threshold, the model parameters of the pre-trained model are adjusted to obtain the updated pre-trained model, and the step of inputting the first text information corresponding to each first audio sample data into the undetermined gating network is executed again until the first target loss value is determined to be greater than the first preset loss threshold is determined, until the current pre-trained model is used as the preset audio synthesis model when the first target loss value is determined to be greater than the first preset loss threshold.

[0121] It should be noted that the first loss function can be a logarithmic loss function, a cross-entropy loss function, or a squared loss function, and the first target loss value can be the result of a weighted sum of the first undetermined loss value and the second undetermined loss value.

[0122] In this way, during fine-tuning, only one (or a few) sub-layers with the highest probability need to be selected for parameter adjustment, and the calculation of model parameters for the remaining N-1 sub-layers can be omitted. That is, the amount of computation is only 1 / N of that of a normal model, which can effectively ensure that a small number of parameters achieve the best results on complex model structures, while also greatly reducing computational resources and improving computational speed.

[0123] Step 303: In response to receiving the second instruction, obtain the target text information.

[0124] The second instruction is triggered after the preset audio synthesis model has been trained, in order to invoke the preset audio synthesis model. The second instruction can be a high-level signal, a low-level signal, an interrupt signal, or other signals in the prior art, which will not be listed here.

[0125] Step 304: Input the target text information into a preset audio synthesis model to obtain the target audio data output by the preset audio synthesis model. The target audio data is the audio data of the specified timbre corresponding to the target text information.

[0126] The preset audio synthesis model includes a gating network and multiple feature extraction networks. Different feature extraction networks are used to extract feature data of different dimensions. The gating network is used to determine the target feature extraction network from the multiple feature extraction networks. The target feature extraction network is used to determine the target audio data corresponding to the target text information.

[0127] The implementation method for this step can be found above. Figure 2 The contents shown in steps 1021 to 1024 are not repeated here.

[0128] The above technical solution determines the target feature extraction network from multiple feature extraction networks by using a gated network in the preset audio synthesis model. Then, the target feature extraction network determines the target audio data corresponding to the target text information. This effectively overcomes the problems of large model structure, small training data volume, and easy overfitting in related technologies. As a result, it can not only improve the realism of synthesized audio, but also significantly reduce the amount of computation required to generate target audio data and improve speech synthesis efficiency.

[0129] Figure 5 This is a flowchart illustrating a training method for a pre-trained model according to an exemplary embodiment of this disclosure; as shown below. Figure 5 As shown, this pre-trained model can be trained in the following way:

[0130] Step 501: Obtain multiple second audio sample data with different timbres, and the second text information corresponding to the second audio sample data.

[0131] Step 502: Using the multiple second audio sample data with different timbres and the second text information corresponding to the second audio sample data as training data, train the preset initial model to obtain the pre-trained model.

[0132] The preset initial model includes an initial gating network and multiple initial feature extraction networks. This step can be achieved through... Figure 6 The steps shown are to be completed. Figure 6 Based on this disclosure Figure 5 The illustrated embodiment presents a flowchart of a training method for a pre-trained model; as shown in the figure. Figure 6 As shown:

[0133] S1, input the second text information corresponding to each second audio sample data into the initial gating network and each initial feature extraction network to obtain the initial feature matrix output by the initial gating network and the initial feature data output by the initial feature extraction network. The elements in the initial feature matrix are used to characterize the initial weights of features of different dimensions.

[0134] S2, based on the initial weights, perform a weighted summation of the multiple initial feature data output by the multiple initial feature extraction networks to obtain the initial audio data.

[0135] S3. Based on the initial audio data and the preset loss function, iteratively train the preset initial model to obtain the pre-trained model.

[0136] The preset loss function includes a first loss function and a second loss function, wherein the second loss function is a loss function that maximizes the preset nuclear norm.

[0137] It should be noted that the first loss function can be a logarithmic loss function, a cross-entropy loss function, or a squared loss function, etc., and the loss function that maximizes the preset nuclear norm is:

[0138]

[0139] Where X is the initial feature matrix and B is the batch size during training.

[0140] In this step, a first specified loss value can be determined based on the first loss function, the initial audio data, and the second audio sample data; a second specified loss value can be determined based on the initial feature matrix and the loss function that maximizes the preset kernel norm; a second target loss value can be determined based on the first specified loss value and the second specified loss value; if the second target loss value is determined to be greater than or equal to the second preset loss threshold, the model parameters of the preset initial model are adjusted to obtain an updated preset initial model, and the step of inputting each second audio sample data and the second text information corresponding to the second audio sample data into the initial gating network and each initial feature extraction network to obtain the initial feature matrix output by the initial gating network and the initial feature data output by the initial feature extraction network is executed again, until the second target loss value is determined to be less than the preset loss threshold, at which point the current preset initial model is used as the pre-trained model.

[0141] It should be noted that the first specified loss value and the second specified loss value can be weighted and summed to obtain the second target loss value.

[0142] The above technical solution, by introducing a loss function that maximizes the pre-defined nuclear norm, effectively ensures the discriminativeness and diversity of the predicted categories by the gating network. This means that multiple dimensional features can be classified during training, thus avoiding wasted dimensions. Consequently, it effectively guarantees the diversity of the discrimination results of the pre-trained audio synthesis model trained based on this pre-trained model.

[0143] Figure 7This is a block diagram illustrating an audio generation apparatus according to an exemplary embodiment of the present disclosure; as shown below. Figure 7 As shown, the audio generation device may include:

[0144] The first acquisition module 701 is configured to acquire target text information;

[0145] The first determining module 702 is configured to input the target text information into a preset audio synthesis model to obtain the target audio data output by the preset audio synthesis model, wherein the target audio data is the audio data of the specified timbre corresponding to the target text information.

[0146] The preset audio synthesis model includes a gating network and multiple feature extraction networks. Different feature extraction networks are used to extract feature data of different dimensions. The gating network is used to determine the target feature extraction network from the multiple feature extraction networks. The target feature extraction network is used to determine the target audio data corresponding to the target text information.

[0147] The above technical solution, by determining at least one target feature extraction network from the multiple feature extraction networks, and determining the target audio data based on the target feature data output by the target feature extraction network, can effectively reduce the amount of computation required for the audio data generation process of a specified timbre while ensuring the fidelity of the audio data, improve the generation efficiency of audio data of a specified timbre, reduce user waiting time during synthesized speech, and thus effectively improve the user experience.

[0148] Optionally, the first determining module 702 is configured to:

[0149] The target text information is input into the gating network to obtain the feature matrix output by the gating network. Different elements in the feature matrix are used to represent the weights of features in different dimensions.

[0150] Based on the feature matrix, at least one target feature extraction network is determined from the plurality of feature extraction networks;

[0151] The target text information is input into the target feature extraction network to obtain the target feature data output by the target feature extraction network.

[0152] The target audio data is determined based on the target feature data.

[0153] Optionally, the first determining module 702 is configured to:

[0154] Based on the feature matrix, determine one or more target feature extraction networks corresponding to the preset number of dimensional features with the largest weights.

[0155] Figure 8 Based on this disclosure Figure 7The illustrated embodiment shows a block diagram of an audio generation apparatus; as shown Figure 8 As shown, the device also includes:

[0156] The second acquisition module 703 is configured to acquire multiple first audio sample data of the specified timbre, and the first text information corresponding to the first audio sample data;

[0157] The second determining module 704 is configured to use multiple first audio sample data and the first text information corresponding to each first audio sample data as training data to train a preset pre-trained model in order to obtain the preset audio synthesis model.

[0158] Optionally, the pre-trained model includes a gated network to be determined and multiple feature extraction networks to be determined, and the second determining module is configured as follows:

[0159] The first text information corresponding to each first audio sample data is input into the undetermined gating network to obtain the undetermined feature matrix output by the undetermined gating network. The elements in the undetermined feature matrix are used to characterize the undetermined weights of features of different dimensions.

[0160] The first undetermined loss value is determined based on the undetermined feature matrix using a loss function that maximizes the preset nuclear norm.

[0161] Based on the undetermined feature matrix, determine one or more specified feature extraction networks corresponding to a preset number of dimensional features with the largest undetermined weights;

[0162] The first text information is input into one or more designated feature extraction networks to obtain designated feature data output by each designated feature extraction network;

[0163] The current specified audio data is determined based on the specified feature data, and the second undetermined loss value of the first loss function is determined based on the specified audio data and the first audio sample data.

[0164] The first target loss value is determined based on the first undetermined loss value and the second undetermined loss value;

[0165] Determine whether the first target loss value is greater than the first preset loss threshold;

[0166] If the first target loss value is greater than the first preset loss threshold, the model parameters of the pre-trained model are adjusted to obtain an updated pre-trained model. The step of inputting the first text information corresponding to each first audio sample data into the undetermined gating network is executed again until the first target loss value is determined to be greater than the first preset loss threshold is determined. Then, if the first target loss value is determined to be greater than the first preset loss threshold, the current pre-trained model is used as the preset audio synthesis model.

[0167] Optionally, the device also includes a model training module 705, configured to:

[0168] Acquire second audio sample data with multiple different timbres, and the second text information corresponding to the second audio sample data;

[0169] The preset initial model is trained using the second audio sample data with multiple different timbres and the second text information corresponding to the second audio sample data as training data to obtain the pre-trained model.

[0170] Optionally, the preset initial model includes an initial gating network and multiple initial feature extraction networks, and the model training module 705 is configured as follows:

[0171] The second text information corresponding to each second audio sample data is input into the initial gating network and each initial feature extraction network to obtain the initial feature matrix output by the initial gating network and the initial feature data output by the initial feature extraction network. The elements in the initial feature matrix are used to characterize the initial weights of features of different dimensions.

[0172] Based on the initial weights, the initial feature data output by the multiple initial feature extraction networks are weighted and summed to obtain the initial audio data;

[0173] The initial model is iteratively trained based on the initial audio data and the preset loss function to obtain the pre-trained model.

[0174] Optionally, the preset loss function includes a first loss function and a second loss function, wherein the second loss function is a loss function that maximizes the preset nuclear norm, and the model training module 705 is configured as follows:

[0175] Based on the first loss function, the initial audio data and the second audio sample data determine a first specified loss value;

[0176] The second specified loss value is determined based on the initial feature matrix and the second loss function;

[0177] A second target loss value is determined based on the first specified loss value and the second specified loss value. If the second target loss value is greater than or equal to the second preset loss threshold, the model parameters of the preset initial model are adjusted to obtain an updated preset initial model. The step of inputting the second text information corresponding to each second audio sample data into the initial gating network and each initial feature extraction network to obtain the initial feature matrix output by the initial gating network and the initial feature data output by the initial feature extraction network is executed again until the second target loss value is determined to be less than the preset loss threshold. Then, the current preset initial model is used as the pre-trained model.

[0178] The above technical solution, by introducing a loss function that maximizes the pre-defined nuclear norm, effectively ensures the discriminativeness and diversity of the predicted categories by the gating network. This means that multiple dimensional features can be classified during training, thus avoiding wasted dimensions. Consequently, it effectively guarantees the diversity of the discrimination results of the pre-trained audio synthesis model.

[0179] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0180] Figure 9 This is a block diagram illustrating an electronic device according to an exemplary embodiment. For example, device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness device, personal digital assistant, etc.

[0181] Reference Figure 9 The device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output interface 812, a sensor component 814, and a communication component 816.

[0182] Processing component 802 typically controls the overall operation of device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the audio generation method described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.

[0183] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of such data include instructions for any application or method operating on device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0184] Power supply component 806 provides power to various components of device 800. Power supply component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to device 800.

[0185] Multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensor may sense not only the boundaries of a touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0186] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.

[0187] Input / output interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0188] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of device 800. For example, sensor assembly 814 may detect the on / off state of device 800, the relative positioning of components such as the display and keypad of device 800, changes in the position of device 800 or a component of device 800, the presence or absence of user contact with device 800, the orientation or acceleration / deceleration of device 800, and temperature changes of device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0189] Communication component 816 is configured to facilitate wired or wireless communication between device 800 and other devices. Device 800 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0190] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the audio generation method described above.

[0191] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of the device 800 to complete the audio generation method described above. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0192] The aforementioned device can be a standalone electronic device or a part of a standalone electronic device. For example, in one embodiment, the device can be an integrated circuit (IC) or a chip, wherein the integrated circuit can be a single IC or a collection of multiple ICs. The chip can include, but is not limited to, the following types: GPU (Graphics Processing Unit), CPU (Central Processing Unit), FPGA (Field Programmable Gate Array), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), and SoC (System on Chip). The aforementioned integrated circuit or chip can be used to execute executable instructions (or code) to implement the aforementioned audio generation method. The executable instructions can be stored in the integrated circuit or chip or obtained from other devices or equipment. For example, the integrated circuit or chip includes a processor, memory, and an interface for communicating with other devices. The executable instructions can be stored in the memory, and when the executable instructions are executed by the processor, the above-described audio generation method is implemented; or, the integrated circuit or chip can receive the executable instructions through the interface and transmit them to the processor for execution to implement the above-described audio generation method.

[0193] In another exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program executable by a programmable device, the computer program having a code portion for performing the above-described audio generation method when executed by the programmable device.

[0194] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of this disclosure. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0195] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. An audio generation method, characterized in that, include: Obtain target text information; The target text information is input into a preset audio synthesis model to obtain the target audio data output by the preset audio synthesis model. The target audio data is the audio data of the specified timbre corresponding to the target text information. The preset audio synthesis model includes a gating network and multiple feature extraction networks. Different feature extraction networks are used to extract feature data of different dimensions. The gating network is used to determine the target feature extraction network from the multiple feature extraction networks. The target feature extraction network is used to determine the target audio data corresponding to the target text information. The gated network is used to output the influence weight of each feature extraction network's corresponding dimension features on the generation of target audio data, and to filter out the dimension with greater influence based on the influence weight, and to use the feature extraction network corresponding to the dimension with greater influence as the target feature extraction network. The preset audio synthesis model can be trained in the following way: Acquire first audio sample data of multiple specified timbres, and first text information corresponding to the first audio sample data; Using multiple first audio sample data and the first text information corresponding to each first audio sample data as training data, a preset pre-trained model is trained to obtain the preset audio synthesis model. The pre-trained model is trained in the following way: Acquire second audio sample data with multiple different timbres, and second text information corresponding to the second audio sample data; The preset initial model is trained using the second audio sample data with multiple different timbres and the second text information corresponding to the second audio sample data as training data to obtain the pre-trained model.

2. The method according to claim 1, characterized in that, The step of inputting the target text information into a preset audio synthesis model to obtain the target audio data output by the preset audio synthesis model includes: The target text information is input into the gating network to obtain the feature matrix output by the gating network. Different elements in the feature matrix are used to represent the weights of features of different dimensions. At least one target feature extraction network is determined from the plurality of feature extraction networks based on the feature matrix; The target text information is input into the target feature extraction network to obtain the target feature data output by the target feature extraction network. The target audio data is determined based on the target feature data.

3. The method according to claim 2, characterized in that, The step of determining at least one target feature extraction network from the plurality of feature extraction networks based on the feature matrix includes: Based on the feature matrix, determine one or more target feature extraction networks corresponding to a preset number of dimensional features with the largest weights.

4. The method according to claim 1, characterized in that, The pre-trained model includes a gated network to be determined and multiple feature extraction networks to be determined. The process of training the pre-trained model using multiple first audio sample data and the first text information corresponding to each first audio sample data as training data to obtain the pre-trained audio synthesis model includes: The first text information corresponding to each first audio sample data is input into the undetermined gating network to obtain the undetermined feature matrix output by the undetermined gating network. The elements in the undetermined feature matrix are used to characterize the undetermined weights of features of different dimensions. The first undetermined loss value is determined based on the undetermined feature matrix using a loss function that maximizes the preset nuclear norm. Based on the undetermined feature matrix, determine one or more designated feature extraction networks corresponding to a preset number of dimensions with the largest undetermined weight, and input the first text information into the one or more designated feature extraction networks to obtain the designated feature data output by each designated feature extraction network; The specified audio data is determined based on the specified feature data, and the second undetermined loss value of the first loss function is determined based on the specified audio data and the first audio sample data. The first target loss value is determined based on the first undetermined loss value and the second undetermined loss value; Determine whether the first target loss value is greater than the first preset loss threshold; If the first target loss value is greater than the first preset loss threshold, the model parameters of the pre-trained model are adjusted to obtain an updated pre-trained model. The step of inputting the first text information corresponding to each first audio sample data into the undetermined gating network is executed again until the first target loss value is determined to be greater than the first preset loss threshold is determined. Then, if the first target loss value is determined to be greater than the first preset loss threshold, the current pre-trained model is used as the preset audio synthesis model.

5. The method according to claim 1, characterized in that, The preset initial model includes an initial gating network and multiple initial feature extraction networks. The preset initial model is trained using the multiple second audio sample data with different timbres and the corresponding second text information as training data to obtain the pre-trained model, including: The second text information corresponding to each second audio sample data is input into the initial gating network and each initial feature extraction network to obtain the initial feature matrix output by the initial gating network and the initial feature data output by the initial feature extraction network. The elements in the initial feature matrix are used to characterize the initial weights of features of different dimensions. The initial feature data output by the multiple initial feature extraction networks are weighted and summed according to the initial weights to obtain the initial audio data; The preset initial model is iteratively trained based on the initial audio data and the preset loss function to obtain the pre-trained model.

6. The method according to claim 5, characterized in that, The preset loss function includes a first loss function and a second loss function, wherein the second loss function is a loss function that maximizes a preset nuclear norm. The step of iteratively training the preset initial model based on the initial audio data and the preset loss function to obtain the pre-trained model includes: Based on the first loss function, a first specified loss value is determined from the initial audio data and the second audio sample data; The second specified loss value is determined based on the initial feature matrix and the second loss function; Determine the second target loss value based on the first specified loss value and the second specified loss value; If the second target loss value is determined to be greater than or equal to the second preset loss threshold, the model parameters of the preset initial model are adjusted to obtain an updated preset initial model. Then, the step of inputting the second text information corresponding to each second audio sample data into the initial gating network and each initial feature extraction network to obtain the initial feature matrix output by the initial gating network and the initial feature data output by the initial feature extraction network is executed again until the second target loss value is determined to be less than the second preset loss threshold. Then, the current preset initial model is used as the pre-trained model.

7. An audio generation apparatus, characterized in that, include: The first acquisition module is configured to acquire target text information; The first determining module is configured to input the target text information into a preset audio synthesis model to obtain target audio data output by the preset audio synthesis model, wherein the target audio data is audio data with a specified timbre corresponding to the target text information. The preset audio synthesis model includes a gating network and multiple feature extraction networks. Different feature extraction networks are used to extract feature data of different dimensions. The gating network is used to determine the target feature extraction network from the multiple feature extraction networks. The target feature extraction network is used to determine the target audio data corresponding to the target text information. The gated network is used to output the influence weight of each feature extraction network's corresponding dimension features on the generation of target audio data, and to filter out the dimension with greater influence based on the influence weight, and to use the feature extraction network corresponding to the dimension with greater influence as the target feature extraction network. The preset audio synthesis model can be trained in the following way: Acquire first audio sample data of multiple specified timbres, and first text information corresponding to the first audio sample data; Using multiple first audio sample data and the first text information corresponding to each first audio sample data as training data, a preset pre-trained model is trained to obtain the preset audio synthesis model. The pre-trained model is trained in the following way: Acquire second audio sample data with multiple different timbres, and second text information corresponding to the second audio sample data; The preset initial model is trained using the second audio sample data with multiple different timbres and the second text information corresponding to the second audio sample data as training data to obtain the pre-trained model.

8. An audio generation device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured as follows: The steps of implementing the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When executed by a processor, the program instructions implement the steps of the method described in any one of claims 1 to 6.

10. A chip, characterized in that, It includes a processor and an interface; the processor is used to read instructions to execute the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Speech synthesis model training method and device, equipment and storage medium

    CN114758645A