Speech synthesis method, speech synthesis model training method and related device

Through Q-former, the voice characteristics of the speaker are extracted and reinforcement learning is combined to generate a speech synthesis model that meets user preferences, solving the pronunciation errors and weird rhythm problems of large-scale synthesized speech, and achieving more natural and accurate speech synthesis.

CN120356452AActive Publication Date: 2025-07-22IFLYTEK CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510848810.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

The existing speech synthesis technology based on large models has problems such as pronunciation errors, weird pronunciation, and slow speech speed, and does not meet the user's timbre and rhythm perception preferences.

Method used

Q-former is used to extract the voice characteristics of the speaker, and the target speech synthesis model is generated through reinforcement learning, and the model parameters are adjusted according to the user's preference for the synthesized speech to ensure that the synthesized speech is in line with the voice characteristics and user preferences of the speaker.

Benefits of technology

The quality of speech synthesis is improved, making the synthesized speech more natural and accurate, conforming to the speaker's voice characteristics and in line with the user's tone and rhythm preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356452A_ABST
    Figure CN120356452A_ABST
Patent Text Reader

Abstract

The invention discloses a speech synthesis method, a speech synthesis model training method and a related device. The method comprises the following steps: acquiring a target text and a target speech to be synthesized; extracting voice features of a speaker in the target voice; based on the voice features of the speaker and the target text, generating a target synthetic voice, the target synthetic voice being a voice having voice features of a reference target voice and having a pronunciation content consistent with the target text; wherein the sound features are extracted by using a Q-former; and / or the target synthetic speech is generated by using a target speech synthesis model, the target speech synthesis model is obtained by performing reinforcement learning on the initial speech synthesis model according to the preference degree of the user for the first sample synthetic speech, and the first sample synthetic speech is generated by using the initial speech synthesis model. In this way, the quality of the target synthetic speech can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a voice synthesis method, a voice synthesis model training method, and related devices. Background Art

[0002] Voice Clone is a very popular direction in the field of voice synthesis in recent years. It aims to construct a voice synthesis system of the recorder by recording a voice segment (usually 5 - 15s) of the target speaker through some common recording devices in life, such as mobile phones, computers, voice recorders, etc.

[0003] Currently, the voice cloning technology based on large models is gradually becoming mainstream, but there are still some problems with the voices synthesized by large models, such as pronunciation errors, strange rhythms, or sudden changes in speech speed. How to solve these problems of synthesized voices is particularly important. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a voice synthesis method, a voice synthesis model training method, and related devices, which can improve the quality of the target synthesized voice.

[0005] To solve the above technical problem, a technical solution adopted by this application is: to provide a voice synthesis method, which includes: obtaining a target text and a target voice to be synthesized; extracting the voice features of the speaker in the target voice; generating a target synthesized voice based on the voice features of the speaker and the target text, where the target synthesized voice is a voice with the same pronunciation content as the target text and referring to the voice characteristics of the target voice; wherein, the voice features are extracted using Q-former; and / or, the target synthesized voice is generated using a target voice synthesis model, and the target voice synthesis model is obtained by performing reinforcement learning on an initial voice synthesis model according to the user's preference degree for the first sample synthesized voice, and the first sample synthesized voice is generated using the initial voice synthesis model.

[0006] To solve the above technical problem, another technical solution adopted by this application is: to provide a training method for a target voice synthesis model, which includes: obtaining a model to be optimized for this reinforcement learning; wherein, the model for the first reinforcement learning is an initial voice synthesis model, and the model for non-first reinforcement learning is the model obtained from the previous reinforcement learning; obtaining each third sample synthesized voice generated by the model to be optimized this time; using the user's preference degree for each third sample synthesized voice to adjust the network parameters of the model to be optimized this time, and repeating the above steps to obtain a target voice synthesis model; wherein, the adjustment direction of the model to be optimized includes: increasing the generation probability of the third sample synthesized voice with a high user preference degree.

[0007] To solve the above technical problems, another technical solution adopted by this application is: to provide a voice synthesis device, including a first acquisition module, an extraction module, and a voice generation module. The first acquisition module is used to acquire the target text and the target voice to be synthesized; the extraction module is used to extract the voice features of the speaker in the target voice; the voice generation module is used to generate a target synthesized voice based on the voice features of the speaker and the target text, and the target synthesized voice is a voice with the same pronunciation content as the target text and the same voice characteristics as the reference target voice; wherein, the voice features are extracted by using Q-former; and / or, the target synthesized voice is generated by using a target voice synthesis model, and the target voice synthesis model is obtained by performing reinforcement learning on the initial voice synthesis model according to the user's preference degree for the first sample synthesized voice, and the first sample synthesized voice is generated by using the initial voice synthesis model.

[0008] To solve the above technical problems, another technical solution adopted by this application is: to provide a target voice synthesis model training device, including: a second acquisition module, a third acquisition module, and a parameter adjustment module. The second acquisition module is used to acquire the model to be optimized for the current reinforcement learning; wherein, the model for the first reinforcement learning is the initial voice synthesis model, and the model for non-first reinforcement learning is the model obtained from the previous reinforcement learning; the third acquisition module is used to acquire each third sample synthesized voice generated by the model to be optimized for the current time; the parameter adjustment module is used to adjust the network parameters of the model to be optimized for the current time by using the user's preference degree for each third sample synthesized voice, and repeat the above steps to obtain the target voice synthesis model; wherein, the adjustment direction of the model to be optimized includes: increasing the generation probability of the third sample synthesized voice with a high user preference degree.

[0009] To solve the above technical problems, another technical solution adopted by this application is: to provide an electronic device, including a memory and a processor coupled to each other, and the memory stores program instructions; the processor is used to execute the program instructions stored in the memory to implement the above method.

[0010] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer-readable storage medium for storing program instructions, and the program instructions can be executed to implement the above method.

[0011] In the above solution, the target synthetic speech is generated based on the voice characteristics of the speaker and the target text. Among them, the voice characteristics are extracted by using Q-former; and / or, the target synthetic speech is generated by using a target speech synthesis model, and the target speech synthesis model is obtained by performing reinforcement learning on an initial speech synthesis model according to the user's preference degree for the first sample synthetic speech. Among them, the method of extracting the speaker's voice characteristics by using Q-former can extract the voice characteristics important for the generation of the target synthetic speech. Furthermore, by using the voice characteristics extracted by Q-former for speech synthesis, the synthesized speech can be made more in line with the speaker's voice characteristics, making the synthesized speech more natural and accurate. In addition, since the target speech synthesis model can be obtained by performing reinforcement learning on the initial speech synthesis model according to the user's preference degree for the synthetic speech, the probability of the target synthetic speech generated by the target speech synthesis model obtained through reinforcement learning conforming to the user's preference is relatively high. Furthermore, it is beneficial to generate the target synthetic speech that conforms to the user's preference. Compared with the method of directly using a speech synthesis model to generate the target synthetic speech based on the input target text and target voice, both the method of using Q-former to extract the speaker's voice characteristics and the method of using the target speech synthesis model obtained by performing reinforcement learning according to the user's preference for the voice in this application can improve the quality of the target synthetic speech. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a schematic flowchart of an embodiment of the speech synthesis method provided by this application; Figure 2 is a schematic structural diagram of an embodiment of the Q-former network provided by this application; Figure 3 is a schematic flowchart of an embodiment of obtaining the target speech synthesis model through reinforcement learning provided by this application; Figure 4 is a schematic flowchart of another embodiment of obtaining the target speech synthesis model through reinforcement learning provided by this application; Figure 5 is a schematic flowchart of an embodiment of the target speech synthesis model training method provided by this application; Figure 6 is a schematic framework diagram of an embodiment of the speech synthesis device provided by this application; Figure 7 is a schematic framework diagram of an embodiment of the target speech synthesis model training device provided by this application; Figure 8 is a schematic framework diagram of an embodiment of the electronic device provided by this application; Figure 9 is a schematic framework diagram of the computer-readable storage medium provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] To make the objectives, technical solutions and effects of this application clearer and more definite, the following further elaborates on this application with reference to the accompanying drawings and by way of examples.

[0014] In addition, if descriptions such as "first" and "second" are involved in the embodiments of this application, these descriptions of "first", "second", etc. are for descriptive purposes only, and should not be construed as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments may be combined with each other, but it must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions results in contradictions or cannot be achieved, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0015] Currently, voice cloning technology based on large models is gradually becoming mainstream, but there are still the following problems: (1) Since large models have no explicit duration constraints and only align speech tokens and text through the attention mechanism inside the large model, the synthesized speech often has problems such as pronunciation errors, strange rhythms, and uneven speech speeds; (2) When large models perform speech synthesis, they model the temporal correlation of speech, and there is diversity in sampling. Sometimes the generated synthesized speech does not conform to the user's perceptual preferences for timbre, rhythm, etc.

[0016] To solve the above problems, this application proposes a speech synthesis method, a speech synthesis model training method, and related devices. Specifically: Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the speech synthesis method provided by this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 1 the shown process sequence. As Figure 1 shown, this embodiment includes: S11: Obtain the target text and target speech to be synthesized.

[0017] S12: Extract the voice features of the speaker in the target speech.

[0018] S13: Generate the target synthesized speech based on the voice features of the speaker and the target text, where the target synthesized speech is a speech that references the voice characteristics of the target speech and has pronunciation content consistent with the target text.

[0019] This embodiment is used to, after obtaining the target speech of the speaker and the target text to be synthesized, convert the target text into speech with the voice characteristics of the speaker by referring to the voice characteristics of the speaker, that is, the target synthesized speech described in the text.

[0020] Among them, the target voice is an audio clip (e.g., 5 - 15 seconds) of a speaker recorded by any electronic device with a recording function. The target text is the text to be converted into the corresponding voice, and the synthesized target synthesized voice is a voice that references the voice characteristics of the target voice and has the same pronunciation content as the target text.

[0021] For example, the target text is "Today the weather is sunny and cloudless", the target voice is an audio clip spoken by the speaker, and the target synthesized voice is a voice that references the voice characteristics in the speaker's audio clip and has the pronunciation content of "Today the weather is sunny and cloudless".

[0022] The reference voice characteristics in this embodiment can be at least some of all the characteristics of the voice. For example, the reference can be all the characteristics of the voice, or the voice characteristics of a part of all the dimensions of the voice. Exemplarily, the reference voice characteristics are, for example, at least one or a combination of at least two of timbre, intonation, speech rate, rhythm, and emotion, etc.

[0023] In an alternative embodiment, the voice features are extracted by using an audio feature extraction network, where the audio feature extraction network is composed of several attention modules, and each attention module includes a self - attention sub - module and a cross - attention sub - module.

[0024] The input data of the self - attention sub - module are several learnable query vectors, and different learnable query vectors are used to extract different - dimensional voice sub - features of the target voice; the self - attention sub - module is used to perform self - attention processing on the input several learnable query vectors to obtain the fusion result of the several learnable query vectors.

[0025] The input data of the cross - attention sub - module include the fusion result output by the self - attention sub - module and the audio features of the target voice. Among them, the cross - attention sub - module is used to respectively use the audio features of the target voice as the key vector and the value vector, use the fusion result as the query vector, and perform cross - attention processing based on the query vector, the key vector, and the value vector to extract the voice features.

[0026] Among them, the audio feature extraction network can be, but is not limited to, a Q - former network, or can also be any other network composed of several attention modules.

[0027] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of an embodiment of the Q - former network provided by this application. To facilitate understanding of the process of extracting voice features by the above - mentioned audio feature extraction network, the process of extracting voice features is briefly described below taking the audio feature extraction network as Q - former as an example.

[0028] First, the Q-former network consists of N attention modules, and each attention module includes a self-attention sub-module and a cross-attention sub-module. Among them, the input data of the self-attention sub-module is a set of discrete learnable query vectors. This set of learnable query vectors includes several different learnable query vectors, and different learnable query vectors are used to extract voice sub-features of different dimensions of the target speech. For example, the learnable query vector Q1 is used to extract the timbre feature in the target speech, the learnable query vector Q2 is used to extract the prosody feature in the target speech, and the learnable query vector Q3 is used to extract the emotional characteristics in the target speech, etc.

[0029] Among them, what features each specific learnable query vector is used to extract is obtained through learning. In other words, the learnable query vector is learnable and is not a fixed value. During the learning process, through backpropagation optimization, it can automatically learn the features important for the generation of the synthesized speech (such as timbre, intonation, or emotion, etc.). Compared with the method of statically extracting features using a pre-trained network, the method of using Q-former to extract voice features in this embodiment can dynamically learn the features important for speech synthesis. Therefore, using the voice features extracted by Q-former for subsequent speech synthesis is beneficial to generating speech that is more in line with the voice characteristics of the speaker and has natural prosody.

[0030] Furthermore, the self-attention sub-module in Q-former is used to perform self-attention interaction on several learnable query vectors to obtain the interaction result of the learnable query vectors, and then use this interaction result as the query vector Q. The audio features of the target speech are used as the value vector V and the key vector K respectively, and cross-attention processing is performed with the audio features of the target speech through the cross-attention sub-module to query the voice characteristics related to each learnable vector from the target speech, and then fuse the queried voice characteristics. Among them, during the above attention processing, the attention mechanism can focus on local key features, which helps to retain more details. Therefore, the above method of using Q-former to extract features can extract richer and more detailed voice characteristics of the speaker from the target speech. Among them, in one embodiment, the audio features of the target speech are obtained by performing voice discretization processing on the target speech to obtain the corresponding speech tokens and then encoding the speech tokens. Of course, in other embodiments, the audio features of the target speech can also be obtained through other methods.

[0031] It should be noted that since Q-former can extract more abundant sound features of the speaker and important sound features for synthesized speech generation from the target speech, the sound features extracted by Q-former can make the generated target synthesized speech features more consistent with the speaker's sound features, which is conducive to generating target synthesized speech that is highly similar to the speaker's sound features.

[0032] At the same time, it should be noted that the method of using Q-former to extract the voice features of the target speaker provided in this embodiment can automatically learn the most important features for synthesized speech generation, so it is conducive to generating target synthesized speech that is highly similar to the speaker's voice features, and can effectively solve the problems of pronunciation errors, strange rhythms, and fast and slow speech speed in synthesized speech caused by the lack of explicit duration constraints in large models, thereby making the generated target synthesized speech more natural and more in line with the speaker's voice characteristics.

[0033] The specific number of learnable query vectors and the data dimension of each learnable query vector can be determined according to the effect of the final synthesized speech. On the basis of meeting the expected effect, the number of learnable query vectors and / or the data dimension of each learnable query vector can be appropriately reduced to avoid introducing too much redundant or even erroneous information, which is very helpful for the naturalness, rhythmic timbre similarity or stability of the synthesized speech generated by the target speech synthesis model.

[0034] In another optional implementation, the target synthesized speech is generated using a target speech synthesis model, wherein the target speech synthesis model is obtained by performing reinforcement learning on an initial speech synthesis model according to the user's preference for the first sample synthesized speech, and the first sample synthesized speech is generated using the initial speech synthesis model.

[0035] Among them, the initial speech synthesis model must undergo several rounds of reinforcement learning before the target speech synthesis model can be learned. Each round of reinforcement learning involves the adjustment of the model network parameters, and the adjustment of each network parameter is essentially based on the model obtained in the previous reinforcement learning.

[0036] For example, the initial speech synthesis model needs to go through n rounds of reinforcement learning to obtain the target speech synthesis model, where the first model to be reinforced is Q0 (i.e. the initial speech synthesis model), Q1 is obtained after the first reinforcement learning, Q1 is obtained after the second reinforcement learning, Q2 is obtained after the third reinforcement learning, and Q2 is obtained after the third reinforcement learning. The above steps are repeated until Q n-1 After the nth reinforcement learning, we get Q n (Target Speech Synthesis Model).

[0037] Among them, the first sample synthetic speech is generated using the initial speech synthesis model. Specifically, the first sample synthetic speech is generated using the models to be reinforced in each round. In this embodiment, the first sample synthetic speech includes a number of second sample synthetic speeches generated by the models to be reinforced in each round.

[0038] In this embodiment, the process of performing reinforcement learning on the initial synthesis model according to the user's preference degree for the first sample synthetic speech to obtain the target speech synthesis model is essentially a process of adjusting the parameters of the model to be reinforced in the corresponding round according to the user's preference degree for the second sample synthetic speeches generated by the models to be reinforced in each round.

[0039] Continuing with the above example, for the first model Q0 to be reinforced, the network parameters of Q0 are adjusted using the user's preference degree for a number of second sample synthetic speeches generated by Q0 to obtain Q1; for the second model Q1 to be reinforced, the network parameters of Q1 are adjusted using the user's preference degree for a number of second sample synthetic speeches generated by Q1 to obtain Q2, and the above operations are repeated until the target speech synthesis model is obtained.

[0040] It should be noted that performing reinforcement learning on the initial synthesis model according to the user's preference degree for the first sample synthetic speech to obtain the target speech synthesis model can make the target synthetic speech generated by the target speech synthesis model more in line with the user's preference for the target synthetic speech. The specific method of performing reinforcement learning on the initial speech synthesis model according to the user's preference degree for the first sample synthetic speech to obtain the target speech synthesis model can refer to the relevant descriptions of the embodiments shown below Figure 3 and Figure 4 the following embodiments.

[0041] In one embodiment, the user's preference degree for the first sample synthetic speech includes at least one of the following preference sub - degrees: prosody naturalness, timbre similarity, and intelligibility. That is, in the process of reinforcement learning, the network parameters can be adjusted based on at least one of the prosody naturalness, timbre similarity, and intelligibility of the first sample synthetic speech, so that the target synthetic speech finally obtained by reinforcement learning can generate a target synthetic speech that is more in line with the user's preference in at least one aspect of prosody naturalness, timbre similarity, and intelligibility.

[0042] Preferably, the preference degree of the first sample synthetic speech can be set to include preference sub - degrees in aspects such as prosody naturalness, timbre similarity, and intelligibility, so that the target synthetic speech generated by the target synthetic speech finally obtained by reinforcement learning meets the user's preference in aspects such as prosody naturalness, timbre similarity, and intelligibility.

[0043] Among them, the preference degree of the user for the first sample synthesized speech is obtained through manual evaluation and / or automated evaluation tools. It should be noted that whether the result is obtained through manual evaluation or by an automated evaluation tool, it is carried out in combination with the user's preference and can characterize the user's preference for the corresponding synthesized speech.

[0044] In an alternative embodiment, the preference degree of the user for the first sample synthesized speech is obtained by the user's annotation based on the preference degree for the first sample synthesized speech after using the initial speech synthesis model to generate the corresponding first sample synthesized speech. For example, the user scores or ranks the first sample synthesized speech in terms of prosody naturalness, timbre similarity, intelligibility, etc.

[0045] In another alternative embodiment, for some objectively measurable data, an automated evaluation tool can be used to assist in scoring to reduce the user's annotation burden.

[0046] In one embodiment, the preference degree of the first sample synthesized speech includes preference sub-degrees for at least one of the following aspects of the first sample synthesized speech: prosody naturalness, timbre similarity, and intelligibility; the automated evaluation tool includes at least one of the following: a voiceprint verification system, a speech recognition system, and a trained reward model.

[0047] Among them, the timbre similarity is used to quantify the degree of closeness in timbre between the first sample synthesized speech and the target speech. In this embodiment, a voiceprint verification system can be used to determine the timbre similarity between the first sample synthesized speech and the target speech, and then the preference degree of the user for the first sample synthesized speech in terms of timbre can be determined based on the timbre similarity. It can be understood that when the user scores the quality of the first sample synthesized speech in terms of timbre, the target speech of the speaker is referred to. Among them, the more similar the timbre of the first sample synthesized speech is to the target speech, the easier it is for the user to give a higher score. Therefore, in this embodiment, the timbre similarity and the preference sub-degree in terms of timbre are set to be positively correlated.

[0048] In one embodiment, a speech recognition system can be used to determine the intelligibility of the first sample synthesized speech. Among them, the intelligibility of speech is used to characterize the degree to which the speech content is correctly recognized by the listener. For example, its quantification method is that if the listener hears 100 words and correctly recognizes 75 of them, the intelligibility is 75%. In an alternative embodiment, the intelligibility can be determined by calculating the WER index, etc. using a speech recognition system (ASR system).

[0049] It can be understood that when the user scores the quality of the first sample synthesized speech in terms of intelligibility, the higher the intelligibility of the first sample synthesized speech, the easier it is for the user to give a higher score. Therefore, in this embodiment, the intelligibility and the preference sub-degree in terms of intelligibility are set to be positively correlated.

[0050] It should be noted that prosodic naturalness is a subjective feeling, and it is difficult to objectively measure prosodic naturalness. Therefore, the preference degree of users for prosodic naturalness can be obtained by means of manual evaluation.

[0051] In some embodiments, a reward model can also be trained in advance using the preference degrees of users for the synthesized speech output by the initial speech synthesis model in various aspects, so that the trained reward model can output a reward value for characterizing the preference degree of users for the first sample synthesized speech.

[0052] Exemplarily, the sample reference speech input into the initial speech synthesis model and the sample synthesized speech output by the initial speech synthesis model are used as the input data of the reward model. The reward model outputs a corresponding reward value based on the input data, and then the cross-entropy loss function is used to train the reward model so that the trained reward model can simulate the preference degree of users for the first sample synthesized speech. Among them, the cross-entropy loss function is expressed as follows:

[0053] In the formula, x: the input sample reference speech, y0 and y1 represent two different sample synthesized speeches generated by the model, i represents the preference label marked by the user (i = 0 means the user thinks y0 is better than y1, and vice versa), The scoring (reward value) of the reward model for the input x and the output yi σ represents the Sigmoid function, which maps the reward difference to the probability interval (0, 1), ( ) represents the preference pair marked by the user, and the annotator will mark which output is better, for example, y0 is better than y1, is the reward difference, that is, the reward difference between the preferred output and the non-preferred output.

[0054] The essence of this loss function is to force the scoring of the reward model to be consistent with the user's preference through cross-entropy loss, so as to provide a reliable reward signal for subsequent reinforcement learning (such as PPO).

[0055] Generally speaking, the initial speech synthesis model is a large model, and the large model can sample and generate different sample synthesized speeches based on the same sample reference speech. Among them, in the process of training the reward model, preference data pairs are first constructed; among them, the preference data pairs include a sample synthesized speech with a higher preference degree (the above-mentioned preferred output) and a sample synthesized speech with a lower preference degree (the above-mentioned non-preferred output); then the constructed preference data pairs are used for the above loss calculation, and then the parameters of the reward model are adjusted according to the loss so that the reward model can output a reward value for characterizing the preference degree of users for the first sample synthesized speech.

[0056] In the above solution, the target synthesized speech is generated based on the voice characteristics of the speaker and the target text. Among them, the voice characteristics are extracted using Q-former; and / or, the target synthesized speech is generated using a target speech synthesis model, and the target speech synthesis model is obtained by performing reinforcement learning on the initial speech synthesis model according to the user's preference degree for the first sample synthesized speech. Among them, the method of using Q-former to extract the speaker's voice characteristics can extract the voice characteristics important for the generation of the target synthesized speech. Furthermore, using the voice characteristics extracted by Q-former for speech synthesis can make the synthesized speech more conform to the speaker's voice characteristics, making the synthesized speech more natural and accurate. In addition, since the target speech synthesis model can be obtained by performing reinforcement learning on the initial speech synthesis model according to the user's preference degree for the synthesized speech, the probability that the target speech synthesis model obtained by reinforcement learning generates the target synthesized speech that conforms to the user's preference is relatively high. Furthermore, it is beneficial to generate the target synthesized speech that conforms to the user's preference. Compared with the method of directly using a speech synthesis model to generate the target synthesized speech based on the input target text and target voice, both the method of using Q-former to extract the speaker's voice characteristics and the method of using the target speech synthesis model obtained by performing reinforcement learning according to the user's preference for the voice in this application can improve the quality of the target synthesized speech.

[0057] In one embodiment, a reinforcement learning algorithm similar to DPO can be used to train the initial speech synthesis model.

[0058] Specifically, please refer to Figure 3 , Figure 3 which is a schematic flowchart of an embodiment of obtaining a target speech synthesis model through reinforcement learning provided by this application. In this embodiment, performing reinforcement learning on the initial speech synthesis model according to the user's preference degree for the first sample synthesized speech to obtain the target speech synthesis model includes: S31: Using the model to be reinforced and learned this time as the model to be optimized, where the model to be reinforced and learned for the first time is the initial speech synthesis model, and the model to be reinforced and learned for non-first time is the model obtained by the previous reinforcement learning; the first sample synthesized speech includes several second sample synthesized speeches generated by the models to be optimized for each time of reinforcement learning.

[0059] As can be seen from the foregoing, the initial speech synthesis model needs to undergo several rounds of reinforcement learning to learn the target speech synthesis model. And when performing each round of reinforcement learning, it involves adjusting the model network parameters. Among them, each adjustment of the network parameters is essentially based on the model obtained from the previous round of reinforcement learning. For example, the initial speech synthesis model needs to undergo n rounds of reinforcement learning to obtain the target speech synthesis model. Among them, the model to be reinforced for the first time is Q0 (i.e., the initial speech synthesis model), Q1 is obtained after the first round of reinforcement learning, Q2 is obtained after Q1 undergoes the second round of reinforcement learning, Q3 is obtained after Q2 undergoes the third round of reinforcement learning, and the above steps are repeated until Q n-1 Undergoes the nth round of reinforcement learning to obtain Q n (the target speech synthesis model).

[0060] In this embodiment, the process of performing reinforcement learning on the initial synthesis model according to the user's preference degree for the first sample synthesized speech to obtain the target speech synthesis model is essentially a process of adjusting the parameters of the model to be reinforced in the corresponding round by using the user's preference degree for the second sample synthesized speech generated by the model to be reinforced in each round.

[0061] For the model to be reinforced in this (this round) round, first, use the model to be reinforced in this round as the model to be optimized in this round. Then, obtain several second sample synthesized speeches generated by the model to be optimized in this round based on the sample reference speech, and obtain the user's preference degree for each second sample synthesized speech in this round. The way to obtain the preference degree can refer to the description in the foregoing, and will not be elaborated here.

[0062] S32: Based on the user's preference degree for each second sample synthesized speech, construct at least one positive and negative speech pair; the preference degree of the positive second sample synthesized speech in the positive and negative speech pair is higher than that of the negative second sample synthesized speech.

[0063] Each of the second sample synthesized speeches described in step S32 is the second sample synthesized speech generated by the model to be optimized in this round. Among them, step S32 is to construct at least one positive and negative speech pair, and each positive and negative speech pair includes a positive second sample synthesized speech (speech positive example) and a negative second sample synthesized speech (speech negative example), and the preference degree of the speech positive example in the positive and negative speech pair is higher than that of the speech negative example.

[0064] S33: Obtain the first generation probability of the model to be optimized for the positive second sample synthesized speech and the second generation probability of the model to be optimized for the negative second sample synthesized speech.

[0065] S34: Based on the first generation probability and the corresponding second generation probability, adjust the network parameters of the model to be optimized, and repeat the above steps to obtain the target speech synthesis model. The adjustment direction of the network parameters of the model to be optimized is: the direction of increasing the first generation probability of the synthesized speech of the positive second sample and decreasing the second generation probability of the synthesized speech of the negative second sample.

[0066] It can be understood that the purpose of reinforcement learning training is to expect that the trained model can output synthesized speech that conforms to user preferences with a relatively high probability. Therefore, during the reinforcement learning training process, after obtaining the first generation probability of the synthesized speech of the positive second sample by the model to be optimized and the second generation probability of the synthesized speech of the negative second sample by the model to be optimized, adjust the network parameters of the model to be optimized in the direction of increasing the first generation probability of the synthesized speech of the positive second sample and decreasing the second generation probability of the synthesized speech of the negative second sample, so that the model after parameter adjustment can increase the first generation probability of the synthesized speech of the positive second sample (positive example speech) and decrease the second generation probability of the synthesized speech of the negative second sample (negative example speech).

[0067] In an embodiment, the difference between the first generation probability of the synthesized speech of the positive second sample and the second generation probability of the synthesized speech of the negative second sample can be used as the objective function, and the network parameters of the model to be optimized are adjusted in the direction of maximizing this objective function.

[0068] In another embodiment, the following method can also be used to adjust the network parameters of the model to be optimized. It specifically includes the following steps: Step 1: Obtain the third generation probability of the synthesized speech of the positive second sample generated by the initial speech synthesis model, and obtain the fourth generation probability of the synthesized speech of the negative second sample generated by the initial speech synthesis model.

[0069] Step 2: Obtain the first ratio between the first generation probability and the third generation probability of the synthesized speech of the positive second sample, and obtain the second ratio between the second generation probability and the fourth generation probability of the synthesized speech of the negative second sample.

[0070] Step 3: Based on the difference between the first ratio of the synthesized speech of the positive second sample and the corresponding second ratio of the synthesized speech of the negative second sample, adjust the network parameters of the model to be optimized.

[0071] This embodiment takes into account that the initial speech synthesis model is trained using a large amount of training data, and the basic stability of the initial speech synthesis model is good. Reinforcement learning is to optimize on the basis of the initial speech synthesis model, and the training data used in reinforcement learning is less. Therefore, during the reinforcement learning process, in order to enable the model of reinforcement learning to maintain the basic stability of the initial speech synthesis model, it is necessary to prevent the model to be optimized from deviating too much from the basic capabilities of the initial speech synthesis model.

[0072] Therefore, in the process of adjusting the network parameters of the model to be optimized in this embodiment, it is necessary to consider the probability difference between the model to be optimized this time and the initial speech synthesis model in generating the second sample synthetic speech, so as to prevent the model to be optimized this time from deviating too much from the original capabilities of the initial speech synthesis model.

[0073] For ease of understanding Figure 3 the steps of the embodiment shown, please refer to the following formula. The following briefly describes the processing of the embodiment shown in combination with the formula: Figure 3 The processing of the embodiment shown is briefly described as follows:

[0074] In the formula, is the model to be optimized, is the initial speech synthesis model, represents the input sample reference speech, represents the generated positive second sample synthetic speech, represents the generated negative second sample synthetic speech, represents the first generation probability of the positive second sample synthetic speech, represents the third generation probability of the initial speech synthesis model generating the positive second sample synthetic speech, represents the first ratio between the first generation probability and the third generation probability of the positive second sample synthetic speech, represents the second generation probability of the negative second sample synthetic speech, represents the fourth generation probability of the initial speech synthesis model generating the negative second sample synthetic speech, represents the second ratio between the second generation probability and the fourth generation probability of the negative second sample synthetic speech, σ represents the activation function, and β is a coefficient.

[0075] Combined with the formula, it can be seen that in step three, based on the difference between the first ratio of the positive second sample synthetic speech and the corresponding second ratio of the negative second sample synthetic speech, the network parameters of the model to be optimized are adjusted. Actually, the adjustment of the network parameters is carried out in the direction of increasing the difference between the first ratio and the second ratio, so as to train the model to be optimized to increase the probability of generating positive example speech and reduce the probability of generating negative example speech.

[0076] In another embodiment, a reinforcement learning algorithm similar to PPO can be used to train the initial speech synthesis model.

[0077] Specifically, please refer to Figure 4 , Figure 4 which is a schematic flowchart of another embodiment of obtaining the target speech synthesis model through reinforcement learning provided by this application. In this embodiment, according to the user's preference degree for the first sample synthetic speech, the initial speech synthesis model is subjected to reinforcement learning to obtain the target speech synthesis model, including: S41: Use the model to be reinforced in this round as the model to be optimized. Among them, the model to be reinforced for the first time is the initial speech synthesis model, and the model to be reinforced for non-first time is the model obtained from the previous reinforcement learning. The first sample synthesized speech includes several second sample synthesized speeches generated by the models to be optimized in each round of reinforcement learning, and the preference degree of the first sample synthesized speech includes the sub-preference degrees of each second sample synthesized speech.

[0078] For specific reference, please refer to the relevant description in part S31, which will not be elaborated here.

[0079] S42: Determine the expected reward corresponding to the model to be optimized based on the sub-preference degrees of each second sample synthesized speech.

[0080] Each second sample synthesized speech described in step S42 is the second sample synthesized speech generated by the model to be optimized in this round. Among them, step S42 is to determine the expected reward corresponding to the model to be optimized in this round by integrating the sub-preference degrees of each second sample synthesized speech.

[0081] Among them, before executing step S42, first obtain a reward model for obtaining the reward values of the first sample synthesized speech. The reward value is used to represent the preference degree of the user for the first sample synthesized speech, and the reward value of the first sample synthesized speech includes the sub-reward values of each second sample synthesized speech. The training method of the reward model can be referred to the relevant description in the previous text, which will not be elaborated here.

[0082] Optionally, the expected reward corresponding to the model to be optimized in this round can be obtained by weighted summing the sub-reward values of each second sample synthesized speech using the generation probabilities of each second sample synthesized speech. Of course, in other embodiments, existing methods for obtaining the expected reward can also be directly adopted.

[0083] S43: Determine the objective function based on the expected reward.

[0084] S44: Adjust the network parameters of the model to be optimized in the direction of maximizing the objective function, and repeat the above steps to obtain the target speech synthesis model. Among them, the value of the objective function is positively correlated with the expected reward.

[0085] In one embodiment, the expected reward can be directly used as the objective function, and then the network parameters of the model to be optimized are adjusted in the direction of maximizing the objective function, and the above steps are repeated until the target speech synthesis model is obtained.

[0086] In another embodiment, first obtain the product of the divergence constraint term and the target coefficient, then use the difference between the expected reward and this product as the objective function, and then adjust the network parameters of the model to be optimized in the direction of maximizing the objective function, and repeat the above steps until the target speech synthesis model is obtained.

[0087] In one embodiment, considering that the reward model is trained with the synthetic speech generated by the initial speech synthesis model, and the initial speech synthesis model is trained with a large amount of training data, the basic stability of the initial speech synthesis model is good. Reinforcement learning is used to optimize based on the initial speech synthesis model, and the training data used in reinforcement learning is less. Therefore, in the process of reinforcement learning, in order to enable the model of reinforcement learning to maintain the basic stability of the initial speech synthesis model, it is necessary to prevent the model to be optimized from deviating too much from the basic capabilities of the initial speech synthesis model. Therefore, in order to prevent the model to be optimized from deviating too far from the initial speech synthesis model, a constraint term can be set to constrain the deviation degree between the model to be optimized and the initial speech synthesis model.

[0088] Exemplarily, please refer to the following formula:

[0089] In the formula, represents the sub-reward value of the second sample synthetic speech output by the reward model, reflecting user preferences, represents the divergence constraint term, and the divergence constraint term is used to characterize the distribution difference between the model to be optimized and the initial speech synthesis model in generating synthetic speech, represents the model to be optimized, represents the initial speech synthesis model, β is the target coefficient, controlling the strength of the divergence constraint term, E represents the expected reward, and θ represents the network parameters of the model to be optimized.

[0090] In a specific embodiment, first obtain the target text and target speech to be synthesized, then use the Q-former network to extract the voice features of the speaker in the target speech, and then use the target speech synthesis model to decode based on the voice features of the speaker and the target text to obtain the target synthetic speech. Among them, in the process of decoding to obtain the target synthetic speech, each speech token is generated in an autoregressive manner, and then all speech tokens are combined to obtain the target synthetic speech.

[0091] Please refer to Figure 5 , Figure 5 is a schematic flowchart of an embodiment of the training method for the target speech synthesis model provided by this application. The training method for the target speech synthesis model in this embodiment includes: S51: Obtain the model to be optimized for the current reinforcement learning; among them, the model for the first reinforcement learning is the initial speech synthesis model, and the model for non-first reinforcement learning is the model obtained from the previous reinforcement learning.

[0092] S52: Obtain each third sample synthetic speech generated by the current model to be optimized.

[0093] S53: Adjust the network parameters of the model to be optimized this time by using the preference degrees of the user for the synthesized voices of each third sample, and repeat the above steps to obtain the target voice synthesis model. Among them, the adjustment direction of the model to be optimized includes: increasing the generation probability of the synthesized voice of the third sample with a high user preference degree.

[0094] For the target voice synthesis model training method in this embodiment, and the method of adjusting the network parameters of the model to be optimized this time by using the preference degrees of the user for the synthesized voices of each third sample, reference can be made to Figure 3 and Figure 4 the method of obtaining the target voice synthesis model by performing reinforcement learning on the initial voice synthesis model by using the preference degrees of the user for the synthesized voices of the first samples shown in the embodiments.

[0095] It should be noted that considering that the small model is limited by the modeling ability, there is a large gap between the generated voice and the natural voice in terms of naturalness and similarity. Therefore, in this embodiment, the above target voice synthesis model is a large model.

[0096] Please refer to Figure 6 , Figure 6 which is a schematic framework diagram of an embodiment of the voice synthesis device provided by the present application. In this embodiment, the voice synthesis device 60 includes a first acquisition module 61, an extraction module 62, and a voice generation module 63. Among them, the first acquisition module 61 is used to acquire the target text and the target voice to be synthesized, the extraction module 62 is used to extract the voice features of the speaker in the target voice, and the voice generation module 63 is used to generate the target synthesized voice based on the voice features of the speaker and the target text. The target synthesized voice is a voice that refers to the voice characteristics of the target voice and has the same pronunciation content as the target text. Among them, the voice features are extracted by using Q-former; and / or, the target synthesized voice is generated by using the target voice synthesis model, and the target voice synthesis model is obtained by performing reinforcement learning on the initial voice synthesis model according to the preference degrees of the user for the synthesized voices of the first samples, and the synthesized voice of the first sample is generated by using the initial voice synthesis model.

[0097] In some embodiments, the preference degrees of the synthesized voice of the first sample include preference sub-degrees for at least one of the following regarding the synthesized voice of the first sample: prosody naturalness, timbre similarity, and intelligibility.

[0098] In some embodiments, each preference sub - degree is obtained through manual evaluation and / or automated evaluation tools; the automated evaluation tools include at least one of the following: a voiceprint verification system, a speech recognition system, and a trained reward model; wherein, the voiceprint verification system is used to determine the timbre similarity between the first sample synthesized speech and the target speech, and the timbre similarity is positively correlated with the corresponding preference sub - degree, and the speech recognition system is used to determine the intelligibility of the first sample synthesized speech, and the intelligibility is positively correlated with the corresponding preference sub - degree; the reward model can determine each preference sub - degree of the user for the first sample synthesized speech.

[0099] In some embodiments, according to the preference degree of the user for the first sample synthesized speech, reinforcement learning is performed on the initial speech synthesis model to obtain the target speech synthesis model, including: using the model to be reinforced this time as the model to be optimized, wherein, the model to be reinforced for the first time is the initial speech synthesis model, and the model to be reinforced for non - first times is the model obtained from the previous reinforcement learning; the first sample synthesized speech includes several second sample synthesized speeches generated by the models to be optimized for each reinforcement learning; based on the preference degree of the user for each second sample synthesized speech, at least one positive - negative speech pair is constructed; the preference degree of the positive second sample synthesized speech in the positive - negative speech pair is higher than that of the negative second sample synthesized speech; obtaining the first generation probability of the model to be optimized for the positive second sample synthesized speech and the second generation probability of the model to be optimized for the negative second sample synthesized speech; based on the first generation probability and the corresponding second generation probability, adjusting the network parameters of the model to be optimized, and repeating the above steps to obtain the target speech synthesis model; wherein, the adjustment direction of the network parameters of the model to be optimized is: the direction of increasing the first generation probability of the positive second sample synthesized speech and decreasing the second generation probability of the negative second sample synthesized speech.

[0100] In some embodiments, based on the first generation probability and the corresponding second generation probability, adjusting the network parameters of the model to be optimized includes: obtaining the third generation probability of the initial speech synthesis model for generating the positive second sample synthesized speech and obtaining the fourth generation probability of the initial speech synthesis model for generating the negative second sample synthesized speech; obtaining the first ratio between the first generation probability and the third generation probability of the positive second sample synthesized speech and obtaining the second ratio between the second generation probability and the fourth generation probability of the negative second sample synthesized speech; based on the difference between the first ratio of the positive second sample synthesized speech and the second ratio of the corresponding negative second sample synthesized speech, adjusting the network parameters of the model to be optimized.

[0101] In some embodiments, to obtain a target speech synthesis model through reinforcement learning on an initial speech synthesis model according to the preference degree of a user for a first sample synthesized speech, the method includes: using the model to be reinforced in this round as the model to be optimized, where the model to be reinforced for the first time is the initial speech synthesis model, and the model to be reinforced for non-first time is the model obtained from the previous reinforcement learning; the first sample synthesized speech includes a plurality of second sample synthesized speeches generated by the models to be optimized in each round of reinforcement learning, and the preference degree of the first sample synthesized speech includes the preference sub-degrees of the second sample synthesized speeches; determining the expected reward corresponding to the model to be optimized based on the preference sub-degrees of the second sample synthesized speeches; determining an objective function based on the expected reward; adjusting the network parameters of the model to be optimized in the direction of maximizing the objective function, and repeating the above steps to obtain the target speech synthesis model; where the value of the objective function is positively correlated with the expected reward.

[0102] In some embodiments, before the speech generation module 63 generates a target synthesized speech based on the voice characteristics of the speaker and the target text, the method further includes: obtaining a reward model for obtaining the reward value of the first sample synthesized speech, where the reward value is used to represent the preference degree of the user for the first sample synthesized speech, and the reward value of the first sample synthesized speech includes the sub-reward values of the second sample synthesized speeches; the expected reward is determined based on the sub-reward values of the second sample synthesized speeches; and / or, determining the objective function based on the expected reward includes: obtaining the product of a divergence constraint term and a target coefficient; the divergence constraint term is used to represent the difference in the probability distributions of the second sample synthesized speeches generated by the model to be optimized in this round and the initial speech synthesis model; using the difference between the expected reward and the product as the objective function.

[0103] In some embodiments, the target speech synthesis model is a large model; and / or, the voice characteristics extracted by the extraction module 62 are obtained by using an audio feature extraction network, and the audio feature extraction network is composed of a plurality of attention modules, and each attention module includes a self-attention sub-module and a cross-attention sub-module; where the self-attention sub-module is used to perform self-attention processing on a plurality of input learnable query vectors to obtain a fusion result of the learnable query vectors, and different learnable query vectors are used to extract voice sub-features of different dimensions of the target speech; the cross-attention sub-module is used to use the audio features of the target speech as key vectors and value vectors respectively, use the fusion result as the query vector, and perform cross-attention processing based on the query vector, key vector and value vector to obtain the voice characteristics.

[0104] Please refer to Figure 7 , Figure 7FIG. 0 is a schematic framework diagram of an embodiment of the target speech synthesis model training device provided by the present application. In this embodiment, the target speech synthesis model training device 70 includes a second acquisition module 71, a third acquisition module 72, and a parameter adjustment module 73. The second acquisition module 71 is configured to acquire a model to be optimized for the current reinforcement learning; wherein, the model for the first reinforcement learning is an initial speech synthesis model, and the model for non-first reinforcement learning is the model obtained from the previous reinforcement learning; the third acquisition module 72 is configured to acquire each third sample synthesized speech generated by the model to be optimized for the current time; the parameter adjustment module 73 is configured to adjust the network parameters of the model to be optimized for the current time by using the preference degree of the user for each third sample synthesized speech, and repeat the above steps to obtain a target speech synthesis model; wherein, the adjustment direction of the model to be optimized includes: increasing the generation probability of the third sample synthesized speech with a high user preference degree.

[0105] Please refer to Figure 8 , Figure 8 FIG. 7 is a schematic framework diagram of an embodiment of the electronic device provided by the present application. In this embodiment, the electronic device 80 includes a memory 81 and a processor 82 which are coupled to each other.

[0106] The memory 81 stores program instructions, and the processor 82 is configured to execute the program instructions stored in the memory 81 to implement the steps of any of the above method embodiments. In a specific implementation scenario, the electronic device 80 may include, but is not limited to: a microcomputer, a server. In addition, the electronic device 80 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited herein.

[0107] Specifically, the processor 82 is configured to control itself and the memory 81 to implement the steps of any of the above embodiments. The processor 82 may also be referred to as a CPU (Central Processing Unit). The processor 82 may be an integrated circuit chip with signal processing capabilities. The processor 82 may also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 82 may be implemented jointly by integrated circuit chips.

[0108] Please refer to Figure 9 , Figure 9It is a schematic framework diagram of the computer-readable storage medium provided by this application. The computer-readable storage medium 90 of the embodiments of this application stores program instructions 91. When the program instructions 91 are executed, they implement the methods provided by any one of the above-mentioned embodiments and any non-conflicting combinations. Among them, the program instructions 91 can form a program file and be stored in the above-mentioned computer-readable storage medium 90 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) can execute all or part of the steps of the methods of various embodiments of this application. The aforementioned computer-readable storage medium 90 includes: various media that can store program codes, such as USB flash drives, external hard drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.

[0109] In the above solution, the target synthetic voice is generated based on the voice characteristics of the speaker and the target text. Among them, the voice characteristics are extracted using Q-former; and / or, the target synthetic voice is generated using the target voice synthesis model, and the target voice synthesis model is obtained by performing reinforcement learning on the initial voice synthesis model according to the user's preference degree for the first sample synthetic voice. Among them, the method of using Q-former to extract the speaker's voice characteristics can extract the voice characteristics important for the generation of the target synthetic voice. Furthermore, using the voice characteristics extracted by Q-former for voice synthesis can make the synthesized voice more conform to the speaker's voice characteristics, making the synthesized voice more natural and accurate. In addition, since the target voice synthesis model can be obtained by performing reinforcement learning on the initial voice synthesis model according to the user's preference degree for the synthetic voice, the probability of the target synthetic voice generated by the target voice synthesis model obtained through reinforcement learning conforming to the user's preference is relatively high. Furthermore, it is beneficial to generate the target synthetic voice that conforms to the user's preference. Compared with the method of directly using the voice synthesis model to generate the target synthetic voice based on the input target text and target voice, both the method of using Q-former to extract the speaker's voice characteristics and the method of using the target voice synthesis model obtained by performing reinforcement learning according to the user's preference for the voice in this application can improve the quality of the target synthetic voice.

[0110] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0111] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0112] In several embodiments provided in the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the apparatuses or units can be in electrical, mechanical or other forms.

[0113] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0114] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0115] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0116] The above is only the embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A voice synthesis method, characterized in that, The method includes: Obtaining a target text and a target voice to be synthesized; Extracting the voice characteristics of the speaker in the target voice; Generating a target synthesized voice based on the voice characteristics of the speaker and the target text, where the target synthesized voice is a voice that refers to the voice characteristics of the target voice and has the same pronunciation content as the target text; Wherein, the voice characteristics are obtained by using Q-former; and / or, the target synthesized voice is generated by using a target voice synthesis model, and the target voice synthesis model is obtained by performing reinforcement learning on an initial voice synthesis model according to the user's preference degree for a first sample synthesized voice, and the first sample synthesized voice is generated by using the initial voice synthesis model.

2. The method according to claim 1, wherein The preference degree of the first sample synthesized voice includes preference sub-degrees for at least one of the following regarding the first sample synthesized voice: prosody naturalness, timbre similarity, and intelligibility.

3. The method according to claim 2, wherein Each of the preference sub-degrees is obtained through manual evaluation and / or an automated evaluation tool; The automated evaluation tool includes at least one of the following: a voiceprint verification system, a speech recognition system, and a trained reward model; Wherein, the voiceprint verification system is used to determine the timbre similarity between the first sample synthesized voice and the target voice, and the timbre similarity is positively correlated with the corresponding preference sub-degree, and the speech recognition system is used to determine the intelligibility of the first sample synthesized voice, and the intelligibility is positively correlated with the corresponding preference sub-degree; the reward model can determine each preference sub-degree of the user for the first sample synthesized voice.

4. The method according to claim 1, characterized in that Performing reinforcement learning on an initial voice synthesis model according to the user's preference degree for a first sample synthesized voice to obtain a target voice synthesis model includes: Taking the model to be reinforced and learned this time as the model to be optimized, where the model to be reinforced and learned for the first time is the initial voice synthesis model, and the model to be reinforced and learned for non-first times is the model obtained from the previous reinforcement learning; the first sample synthesized voice includes a plurality of second sample synthesized voices generated by the model to be optimized for each time of reinforcement learning; Constructing at least one positive and negative voice pair based on the user's preference degree for each of the second sample synthesized voices; the preference degree of the positive second sample synthesized voice in the positive and negative voice pair is higher than that of the negative second sample synthesized voice; Obtaining a first generation probability of the model to be optimized for the positive second sample synthesized voice and a second generation probability of the model to be optimized for the negative second sample synthesized voice; Adjusting the network parameters of the model to be optimized based on the first generation probability and the corresponding second generation probability, and repeating the above steps to obtain the target voice synthesis model; wherein, the adjustment direction of the network parameters of the model to be optimized is: the direction of increasing the first generation probability of the positive second sample synthesized voice and decreasing the second generation probability of the negative second sample synthesized voice.

5. The method according to claim 4, wherein Adjusting the network parameters of the model to be optimized based on the first generation probability and the corresponding second generation probability includes: Obtain a third generation probability that the initial speech synthesis model generates the positive second sample synthetic speech, and obtain a fourth generation probability that the initial speech synthesis model generates the negative second sample synthetic speech; Obtain a first ratio between the first generation probability and the third generation probability of the positive second sample synthetic speech, and obtain a second ratio between the second generation probability and the fourth generation probability of the negative second sample synthetic speech; Based on the difference between the first ratio of the positive second sample synthetic speech and the second ratio of the corresponding negative second sample synthetic speech, adjust the network parameters of the model to be optimized.

6. The method according to claim 1, wherein Performing reinforcement learning on the initial speech synthesis model according to the user's preference degree for the first sample synthetic speech to obtain a target speech synthesis model, including: Taking the model to be reinforced in this time as the model to be optimized, wherein the model to be reinforced for the first time is the initial speech synthesis model, and the model to be reinforced for non-first time is the model obtained by the previous reinforcement learning; the first sample synthetic speech includes a plurality of second sample synthetic speeches generated by the model to be optimized for each time of reinforcement learning, and the preference degree of the first sample synthetic speech includes the sub-preference degrees of each second sample synthetic speech; Based on the sub-preference degrees of each second sample synthetic speech, determine the expected reward corresponding to the model to be optimized; Based on the expected reward, determine the objective function; Adjust the network parameters of the model to be optimized in the direction of maximizing the objective function, and repeat the above steps to obtain the target speech synthesis model; wherein, the value of the objective function is positively correlated with the expected reward.

7. The method according to claim 6, wherein Before generating the target synthetic speech based on the voice feature of the speaker and the target text, the method further includes: Obtain a reward model for obtaining the reward value of the first sample synthetic speech, the reward value being used to characterize the user's preference degree for the first sample synthetic speech, and the reward value of the first sample synthetic speech including the sub-reward values of each second sample synthetic speech; The expected reward is determined based on the sub-reward values of each second sample synthetic speech; And / or, the determining the objective function based on the expected reward includes: Obtain the product of the divergence constraint term and the target coefficient; the divergence constraint term is used to characterize the probability distribution difference between the model to be optimized in this time and the initial speech synthesis model for generating each second sample synthetic speech; Taking the difference between the expected reward and the product as the objective function.

8. The method according to claim 1, characterized in that The target speech synthesis model is a large model; And / or, the voice feature is extracted by using an audio feature extraction network, and the audio feature extraction network is composed of a plurality of attention modules, and each attention module includes a self-attention sub-module and a cross-attention sub-module; wherein, The self-attention sub-module is used to perform self-attention processing on a plurality of input learnable query vectors to obtain a fusion result of the plurality of learnable query vectors, and different learnable query vectors are used to extract different-dimensional voice sub-features of the target voice; The cross-attention sub-module is used to use the audio features of the target speech as the key vector and the value vector respectively, use the fusion result as the query vector, and perform cross-attention processing based on the query vector, key vector and value vector to obtain the voice features.

9. A method for training a target speech synthesis model, characterized in that, The method includes: Obtaining an optimization model to be optimized for the current reinforcement learning; wherein, the model for the first reinforcement learning is the initial speech synthesis model, and the model for non-first reinforcement learning is the model obtained from the previous reinforcement learning; Obtaining each third-sample synthesized speech generated by the optimization model to be optimized this time; Using the preference degree of the user for each third-sample synthesized speech to adjust the network parameters of the optimization model to be optimized this time, and repeating the above steps to obtain the target speech synthesis model; wherein, the adjustment direction of the optimization model includes: increasing the generation probability of the third-sample synthesized speech with a high user preference degree.

10. A voice synthesis device, characterized in that, The device includes: A first acquisition module, configured to acquire a target text and a target speech to be synthesized; An extraction module, configured to extract the voice features of the speaker in the target speech; A speech generation module, configured to generate a target synthesized speech based on the voice features of the speaker and the target text, where the target synthesized speech is a speech that refers to the voice characteristics of the target speech and has the same pronunciation content as the target text; Wherein, the voice features are extracted by using Q-former; and / or, the target synthesized speech is generated by using a target speech synthesis model, and the target speech synthesis model is obtained by performing reinforcement learning on the initial speech synthesis model according to the preference degree of the user for the first-sample synthesized speech, and the first-sample synthesized speech is generated by using the initial speech synthesis model.

11. An apparatus for training a target speech synthesis model, characterized in that, The device includes: A second acquisition module, configured to acquire an optimization model to be optimized for the current reinforcement learning; wherein, the model for the first reinforcement learning is the initial speech synthesis model, and the model for non-first reinforcement learning is the model obtained from the previous reinforcement learning; A third acquisition module, configured to acquire each third-sample synthesized speech generated by the optimization model to be optimized this time; A parameter adjustment module, configured to use the preference degree of the user for each third-sample synthesized speech to adjust the network parameters of the optimization model to be optimized this time, and repeating the above steps to obtain the target speech synthesis model; wherein, the adjustment direction of the optimization model includes: increasing the generation probability of the third-sample synthesized speech with a high user preference degree.

12. An electronic device, characterized in that, Including a memory and a processor coupled to each other, The memory stores program instructions; The processor is configured to execute the program instructions stored in the memory to implement the method according to any one of claims 1-9.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions that can be run by a processor, and the program instructions can be executed by the processor to implement the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Speech synthesis method and device, equipment and storage medium

    CN116612742A

  • Deep synthesis audio detection method, system and product combined with large language model

    CN117577120A

  • Audio analysis method and device, storage medium and electronic equipment

    CN117727301A

  • Driver risk identification method and device, terminal equipment and readable storage medium

    CN118541685A

  • Speech synthesis method and device

    CN119152837A