Method, device, and storage medium for speech synthesis
By introducing the selection and combination functions of targeted pronunciation sound and style in the pronunciation synthesis model, and using the pitch speed determination model, feature fusion model and spectrum synthesis model, the problem that the pronunciation synthesis model can only synthesize the fixed speaker's pronunciation audio in the prior art is solved, achieving a more flexible and realistic pronunciation synthesis effect.
Patent Information
- Application Number
- CN202211529185.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-30
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-11-30
AI Technical Summary
The existing speech synthesis model can only synthesize the aloud audio of a fixed speaker (reading style and pronunciation aloud) and lacks a solution to arbitrarily select and combine the timbre and style.
By determining the target text, target reading timbre and target reading style, and using pre-trained pitch speech speed to determine the model, feature fusion model and spectrum synthesis model, generate tone information that combines pronunciation characteristics and target reading timbre, and finally output the reading audio corresponding to the target text.
The arbitrary selection and combination of the target reading sound and the target reading style of the reading sound of the aloud audio is realized. The generated reading sound incorporates the reading style characteristics of the target reading style and the tone information of the target reading sound, which improves the flexibility and authenticity of speech synthesis.
Smart Images

Figure CN115862594B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech synthesis, and particularly to a method, device, and storage medium for speech synthesis. Background Art
[0002] With the development of products such as voice assistants, intelligent navigation, intelligent customer service, and e-books, text-to-speech (TTS), also known as speech synthesis, has become increasingly common in daily life.
[0003] In related technologies, the reading audio of multiple speakers reading text is used as a sample and an identifier is assigned to each speaker for training. In this way, a speech synthesis model that can output the reading audio of multiple different speakers can be obtained. In actual applications, when the text and the speaker identifier are input into the speech synthesis model, the reading audio of the corresponding speaker reading the text can be obtained.
[0004] In the above technologies, during the model training process, the reading style (the reading style can include characteristics such as speech rate and pitch) of the speaker reading the text is generally fixed. If the reading style changes, it will cause the audio output by the model to be chaotic. It can be seen that the speech synthesis models in related technologies can only synthesize the reading audio of a fixed speaker (with a fixed reading style and pronunciation timbre), and there is a lack of a solution for arbitrarily selecting both the timbre and the style to synthesize the reading audio. Summary of the Invention
[0005] Embodiments of this application provide a method, device, and storage medium for speech synthesis, which can solve the problem of low speech synthesis efficiency. The technical solutions are as follows:
[0006] In a first aspect, a method for speech synthesis is provided. The method includes:
[0007] Determine a target text, and determine the target pronunciation timbre and target reading style of the target text;
[0008] Input the text pronunciation features of the target text into a pre-trained first pitch and speech rate determination model, and the first pitch and speech rate determination model outputs the first reading style features of the target text when read in the target reading style;
[0009] Input the text pronunciation features of the target text into a pre-trained second pitch and speech rate determination model, and the second pitch and speech rate determination model outputs the second reading style features of the target text when read in the target pronunciation timbre;
[0010] Input the text pronunciation features of the target text, the first reading style features, and the second reading style features into a pre-set feature fusion model, and the feature fusion model outputs the fused pronunciation features;
[0011] Input the timbre information fusing the pronunciation feature and the target pronunciation timbre into a pre-trained spectral synthesis model, and output the pronunciation audio corresponding to the target text by the spectral synthesis model.
[0012] In a possible implementation manner, the inputting the text pronunciation feature of the target text into a pre-trained first pitch and speech rate determination model includes:
[0013] Input the text pronunciation feature of the target text into the first pitch and speech rate determination model corresponding to the pre-trained first pronunciation style.
[0014] In a possible implementation manner, the first pitch and speech rate determination model and the second pitch and speech rate determination model are the same pitch and speech rate determination model, and the pitch and speech rate determination model is trained by pronunciation audio samples of multiple pronunciation timbres under multiple pronunciation styles, and each pronunciation audio sample has its corresponding pronunciation timbre code;
[0015] The inputting the text pronunciation feature of the target text into a pre-trained first pitch and speech rate determination model, and the inputting the text pronunciation feature of the target text into a pre-trained second pitch and speech rate determination model include:
[0016] Input the text pronunciation feature of the target text and the pronunciation timbre code corresponding to the target pronunciation style into the pitch and speech rate determination model, and input the text pronunciation feature of the target text and the pronunciation timbre code corresponding to the target pronunciation timbre into the pitch and speech rate determination model.
[0017] In a possible implementation manner, the first pronunciation style feature includes a first pitch feature and a first speech rate feature; the second pronunciation style feature includes a second pitch feature and a second speech rate feature;
[0018] The inputting the text pronunciation feature of the target text, the first pronunciation style feature and the second pronunciation style feature into a pre-set feature fusion model, and outputting the fused pronunciation feature by the feature fusion model includes:
[0019] Generate a fused pitch feature according to the first pitch feature and the second pitch feature;
[0020] Generate a fused speech rate feature according to the first speech rate feature and the second speech rate feature;
[0021] Input the fused pitch feature, the fused speech rate feature and the text pronunciation feature into a pre-set feature fusion model, and output the fused pronunciation feature by the feature fusion model.
[0022] In a possible implementation, the first pitch feature includes a plurality of first fundamental frequencies, where the first fundamental frequencies are the fundamental frequencies of each audio frame when the target text is read in the target reading style, and the second pitch feature includes a plurality of second fundamental frequencies, where the second fundamental frequencies are the fundamental frequencies of each audio frame when the target text is read in the target reading voice color;
[0023] Generating a fused pitch feature based on the first pitch feature and the second pitch feature includes:
[0024] Determine the average value of the plurality of second fundamental frequencies, determine the difference between each of the first fundamental frequencies and the average value, and form a third pitch feature with the plurality of differences;
[0025] Perform weighted summation on the third pitch feature and the second pitch feature according to a first weight corresponding to the third pitch feature and a second weight corresponding to the second pitch feature to obtain a fused pitch feature, where the first weight is greater than the second weight.
[0026] In a possible implementation, the first pitch feature includes a plurality of first fundamental frequencies, where the first fundamental frequencies are the fundamental frequencies of each audio frame when the target text is read in the target reading style, and the second pitch feature includes a plurality of second fundamental frequencies, where the second fundamental frequencies are the fundamental frequencies of each audio frame when the target text is read in the target reading voice color;
[0027] Determining a fused pitch feature according to the first pitch feature and the second pitch feature includes:
[0028] Calculate the average value of the plurality of first fundamental frequencies to obtain a first average value, calculate the average value of the plurality of second fundamental frequencies to obtain a second average value, and determine the ratio of the variance of the second pitch feature to the variance of the first pitch feature;
[0029] Determine the difference between each of the first fundamental frequencies and the first average value, determine the product of each difference and the ratio, determine the sum value of each product and the second average value, and use the sequence composed of the plurality of sum values as the fused pitch feature.
[0030] In a possible implementation, the first speech rate feature includes the first audio duration of each phoneme when the target text is read in the target reading style, and the second speech rate feature includes the second audio duration of each phoneme when the target text is read in the target reading voice color;
[0031] Determining a fused speech rate feature according to the first speech rate feature and the second speech rate feature includes:
[0032] Weight the first speech rate feature and the second speech rate feature according to the weights corresponding to the first speech rate feature and the weights corresponding to the second speech rate feature, and perform weighted summation to obtain a fused speech rate feature, where the weight corresponding to the first speech rate feature is greater than the weight corresponding to the second speech rate feature.
[0033] In a possible implementation, the inputting the fused pitch feature, the fused speech rate feature, and the text pronunciation feature into a pre-set feature fusion model includes:
[0034] Determine the average value and variance of multiple fused fundamental frequencies included in the fused pitch feature;
[0035] Determine the difference between each of the fused fundamental frequencies and the average value, determine the value obtained by multiplying each of the differences by a preset control parameter and dividing by the variance, determine the sum value of each of the values and the average value, and use the sequence composed of multiple sum values as the adjusted fused pitch feature;
[0036] Input the adjusted fused pitch feature, the fused speech rate feature, and the text pronunciation feature into a pre-set feature fusion model.
[0037] In a possible implementation, the method further includes:
[0038] Determine the phoneme sequence corresponding to the target text;
[0039] Input the phoneme sequence into a pre-trained encoder, and output the text pronunciation feature of the target text by the encoder.
[0040] In a possible implementation, the spectral synthesis model includes a decoder and a vocoder;
[0041] The inputting the fused pronunciation feature and the timbre information of the target pronunciation timbre into a pre-trained spectral synthesis model, and outputting the pronunciation audio corresponding to the target text by the spectral synthesis model includes:
[0042] Input the fused pronunciation feature and the timbre information of the target pronunciation timbre into the decoder, and output Mel spectrum features by the decoder;
[0043] Input the Mel spectrum features into the vocoder, and output the pronunciation audio corresponding to the target text by the vocoder.
[0044] In a second aspect, a voice synthesis device is provided, and the device includes:
[0045] A determination module, configured to determine a target text, and determine the target pronunciation timbre and target pronunciation style of the target text;
[0046] An input module, configured to input the text pronunciation features of the target text into a pre-trained first pitch and speech rate determination model, and output, by the first pitch and speech rate determination model, the first reading style features when the target text is read in the target reading style; and is further configured to input the text pronunciation features of the target text into a pre-trained second pitch and speech rate determination model, and output, by the second pitch and speech rate determination model, the second reading style features when the target text is read in the target reading voice color.
[0047] A fusion module, configured to input the text pronunciation features of the target text, the first reading style features, and the second reading style features into a pre-set feature fusion model, and output, by the feature fusion model, the fused pronunciation features.
[0048] A synthesis module, configured to input the fused pronunciation features and the voice color information of the target reading voice color into a pre-trained spectral synthesis model, and output, by the spectral synthesis model, the reading audio corresponding to the target text.
[0049] In a possible implementation manner, the input module is configured to:
[0050] Input the text pronunciation features of the target text into a first pitch and speech rate determination model corresponding to a pre-trained first reading style.
[0051] In a possible implementation manner, the first pitch and speech rate determination model and the second pitch and speech rate determination model are the same pitch and speech rate determination model, and the pitch and speech rate determination model is trained with reading audio samples of multiple reading voice colors in multiple reading styles, and each reading audio sample has its corresponding reading voice color code;
[0052] The input module is configured to input the text pronunciation features of the target text and the reading voice color code corresponding to the target reading style into the pitch and speech rate determination model, and input the text pronunciation features of the target text and the reading voice color code corresponding to the target reading voice color into the pitch and speech rate determination model.
[0053] In a possible implementation manner, the first reading style features include a first pitch feature and a first speech rate feature; the second reading style features include a second pitch feature and a second speech rate feature;
[0054] The fusion module is configured to generate a fused pitch feature according to the first pitch feature and the second pitch feature;
[0055] Generate a fused speech rate feature according to the first speech rate feature and the second speech rate feature;
[0056] Input the fused pitch feature, the fused speech rate feature, and the text pronunciation feature into a pre-set feature fusion model, and the feature fusion model outputs a fused pronunciation feature.
[0057] In a possible implementation, the first pitch feature includes multiple first fundamental frequencies, where the first fundamental frequencies are the fundamental frequencies of each audio frame when the target text is read in the target reading style, and the second pitch feature includes multiple second fundamental frequencies, where the second fundamental frequencies are the fundamental frequencies of each audio frame when the target text is read in the target reading voice color.
[0058] The fusion module is configured to determine the average value of the multiple second fundamental frequencies, determine the difference between each first fundamental frequency and the average value, and form a third pitch feature with the multiple differences.
[0059] Perform weighted summation on the third pitch feature and the second pitch feature according to the first weight corresponding to the third pitch feature and the second weight corresponding to the second pitch feature to obtain a fused pitch feature, where the first weight is greater than the second weight.
[0060] In a possible implementation, the first pitch feature includes multiple first fundamental frequencies, where the first fundamental frequencies are the fundamental frequencies of each audio frame when the target text is read in the target reading style, and the second pitch feature includes multiple second fundamental frequencies, where the second fundamental frequencies are the fundamental frequencies of each audio frame when the target text is read in the target reading voice color.
[0061] The fusion module is configured to calculate the average value of the multiple first fundamental frequencies to obtain a first average value, calculate the average value of the multiple second fundamental frequencies to obtain a second average value, and determine the ratio of the variance of the second pitch feature to the variance of the first pitch feature.
[0062] Determine the difference between each first fundamental frequency and the first average value, determine the product of each difference and the ratio, determine the sum value of each product and the second average value, and use the sequence composed of the multiple sum values as the fused pitch feature.
[0063] In a possible implementation, the first speech rate feature includes the first audio duration of each phoneme when the target text is read in the target reading style, and the second speech rate feature includes the second audio duration of each phoneme when the target text is read in the target reading voice color.
[0064] The fusion module is configured to perform weighted summation on the first speech rate feature and the second speech rate feature according to the weight corresponding to the first speech rate feature and the weight corresponding to the second speech rate feature, so as to obtain a fused speech rate feature, where the weight corresponding to the first speech rate feature is greater than the weight corresponding to the second speech rate feature.
[0065] In a possible implementation, the fusion module is configured to:
[0066] Determine the average value and variance of multiple fused fundamental frequencies included in the fused pitch feature;
[0067] Determine the difference between each of the fused fundamental frequencies and the average value, determine the value obtained by multiplying each of the differences by a preset control parameter and dividing by the variance, determine the sum value of each of the values and the average value, and use the sequence composed of multiple sum values as the adjusted fused pitch feature;
[0068] Input the adjusted fused pitch feature, the fused speech rate feature, and the text pronunciation feature into a preset feature fusion model.
[0069] In a possible implementation, the determination module is further configured to:
[0070] Determine the phoneme sequence corresponding to the target text;
[0071] Input the phoneme sequence into a pre-trained encoder, and the encoder outputs the text pronunciation feature of the target text.
[0072] In a possible implementation, the spectrum synthesis model includes a decoder and a vocoder;
[0073] The synthesis module is configured to input the fused pronunciation feature and the timbre information of the target reading timbre into the decoder, and the decoder outputs Mel spectrum features;
[0074] Input the Mel spectrum features into the vocoder, and the vocoder outputs the reading audio corresponding to the target text.
[0075] In a third aspect, a computer device is provided. The computer device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the method for speech synthesis as described in the first aspect above.
[0076] In a fourth aspect, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by the processor to implement the method for speech synthesis as described in the first aspect above.
[0077] In a fifth aspect, a computer program product is provided. The computer program product includes computer program code, and when the computer program code is executed by a computer device, the computer device executes the method according to the first aspect and its possible implementation manners.
[0078] The beneficial effects brought by the technical solutions provided in the embodiments of this application at least include:
[0079] In the embodiments of this application, a user can select a target pronunciation timbre, a target pronunciation style, and a target text of the pronunciation audio, and then use a pitch and speech rate determination model, a feature fusion model, and a spectrum synthesis model to synthesize the pronunciation audio. The synthesized pronunciation audio incorporates the pronunciation style features of the target pronunciation style and also incorporates the timbre information of the target pronunciation timbre. During the synthesis process, the target pronunciation timbre and the target pronunciation style can be arbitrarily selected and combined. It can be seen that the embodiments of this application provide a solution that can arbitrarily select both the timbre and the style to synthesize the pronunciation audio. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0081] Figure 1 is a schematic structural diagram of a computer device provided by the embodiments of this application;
[0082] Figure 2 is a flowchart of a speech synthesis method provided by the embodiments of this application;
[0083] Figure 3 is an algorithm flowchart of a speech synthesis method provided by the embodiments of this application;
[0084] Figure 4 is a flowchart of a method for generating fused pronunciation features provided by the embodiments of this application;
[0085] Figure 5 is a flowchart of a method for adjusting fused pronunciation features provided by the embodiments of this application;
[0086] Figure 6 is a flowchart of a model training provided by the embodiments of this application;
[0087] Figure 7 is an algorithm flowchart of a model training provided by the embodiments of this application;
[0088] Figure 8It is a structural diagram of a voice synthesis device provided by an embodiment of the present application. Detailed implementation manners
[0089] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0090] A voice synthesis method provided by an embodiment of the present application. The execution subject of this method can be a computer device, and this computer device can be a server or a terminal. This computer device can be a single server or a server group. If it is a single server, this server can be responsible for all the processing in the following solutions. If it is a server group, different servers in the server group can be responsible for different processing in the following solutions respectively. The specific processing allocation can be arbitrarily set by technicians according to actual needs and will not be elaborated here.
[0091] This server can be the background server of a certain application. This application can be an application or website with text-to-speech (voice synthesis) function. This application or website can be a voice synthesis application or website, etc. In the embodiments of the present application, the application is taken as a voice synthesis program as an example to elaborate the solution in detail. Other situations are similar and will not be elaborated in this embodiment.
[0092] Figure 1 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. From the perspective of hardware composition, the structure of the computer device can be as Figure 1 shown, including a processor 110, a memory 120, and a communication component 130.
[0093] The processor 110 can be a central processing unit (CPU) or a system on chip (SoC), etc. The processor 110 can be used to obtain text pronunciation features, generate fused pronunciation features, generate a reading audio corresponding to the target text, etc.
[0094] The memory 120 can include various volatile memories or non-volatile memories, such as a solid-state disk (SSD), a dynamic random access memory (DRAM), etc. The memory 120 can be used to store the initial data, intermediate data, and result data used in the voice synthesis process. For example, the text to be converted into a reading audio, the synthesized reading audio, etc.
[0095] The communication component 130 can be a wired network connector, ultra wide band (UWB), wireless fidelity (WiFi) module, Bluetooth module, cellular network communication module, etc. The communication component 130 can be used for data transmission with other devices, which can be servers or terminals, etc. For example, receiving text to be converted into spoken audio and sending the synthesized spoken audio, etc.
[0096] When performing audio synthesis using the target pronunciation timbre and target pronunciation style, in order to make the speech more natural and realistic, a small amount of pronunciation style features of the target pronunciation timbre can be incorporated into the target pronunciation style, and then audio synthesis is performed. The corresponding processing flow can be as Figure 2 shown, and the overall algorithm architecture is as Figure 3 shown.
[0097] 201, determine the target text, and determine the target pronunciation timbre and target pronunciation style of the target text.
[0098] The target text can be a word, phrase, or single text sentence, or multiple text sentences, etc. The target text can be randomly obtained from a text library, or can also be obtained by web crawling, or can also be obtained by manual input. The user can manually set the target pronunciation timbre and target pronunciation style in the timbre option and style option. The target pronunciation style can be considered as the style used by the first person to read the target text, and the target pronunciation timbre can be considered as the pronunciation timbre of the second person. The first person and the second person can be the same person or different persons.
[0099] For example, the acquisition method of the target text in a voice assistant can be: the server or terminal inputs the question text into a question-and-answer model to obtain the answer text as the target text. Another example, the acquisition method of the target text in intelligent navigation can be: the server or terminal generates route prompt text in real time according to the located position and route as the target text. Another example, the acquisition method of the target text in intelligent customer service can be: after receiving the instruction or question from the terminal, the server obtains the reply text of the corresponding instruction or question in the text library as the target text. Another example, in the process of generating audiobook audio, the acquisition method of the target text can be: all the text in the e-book can be used as the target text, or a certain paragraph or sentence of the text can be selected separately as the target text.
[0100] After the computer device obtains the target text, it first converts the target text into a corresponding phoneme sequence. Then, the phoneme sequence corresponding to the target text is input into the encoder, and the text pronunciation feature of the target text is output. The text pronunciation feature can also be called text encoding.
[0101] 202. Input the text pronunciation features of the target text into a pre-trained first pitch and speech rate determination model, and the first pitch and speech rate determination model outputs the first reading style features when the target text is read in the target reading style.
[0102] Among them, the pitch and speech rate determination model can be a machine learning model, such as a convolutional neural network, a recurrent neural network, and a recursive neural network, etc. The pitch and speech rate determination model is trained with the reading audio samples of various reading timbres in various reading styles, and each reading audio sample has its corresponding reading timbre code. The reading timbre code can be a code number, and each code number is assigned to a timbre and corresponds to a timbre. Each person has their own unique timbre, that is to say, each code number can correspond to a person. In addition, it can be set that each person only corresponds to one reading style, so that each code number can correspond to one reading style.
[0103] For example, it can be set that the person code number of the first person to which the reading style belongs is 0001, that is, the reading timbre code is 0001; or, it can be set that the person code number of the first person to which the reading style belongs is speaker1, that is, the reading timbre code is speaker1.
[0104] The training process of the pitch and speech rate determination model will be described in detail in the following content. The first reading style features include the first pitch feature and the first speech rate feature. The pitch feature can be the fundamental frequency of each audio frame in the reading audio corresponding to the target text, and the speech rate feature can be the audio duration of each phoneme in the reading audio corresponding to the target text.
[0105] The determination method of the first reading style features can be various. The following describes several feasible determination methods:
[0106] Method 1: Train a pitch and speech rate determination model that can be used for each reading style. The computer device can input the text pronunciation features of the target text and the reading timbre code corresponding to the target reading style into the pitch and speech rate determination model, and the pitch and speech rate determination model outputs the first reading style features when the target text is read in the target reading style.
[0107] Method 2: Train multiple pitch and speech rate determination models for different reading styles. The computer device can input the text pronunciation features into the first pitch and speech rate determination model corresponding to the target reading style, and can output the first reading style features when the target text is read in the target reading style.
[0108] 203. Input the text pronunciation features of the target text into a pre-trained second pitch and speech rate determination model, and the second pitch and speech rate determination model outputs the second reading style features when the target text is read in the target reading timbre.
[0109] Among them, the processing method of step 203 is the same as that of step 202. For the relevant description content of step 202, please refer to it, and it will not be elaborated here.
[0110] The first pitch and speech rate determination model and the second pitch and speech rate determination model can be the same pitch and speech rate determination model. The pitch and speech rate determination model is trained by using the reading audio samples of multiple reading voices in multiple reading styles, and each reading audio sample has its corresponding reading voice encoding. The first reading style feature includes the first pitch feature and the first speech rate feature, and the second reading style feature includes the second pitch feature and the second speech rate feature. The determined first reading style feature and second reading style feature are used to generate the fused pronunciation feature subsequently. The second reading style feature and the target reading voice both belong to the second person.
[0111] 204. Input the text pronunciation feature, the first reading style feature, and the second reading style feature of the target text into a pre-set feature fusion model, and the feature fusion model outputs the fused pronunciation feature.
[0112] Among them, the feature fusion model can be a machine learning model, such as a convolutional neural network, a recurrent neural network, and a recursive neural network, etc. The feature fusion model can be pre-set, and the setting process will be described in detail in the following content. The feature fusion model can be set when training the pitch and speech rate determination model, or it can also be set separately. The fused pronunciation feature incorporates the pronunciation characteristics of the target text and the reading style features of the first person and the second person.
[0113] The computer device can input the text pronunciation feature, the first reading style feature, and the second reading style feature of the target text into the feature fusion model to obtain the fused pronunciation feature. Since the text pronunciation feature can also be called text encoding, the fused pronunciation feature can be considered as the text encoding fused with the first reading style feature and the second reading style feature (specifically, the first pitch feature, the first speech rate feature, the second pitch feature, and the second speech rate feature).
[0114] There can be various methods for generating the fused pronunciation feature. Several feasible determination methods are described below:
[0115] Method 1: The computer device can input the text pronunciation feature, the first pitch feature, the first speech rate feature, the second pitch feature, and the second speech rate feature into the feature fusion model together to output the fused pronunciation feature.
[0116] Method 2: The computer device can separately perform pitch feature fusion and speech rate feature fusion, and then input the obtained fused pitch feature, fused speech rate feature, and text pronunciation feature into the feature fusion model to output the fused pronunciation feature. The specific methods of pitch feature fusion and speech rate feature fusion will be described in detail in the following content.
[0117] 205, input the timbre information that fuses the pronunciation features and the target pronunciation timbre into a pre-trained spectral synthesis model, and the spectral synthesis model outputs the pronunciation audio corresponding to the target text.
[0118] Among them, the spectral synthesis model is a decoder and a vocoder. The timbre information can be a code uniquely corresponding to the target pronunciation timbre. This code can be the code assigned to the target pronunciation timbre, that is, the timbre code, or the code assigned to the person to which the target pronunciation timbre belongs, that is, the person code. The timbre information input in this step can be regarded as the timbre information of the second person to which the target pronunciation timbre belongs. The second person and the first person to which the target pronunciation style belongs can be different persons or the same person. The timbre information can also be a timbre feature vector.
[0119] The computer device can input the timbre information that fuses the pronunciation features and the target pronunciation timbre into the decoder, and the decoder outputs the Mel spectrum features. Then, the computer device can input the obtained Mel spectrum features into the vocoder, and can output the pronunciation audio corresponding to the target text. The pronunciation audio generated in the above manner incorporates the target pronunciation timbre and the target pronunciation style.
[0120] After obtaining the pronunciation audio, the subsequent processing also varies according to different usage scenarios. For example, after the voice assistant obtains the pronunciation audio corresponding to the target text, it plays it directly; the intelligent navigation terminal plays the pronunciation audio in real time for driving prompts; the intelligent customer service plays the pronunciation audio during the call to answer the user's questions; in the application scenario of the e-book, the generated pronunciation audio is stored and used as an audiobook for the user to download.
[0121] In the embodiments of the present application, the user can select the target pronunciation timbre, the target pronunciation style, and the target text of the pronunciation audio, and then use the pitch and speech rate determination model, the feature fusion model, and the spectral synthesis model to synthesize the pronunciation audio. The synthesized pronunciation audio incorporates the pronunciation style features of the target pronunciation style and also incorporates the timbre information of the target pronunciation timbre. During the synthesis process, the target pronunciation timbre and the target pronunciation style can be arbitrarily selected and combined. It can be seen that the embodiments of the present application provide a solution that can arbitrarily select both the timbre and the style to synthesize the pronunciation audio.
[0122] The processing flow of the method for generating the fused pronunciation features provided by the embodiments of the present application can be as Figure 4 shown, including the following steps:
[0123] 401, generate the fused pitch feature according to the first pitch feature and the second pitch feature.
[0124] Among them, the first pitch feature includes a plurality of first fundamental frequencies F0 1The first fundamental frequency is the fundamental frequency of each audio frame when the target text is read aloud in the target reading style, and the second pitch feature includes a plurality of second fundamental frequencies F0 2 , the second fundamental frequency is the fundamental frequency of each audio frame when the target text is read aloud with the target reading timbre. Based on the above introduction, the target reading style corresponds to the first character, and the target reading timbre corresponds to the second character. When each character is set to correspond to only one reading style, the audio frame obtained when reading with the target reading timbre will have the reading style of the second character. Therefore, the first fundamental frequency corresponds to the reading style of the first character, and the second fundamental frequency corresponds to the reading style of the second character.
[0125] There are many ways to implement pitch feature fusion. Here are some feasible fusion methods:
[0126] Method 1: Determine multiple second fundamental frequencies F0 2 The average value of 2 , determine each first fundamental frequency F0 1 Respectively with the mean 2 The difference is calculated as shown in formula 1 to obtain multiple third fundamental frequencies F0 3 , this plurality of third fundamental frequencies F0 3 Constituent third pitch characteristic.
[0127] F0 1 -mean 2 =F0 3 ………………Formula 1
[0128] According to the third fundamental frequency F0 3 The corresponding first weight x 1 and the second fundamental frequency F0 2 The corresponding second weight x 2 , for the third fundamental frequency F0 3 and the second fundamental frequency F0 2 A weighted sum is performed to obtain multiple fused fundamental frequencies F0 as shown in Formula 2, and the multiple fused fundamental frequencies F0 constitute a fused pitch feature. The first weight is greater than the second weight, and the relationship between the first weight and the second weight can be shown in Formula 3. For example, the first weight is 0.8 and the second weight is 0.2.
[0129] x 1 ×F0 1 +x 2 ×mean 2 =F0………………Formula 2
[0130] 1-x 1 =x 2 ………………Formula 3
[0131] Method 2: Determine multiple first fundamental frequencies F01 The mean value of 1 and the variance sigma 1 , multiple second fundamental frequencies F0 2 The mean value of 2 and the variance sigma 2 , and the second fundamental frequency F0 is calculated as shown in Equation 4 2 The variance sigma of 2 and the first fundamental frequency F0 1 The variance sigma of 1 The ratio y.
[0132]
[0133] Determine each first fundamental frequency F0 1 The difference from the mean value respectively 1 , determine the product of each difference and the ratio y respectively, determine the sum of each product and the mean value respectively 2 , and calculate multiple fused fundamental frequencies F0 as shown in Equation 5. These multiple fused fundamental frequencies F0 constitute the fused pitch feature.
[0134] (F0 1 - mean 1 ) × y + mean 2 = F0………………Equation 5
[0135] 402, generate a fused speech rate feature according to the first speech rate feature and the second speech rate feature.
[0136] Among them, the first speech rate feature includes the first audio duration t of each phoneme when the target text is read in the target reading style 1 , and the second speech rate feature includes the second audio duration t of each phoneme when the target text is read in the target reading voice color 2 .
[0137] According to the weight x corresponding to the first audio duration t 1 and the weight x corresponding to the second audio duration t 3 , perform weighted summation on the first audio duration t 2 and the second audio duration t 4 . Calculate multiple fused durations t as shown in Formula 6. These multiple fused durations t constitute the fused speech rate feature. Among them, the weight corresponding to the first audio duration t 1 is greater than the weight corresponding to the second audio duration t 2 . The relationship between x 1 and x 1 can be as shown in Equation 7. 3 and x 4 The relationship can be as shown in Equation 7.
[0138] x 3 ×t 1 +x 4 ×t 2 =t………………Equation 6
[0139] 1 - x 3 =x 4 ………………Equation 7
[0140] 403. Input the fused pitch feature, the fused speech rate feature, and the text pronunciation feature into a pre - set feature fusion model, and the feature fusion model outputs a fused pronunciation feature.
[0141] Among them, the fused pitch feature includes multiple fused fundamental frequencies, and the fused speech rate feature includes multiple fused durations.
[0142] The computer device can input the fused pitch feature, the fused speech rate feature, and the text pronunciation feature into the feature fusion model to output a fused pronunciation feature.
[0143] There is no necessary chronological order between the above steps 401 and 402. Step 401 can be prior, step 402 can be prior, or they can be executed simultaneously.
[0144] The embodiment of the present application can also adjust the above - mentioned fused pitch feature to adjust the tone strength of the reading audio. Correspondingly, the processing of the above step 403 can be as Figure 5 shown, including the following steps:
[0145] 501. Determine the average value and variance of the multiple fused fundamental frequencies included in the fused pitch feature.
[0146] Based on multiple fused fundamental frequencies F0, determine the average value mean and variance sigma of the multiple fused fundamental frequencies F0.
[0147] 502. Generate an adjusted fused pitch feature according to the multiple fused fundamental frequencies, the average value, the variance, and a preset control parameter.
[0148] Determine the difference between each fused fundamental frequency F0 and the average value mean respectively, determine the value obtained by multiplying each difference by the preset control parameter a and dividing by the variance sigma respectively, and calculate the sum value F0′ of each value and the average value as shown in Equation 8. These multiple sum values F0′ constitute the adjusted fused pitch feature.
[0149]
[0150] Among them, the preset control parameter can be set manually. The larger the preset control parameter is, the larger the variance of the pitch feature adjusted according to the preset control parameter is, and the greater the pitch fluctuation of the generated speech is. For example, the synthesized speech is used for reading e-books. In an e-book, there is no need for a large pitch fluctuation at the "narration" of the novel text, but a larger pitch fluctuation is required at the "dialogue". Therefore, different preset control parameters can be set for the "narration" text and the "dialogue" text. The preset control parameter of the "narration" text can be set to 0.8, and the preset control parameter of the "dialogue" text can be set to 1.2.
[0151] 503. Input the adjusted fused pitch feature, fused speech rate feature, and text pronunciation feature into a preset feature fusion model, and the feature fusion model outputs a fused pronunciation feature.
[0152] The computer device can input the adjusted fused pitch feature, fused speech rate feature, and text pronunciation feature into the feature fusion model and output a fused pronunciation feature.
[0153] By using the method of the embodiment of the present application to adjust the fused pitch feature, the variance of the fundamental frequency in the fused pitch feature can be adjusted, thereby adjusting the strength of the tone of the reading audio. For example, by adjusting the fused pitch feature to increase the variance of the fundamental frequency in the fused pitch feature, the pitch fluctuation of the reading audio synthesized with the fused pitch feature is increased, that is, the sound fluctuation of the reading audio is more obvious, and the more obvious the sound fluctuation is, the stronger the tone is.
[0154] In the embodiment of the present application, the pitch and speech rate determination model is the pitch and speech rate determination model that can be used for each reading style in step 202. In the embodiment of the present application, the model training process can be as Figure 6 shown, including the following steps:
[0155] 601. Obtain a sample text, obtain the reading audio of the third person reading the sample text in the third reading style as a reference reading audio, and assign a reading voice color code to the reference reading audio.
[0156] Among them, multiple reference reading audios of different styles are required for training. The reference reading audios belonging to the same style are recorded by the same person. The reading voice color code can be a person code, and the reading voice color code can be used to indicate the reading style corresponding to the person to the pitch and speech rate determination model and the reading voice color corresponding to the person to the spectral synthesis model.
[0157] 602. Obtain the text pronunciation feature of the sample text.
[0158] The processing method of step 702 is the same as that of step 201. For the relevant description content of step 201, reference can be made, and details are not described here.
[0159] 603. Determine the reading style features of the reference reading audio as the reference reading style features for training.
[0160] The reference reading style features include reference pitch features and reference speech rate features. The computer device extracts the fundamental frequency in the reference reading audio and uses the fundamental frequency of each audio frame in the reference reading audio as the reference pitch features. Among them, the extraction method of the fundamental frequency can be the time domain method, the frequency domain method, etc.
[0161] The computer device determines the reference phoneme sequence corresponding to the sample text according to the sample text. The computer device identifies the number of frames and the audio frame duration corresponding to each phoneme in the reference reading audio according to the reference phoneme sequence, obtains the audio duration of each phoneme in the reference reading audio, and uses the sequence composed of the audio durations of each phoneme in the reference reading audio as the reference speech rate features.
[0162] 604. Input the text pronunciation features and the pronunciation color coding of the reference reading audio into the pitch and speech rate determination model to be trained, and the pitch and speech rate determination model to be trained outputs the predicted reading style features when the sample text is read in the third reading style.
[0163] Among them, the predicted reading style features include predicted pitch features and predicted speech rate features.
[0164] 605. Input the predicted reading style features and the text pronunciation features of the sample text into the pre-set feature fusion model, and the pre-set feature fusion model outputs the predicted fusion pronunciation features.
[0165] Among them, the parameters in the feature fusion model are pre-set by technicians according to experience.
[0166] 606. Input the predicted fusion pronunciation features and the pronunciation color information of the third person's reading audio into the spectrum synthesis model, and the spectrum synthesis model to be trained outputs the predicted reading audio corresponding to the sample text.
[0167] Among them, the spectrum synthesis model can be a model that has been trained. The pronunciation color information can be the pronunciation color coding of the reference reading audio.
[0168] The processing methods of steps 604-606 are the same as those of steps 202-204. For the relevant description content of steps 202-204, please refer to it, and it will not be elaborated here.
[0169] 607. Input the predicted reading style features and the reference reading style features into the first loss function, and the first loss function outputs the first loss value. Input the predicted reading audio and the reference reading audio into the second loss function, and the second loss function outputs the second loss value.
[0170] Among them, the first loss function and the second loss function can be corresponding squared loss functions, logarithmic loss functions, exponential loss functions, and so on.
[0171] 608. Adjust the parameters of the pitch and speech rate determination model to be trained and the preset feature fusion model according to the first loss value and the second loss value.
[0172] 609. If the training end condition is satisfied, determine the pitch and speech rate determination model after parameter adjustment as the trained pitch and speech rate determination model, and determine the feature fusion model after parameter adjustment as the set feature fusion model. If the training end condition is not satisfied, obtain other sample texts and re-execute the above process.
[0173] There can be many choices for the training end condition. The following are several examples:
[0174] Condition 1: Reach the specified number of training times. Condition 2: Each loss value is less than the specified value. Condition 3: Each loss value no longer has a decreasing trend. Condition 4: Use a certain number of samples to verify the accuracy of each model after parameter adjustment. Compare the predicted pitch features with the reference pitch features, compare the predicted speech rate features with the reference speech rate features, and compare the predicted reading audio with the reference reading audio. The matching degrees all reach the specified value. In the embodiments of the present application, the pitch and speech rate determination model training and the feature fusion model setting are carried out simultaneously as an example. The feature fusion model can be pre-set by technicians alone without parameter adjustment.
[0175] In Figure 7 the internal structures of the pitch and speech rate determination model, the feature fusion model, and the spectral synthesis model are shown, and the data transmission relationships between the models in the above process are shown. In the embodiments of the present application, the specific structure of the pitch and speech rate determination model may include two convolutional layers and a fully connected layer. The convolutional layer is composed of a rectified linear units (ReLU) activation layer, a one-dimensional convolutional layer, a normalization layer, and a dropout layer. The specific structure of the feature fusion model may include a one-dimensional convolutional layer and a positional encoding layer. The specific structure of the spectral synthesis model may include a one-dimensional transposed convolutional layer, a one-dimensional dilated convolutional layer, a gated activation layer, a 1×1 convolutional layer, and a one-dimensional convolutional layer. The number of one-dimensional dilated convolutional layers of the spectral synthesis model can be set according to actual requirements or experimental effects.
[0176] In the embodiments of the present application, the user can select the target pronunciation timbre, the target pronunciation style, and the target text of the audio to be read aloud, and then use the pitch and speech rate determination model, the feature fusion model, and the spectral synthesis model to synthesize the audio for reading aloud. The synthesized audio for reading aloud incorporates the pronunciation style features of the target pronunciation style and also incorporates the timbre information of the target pronunciation timbre. During the synthesis process, the target pronunciation timbre and the target pronunciation style can be arbitrarily selected and combined. It can be seen that the embodiments of the present application provide a solution that can arbitrarily select both the timbre and the style to synthesize the audio for reading aloud.
[0177] All of the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present disclosure, which will not be elaborated herein one by one.
[0178] Based on the same technical concept, the embodiments of the present application also provide a voice synthesis device, which is applied to the computer device in the above embodiments, such as Figure 8 As shown, the device includes:
[0179] A determination module 810, configured to determine the target text, and determine the target pronunciation timbre and the target pronunciation style of the target text. Specifically, it can implement the determination function in step 201 above, as well as other implicit steps.
[0180] An input module 820, configured to input the text pronunciation features of the target text into a pre-trained first pitch and speech rate determination model, and output the first pronunciation style features when the target text is read aloud in the target pronunciation style by the first pitch and speech rate determination model. It is also configured to input the text pronunciation features of the target text into a pre-trained second pitch and speech rate determination model, and output the second pronunciation style features when the target text is read aloud in the target pronunciation timbre by the second pitch and speech rate determination model. Specifically, it can implement the input function in steps 202-203 above, as well as other implicit steps.
[0181] A fusion module 830, configured to input the text pronunciation features, the first pronunciation style features, and the second pronunciation style features of the target text into a pre-set feature fusion model, and output the fused pronunciation features by the feature fusion model. Specifically, it can implement the fusion function in step 204 above, as well as other implicit steps.
[0182] A synthesis module 840, configured to input the fused pronunciation features and the timbre information of the target pronunciation timbre into a pre-trained spectral synthesis model, and output the audio for reading aloud corresponding to the target text by the spectral synthesis model. Specifically, it can implement the synthesis function in step 205 above, as well as other implicit steps.
[0183] In a possible implementation manner, the input module 820 is configured to input the text pronunciation features of the target text into the first pitch and speech rate determination model corresponding to the pre-trained first pronunciation style.
[0184] In a possible implementation, the first pitch and speech rate determination model and the second pitch and speech rate determination model are the same pitch and speech rate determination model, which is trained with reading audio samples of multiple reading voices in multiple reading styles, and each reading audio sample has its corresponding reading voice encoding.
[0185] The input module 820 is configured to input the text pronunciation features of the target text and the reading voice encoding corresponding to the target reading style into the pitch and speech rate determination model, and input the text pronunciation features of the target text and the reading voice encoding corresponding to the target reading voice into the pitch and speech rate determination model.
[0186] In a possible implementation, the first reading style feature includes a first pitch feature and a first speech rate feature, and the second reading style feature includes a second pitch feature and a second speech rate feature.
[0187] The fusion module 830 is configured to generate a fused pitch feature according to the first pitch feature and the second pitch feature, generate a fused speech rate feature according to the first speech rate feature and the second speech rate feature, and input the fused pitch feature, the fused speech rate feature, and the text pronunciation features into a pre-set feature fusion model, and the feature fusion model outputs the fused pronunciation features. Specifically, it can implement the fusion function in step 204 above, as well as other implicit steps.
[0188] In a possible implementation, the first pitch feature includes a plurality of first fundamental frequencies, and the first fundamental frequency is the fundamental frequency of each audio frame when the target text is read in the target reading style. The second pitch feature includes a plurality of second fundamental frequencies, and the second fundamental frequency is the fundamental frequency of each audio frame when the target text is read in the target reading voice.
[0189] The fusion module 830 is configured to determine the average value of the plurality of second fundamental frequencies, determine the difference between each first fundamental frequency and the average value, and form a third pitch feature with the plurality of differences. According to the first weight corresponding to the third pitch feature and the second weight corresponding to the second pitch feature, the third pitch feature and the second pitch feature are weighted and summed to obtain the fused pitch feature, where the first weight is greater than the second weight. Specifically, it can implement the fusion function in step 401 above, as well as other implicit steps.
[0190] In a possible implementation, the first pitch feature includes a plurality of first fundamental frequencies, and the first fundamental frequency is the fundamental frequency of each audio frame when the target text is read in the target reading style. The second pitch feature includes a plurality of second fundamental frequencies, and the second fundamental frequency is the fundamental frequency of each audio frame when the target text is read in the target reading voice.
[0191] The fusion module 830 is used to calculate the average value of multiple first fundamental frequencies to obtain a first average value, calculate the average value of multiple second fundamental frequencies to obtain a second average value, and determine the ratio of the variance of the second pitch feature to the variance of the first pitch feature. Determine the difference between each first fundamental frequency and the first average value respectively, determine the product of each difference and the ratio respectively, determine the sum value of each product and the second average value respectively, and use the sequence composed of multiple sum values as the fused pitch feature. Specifically, it can implement the fusion function in step 401 above, as well as other implicit steps.
[0192] In a possible implementation manner, the first speech rate feature includes the first audio duration of each phoneme when the target text is read in the target reading style, and the second speech rate feature includes the second audio duration of each phoneme when the target text is read in the target reading timbre.
[0193] The fusion module 830 is used to perform weighted summation on the first speech rate feature and the second speech rate feature according to the weight corresponding to the first speech rate feature and the weight corresponding to the second speech rate feature to obtain a fused speech rate feature, where the weight corresponding to the first speech rate feature is greater than the weight corresponding to the second speech rate feature. Specifically, it can implement the fusion function in step 402 above, as well as other implicit steps.
[0194] In a possible implementation manner, the fusion module 830 is used to determine the average value and variance of multiple fused fundamental frequencies included in the fused pitch feature. Determine the difference between each fused fundamental frequency and the average value respectively, determine the value obtained by multiplying each difference by a preset control parameter and dividing by the variance respectively, determine the sum value of each value and the average value respectively, and use the sequence composed of multiple sum values as the adjusted fused pitch feature. Input the adjusted fused pitch feature, the fused speech rate feature, and the text pronunciation feature into a pre-set feature fusion model. Specifically, it can implement the fusion functions in steps 501 - 503 above, as well as other implicit steps.
[0195] In a possible implementation manner, the determination module 810 is further used to determine the phoneme sequence corresponding to the target text. Input the phoneme sequence into a pre-trained encoder, and the encoder outputs the text pronunciation feature of the target text. Specifically, it can implement the determination function in step 201 above, as well as other implicit steps.
[0196] In a possible implementation manner, the spectral synthesis model includes a decoder and a vocoder. The synthesis module 840 is used to input the fused pronunciation feature and the timbre information of the target reading timbre into the decoder, and the decoder outputs Mel spectrum features. Input the Mel spectrum features into the vocoder, and the vocoder outputs the reading audio corresponding to the target text. Specifically, it can implement the synthesis function in step 205 above, as well as other implicit steps.
[0197] In the embodiments of the present application, the user can select the target pronunciation timbre, target pronunciation style, and target text of the recited audio, and then use the pitch and speech rate determination model, feature fusion model, and spectral synthesis model to synthesize the recited audio. The synthesized recited audio incorporates the pronunciation style features of the target pronunciation style and also incorporates the timbre information of the target pronunciation timbre. During the synthesis process, the target pronunciation timbre and target pronunciation style can be arbitrarily selected and combined. It can be seen that the embodiments of the present application provide a solution that can arbitrarily select both the timbre and style to synthesize the recited audio.
[0198] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0199] It should be noted that when the voice synthesis device provided in the above embodiments synthesizes voice, only the above-mentioned division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the voice synthesis device provided in the above embodiments and the method embodiments of voice synthesis belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be elaborated here.
[0200] In an exemplary embodiment, a computer-readable storage medium is further provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the voice synthesis method in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), random access memory (RAM), compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0201] In an exemplary embodiment, a computer program product is further provided. The computer program product includes at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the voice synthesis method.
[0202] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0203] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0204] It should be noted that the information involved in the present application (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) are all authorized by users or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions. For example, the benchmark reading audio involved in the present application is obtained under full authorization.
Claims
1. A method for speech synthesis, characterized in that, the method comprises: determining a target text, and determining a target pronunciation timbre and a target pronunciation style of the target text; inputting the text pronunciation features of the target text and the pronunciation timbre codes corresponding to the target pronunciation style into a pitch and speech rate determination model, and outputting, by the pitch and speech rate determination model, a first pronunciation style feature when the target text is read in the target pronunciation style, where the pitch and speech rate determination model is trained with reading audio samples of multiple pronunciation timbres in multiple pronunciation styles, and each reading audio sample has its corresponding pronunciation timbre code; inputting the text pronunciation features of the target text and the pronunciation timbre codes corresponding to the target pronunciation timbre into the pitch and speech rate determination model, and outputting, by the pitch and speech rate determination model, a second pronunciation style feature when the target text is read in the target pronunciation timbre; inputting the text pronunciation features, the first pronunciation style feature, and the second pronunciation style feature of the target text into a pre-set feature fusion model, and outputting, by the feature fusion model, a fused pronunciation feature; inputting the fused pronunciation feature and the timbre information of the target pronunciation timbre into a pre-trained spectral synthesis model, and outputting, by the spectral synthesis model, a reading audio corresponding to the target text.
2. The method according to claim 1, characterized in that, the first pronunciation style feature includes a first pitch feature and a first speech rate feature; the second pronunciation style feature includes a second pitch feature and a second speech rate feature; the step of inputting the text pronunciation features, the first pronunciation style feature, and the second pronunciation style feature of the target text into a pre-set feature fusion model, and outputting, by the feature fusion model, a fused pronunciation feature includes: generating a fused pitch feature according to the first pitch feature and the second pitch feature; generating a fused speech rate feature according to the first speech rate feature and the second speech rate feature; inputting the fused pitch feature, the fused speech rate feature, and the text pronunciation features into a pre-set feature fusion model, and outputting, by the feature fusion model, a fused pronunciation feature.
3. The method according to claim 2, characterized in that, the first pitch feature includes a plurality of first fundamental frequencies, the first fundamental frequencies being the fundamental frequencies of each audio frame when the target text is read in the target pronunciation style, the second pitch feature includes a plurality of second fundamental frequencies, the second fundamental frequencies being the fundamental frequencies of each audio frame when the target text is read in the target pronunciation timbre; the step of generating a fused pitch feature according to the first pitch feature and the second pitch feature includes: determining an average value of the plurality of second fundamental frequencies, determining a difference between each first fundamental frequency and the average value, and forming a third pitch feature with the plurality of differences; performing weighted summation on the third pitch feature and the second pitch feature according to a first weight corresponding to the third pitch feature and a second weight corresponding to the second pitch feature to obtain a fused pitch feature, where the first weight is greater than the second weight.
4. The method according to claim 2, It is characterized in that the first pitch feature includes a plurality of first fundamental frequencies, which are the fundamental frequencies of each audio frame when the target text is read in the target reading style, and the second pitch feature includes a plurality of second fundamental frequencies, which are the fundamental frequencies of each audio frame when the target text is read in the target reading voice color; generating a fused pitch feature according to the first pitch feature and the second pitch feature includes: calculating an average value of the plurality of first fundamental frequencies to obtain a first average value, calculating an average value of the plurality of second fundamental frequencies to obtain a second average value, and determining a ratio of the variance of the second pitch feature to the variance of the first pitch feature; determining the difference between each of the first fundamental frequencies and the first average value, determining the product of each of the differences and the ratio, determining the sum value of each of the products and the second average value, and using the sequence composed of the plurality of sum values as the fused pitch feature.
5. The method according to claim 2, It is characterized in that the first speech rate feature includes the first audio duration of each phoneme when the target text is read in the target reading style, and the second speech rate feature includes the second audio duration of each phoneme when the target text is read in the target reading voice color; generating a fused speech rate feature according to the first speech rate feature and the second speech rate feature includes: performing weighted summation on the first speech rate feature and the second speech rate feature according to the weight corresponding to the first speech rate feature and the weight corresponding to the second speech rate feature to obtain a fused speech rate feature, wherein the weight corresponding to the first speech rate feature is greater than the weight corresponding to the second speech rate feature.
6. The method according to claim 2, It is characterized in that inputting the fused pitch feature, the fused speech rate feature, and the text pronunciation feature into a pre-set feature fusion model includes: determining the average value and variance of the plurality of fused fundamental frequencies included in the fused pitch feature; determining the difference between each of the fused fundamental frequencies and the average value, determining the value obtained by multiplying each of the differences by a preset control parameter and dividing by the variance, determining the sum value of each of the values and the average value, and using the sequence composed of the plurality of sum values as the adjusted fused pitch feature; inputting the adjusted fused pitch feature, the fused speech rate feature, and the text pronunciation feature into a pre-set feature fusion model.
7. The method according to any one of claims 1-6, It is characterized in that the spectrum synthesis model includes a decoder and a vocoder; inputting the fused pronunciation feature and the timbre information of the target reading voice color into a pre-trained spectrum synthesis model, and outputting the reading audio corresponding to the target text by the spectrum synthesis model includes: inputting the fused pronunciation feature and the timbre information of the target reading voice color into the decoder, and outputting Mel spectrum features by the decoder; inputting the Mel spectrum features into the vocoder, and outputting the reading audio corresponding to the target text by the vocoder.
8. A computer device, It is characterized in that The computer device includes a memory and a processor, where the memory is used to store computer instructions; the processor executes the computer instructions stored in the memory so that the computer device executes the method described in any one of claims 1-7 above.
9. A computer-readable storage medium, characterized in that the computer-readable storage medium stores computer program code, and in response to the computer program code being executed by a computer device, the computer device executes the method described in any one of claims 1-7 above.
Citation Information
Patent Citations
Voice style migration method and device, readable medium and electronic equipment
CN112927674A