Artificial intelligence-based speech synthesis method, device, computer equipment and medium
By extracting and fusing features of the target user's reference speech spectrum and phonemes, the speech synthesis result for the target user is generated, which solves the problem of large timbre differences under zero speech samples or lightweight speech samples and optimizes the speech synthesis effect.
Patent Information
- Application Number
- CN202210816256.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-12
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-07-12
AI Technical Summary
In the case of zero or lightweight voice samples, the timbre of the synthesized voice is quite different from that of the target user, resulting in poor speech synthesis effect.
By obtaining the target user's reference speech spectrum and target speech phonemes, the trained spectrum encoder, phoneme encoder, recognition encoder, user representation predictor and spectrum decoder are used to extract and fuse features to generate the target user's speech synthesis result and reduce timbre differences.
The speech synthesis effect has been optimized, reducing the difference between the timbre of the synthesized speech and the user's own voice, and improving the quality of speech synthesis.
Smart Images

Figure CN115019769B_ABST
Abstract
Description
Technical Field
[0001] The present invention is applicable to the field of speech synthesis technology, and in particular relates to a speech synthesis method, apparatus, computer equipment and medium based on artificial intelligence. Background Art
[0002] Speech synthesis is a technology that converts text information into speech information, that is, converting text information into any audible speech, involving multiple disciplines such as acoustics, linguistics and computer science.
[0003] With the development of deep learning technology, most of the current mainstream end-to-end speech synthesis systems use the attention mechanism to implicitly learn the alignment relationship between text and speech. At the same time, they adopt an autoregressive speech generation model, requiring the generation of the subsequent speech frame to use the previous speech frame as input, and have strong front-end dependency and temporal sequence for the speech frames. Therefore, they have high requirements for the data volume and quality of speech samples. When speech synthesis is based on zero speech samples or lightweight speech samples, the timbre of the synthesized speech is quite different from the user's own timbre, resulting in poor speech synthesis effect.
[0004] Therefore, in the field of speech synthesis technology, how to reduce the difference between the timbre of the synthesized speech and the timbre of the target user with zero speech samples or lightweight speech samples and improve the speech synthesis effect has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, the embodiments of the present invention provide an artificial intelligence-based speech synthesis method, apparatus, computer equipment and medium to solve the problem in the prior art that the timbre of the synthesized speech is significantly different from the timbre of the target user when there are zero speech samples or lightweight speech samples.
[0006] In a first aspect, an embodiment of the present invention provides a speech synthesis method, the speech synthesis method comprising:
[0007] Obtaining a reference speech spectrum and target speech phonemes of a target user, and processing the reference speech spectrum and the target speech phonemes based on a trained speech synthesis model, wherein the speech synthesis model includes a trained spectrum encoder, a trained phoneme encoder, a trained recognition encoder, a trained user representation predictor, and a trained spectrum decoder; the processing includes:
[0008] Inputting the reference speech spectrum into the trained spectrum encoder to obtain reference timbre content features, and inputting the target speech phonemes into the trained phoneme encoder to obtain target content features;
[0009] Inputting the reference timbre content feature and the target content feature into the trained recognition encoder to obtain the target timbre content feature;
[0010] Sampling the target timbre content feature, and inputting the sampling result into the trained user representation predictor to obtain the user identity content feature;
[0011] The target timbre content feature and the user identity content feature are subjected to feature fusion, and the obtained fusion feature is input into the trained spectrum decoder to obtain a speech synthesis result of the target user.
[0012] In a second aspect, an embodiment of the present invention provides a speech synthesis device, the speech synthesis device comprising:
[0013] Data acquisition module: used to obtain the reference speech spectrum and target speech phonemes of the target user;
[0014] a spectrum encoder, configured to input the reference speech spectrum and output reference timbre content features;
[0015] A phoneme encoder, configured to input the target speech phonemes and output target content features;
[0016] an identification encoder, configured to input the reference timbre content feature and the target content feature, and output the target timbre content feature;
[0017] A user characterization predictor, configured to sample the target timbre content feature and output a user identity content feature based on the sampled result;
[0018] Spectrum decoder: used to fuse the target timbre content features and the user identity content features, and output the speech synthesis result of the target user based on the obtained fusion features.
[0019] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the speech synthesis method as described in the first aspect when executing the computer program.
[0020] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the speech synthesis method as described in the first aspect is implemented.
[0021] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: by obtaining the reference speech spectrum and target speech phonemes of the target user, inputting the reference speech spectrum into a trained spectrum encoder to obtain reference timbre content features, inputting the target speech phonemes into a trained phoneme encoder to obtain target content features, and then inputting the reference timbre content features and the target content features into a trained recognition encoder to obtain target timbre content features, by sampling the target timbre content features, inputting the sampling results into a trained user representation predictor to obtain user identity content features, and then performing feature fusion of the target timbre content features and the user identity content features, and inputting the fused features into a trained spectrum decoder to obtain the speech synthesis result of the target user, and characterizing the timbre and target content of the target user by using the fused features obtained by fusing the one-to-one corresponding target timbre content features and the user identity content features to obtain the speech synthesis result of the target user, thereby reducing the difference between the timbre of the synthesized speech and the timbre of the user itself and optimizing the speech synthesis effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0023] Figure 1 This is a schematic diagram of an application environment of a speech synthesis method provided in Example 1 of the present invention;
[0024] Figure 2 This is a flow chart of a speech synthesis method provided in Example 1 of the present invention;
[0025] Figure 3 This is a structural diagram of a speech synthesis device provided in Embodiment 2 of the present invention;
[0026] Figure 4 This is a structural diagram of a computer device provided in Example 3 of the present invention. DETAILED DESCRIPTION
[0027] In the following description, specific details such as particular system structures and techniques are provided for purposes of illustration, not limitation, to facilitate a thorough understanding of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present invention with unnecessary detail.
[0028] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0029] It will also be understood that the term "and / or" used in the present description and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0030] As used in the present specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0031] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0032] References to "one embodiment" or "some embodiments" in the present specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present invention. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0033] Embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.
[0034] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0035] It should be understood that the order of execution of the steps in the following embodiments does not necessarily mean the order in which they are executed. The order in which each process is executed should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0036] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.
[0037] A speech synthesis method provided in the first embodiment of the present invention can be applied in the following situations: Figure 1 In an application environment, a client communicates with a server. Clients include, but are not limited to, PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, and personal digital assistants (PDAs). The server can be implemented as a standalone server or a server cluster consisting of multiple servers.
[0038] See also Figure 2 , is a flow chart of a speech synthesis method provided by the first embodiment of the present invention, the above-mentioned speech synthesis method can be applied to Figure 1 In the client, the speech synthesis method may include the following steps:
[0039] Step S201: Obtain a reference speech spectrum and target speech phonemes of a target user, and process the reference speech spectrum and target speech phonemes based on a trained speech synthesis model, wherein the speech synthesis model includes a trained spectrum encoder, a trained phoneme encoder, a trained recognition encoder, a trained user representation predictor, and a trained spectrum decoder.
[0040] Among them, the reference speech spectrum of the target user can be obtained based on the target user's actual pronunciation of the reference text. The reference text can be a sentence or a word to ensure a lightweight speech sample or a near-zero speech sample. The reference speech spectrum contains both the timbre information of the target user and the content information of the reference text.
[0041] The target speech phonemes can be obtained by searching the target speech text according to the pronunciation dictionary. The target speech text can be a sentence, a word, or a paragraph, so that the speech spectrum generated by the target user contains the content information of the target speech phonemes.
[0042] The encoders used in the speech synthesis model are all pre-trained encoders, including a trained spectrum encoder, a trained phoneme encoder, a trained recognition encoder, a trained user representation predictor, and a trained spectrum decoder. They are used to extract and process features of the reference speech spectrum and target speech phonemes to obtain the speech synthesis results for the target user.
[0043] Optionally, during speech synthesis model training, a temporary embedding layer and a pre-trained temporary encoder are added to the speech synthesis model, with the sample speech spectrum, sample speech phonemes, and sample user ID of the sample user as training samples, and the real sample spectrum as training labels;
[0044] The training process of the speech synthesis model includes:
[0045] Input the sample speech spectrum into the spectrum encoder for feature extraction to obtain the sample spectrum features;
[0046] Input the sample speech phonemes into the phoneme encoder for feature extraction to obtain the sample phoneme features;
[0047] Input the sample user number into the temporary embedding layer for feature extraction to obtain the sample embedding vector;
[0048] Fusing the sample spectrum feature and the sample phoneme feature, and inputting the obtained first fusion feature into the recognition encoder to obtain the sample timbre content feature;
[0049] Perform feature fusion on the sample embedding vector and the sample phoneme feature, and input the obtained second fusion feature into the pre-trained temporary encoder to obtain the sample identity content feature;
[0050] Perform Gaussian sampling on the sample timbre content features, input the sampling results into the user representation predictor, and obtain the sample prediction number;
[0051] Multiply the sampling result by the sample user number to obtain the fused sample spectrum feature;
[0052] The fused sample spectrum features are input into the spectrum decoder to obtain the predicted speech spectrum.
[0053] Among them, the training samples of the speech synthesis model include the sample user's sample speech spectrum, sample speech phonemes and sample user number, and the training label is the real sample spectrum. Among them, the training set contains a large number of sample users. The sample speech spectrum can be obtained based on the sample user's actual pronunciation of the sample reference text. The sample reference text can be a sentence or a word. The sample speech spectrum contains both the sample user's timbre information and the content information of the sample reference text.
[0054] The sample speech phonemes can be obtained by searching the sample speech text into phonemes according to the pronunciation dictionary. The sample speech text can be a sentence, a word, or a paragraph to indicate that the real sample spectrum of the sample user contains content information of the sample speech phonemes.
[0055] The sample user number is the identity authentication information of the sample user, and the sample user number, the sample user timbre and the sample user have a one-to-one correspondence.
[0056] The real sample spectrum is obtained based on the target user's actual pronunciation of the sample speech text, and includes both the timbre information of the sample user and the content information of the sample speech text.
[0057] When training the speech synthesis model, a temporary embedding layer and a pre-trained temporary encoder are added to the speech synthesis model. Therefore, the speech synthesis model during training includes a spectrum encoder, a phoneme encoder, a temporary embedding layer, a recognition encoder, a temporary encoder, a user representation predictor, and a spectrum decoder. The specific training process of the speech synthesis model is as follows:
[0058] First, the sample speech spectrum is input into the spectrum encoder for feature extraction to obtain sample spectrum features, which are used to represent the timbre information of the sample user and the content information of the sample reference text. The sample speech phonemes are input into the phoneme encoder for feature extraction to obtain sample phoneme features, which are used to represent the content information of the sample speech phonemes. The sample user number is input into the temporary embedding layer for feature extraction to obtain a sample embedding vector, which is used to represent the identity authentication information of the sample user, wherein the timbre information and identity authentication information of the same sample user correspond one to one.
[0059] Secondly, the sample spectrum features and the sample phoneme features are feature fused to obtain a first fused feature, and the first fused feature is input into the recognition encoder to obtain a sample timbre content feature, which is used to represent the timbre information of the sample user, the content information of the sample reference text, and the content information of the sample speech phonemes. The sample embedding vector and the sample phoneme features are feature fused to obtain a second fused feature, and the second fused feature is input into a pre-trained temporary encoder to obtain a sample identity content feature, which is used to represent the identity authentication information of the sample user and the content information of the sample speech phonemes.
[0060] Again, Gaussian sampling is performed on the sample timbre content features, and the sampling results are input into the user characterization predictor to obtain a sample prediction number, which is used to characterize the identity authentication information of the predicted sample user. At the same time, the sampling result is multiplied by the sample user number to obtain a fused sample spectrum feature, which is used to obtain a fused sample spectrum feature based on the fusion of the sample user's timbre information and the sample user's identity authentication information, thereby improving the representation of the sample user's timbre.
[0061] Finally, the fused sample spectrum features are input into the spectrum decoder to obtain the predicted speech spectrum, which is used to characterize the timbre information of the predicted sample user and the content information of the predicted sample speech text, so as to generate the predicted speech spectrum through the fused sample spectrum features with a stronger representation degree, reduce the difference between the timbre of the generated predicted speech spectrum and the timbre of the real sample spectrum of the sample user, and optimize the speech synthesis effect of the speech synthesis model.
[0062] During the training process of the speech synthesis model, in order to ensure the quality of model training, a loss function is calculated based on the obtained sample timbre content features, sample identity content features, sample prediction number, sample user number, predicted speech spectrum and actual sample spectrum during the training process. Based on the loss function, the parameters of the spectrum encoder, phoneme encoder, recognition encoder, user representation predictor and spectrum decoder are updated by the gradient descent method to obtain a trained spectrum encoder, a trained phoneme encoder, a trained recognition encoder, a trained user representation predictor and a trained spectrum decoder, thereby achieving the purpose of generating high-quality speech synthesis results for the target user based on lightweight speech samples or near-zero speech samples.
[0063] Optionally, the loss function includes:
[0064] Relative entropy is calculated based on the sample timbre content characteristics and the sample identity content characteristics. Based on the relative entropy, the parameters of the spectrum encoder, phoneme encoder and recognition encoder are updated by the gradient descent method until the relative entropy converges, and the pre-trained spectrum encoder, pre-trained phoneme encoder and pre-trained recognition encoder are obtained.
[0065] Among them, the sample timbre content features exist in the form of feature distribution, representing the timbre information of the sample user, the content information of the sample reference text, and the content information of the sample speech phonemes. The sample identity content features also exist in the form of feature distribution, representing the identity authentication information of the sample user and the content information of the sample speech phonemes, and the timbre information of the sample user corresponds one-to-one to the identity authentication information of the sample user.
[0066] Therefore, the relative entropy between the sample timbre content features and the sample identity content features is calculated, and the relative entropy is used to characterize the similarity between the sample timbre content features and the sample identity content features. The smaller the relative entropy, the closer the sample timbre content features and the sample identity content features are, indicating that the sample timbre content features have a smaller representation of the content information of the sample reference text. The larger the relative entropy, the greater the difference between the sample timbre content features and the sample identity content features, indicating that the sample timbre content features have a greater representation of the content information of the sample reference text.
[0067] After updating the parameters of the spectrum encoder, phoneme encoder and recognition encoder by the gradient descent method, the relative entropy between the sample timbre content features and the sample identity content features is recalculated until the relative entropy converges, so that the sample timbre content features are as close as possible to the sample identity content features, so as to ensure that the spectrum encoder can filter out the content information of the sample reference text and learn the timbre information of the sample user as much as possible, so as to improve the spectrum encoder's representation of the sample user's timbre information.
[0068] In one embodiment, the sample timbre content feature is recorded as P(x), the sample identity content feature is recorded as Q(x), and the relative entropy between the sample timbre content feature P(x) and the sample identity content feature Q(x) is recorded as D KL (P‖Q), then:
[0069] D KL (P∥Q)=E[logP(x)-logQ(x)]
[0070] where P(x) is the sample timbre content feature, Q(x) is the sample identity content feature, and E[logP(x)-logQ(x)] is the expected logarithmic difference between the sample timbre content feature P(x) and the sample identity content feature Q(x).
[0071] Relative entropy DKL The smaller (P‖Q), the closer the sample timbre content feature P(x) is to the sample identity content feature Q(x), indicating that the sample timbre content feature P(x) has a smaller representation degree of the content information of the sample reference text. Then, after updating the parameters of the spectrum encoder, phoneme encoder, and recognition encoder by the gradient descent method, the sample timbre content feature after parameter update is obtained, which is recorded as P(x)′ and the sample identity content feature Q(x)′. The relative entropy D between the sample timbre content feature P(x)′ and the sample identity content feature Q(x)′ is recalculated. KL (P‖Q)′, until the relative entropy converges, to ensure that the spectrum encoder can filter out the content information of the sample reference text and learn the timbre information of the sample user as much as possible, so as to improve the spectrum encoder's representation of the sample user's timbre information.
[0072] Optionally, the loss function includes:
[0073] The first mean square error loss is calculated based on the prediction number and the sample user number. Based on the first mean square error loss, the parameters of the user representation predictor are updated by the gradient descent method until the first mean square error loss converges to obtain the pre-trained user representation prediction.
[0074] Among them, the predicted number represents the identity authentication information of the predicted sample user, the sample user number is the identity authentication information of the sample user, and the first mean square error loss between the predicted number and the sample user number is calculated. The first mean square error loss is used to represent the similarity between the predicted number and the sample user number. The smaller the first mean square error loss, the closer the predicted number and the sample user number are, which means that the user representation predictor has a better representation effect on the sample timbre content characteristics. The larger the first mean square error loss, the greater the difference between the predicted number and the sample user number, which means that the user representation predictor has a worse representation effect on the sample timbre content characteristics.
[0075] Therefore, after updating the parameters of the user representation predictor by the gradient descent method, the first mean square error loss between the predicted number and the sample user number is recalculated until the first mean square error loss converges, so that the predicted number is as close as possible to the sample user number, thereby improving the degree of representation of the sample timbre content characteristics by the user representation predictor.
[0076] In one embodiment, the prediction number is recorded as Y, the sample user number is recorded as B, and the first mean square error loss between the prediction number Y and the sample user number B is recorded as S1. The first mean square error loss S1 is:
[0077]
[0078] Where Y is the predicted number and B is the sample user number.
[0079] The smaller the first mean square error loss S1 between the prediction number Y and the sample user number B, the better the representation effect of the user representation predictor on the sample timbre content characteristics. After updating the parameters of the user representation predictor by the gradient descent method, the prediction number Y′ and sample user number B′ after parameter update are obtained, and the first mean square error loss S1′ between the updated prediction number Y′ and sample user number B′ is recalculated until the first mean square error loss converges to improve the representation degree of the sample timbre content characteristics of the user representation predictor.
[0080] Optionally, the loss function includes:
[0081] The second mean square error loss is calculated based on the predicted speech spectrum and the real sample spectrum. Based on the second mean square error loss, the parameters of the spectrum decoder are updated by the gradient descent method until the second mean square error loss converges to obtain a pre-trained spectrum decoder.
[0082] Among them, the predicted speech spectrum represents the timbre information of the predicted sample user and the content information of the predicted sample speech text, the real sample spectrum is the timbre information of the sample user and the content information of the sample speech text, and the second mean square error loss between the predicted speech spectrum and the real sample spectrum is calculated. The second mean square error loss is used to represent the similarity between the predicted speech spectrum and the real sample spectrum. The smaller the second mean square error loss, the closer the predicted speech spectrum and the real sample spectrum are, which means that the spectrum decoder has a better representation effect on the fusion features of the sample timbre content features and the sample user number. The larger the second mean square error loss, the greater the difference between the predicted speech spectrum and the real sample spectrum, which means that the spectrum decoder has a worse representation effect on the fusion features of the sample timbre content features and the sample user number.
[0083] Therefore, after updating the parameters of the spectrum decoder by the gradient descent method, the second mean square error loss between the predicted speech spectrum and the true sample spectrum is recalculated until the second mean square error loss converges, so that the predicted speech spectrum is as close as possible to the true sample spectrum, so as to improve the spectrum decoder's representation of the sample timbre content features and the content information of the sample speech phonemes, reduce the difference between the timbre of the synthesized speech and the user's own timbre, and optimize the speech synthesis effect of the speech synthesis model.
[0084] In one embodiment, the predicted speech spectrum is recorded as J, the real sample spectrum is recorded as Z, and the second mean square error loss between the predicted speech spectrum J and the real sample spectrum Z is recorded as S2. The second mean square error loss S2 is:
[0085]
[0086] Where J is the predicted speech spectrum and Z is the true sample spectrum.
[0087] The smaller the second mean square error loss S2 between the predicted speech spectrum J and the true sample spectrum Z, the closer the predicted speech spectrum J and the true sample spectrum Z are, indicating that the spectrum decoder has a better representation effect on the fusion features of the sample timbre content features and the sample user number. After updating the parameters of the spectrum decoder by the gradient descent method, the predicted speech spectrum J′ and the true sample spectrum Z′ after parameter update are obtained, and the second mean square error loss S2′ between the updated predicted speech spectrum J′ and the true sample spectrum Z′ is recalculated until the second mean square error loss converges, so as to improve the representation degree of the spectrum decoder on the sample timbre content features and the content information of the sample speech phonemes, reduce the difference between the timbre of the synthesized speech and the timbre of the user itself, and optimize the speech synthesis effect of the speech synthesis model.
[0088] According to the training process of the above-mentioned speech synthesis model, the training of the spectrum encoder, phoneme encoder, recognition encoder, user representation predictor and spectrum decoder is completed to obtain a trained spectrum encoder, a trained phoneme encoder, a trained recognition encoder, a trained user representation predictor and a trained spectrum decoder, which are used to generate speech synthesis results for the target user based on lightweight speech samples or near-zero speech samples.
[0089] Step S202: Input the reference speech spectrum into a trained spectrum encoder to obtain reference timbre content features, and input the target speech phonemes into a trained phoneme encoder to obtain target content features.
[0090] Among them, the reference speech spectrum can be obtained based on the target user's actual pronunciation of the reference text. The reference text can be a sentence or a word to ensure lightweight speech samples or close to zero speech samples. A trained spectrum encoder in the speech synthesis model is obtained, and the reference speech spectrum is input into the trained spectrum encoder to obtain a reference timbre content feature. The reference timbre content feature is used to represent the timbre information of the target user and the content information of the reference text.
[0091] The target speech phonemes can be obtained by completing the text-to-phoneme search according to the pronunciation dictionary based on the target speech text. The target speech text can be a sentence, a word, or a paragraph. The trained phoneme encoder in the speech synthesis model is obtained, and the target speech phonemes are input into the trained phoneme encoder to obtain the target content features. The target content features are used to represent the content information corresponding to the target speech factors.
[0092] Step S203: Input the reference timbre content feature and the target content feature into the trained recognition encoder to obtain the target timbre content feature.
[0093] Among them, the reference timbre content feature represents the timbre information of the target user and the content information of the reference text, the target content feature represents the content information corresponding to the target speech factor, and a trained recognition encoder in the speech synthesis model is obtained. The reference timbre content feature and the target content feature are input into the trained recognition encoder together to obtain the target timbre content feature. The trained recognition encoder has filtered out the content information of the reference text in the target timbre content feature. Therefore, the target timbre content feature represents the timbre information of the target user and the content information of the sample speech phonemes.
[0094] Step S204: sampling the target timbre content features, and inputting the sampling results into the trained user representation predictor to obtain the user identity content features.
[0095] Among them, the target timbre content feature exists in the form of feature distribution, which represents the timbre information of the sample user and the content information of the sample speech phonemes. The target timbre content feature is sampled to obtain the sampling result, and the trained user representation predictor in the speech synthesis model is obtained. The obtained sampling result is input into the trained user representation predictor to obtain the user identity content feature, and the user identity content feature is the same as the target timbre content feature, and can both represent the timbre information of the target user and the content information of the sample speech phonemes.
[0096] Step S205 , performing feature fusion on the target timbre content features and the user identity content features, and inputting the obtained fusion features into the trained spectrum decoder to obtain the speech synthesis result of the target user.
[0097] Among them, the target timbre content feature and the user identity content feature both represent the timbre information of the target user and the content information of the sample speech phonemes. The target timbre content feature and the user identity content feature are fused to obtain a fusion feature. The fusion feature is obtained by fusion of features of different representation forms, and has a stronger representation degree of the target user's timbre information and the content information of the sample speech phonemes.
[0098] Obtain the trained spectrum decoder in the speech synthesis model, input the fusion feature into the trained spectrum decoder, and obtain the speech synthesis result of the target user.
[0099] The embodiment of the present invention obtains a reference speech spectrum and target speech phonemes of a target user, inputs the reference speech spectrum into a trained spectrum encoder to obtain a reference timbre content feature, inputs the target speech phonemes into a trained phoneme encoder to obtain a target content feature, then inputs the reference timbre content feature and the target content feature into a trained recognition encoder to obtain a target timbre content feature, samples the target timbre content feature, inputs the sampling result into a trained user representation predictor to obtain a user identity content feature, then fuses the target timbre content feature and the user identity content feature, inputs the fused feature into a trained spectrum decoder to obtain a speech synthesis result for the target user, characterizes the timbre and target content of the target user by using the fused feature obtained by fusing the one-to-one corresponding target timbre content feature and the user identity content feature to obtain a speech synthesis result for the target user, thereby reducing the difference between the timbre of the synthesized speech and the timbre of the user itself and optimizing the speech synthesis effect.
[0100] Corresponding to the speech synthesis method of the above embodiment, Figure 3 A structural block diagram of a speech synthesis device provided in a second embodiment of the present invention is given. For ease of description, only the parts related to the embodiment of the present invention are shown.
[0101] See also Figure 3 , the speech synthesis device comprises:
[0102] Data acquisition module 31: used to obtain the reference speech spectrum and target speech phonemes of the target user;
[0103] The spectrum encoder 32 is used to input a reference speech spectrum and output a reference timbre content feature;
[0104] The phoneme encoder 33 is used to input the target speech phonemes and output the target content features;
[0105] The recognition encoder 34 is used to input the reference timbre content feature and the target content feature and output the target timbre content feature;
[0106] A user characterization predictor 35 is used to sample the target timbre content feature and output the user identity content feature based on the obtained sampling result;
[0107] Spectrum decoder 36: used to fuse the target timbre content features and the user identity content features, and output the speech synthesis result of the target user according to the obtained fusion features.
[0108] Optionally, the speech synthesis device further includes:
[0109] The spectrum encoder 32 is used to input the sample speech spectrum and output the sample spectrum features during model training;
[0110] The phoneme encoder 33 is used to input sample speech phonemes and output sample phoneme features during model training;
[0111] Temporary embedding layer, used to input sample user ID and output sample embedding vector during model training;
[0112] The recognition encoder 34 is used to fuse the sample spectrum features and the sample phoneme features during model training, and output the sample timbre content features based on the obtained first fusion features;
[0113] A temporary encoder is used to fuse the sample embedding vector and the sample phoneme feature during model training, and output the sample identity content feature based on the obtained second fused feature;
[0114] The user representation predictor 35 is used to sample the sample timbre content features during model training and output a sample prediction number based on the obtained sampling results;
[0115] The feature fusion module is used to multiply the sampling result by the sample user number during model training to obtain the fused sample spectrum feature;
[0116] The spectrum decoder 36 is used to input the fused sample spectrum features and output the predicted speech spectrum during model training.
[0117] The parameter update module is used to calculate the loss function according to the sample timbre content features, sample identity content features, sample prediction number, sample user number, predicted speech spectrum and true sample spectrum during model training, and update the parameters of the spectrum encoder, phoneme encoder, recognition encoder, user representation predictor and spectrum decoder through the gradient descent method based on the loss function.
[0118] Optionally, the parameter updating module includes:
[0119] A relative entropy calculation submodule is used to calculate the parameters of the spectrum encoder, phoneme encoder and recognition encoder based on the sample timbre content characteristics and the sample identity content characteristics, and update the parameters of the spectrum encoder, phoneme encoder and recognition encoder by gradient descent method based on the relative entropy until the relative entropy converges, thereby obtaining a pre-trained spectrum encoder, a pre-trained phoneme encoder and a pre-trained recognition encoder;
[0120] The first mean square error loss submodule is used to calculate the predicted number and the sample user number, and update the parameters of the user representation predictor by gradient descent based on the first mean square error loss until the first mean square error loss converges to obtain the pre-trained user representation prediction;
[0121] The second mean square error loss submodule is used to calculate the predicted speech spectrum and the real sample spectrum. Based on the second mean square error loss, the parameters of the spectrum decoder are updated by the gradient descent method until the second mean square error loss converges to obtain a pre-trained spectrum decoder.
[0122] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0123] Figure 4 This is a schematic diagram of the structure of a computer device provided in the third embodiment of the present invention. Figure 4 As shown, the computer device of this embodiment includes: at least one processor ( Figure 4 Only one is shown), a memory, and a computer program stored in the memory and executable on at least one processor, wherein when the processor executes the computer program, the steps in any of the above-mentioned embodiments of the speech synthesis method are implemented.
[0124] The computer device may include, but is not limited to, a processor and a memory. It will be understood by those skilled in the art that Figure 4 The above is merely an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include a network interface, a display screen, and an input device.
[0125] The processor may be a CPU, or other general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. A general-purpose processor may be a microprocessor, or any conventional processor.
[0126] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be the hard disk of the computer device, and in other embodiments, it can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Furthermore, the memory can also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, boot loaders (BootLoader), data, and other programs, such as the program code of the computer program. The memory can also be used to temporarily store data that has been output or is about to be output.
[0127] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned method embodiment. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include at least: any entity or device capable of carrying computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0128] The present invention may implement all or part of the processes in the above-mentioned method embodiments, and may also be completed through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0129] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0130] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0131] In the embodiments provided by the present invention, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0132] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0133] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A speech synthesis method, characterized in that: The speech synthesis method comprises: Obtaining a reference speech spectrum and target speech phonemes of a target user, and processing the reference speech spectrum and the target speech phonemes based on a trained speech synthesis model, wherein the speech synthesis model includes a trained spectrum encoder, a trained phoneme encoder, a trained recognition encoder, a trained user representation predictor, and a trained spectrum decoder; the processing includes: Inputting the reference speech spectrum into the trained spectrum encoder to obtain reference timbre content features, and inputting the target speech phonemes into the trained phoneme encoder to obtain target content features; Inputting the reference timbre content feature and the target content feature into the trained recognition encoder to obtain the target timbre content feature; Sampling the target timbre content feature, and inputting the sampling result into the trained user representation predictor to obtain the user identity content feature; The target timbre content feature and the user identity content feature are subjected to feature fusion, and the obtained fusion feature is input into the trained spectrum decoder to obtain a speech synthesis result of the target user.
2. The speech synthesis method according to claim 1, wherein: When training the speech synthesis model, a temporary embedding layer and a pre-trained temporary encoder are added to the speech synthesis model, a sample speech spectrum, a sample speech phoneme, and a sample user ID of a sample user are used as training samples, and a real sample spectrum is used as a training label; The training process of the speech synthesis model includes: Inputting the sample speech spectrum into the spectrum encoder for feature extraction to obtain sample spectrum features; Inputting the sample speech phonemes into the phoneme encoder for feature extraction to obtain sample phoneme features; Inputting the sample user number into the temporary embedding layer for feature extraction to obtain a sample embedding vector; Performing feature fusion on the sample spectrum feature and the sample phoneme feature, and inputting the obtained first fusion feature into the recognition encoder to obtain the sample timbre content feature; Performing feature fusion on the sample embedding vector and the sample phoneme feature, and inputting the obtained second fused feature into the pre-trained temporary encoder to obtain the sample identity content feature; Performing Gaussian sampling on the sample timbre content features, inputting the sampling results into the user representation predictor to obtain a sample prediction number; Multiplying the sampling result by the sample user number to obtain a fused sample spectrum feature; Inputting the fused sample spectrum features into the spectrum decoder to obtain a predicted speech spectrum; The speech synthesis model training process also includes: A loss function is calculated based on the sample timbre content features, the sample identity content features, the sample prediction number, the sample user number, the predicted speech spectrum and the true sample spectrum. Based on the loss function, the parameters of the spectrum encoder, the phoneme encoder, the recognition encoder, the user representation predictor and the spectrum decoder are updated by the gradient descent method.
3. The speech synthesis method according to claim 2, wherein: The loss function includes: Relative entropy, the relative entropy is calculated based on the sample timbre content characteristics and the sample identity content characteristics. Based on the relative entropy, the parameters of the spectrum encoder, the phoneme encoder, and the recognition encoder are updated by gradient descent method until the relative entropy converges, thereby obtaining a pre-trained spectrum encoder, a pre-trained phoneme encoder, and a pre-trained recognition encoder.
4. The speech synthesis method according to claim 3, wherein: The loss function also includes: A first mean square error loss is calculated based on the prediction number and the sample user number. Based on the first mean square error loss, the parameters of the user representation predictor are updated by a gradient descent method until the first mean square error loss converges, thereby obtaining a pre-trained user representation predictor.
5. The speech synthesis method according to claim 4, characterized in that The loss function also includes: A second mean square error loss is calculated based on the predicted speech spectrum and the true sample spectrum. Based on the second mean square error loss, the parameters of the spectrum decoder are updated by the gradient descent method until the second mean square error loss converges, thereby obtaining a pre-trained spectrum decoder.
6. A speech synthesis device, characterized in that: The speech synthesis device comprises: Data acquisition module: used to obtain the reference speech spectrum and target speech phonemes of the target user; a spectrum encoder, configured to input the reference speech spectrum and output reference timbre content features; A phoneme encoder, configured to input the target speech phonemes and output target content features; an identification encoder, configured to input the reference timbre content feature and the target content feature, and output the target timbre content feature; A user characterization predictor, configured to sample the target timbre content feature and output a user identity content feature based on the sampled result; Spectrum decoder: used to fuse the target timbre content features and the user identity content features, and output the speech synthesis result of the target user based on the obtained fusion features.
7. The speech synthesis device according to claim 6, characterized in that The speech synthesis device further comprises: Temporary embedding layer, used to input sample user ID and output sample embedding vector during model training; A temporary encoder is used to perform feature fusion on the sample embedding vector and the sample phoneme features extracted by the phoneme encoder during model training, and output the sample identity content features based on the obtained fusion features; The speech synthesis device also includes: The spectrum encoder is used to input a sample speech spectrum and output sample spectrum features during model training; The phoneme encoder is used to input sample speech phonemes and output sample phoneme features during model training; The recognition encoder is used to fuse the sample spectrum features and the sample phoneme features during model training, and output the sample timbre content features based on the obtained first fusion features; The user representation predictor is used to sample the sample timbre content features during model training and output a sample prediction number based on the obtained sampling results; The feature fusion module is used to multiply the sampling result by the sample user number during model training to obtain the fused sample spectrum feature; The spectrum decoder is used to input the fused sample spectrum features and output the predicted speech spectrum during model training; The parameter update module is used to calculate the loss function according to the sample timbre content features, sample identity content features, sample prediction number, sample user number, predicted speech spectrum and true sample spectrum during model training, and update the parameters of the spectrum encoder, phoneme encoder, recognition encoder, user representation predictor and spectrum decoder through the gradient descent method based on the loss function.
8. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the speech synthesis method according to any one of claims 1 to 5 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the speech synthesis method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Audio generation method and device, storage medium and electronic equipment
CN113205793A
Speech synthesis model training method, speech synthesis method and related device
CN114187891A