Speech synthesis method, computer device and computer readable storage medium
By obtaining and processing emotional information in speech synthesis technology, and generating emotional vector sequences with length matching phoneme sequences to be synthesized texts, the problem that synthetic speech in the prior art is difficult to express emotions is solved, and a richer emotional expression is achieved.
Patent Information
- Application Number
- CN202310163997.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-15
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-02-15
AI Technical Summary
Existing speech synthesis technologies are difficult to display the emotional information of synthesized speech, especially in the intonation characteristics, and cannot express emotional diversity.
By obtaining the pending information about the text to be synthesized, setting the emotional category and setting the emotional intensity, the emotional base vector matching the emotional category is obtained, and the target emotional vector is determined based on the emotional intensity and emotional base vector. Then, based on the target emotion vector generation sequences of the phoneme sequences of the text to be synthesized, the speech synthesis process is performed in combination with the text to be synthesized, and a synthetic speech that can display the set emotional categories and emotional intensity is generated.
The synthetic voice can show the setting of emotional categories and setting emotional intensity, making the emotional information of synthetic voice richer and more diverse, and meeting the need to express emotions in speech synthesis.
Smart Images

Figure CN116072099B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a speech synthesis method, a computer device, and a computer-readable storage medium. Background Art
[0002] Speech synthesis technology, also known as text-to-speech technology, can convert any text information into standard and fluent speech. Today, speech synthesis technology is widely used in many fields such as smart speakers, map navigation, and voice assistants. With the development of deep learning technology, end-to-end speech synthesis systems have made significant progress. However, the intonation characteristics of synthesized speech are usually relatively fixed and cannot show emotional information. Summary of the invention
[0003] The embodiments of the present application provide a speech synthesis method, a computer device, and a computer-readable storage medium, which can use text to synthesize speech that can express set emotions and set emotion intensity, thereby enriching the emotional information of the synthesized speech.
[0004] On the one hand, the present application provides a speech synthesis method, the method comprising:
[0005] Acquiring information to be processed, wherein the information to be processed includes text to be synthesized, a set emotion category, and a set emotion intensity;
[0006] Acquire a first emotion basis vector that matches the set emotion category, and determine a target emotion vector according to the set emotion intensity and the first emotion basis vector;
[0007] Generate a first emotion vector sequence according to the target emotion vector, wherein the length of the first emotion vector sequence matches the length of the phoneme sequence of the text to be synthesized;
[0008] Speech synthesis processing is performed based on the text to be synthesized and the first emotion vector sequence to obtain a first synthesized speech; the emotion category displayed by the first synthesized speech matches the set emotion category, and the emotion intensity displayed by the first synthesized speech matches the set emotion intensity.
[0009] In one aspect, the present application provides a speech synthesis device, the device comprising:
[0010] An acquisition unit, used for acquiring information to be processed, wherein the information to be processed includes a text to be synthesized, a set emotion category and a set emotion intensity;
[0011] A processing unit, acquiring a first emotion basis vector matching the set emotion category, and determining a target emotion vector according to the set emotion intensity and the first emotion basis vector;
[0012] The processing unit is further used to generate a first emotion vector sequence according to the target emotion vector, wherein the length of the first emotion vector sequence matches the length of the phoneme sequence of the text to be synthesized;
[0013] The processing unit is also used to perform speech synthesis processing based on the text to be synthesized and the first emotion vector sequence to obtain a first synthesized speech; the emotion category displayed by the first synthesized speech matches the set emotion category, and the emotion intensity displayed by the first synthesized speech matches the set emotion intensity.
[0014] On the one hand, an embodiment of the present application provides a computer device, including: a processor and a memory, the memory storing executable program code, and the processor being used to call the executable program code to implement the speech synthesis method provided in the embodiment of the present application.
[0015] Accordingly, an embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is executed on a computer, the computer implements the speech synthesis method provided in the embodiment of the present application.
[0016] Accordingly, the embodiment of the present application also provides a computer program product or a computer program, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device implements the speech synthesis method provided in the embodiment of the present application.
[0017] In the present application, information to be processed including a text to be synthesized, a set emotion category and a set emotion intensity is obtained, an emotion base vector matching the set emotion category is obtained, and a target emotion vector is determined according to the set emotion intensity and the emotion base vector; an emotion vector sequence whose length matches the length of the phoneme sequence of the text to be synthesized is generated according to the target emotion vector, and speech synthesis processing is performed according to the text to be synthesized and the emotion vector sequence to obtain synthesized speech. The speech synthesis method provided in the present application adds an emotion vector sequence including an emotion vector and an emotion intensity during speech synthesis, so that the speech synthesized according to the text can show the set emotion category and the set emotion intensity, thereby enriching the emotional information of the synthesized speech on the basis of synthesizing the text into speech. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0019] Figure 1 Schematic diagram of a system architecture applicable to the speech synthesis method provided in the embodiment of the present application;
[0020] Figure 2 It is a flowchart of a speech synthesis method provided in an embodiment of the present application;
[0021] Figure 3 is a schematic diagram of an emotional integration method provided in an embodiment of the present application;
[0022] Figure 4 is a structural schematic diagram of an acoustic synthesis model provided in an embodiment of the present application;
[0023] Figure 5 It is a structural diagram of a sentiment information extraction model provided in an embodiment of the present application;
[0024] Figure 6 It is a flowchart of a model training method provided in an embodiment of the present application;
[0025] Figure 7 It is a structural diagram of an emotion feature extraction network provided in an embodiment of the present application;
[0026] Figure 8 is a structural diagram of an emotion representation network provided in an embodiment of the present application;
[0027] Fig. 9 is a structural schematic diagram of a speech synthesis device provided in an embodiment of the present application;
[0028] Fig.10 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0030] It should be noted that the descriptions of "first", "second", etc. involved in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the technical features defined as "first" or "second" may explicitly or implicitly include at least one of the features.
[0031] The embodiment of the present application provides a speech synthesis method, which can use text to synthesize speech that can show set emotions and set emotional intensity, so as to achieve the effect of enriching the emotional information of the synthesized speech. The speech synthesis method provided in the embodiment of the present application can be applied to the field of artificial intelligence. Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. And an important research direction in the field of artificial intelligence is speech synthesis technology.
[0032] The speech synthesis method provided in the embodiment of the present application can be Figure 1 The speech synthesis device 101 shown is implemented. In the process of the speech synthesis device 101 synthesizing speech from text, the speech synthesis device 101 obtains the information to be processed including the text to be synthesized, the set emotion category and the set emotion intensity, and obtains the emotion basis vector matching the set emotion category. The emotion basis vector is vector data representing a specific emotion category. Each emotion basis vector represents a specific emotion category. The target emotion vector is determined based on the emotion basis vector and the set emotion intensity, and the target emotion vector contains the emotion features of the set emotion category and the set emotion intensity. The target emotion vector generates an emotion vector sequence whose length matches the length of the phoneme sequence of the text to be synthesized, and the emotion vector sequence is subjected to speech synthesis processing with the text to be synthesized to obtain synthesized speech. The synthesized speech can show the set emotion category and the set emotion intensity, enriching the emotional information of the synthesized speech.
[0033] In one embodiment, the speech synthesis device 101 may be a terminal device. The terminal device may obtain the information to be processed, including the text to be processed, the set emotion category and the set emotion intensity, input by the user through human-computer interaction with the user, and perform speech synthesis processing to obtain a synthesized speech that can display the set emotion category and the set emotion intensity. The terminal device may be a smart home appliance, a handheld device with wireless communication function (such as a smart phone, a tablet computer), a computing device (such as a personal computer (PC), a vehicle-mounted terminal, an intelligent voice interaction device, a wearable device or other smart device, but is not limited thereto.
[0034] In another embodiment, the speech synthesis device 101 may be a server. The terminal device obtains user demand information by human-computer interaction with the user, and then processes the user demand information to obtain information to be processed including text to be processed, set emotion categories, and set emotion intensity. The server obtains the information to be processed from the terminal device, and performs speech synthesis processing on it to obtain a synthesized speech that can display the set emotion category and the set emotion intensity. The server then sends the synthesized speech to the terminal device, and finally the terminal device conveys the synthesized speech to the user. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0035] It is understandable that the system architecture diagram applicable to the speech synthesis method described in the embodiment of the present application is to more clearly illustrate the technical solution of the embodiment of the present application, and does not constitute a limitation on the speech synthesis method provided in the embodiment of the present application. Figure 1 The number of speech synthesis devices 101 in the embodiment is only for illustration. Any number of servers and terminal devices may be configured according to the business implementation requirements. Moreover, with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0036] See also Figure 2 , which is a flowchart of a speech synthesis method provided by an exemplary embodiment of the present application. The present application embodiment applies the method to Figure 1 Taking the speech synthesis device shown in the figure as an example, the speech synthesis method may include but is not limited to the following steps:
[0037] S201, obtaining information to be processed, wherein the information to be processed includes text to be synthesized, set emotion category and set emotion intensity.
[0038] In an embodiment of the present application, the information to be processed includes the text to be synthesized, the set emotion category and the set emotion intensity corresponding to the set emotion category. The text to be synthesized can be a sentence or a word, or a paragraph or an article. The method for the speech synthesis device to obtain the information to be processed can be to obtain the information to be processed from other devices, or to obtain the information to be processed by interacting with the user, or to obtain the information to be processed from a database. The set emotion intensity is used to indicate the proportion of the corresponding set emotion category in all set emotion categories.
[0039] In one embodiment, emotions include but are not limited to six basic emotion categories: anger, disgust, fear, happiness, sadness, and surprise. The speech synthesis device obtains information to be processed, for example: the text to be processed is "Oh my God! How come you are here", the emotion categories are set to "surprise" and "happy", and the emotion intensity is set to 0.7 for "surprise" and 0.3 for "happy". It can be seen that the set emotion category can include one or more, and correspondingly, the set emotion intensity can be the intensity ratio corresponding to each set emotion category.
[0040] S202: Obtain a first emotion basis vector that matches the set emotion category.
[0041] In an embodiment of the present application, the first emotion base vector is vector data representing a specific emotion category and matches the set emotion category. For example, if the emotion category is set to "anger", the first emotion base vector is vector data corresponding to the set emotion category "anger".
[0042] In one embodiment, a first emotion basis vector matching a set emotion category can be determined from a plurality of baseline emotion basis vectors. The baseline emotion basis vector is vector data representing a basic emotion category, and there can be multiple baseline emotion basis vectors. For example, assuming that there are three basic emotion categories, namely, the "happy" emotion category, the "sad" emotion category, and the "angry" emotion category, there can be three baseline emotion basis vectors, namely, a baseline emotion basis vector representing "happy", a baseline emotion basis vector representing "sad", and a baseline emotion basis vector representing "angry"; assuming that the set emotion category is "sad", the baseline emotion basis vector representing "sad" can be determined as the first emotion basis vector from the above three baseline emotion basis vectors according to the set emotion category.
[0043] S203: Determine a target emotion vector according to the set emotion intensity and the first emotion basis vector.
[0044] In an embodiment of the present application, the first emotion vector corresponding to the set emotion category and the set emotion intensity are calculated to obtain a target emotion vector, and the target emotion vector includes the emotion characteristics of the set emotion category and the set emotion intensity.
[0045] In one embodiment, the number of set emotion categories is M, the number of set emotion intensities is M, each set emotion category corresponds to a set emotion intensity, the number of first emotion basis vectors is M, M is a positive integer greater than 1, and each set emotion category corresponds to a first emotion basis vector, then the step of determining the target emotion vector according to the set emotion intensity and the first emotion basis vector may be: weighting the set emotion intensities matching the M set emotion categories and the first emotion basis vectors matching the M set emotion categories to obtain M weighted emotion basis vectors corresponding to the M set emotion categories; adding the M weighted emotion basis vectors corresponding to the M set emotion categories to obtain the target emotion vector.
[0046] In the embodiment of the present application, the weighted sum formula of setting the emotion category and setting the emotion intensity can be as follows (1):
[0047]
[0048] Where E represents the target emotion vector, T i represents the first emotion basis vector corresponding to the i-th set emotion category, ω i Represents the set emotion intensity corresponding to the i-th set emotion category. For example: M is 2, the first set emotion category is "happy", and the set emotion intensity of the set emotion category "happy" is 0.4; the second set emotion category is "surprised", and the set emotion intensity of the set emotion category "surprised" is 0.6; then the first emotion basis vector corresponding to the set emotion category "happy" is weighted with the set emotion intensity (0.4), the first emotion basis vector corresponding to the set emotion category "surprised" is weighted with the set emotion intensity (0.6), and the two weighted emotion vectors are summed to obtain the target emotion vector.
[0049] In one embodiment, the M weighted emotion basis vectors corresponding to the M set emotion categories can be calculated in a nonlinear manner to obtain a target emotion vector. In one implementation, the M weighted emotion basis vectors are input into a preset mapping relationship to obtain a mapped weighted emotion basis vector. The mapped weighted emotion basis vectors are then summed to obtain a target emotion vector.
[0050] In one embodiment, there is only one set emotion category and only one set emotion intensity matching the set emotion category. The step of determining the target emotion vector according to the set emotion intensity and the first emotion base vector may include but is not limited to the following steps 1 and 2:
[0051] 1. Obtain a neutral emotion basis vector, and determine a weight parameter of the neutral emotion basis vector according to the set emotion intensity; perform weighted processing on the neutral emotion basis vector according to the weight parameter to obtain a weighted neutral emotion basis vector.
[0052] In the embodiment of the present application, the neutral emotion basis vector is the emotion basis vector corresponding to the emotion type when the emotion intensity of all emotion types is the weakest, that is, the emotion intensity is 0. When there is only one set emotion category, all baseline emotion basis vectors except the first emotion basis vector corresponding to the set emotion category are neutral emotion basis vectors. The weighted neutral emotion basis vector is used to represent the distance between the target emotion vector and the neutral emotion basis vector. The closer the distance to the neutral emotion basis vector is, the smaller the set emotion intensity is, and the farther the distance to the neutral emotion basis vector is, the greater the set emotion intensity is.
[0053] 2. Perform weighted processing on the first emotion basis vector according to the set emotion intensity to obtain a weighted first emotion basis vector; perform addition processing on the weighted neutral emotion basis vector and the weighted first emotion basis vector to obtain a target emotion vector.
[0054] In the embodiment of the present application, the weighted first emotion basis vector is used to represent the distance between the target emotion vector and the first emotion basis vector. The weighted neutral emotion basis vector and the weighted first emotion basis vector are added to obtain the target emotion vector. In one embodiment, taking "happy" as an example of a set emotion category, the calculation formula for the target emotion vector of the set emotion category is as follows (2):
[0055] E=αT hap +(1-)T neu (2)
[0056] Among them, E represents the target emotion vector of the set emotion category, α represents the set emotion intensity of the set emotion category, T hap Represents the first emotion basis vector corresponding to the set emotion category, T neu Represents the neutral emotion basis vector corresponding to the set emotion category, and 1-α represents the emotion intensity of the neutral emotion basis vector corresponding to the set emotion category. In one embodiment, the "1" in the above 1-α can also be other parameters (for example: 2, 3, etc.), and is used to subtract from α to determine the emotion intensity of the neutral emotion basis vector corresponding to the set emotion category. The method provided in the embodiment of the present application can not only realize the fusion of multiple emotion categories, but also realize the distinction between a single emotion category and a neutral emotion category, and make the emotion contained in the final synthesized speech different from the neutral emotion.
[0057] In one embodiment, the steps of obtaining the first emotion base vector matching the set emotion category and determining the target emotion vector according to the set emotion intensity and the first emotion base vector may be performed by an emotion coloring module.
[0058] See also Figure 3 , which is a schematic diagram of an emotional fusion method provided by an exemplary embodiment of the present application, Figure 3 The emotional fusion method shown can be implemented by the emotional coloring module. Figure 3 As shown in the figure, the emotion is divided into six basic emotion categories: anger, disgust, fear, joy, sadness, and surprise. The emotion coloring module contains the reference emotion basis vectors corresponding to the above six basic emotion categories. Assume that the set emotion category includes two basic emotion categories of "surprise" and "joy". Then, the first emotion basis vector representing the two emotion categories of "surprise" and "joy" can be determined from the six reference emotion basis vectors according to the two set emotion categories (the first emotion basis vector is the vector data representing a specific emotion category). The first emotion basis vector representing the two emotion categories of "surprise" and "joy" is multiplied by the set emotion intensity matching the two set emotion categories, respectively, to obtain the weighted emotion basis vector representing the two emotion categories of "surprise" and "joy", and finally the weighted emotion basis vectors representing the two emotion categories of "surprise" and "joy" are added to obtain the target emotion vector representing the emotion category of "surprise". The emotion coloring module enables the target emotion vector to include multiple set emotion categories and multiple set emotion intensities, so that the final synthesized speech can contain multiple emotions.
[0059] S204: Generate a first emotion vector sequence according to the target emotion vector, wherein the length of the first emotion vector sequence matches the length of the phoneme sequence of the text to be synthesized.
[0060] In the embodiment of the present application, the length of the first emotion vector sequence matches the length of the phoneme sequence of the text to be synthesized, so as to facilitate the combination of the phoneme sequence of the text to be synthesized with the first emotion vector sequence.
[0061] In one embodiment, the information to be processed may also include emotion intensity change information, and the emotion intensity change information is used to indicate the change of the emotion vector sequence. Then the step of generating the first emotion vector sequence according to the target emotion vector may be:
[0062] 1. Obtain the length of the phoneme sequence of the text to be synthesized; expand the target emotion vector according to the length to obtain an initial emotion vector sequence.
[0063] In an embodiment of the present application, the target emotion vector is expanded to an initial emotion vector sequence that matches the length of the phoneme sequence of the text to be synthesized according to the length of the phoneme sequence of the text to be synthesized.
[0064] 2. According to the emotion intensity change information, specific vector features of all or part of the emotion vectors in the initial emotion vector sequence are adjusted to obtain a first emotion vector sequence.
[0065] In the embodiment of the present application, the specific vector feature is a vector feature used to indicate the intensity of emotion. The emotion intensity change information is the emotion vector sequence indication information, which can be used to indicate the change of the emotion intensity in the emotion vector sequence. For example, if the emotion intensity change information indicates that "the emotion intensity gradually weakens", the specific vector feature in the initial emotion vector sequence (i.e., the vector feature of the emotion intensity) changes accordingly. The specific vector features of all the emotion vectors in the initial emotion vector sequence are adjusted according to the emotion intensity change information to obtain a first emotion vector sequence.
[0066] In one embodiment, the emotion intensity change information can be used to indicate the change mode of the phoneme sequence of the text to be synthesized. For example, the emotion intensity change information is "starting from the first phoneme in the phoneme sequence, the emotion intensity changes from strong to weak." Then the emotion intensity of all emotion vectors in the initial emotion vector sequence is adjusted to obtain the first emotion vector sequence. The method provided in the embodiment of the present application realizes emotion control at the phoneme level in synthesized speech and enhances the emotional expressiveness of synthesized speech.
[0067] In one embodiment, the emotion intensity change information can be used to indicate the change mode of the sentence of the text to be synthesized. For example, the emotion intensity change information is "the emotion intensity of the second half of a sentence is enhanced, and the rest remains unchanged". Then the emotion intensity of the emotion vector of the second half of the initial emotion vector sequence is adjusted to obtain the first emotion vector sequence.
[0068] S205. Perform speech synthesis processing according to the text to be synthesized and the first emotion vector sequence to obtain a first synthesized speech; the emotion category displayed by the first synthesized speech matches the set emotion category, and the emotion intensity displayed by the first synthesized speech matches the set emotion intensity.
[0069] In an embodiment of the present application, the text to be synthesized and the first emotion vector sequence are subjected to speech synthesis processing. Since the first emotion vector sequence contains emotion features of a set emotion category and a set emotion intensity, the emotion category exhibited by the obtained first synthesized speech matches the set emotion category, and the emotion intensity exhibited by the first synthesized speech matches the set emotion intensity. The method provided in an embodiment of the present application takes into account the influence of different emotion categories on synthesized emotional speech, introduces a set emotion intensity, so that the emotion representations between the same emotion categories will be different based on the different set emotion intensities, so that the synthesized speech can contain rich emotional information.
[0070] In one embodiment, an acoustic synthesis model can be used to perform speech synthesis processing according to the text to be synthesized and the first emotion vector sequence to obtain a first synthesized speech. The specific implementation method can be: input the text to be synthesized and the first emotion vector sequence into the target acoustic synthesis model for speech synthesis processing to obtain the first synthesized speech. Among them, the target acoustic synthesis model is obtained by training in combination with the emotion information extraction model; during the training process, the emotion information extraction model is used to process the sample speech to determine the second emotion vector sequence, and the initial acoustic synthesis model is used to perform speech synthesis processing according to the training text corresponding to the second emotion vector sequence and the sample speech to obtain the second synthesized speech; the target acoustic synthesis model is obtained by adjusting the model parameters of the initial acoustic synthesis model using the first loss parameter determined according to the sample speech and the second synthesized speech.
[0071] See also Figure 4 , which is a schematic diagram of the structure of an acoustic synthesis model provided by an exemplary embodiment of the present application. Figure 4 As shown, the acoustic synthesis model mainly includes two parts: an acoustic module and a vocoding unit, wherein the acoustic module includes a text-to-phoneme and phoneme embedding unit, a phoneme encoding unit, a variable adaptation unit, and a decoding unit. After the text to be synthesized is input into the acoustic synthesis model, it first passes through the text-to-phoneme and phoneme embedding unit in the acoustic module; the output data of the text-to-phoneme and phoneme embedding unit is combined with the position code and input into the phoneme encoding unit to obtain the phoneme sequence of the text to be synthesized; then the phoneme sequence of the text to be synthesized is added to the emotion vector sequence containing the emotion features of the set emotion category and the set emotion intensity to obtain the phoneme sequence containing the emotion vector. The phoneme sequence containing the emotion vector is input into the variable information adaptation unit for processing; the output data of the variable information adaptation unit is combined with the position code and input into the decoding unit for processing to obtain the mel spectrum of the text to be synthesized. The mel spectrum of the text to be synthesized is input into the vocoding unit for synthesis to obtain the synthesized speech. The emotion category displayed by the synthesized speech matches the set emotion category, and the emotion intensity displayed by the synthesized speech matches the set emotion intensity. The acoustic synthesis model proposed in the embodiment of the present application can express the differences in the intensity of different emotions between the same emotion categories, and can also express many different emotions by combining multiple different emotion categories.
[0072] It should be noted that the acoustic synthesis model provided in the above embodiment is only one possible structure of the acoustic synthesis model of the present application. Other end-to-end acoustic synthesis models that can realize speech synthesis are included in the coverage of the present application.
[0073] In the embodiment of the present application, the acoustic synthesis model is obtained by training in combination with the emotion information extraction model. Figure 5As shown, the emotion information extraction model may include a spectrum extraction network 501, an emotion feature extraction network 502 and an emotion representation network 503. The output of the spectrum extraction network 501 is connected to the input of the emotion feature extraction network 502, and the output of the emotion feature extraction network 502 is connected to the input of the emotion representation network 503. During model training, the emotion information extraction model processes the input sample speech to determine the emotion vector sequence. The initial acoustic synthesis model is trained in combination with the emotion vector sequence determined by the emotion information extraction model to obtain the target acoustic synthesis model.
[0074] The training of the target acoustic synthesis model can be performed by a model training device, which can be the speech synthesis device in the foregoing embodiment, or can be other computing devices different from the speech synthesis device in the foregoing embodiment.
[0075] Based on the above embodiments, the beneficial effects of the present application are: the speech synthesis method provided by the present application can fuse multiple emotion categories according to the set emotion category and the set emotion intensity, so that the synthesized speech can include emotions beyond the basic emotion categories; at the same time, for a single emotion category, the method provided by the present application can also reflect the difference between it and the neutral emotion basis vector, so that the emotion representation space of the synthesized speech is richer.
[0076] See also Figure 6 , which shows a flow chart of a model training method provided in an embodiment of the present application. The steps of the model training method include but are not limited to:
[0077] S601. Acquire a training data pair, wherein the training data pair includes a training text, a sample speech corresponding to the training text, and a sample emotion category corresponding to the sample speech.
[0078] In the embodiment of the present application, there are multiple groups of training data pairs, and each group of training data pairs is not completely the same. The training text can be a sentence or a word, or a paragraph or a chapter.
[0079] S602: Input the sample speech into the spectrum extraction network for processing to obtain a mel spectrum of the sample speech.
[0080] In the embodiment of the present application, the spectrum extraction network in the emotion information extraction model is used to convert speech into mel spectrum, which is a mathematical representation of speech, and is convenient for the emotion feature extraction network to extract emotion features.
[0081] S603: Input the mel spectrum into the emotion feature extraction network for processing to obtain reference emotion features of the sample speech.
[0082] In the embodiment of the present application, the mel spectrum is input into the emotion feature extraction network in the emotion information extraction model for processing to obtain the reference emotion feature of the sample speech, which is used to represent the emotion feature contained in the mel spectrum of the sample speech. The reference emotion feature of the sample speech can be used to train the acoustic synthesis model.
[0083] See also Figure 7 , which is a schematic diagram of the structure of an emotion feature extraction network proposed in an embodiment of the present application. The emotion feature extraction network mainly includes three parts, namely a convolution unit, a gate control loop unit and a second activation function unit. Among them, the convolution unit includes N convolution layers (N is a positive integer greater than 1), and each convolution layer includes a one-dimensional convolution layer, a normalization function and a first activation function. The Mel spectrum of the sample speech is input into the convolution unit of the emotion feature extraction network for processing. After being processed by the convolution unit, the potential relationship between frames in the Mel spectrum of the sample speech can be extracted. Then, the output data of the convolution unit is input into the gate control loop unit for processing, and finally the output data of the gate control loop unit is input into the first activation function unit for calculation to obtain the reference emotion feature of the sample speech. In one embodiment, the reference Mel spectrum includes an R frame. After the R frame reference Mel spectrum is input into the convolution unit for calculation, an R frame reference Mel spectrum containing a potential relationship representation between frames is obtained. The output result of the convolution unit is then input into the gate control loop unit for processing, and the operation result of the gate control loop unit is input into the second activation function unit for processing to obtain the global emotional features of the R frame reference mel spectrum. The structure of the emotional feature extraction network provided in the embodiment of the present application can extract the global emotional features of the sample speech, can better avoid the mixing of sample information, and reduce the influence of the sample speech on the synthetic speech effect.
[0084] It should be noted that the above Figure 7 The emotional feature extraction network structure shown is only one possible structure provided by the embodiment of the present application, and other structures that can realize the functions of the emotional feature extraction network are also included in the scope of the present application. For example, the emotional feature extraction network is constructed based on Convolutional Neural Networks (CNN) to realize the above functions.
[0085] S604: Input the reference emotion feature into the emotion representation network for processing, and determine a second emotion basis vector matching the reference emotion feature from a plurality of initial emotion basis vectors.
[0086] In the embodiment of the present application, the initial emotion basis vector is vector data representing the basic emotion category, and each initial emotion basis vector corresponds to an emotion category. When there are multiple emotion categories, there are also multiple corresponding initial emotion basis vectors. The step of determining the second emotion basis vector matching the reference emotion feature from multiple initial emotion basis vectors by the emotion representation network can be:
[0087] 1. Input the reference emotion feature into the emotion representation network for processing, and determine the similarity between the reference emotion feature and each of the initial emotion basis vectors.
[0088] In the embodiment of the present application, the similarity between the reference emotion feature and each initial emotion basis vector can indicate which initial emotion basis vector the reference emotion feature is most similar to.
[0089] 2. Determine a second emotion basis vector that matches the reference emotion feature from the multiple initial emotion basis vectors based on the similarity.
[0090] In the embodiment of the present application, a second emotion basis vector matching the reference emotion feature is determined based on the similarity, and the second emotion basis vector can represent the emotion feature included in the reference emotion feature.
[0091] In one embodiment, the initial emotion basis vector is randomly initialized vector data, which cannot represent the corresponding emotion category, so it is necessary to adjust multiple initial emotion basis vectors to obtain multiple benchmark emotion basis vectors that can represent the characteristic emotion category. The benchmark emotion basis vector is vector data representing a specific emotion category. The process of adjusting the initial emotion basis vector to obtain the benchmark emotion basis vector may include but is not limited to:
[0092] 1. Determine a second loss parameter according to the sample emotion category corresponding to the sample speech and the similarity between the reference emotion feature and each of the initial emotion basis vectors.
[0093] In the embodiment of the present application, the sample emotion category corresponding to the sample speech in the training data pair represents the emotion category contained in the reference emotion feature. The size of the second loss parameter indicates the degree to which the initial emotion basis vector can represent the emotion category corresponding to the initial emotion basis vector; the smaller the value of the second loss parameter, the more accurately the initial emotion basis vector can represent the emotion category corresponding to it; the larger the value of the second loss parameter, the less accurately the initial emotion basis vector can represent the emotion category corresponding to it.
[0094] 2. Adjust the multiple initial emotion basis vectors according to the second loss parameter to obtain the multiple benchmark emotion basis vectors.
[0095] In an embodiment of the present application, multiple initial emotion basis vectors are adjusted according to the second loss parameter to obtain benchmark emotion basis vectors that can represent the emotion categories corresponding to them.
[0096] It should be noted that multiple initial emotion basis vectors need to be adjusted multiple times according to the second loss parameter to obtain multiple reference emotion basis vectors. The embodiment of the present application only describes one training process among multiple trainings, and the rest of the training processes are similar to the training processes described above.
[0097] S605: Generate a second emotion vector sequence according to the second emotion basis vector.
[0098] In the embodiment of the present application, since the second emotion base vector is determined based on the reference emotion feature corresponding to the training text, the second emotion vector sequence generated based on the second emotion base vector contains the emotion feature in the reference emotion feature, and the length of the second emotion vector sequence matches the length of the phoneme sequence of the training text. In the subsequent speech synthesis process, the second emotion vector sequence can add specific emotion features to the synthesized speech.
[0099] In one embodiment, the process of obtaining the second emotion vector sequence according to the reference emotion feature and training the initial emotion base vector according to the sample emotion category corresponding to the sample speech can be implemented by an emotion representation network. Figure 8 , which is a schematic diagram of the structure of an emotion representation network proposed in an embodiment of the present application. The emotion representation network mainly includes a multi-head attention mechanism unit and a cross entropy unit. When the emotion representation network receives the input reference emotion feature, a similarity weight calculation is performed according to a plurality of randomly initialized initial emotion basis vectors to obtain the similarity between the reference emotion feature and each of the initial emotion basis vectors. According to the similarity, a second emotion basis vector matching the reference emotion feature is determined from a plurality of initial emotion basis vectors; and then a second emotion vector sequence is determined according to the second emotion vector sequence. According to the similarity between the reference emotion feature and each of the initial emotion basis vectors and the sample emotion category corresponding to the sample speech, a cross entropy calculation is performed in the cross entropy unit to obtain a second loss parameter. Finally, the initial emotion basis vector is adjusted according to the second loss parameter to finally obtain a baseline emotion basis vector.
[0100] In one embodiment, the types of basic emotions include but are not limited to six, namely anger, disgust, fear, joy, sadness, and surprise. The emotion basis vectors corresponding to the six basic emotion categories are randomly initialized to obtain the initial emotion basis vector, which is the vector data representing the basic emotion category. The emotion representation network receives the input reference emotion feature, and inputs it and the initial emotion basis vector into the attention mechanism unit for similarity weight calculation to obtain the similarity between the reference emotion feature and the six initial emotion basis vectors. According to these similarities, the second emotion basis vector matching the reference emotion feature is determined from the six initial emotion basis vectors. For example: the sample emotion category corresponding to the sample speech is "surprise", and the second emotion basis vector determined according to the reference emotion feature corresponding to the sample speech is the emotion basis vector representing "anger", then it can be seen that the second emotion basis vector at this time cannot fully and accurately represent the emotion category in the sample speech, and the emotion basis vector needs to be adjusted. Therefore, the second emotion basis vector and the sample emotion category corresponding to the sample speech are cross-entropy calculated to obtain the second loss parameter, and the initial emotion basis vector is adjusted according to the second loss parameter so that the initial emotion basis vector can accurately represent the basic emotion category. The emotion representation network proposed in the embodiment of the present application adjusts the initial emotion basis vector by using the second loss parameter to obtain a reference emotion basis vector, so that the reference emotion basis vector can accurately represent the basic emotion category. When the model is used, the first emotion basis vector is determined from multiple reference emotion basis vectors according to the set emotion category, and the first emotion basis vector is used for synthesized speech, so the reference emotion basis vector is conducive to the synthesized speech containing rich and accurate emotion categories.
[0101] S606: Input the training text and the second emotion vector sequence into the initial acoustic synthesis model for speech synthesis processing to obtain a second synthesized speech.
[0102] In the embodiment of the present application, the second emotion vector sequence includes the emotion features extracted from the sample speech by the emotion information extraction model. The training text and the second emotion vector sequence are input into the initial acoustic synthesis model for speech synthesis processing to obtain the second synthesized speech, and the second synthesized speech includes some or all of the emotion features in the sample speech. The structure of the initial acoustic synthesis model can refer to the above Figure 4 The structure shown.
[0103] S607. Determine a first loss parameter according to the sample speech and the second synthesized speech, and adjust model parameters of the initial acoustic synthesis model according to the first loss parameter to obtain the target acoustic synthesis model.
[0104] In the embodiment of the present application, the first loss parameter determined based on the sample speech and the second synthesized speech can represent the degree of difference between the second synthesized speech and the sample speech. The larger the first loss parameter, the greater the difference between the sample speech and the second synthesized speech, and the smaller the first loss parameter, the smaller the difference between the sample speech and the second synthesized speech. The model parameters of the initial acoustic synthesis model are adjusted according to the first loss parameter. When the first loss parameter determined by the sample speech and the second synthesized speech is less than the set threshold, it indicates that the training of the initial acoustic synthesis model is completed, and the target acoustic synthesis model is obtained.
[0105] It should be noted that the above model training process will be repeated multiple times until the first loss parameter determined according to the sample speech and the second synthesized speech and the second loss parameter determined according to the sample emotion category, reference emotion features and initial emotion basis vector meet the preset requirements, and the model training process will stop.
[0106] Based on the above embodiments, the beneficial effects of the present application are: in the model training process, the present application proposes to use training data that does not contain emotion intensity (i.e., sample speech, sample emotion category) to perform model training, and use the emotion feature extraction network and the emotion representation network to train the initial acoustic model and the initial emotion basis vector to obtain the target acoustic synthesis model and the baseline emotion basis vector representing the basic emotion category. In the reasoning process, the present application proposes to use the emotion coloring module to perform emotion fusion on the set emotion intensity and the set emotion category, so as to obtain emotion categories other than the basic emotion categories. Then, the emotion vector sequence containing the fused emotion category and the target acoustic synthesis model are used to synthesize speech containing multiple emotion categories and multiple emotion intensities, so that the synthesized speech contains rich emotion information.
[0107] See also Fig. 9 , which is a schematic diagram of the structure of a speech synthesis device provided in an embodiment of the present application. Fig. 9 The speech synthesis device shown may specifically include:
[0108] An acquisition unit 901 is used to acquire information to be processed, wherein the information to be processed includes a text to be synthesized, a set emotion category, and a set emotion intensity;
[0109] The processing unit 902 is used to obtain a first emotion basis vector that matches the set emotion category, and determine a target emotion vector according to the set emotion intensity and the first emotion basis vector;
[0110] The processing unit 902 is further configured to generate a first emotion vector sequence according to the target emotion vector, wherein the length of the first emotion vector sequence matches the length of the phoneme sequence of the text to be synthesized;
[0111] The processing unit 902 is also used to perform speech synthesis processing based on the text to be synthesized and the first emotion vector sequence to obtain a first synthesized speech; the emotion category displayed by the first synthesized speech matches the set emotion category, and the emotion intensity displayed by the first synthesized speech matches the set emotion intensity.
[0112] In one embodiment, when the processing unit 902 is used to perform speech synthesis processing according to the text to be synthesized and the first emotion vector sequence to obtain the first synthesized speech, it is specifically used to: input the text to be synthesized and the first emotion vector sequence into a target acoustic synthesis model for speech synthesis processing to obtain the first synthesized speech; wherein the target acoustic synthesis model is trained by the second emotion vector sequence corresponding to the sample speech and the training text corresponding to the sample speech, and the second emotion vector sequence is obtained by processing the sample speech by an emotion information extraction model.
[0113] In one embodiment, the emotion information extraction model includes a spectrum extraction network, an emotion feature extraction network and an emotion representation network, and the processing unit 902 is further used to: input the sample speech into the spectrum extraction network for processing to obtain the mel spectrum of the sample speech, and input the mel spectrum into the emotion feature extraction network for processing to obtain the reference emotion feature of the sample speech; input the reference emotion feature into the emotion representation network for processing, and the emotion representation network determines a second emotion basis vector matching the reference emotion feature from multiple initial emotion basis vectors, and generates a second emotion vector sequence based on the second emotion basis vector; wherein the length of the second emotion vector sequence matches the length of the phoneme sequence of the training text, and each of the initial emotion basis vectors corresponds to an emotion category; input the training text and the second emotion vector sequence into the initial acoustic synthesis model for speech synthesis processing to obtain a second synthesized speech; determine a first loss parameter based on the sample speech and the second synthesized speech, and adjust the model parameters of the initial acoustic synthesis model based on the first loss parameter to obtain the target acoustic synthesis model.
[0114] In one embodiment, when the processing unit 902 is used to input the reference emotion feature into the emotion representation network for processing, and the emotion representation network determines a second emotion basis vector matching the reference emotion feature from multiple initial emotion basis vectors, it is specifically used to: input the reference emotion feature into the emotion representation network for processing, and the emotion representation network determines the similarity between the reference emotion feature and each of the initial emotion basis vectors; and the emotion representation network determines a second emotion basis vector matching the reference emotion feature from the multiple initial emotion basis vectors based on the similarity.
[0115] In one embodiment, the processing unit 902 is further used to: determine a second loss parameter based on the sample emotion category corresponding to the sample speech and the similarity between the reference emotion feature and each of the initial emotion basis vectors; adjust the multiple initial emotion basis vectors based on the second loss parameter to obtain the multiple benchmark emotion basis vectors; wherein the set emotion categories are M, M is a positive integer greater than 1, and the processing unit 902 is specifically used to: determine, for a target set emotion category, a first emotion basis vector matching the target set emotion category from the multiple benchmark emotion basis vectors; the target set emotion category is any one of the M set emotion categories, and different set emotion categories correspond to different first emotion basis vectors.
[0116] In one embodiment, the number of set emotion categories is M, the number of set emotion intensities is M, each of the set emotion categories corresponds to a set emotion intensity, the number of first emotion basis vectors is M, each of the set emotion categories corresponds to a first emotion basis vector, where M is a positive integer greater than 1, and the above-mentioned processing unit 902, when used to determine the target emotion vector according to the set emotion intensity and the first emotion basis vector, is specifically used for: for the target set emotion category, according to the matching set emotion intensity among the M set emotion intensities that matches the target set emotion category, weighting the first emotion basis vector that matches the target set emotion category to obtain the weighted emotion basis vector corresponding to the target set emotion category; the target set emotion category is any one of the M set emotion categories; the weighted emotion basis vectors corresponding to each of the target set emotion categories are added to obtain the target emotion vector.
[0117] In one embodiment, when the processing unit 902 is used to determine the target emotion vector according to the set emotion intensity and the first emotion basis vector, it is specifically used to: obtain a neutral emotion basis vector, and determine a weight parameter of the neutral emotion basis vector according to the set emotion intensity; then weight the neutral emotion basis vector according to the weight parameter to obtain a weighted neutral emotion basis vector; weight the first emotion basis vector according to the set emotion intensity to obtain a weighted first emotion basis vector; and add the weighted neutral emotion basis vector and the weighted first emotion basis vector to obtain a target emotion vector.
[0118] In one embodiment, the above-mentioned processing unit 902 is used to generate a first emotion vector sequence according to the target emotion vector; when the information to be processed also includes emotion intensity change information, it is specifically used to: first obtain the length of the phoneme sequence of the text to be synthesized; expand the target emotion vector according to the length to obtain an initial emotion vector sequence; adjust the specific vector features of all or part of the emotion vectors in the initial emotion vector sequence according to the emotion intensity change information to obtain a first emotion vector sequence; wherein the specific vector features are vector features used to indicate emotion intensity.
[0119] It should be noted that the functions of the various functional modules of the speech synthesis device in the embodiment of the present application can be specifically implemented according to the method in the above embodiment. The specific implementation process can refer to the relevant description of the above method embodiment, which will not be repeated here.
[0120] Based on the above embodiments, the beneficial effects of the present application are: the speech synthesis method provided by the present application can fuse multiple emotion categories according to the set emotion category and the set emotion intensity, so that the synthesized speech can include emotions beyond the basic emotion categories; at the same time, for a single emotion category, the method provided by the present application can also reflect the difference between it and the neutral emotion basis vector, so that the emotion representation space of the synthesized speech is richer.
[0121] See also Fig.10 , which is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Fig.10 The computer device in the embodiment of the present application shown may include: a processor 1001 and a storage device 1002. The processor 1001 and the storage device 1002 may exchange data.
[0122] The above-mentioned storage device 1002 may include a volatile memory (volatile memory), such as a random-access memory (RAM); the storage device 1002 may also include a non-volatile memory (non-volatile memory), such as a flash memory, a solid-state drive (SSD), etc.; the above-mentioned storage device 1002 may also include a combination of the above-mentioned types of memory.
[0123] The processor 1001 may be a central processing unit (CPU). In one embodiment, the storage device 1002 is used to store program instructions, and the processor 1001 may call the program instructions to implement the following operations:
[0124] Acquiring information to be processed, wherein the information to be processed includes text to be synthesized, a set emotion category, and a set emotion intensity;
[0125] Acquire a first emotion basis vector that matches the set emotion category, and determine a target emotion vector according to the set emotion intensity and the first emotion basis vector;
[0126] Generate a first emotion vector sequence according to the target emotion vector, wherein the length of the first emotion vector sequence matches the length of the phoneme sequence of the text to be synthesized;
[0127] Speech synthesis processing is performed based on the text to be synthesized and the first emotion vector sequence to obtain a first synthesized speech; the emotion category displayed by the first synthesized speech matches the set emotion category, and the emotion intensity displayed by the first synthesized speech matches the set emotion intensity.
[0128] In one embodiment, when the processor 1001 is used to perform speech synthesis processing according to the text to be synthesized and the first emotion vector sequence to obtain the first synthesized speech, it is specifically used to: input the text to be synthesized and the first emotion vector sequence into a target acoustic synthesis model for speech synthesis processing to obtain the first synthesized speech; wherein the target acoustic synthesis model is trained by the second emotion vector sequence corresponding to the sample speech and the training text corresponding to the sample speech, and the second emotion vector sequence is obtained by processing the sample speech by an emotion information extraction model.
[0129] In one embodiment, the emotion information extraction model includes a spectrum extraction network, an emotion feature extraction network and an emotion representation network, and the processor 1001 is further used to: input the sample speech into the spectrum extraction network for processing to obtain the mel spectrum of the sample speech, and input the mel spectrum into the emotion feature extraction network for processing to obtain the reference emotion feature of the sample speech; input the reference emotion feature into the emotion representation network for processing, and the emotion representation network determines a second emotion basis vector matching the reference emotion feature from multiple initial emotion basis vectors, and generates a second emotion vector sequence based on the second emotion basis vector; wherein the length of the second emotion vector sequence matches the length of the phoneme sequence of the training text, and each of the initial emotion basis vectors corresponds to an emotion category; input the training text and the second emotion vector sequence into the initial acoustic synthesis model for speech synthesis processing to obtain a second synthesized speech; determine a first loss parameter based on the sample speech and the second synthesized speech, and adjust the model parameters of the initial acoustic synthesis model based on the first loss parameter to obtain the target acoustic synthesis model.
[0130] In one embodiment, when the processor 1001 is used to input the reference emotion feature into the emotion representation network for processing, and the emotion representation network determines a second emotion basis vector matching the reference emotion feature from multiple initial emotion basis vectors, the processor 1001 is specifically used to: input the reference emotion feature into the emotion representation network for processing, and the emotion representation network determines the similarity between the reference emotion feature and each of the initial emotion basis vectors; and the emotion representation network determines a second emotion basis vector matching the reference emotion feature from the multiple initial emotion basis vectors based on the similarity.
[0131] In one embodiment, the processor 1001 is further used to: determine a second loss parameter based on the sample emotion category corresponding to the sample speech and the similarity between the reference emotion feature and each of the initial emotion basis vectors; adjust the multiple initial emotion basis vectors based on the second loss parameter to obtain the multiple benchmark emotion basis vectors; wherein, the emotion categories are set to M, M is a positive integer greater than 1, and when the processor 1001 obtains the first emotion basis vector matching the set emotion category, it is specifically used to: determine, for a target set emotion category, a first emotion basis vector matching the target set emotion category from the multiple benchmark emotion basis vectors; the target set emotion category is any one of the M set emotion categories, and the first emotion basis vectors corresponding to different set emotion categories are different.
[0132] In one embodiment, the number of set emotion categories is M, the number of set emotion intensities is M, each of the set emotion categories corresponds to a set emotion intensity, the number of first emotion base vectors is M, each of the set emotion categories corresponds to a first emotion base vector, and M is a positive integer greater than 1. When the processor 1001 is used to determine the target emotion vector according to the set emotion intensity and the first emotion base vector, it is specifically used to: for the target set emotion category, according to the matching set emotion intensity among the M set emotion intensities that matches the target set emotion category, weight the first emotion base vector that matches the target set emotion category to obtain the weighted emotion base vector corresponding to the target set emotion category; the target set emotion category is any one of the M set emotion categories; the weighted emotion base vectors corresponding to each of the target set emotion categories are added to obtain the target emotion vector.
[0133] In one embodiment, when the processor 1001 is used to determine the target emotion vector according to the set emotion intensity and the first emotion basis vector, it is specifically used to: obtain a neutral emotion basis vector, and determine a weight parameter of the neutral emotion basis vector according to the set emotion intensity; weight the neutral emotion basis vector according to the weight parameter to obtain a weighted neutral emotion basis vector; weight the first emotion basis vector according to the set emotion intensity to obtain a weighted first emotion basis vector; add the weighted neutral emotion basis vector and the weighted first emotion basis vector to obtain a target emotion vector.
[0134] In one embodiment, the above-mentioned processor 1001 is used to generate a first emotion vector sequence according to the target emotion vector; when the information to be processed also includes emotion intensity change information, it is specifically used to obtain the length of the phoneme sequence of the text to be synthesized; the target emotion vector is expanded according to the length to obtain an initial emotion vector sequence; specific vector features of all or part of the emotion vectors in the initial emotion vector sequence are adjusted according to the emotion intensity change information to obtain a first emotion vector sequence; wherein the specific vector feature is a vector feature used to indicate emotion intensity.
[0135] The processor 1001 and the storage device 1002 described in the embodiments of the present application can execute the embodiments of the present application. Figure 2 or Figure 6 The implementation method described in the relevant embodiments of the provided speech synthesis method will not be repeated here.
[0136] Based on the above embodiments, the beneficial effects of the present application are: the speech synthesis method provided by the present application can fuse multiple emotion categories according to the set emotion category and the set emotion intensity, so that the synthesized speech can include emotions beyond the basic emotion categories; at the same time, for a single emotion category, the method provided by the present application can also reflect the difference between it and the neutral emotion basis vector, so that the emotion representation space of the synthesized speech is richer.
[0137] In the several embodiments provided in the present application, it should be understood that the disclosed methods, devices and systems can be implemented in other ways. For example, the device embodiments described above are merely schematic; for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0138] The present application provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, the computer program including program instructions, and when the processor executes the program instructions, the above-mentioned Figure 2 , Figure 6 The speech synthesis method in the corresponding embodiment will not be described here. For technical details not disclosed in the computer-readable storage medium embodiment involved in the present application, please refer to the description of the method embodiment of the present application. As an example, the program instructions can be executed by steps on a computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed in multiple locations and interconnected by a communication network. Multiple computer devices distributed in multiple locations and interconnected by a communication network can constitute a blockchain system.
[0139] The present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device can implement the above-mentioned Figure 2 , Figure 6 The speech synthesis method in the corresponding embodiment will therefore not be described in detail here.
[0140] It should be noted that, for the above-mentioned various method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0141] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the above-mentioned program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The above-mentioned storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).
[0142] The above disclosure is only part of the embodiments of the present application, which certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A speech synthesis method, characterized in that: The method comprises: Acquire information to be processed, wherein the information to be processed includes text to be synthesized, set emotion category, set emotion intensity and emotion intensity change information; Acquire a first emotion basis vector that matches the set emotion category, and determine a target emotion vector according to the set emotion intensity and the first emotion basis vector; Acquire the length of the phoneme sequence of the text to be synthesized; perform expansion processing on the target emotion vector according to the length to obtain an initial emotion vector sequence; adjust specific vector features of all or part of the emotion vectors in the initial emotion vector sequence according to the emotion intensity change information to obtain a first emotion vector sequence; wherein the specific vector feature is a vector feature for indicating emotion intensity, and the length of the first emotion vector sequence matches the length of the phoneme sequence of the text to be synthesized; Speech synthesis processing is performed according to the text to be synthesized and the first emotion vector sequence to obtain a first synthesized speech; the emotion category displayed by the first synthesized speech matches the set emotion category, and the emotion intensity displayed by the first synthesized speech matches the set emotion intensity.
2. The method according to claim 1, characterized in that The step of performing speech synthesis processing according to the text to be synthesized and the first emotion vector sequence to obtain a first synthesized speech includes: The text to be synthesized and the first emotion vector sequence are input into a target acoustic synthesis model for speech synthesis processing to obtain a first synthesized speech; wherein the target acoustic synthesis model is trained by a second emotion vector sequence corresponding to a sample speech and a training text corresponding to the sample speech, and the second emotion vector sequence is obtained by processing the sample speech with an emotion information extraction model.
3. The method according to claim 2, characterized in that The emotion information extraction model includes a spectrum extraction network, an emotion feature extraction network and an emotion representation network, and the method further includes: Inputting the sample speech into the spectrum extraction network for processing to obtain a mel spectrum of the sample speech, and inputting the mel spectrum into the emotion feature extraction network for processing to obtain a reference emotion feature of the sample speech; The reference emotion feature is input into the emotion representation network for processing, and the emotion representation network determines a second emotion basis vector matching the reference emotion feature from a plurality of initial emotion basis vectors, and generates a second emotion vector sequence according to the second emotion basis vector; wherein the length of the second emotion vector sequence matches the length of the phoneme sequence of the training text, and each of the initial emotion basis vectors corresponds to an emotion category; Inputting the training text and the second emotion vector sequence into an initial acoustic synthesis model for speech synthesis processing to obtain a second synthesized speech; A first loss parameter is determined according to the sample speech and the second synthesized speech, and model parameters of the initial acoustic synthesis model are adjusted according to the first loss parameter to obtain the target acoustic synthesis model.
4. The method according to claim 3, wherein the step of inputting the reference emotion feature into the emotion representation network for processing, and determining by the emotion representation network a second emotion basis vector matching the reference emotion feature from a plurality of initial emotion basis vectors, comprises: Inputting the reference emotion feature into the emotion representation network for processing, and determining the similarity between the reference emotion feature and each of the initial emotion basis vectors by the emotion representation network; The emotion representation network determines, according to the similarity, a second emotion basis vector that matches the reference emotion feature from the multiple initial emotion basis vectors.
5. The method according to claim 4, characterized in that The method further comprises: Determining a second loss parameter according to the sample emotion category corresponding to the sample speech and the similarity between the reference emotion feature and each of the initial emotion basis vectors; Adjusting the multiple initial emotion basis vectors according to the second loss parameter to obtain multiple benchmark emotion basis vectors; The number of set emotion categories is M, where M is a positive integer greater than 1; and obtaining a first emotion basis vector matching the set emotion category includes: For a target set emotion category, a first emotion basis vector matching the target set emotion category is determined from the multiple baseline emotion basis vectors; the target set emotion category is any one of the M set emotion categories, and different set emotion categories correspond to different first emotion basis vectors.
6. The method according to any one of claims 1 to 5, characterized in that: The number of the set emotion categories is M, the number of the set emotion strengths is M, each of the set emotion categories corresponds to a set emotion strength, the number of the first emotion basis vectors is M, each of the set emotion categories corresponds to a first emotion basis vector, and M is a positive integer greater than 1; The determining the target emotion vector according to the set emotion intensity and the first emotion base vector comprises: For the target setting emotion category, according to the matching set emotion intensity that matches the target setting emotion category among the M set emotion intensities, a first emotion basis vector that matches the target setting emotion category is weighted to obtain a weighted emotion basis vector corresponding to the target setting emotion category; the target setting emotion category is any one of the M set emotion categories; The weighted emotion basis vectors corresponding to each of the target setting emotion categories are added together to obtain a target emotion vector.
7. The method according to any one of claims 1 to 5, characterized in that: The determining the target emotion vector according to the set emotion intensity and the first emotion base vector comprises: Obtaining a neutral emotion basis vector, and determining a weight parameter of the neutral emotion basis vector according to the set emotion intensity; Performing weighted processing on the neutral sentiment basis vector according to the weight parameter to obtain a weighted neutral sentiment basis vector; Performing weighted processing on the first emotion basis vector according to the set emotion intensity to obtain a weighted first emotion basis vector; The weighted neutral emotion basis vector and the weighted first emotion basis vector are added together to obtain a target emotion vector.
8. A computer device, characterized in that: include: A processor and a memory, wherein the memory stores executable program code, and the processor is used to call the executable program code to implement the speech synthesis method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, which, when executed on a computer, enable the computer to implement the speech synthesis method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Speech synthesis method and device, computer equipment and storage medium
CN115376486A