Speech synthesis method and system, electronic device, and storage medium
Through the conditional dynamic convolutional neural network injecting tone representation, the problem of computing power cost and real-time reduction caused by the improvement of tone diversity in the existing technology is solved, and efficient and robust speech synthesis effect is achieved.
Patent Information
- Application Number
- PCT/CN2025/070132
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-01
- Filing Date
- 2025-01-02
- Publication Date
- 2025-09-04
AI Technical Summary
In the prior art, when personalized speech synthesis models improve timbre diversity, they lead to an increase in computing power costs and a decrease in real-time performance, making it difficult to improve model expressiveness and robustness without large-scale increase in model parameters.
The conditional dynamic convolution neural network is used to construct the first convolution neural network by obtaining tone representation, and the convolution operation is performed using tone representation and text content representation, and splicing and time extension along the time axis, injecting tone information to enhance voice information.
Without large-scale addition of model parameters, the expressiveness and robustness of the speech synthesis model are improved, and more efficient tone modeling is achieved.
Smart Images

Figure CN2025070132_04092025_PF_FP_ABST
Abstract
Description
Speech synthesis method and system, electronic device and storage medium
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on March 1, 2024, with application number "202410238607.7" and invention name "Speech Synthesis Method and System, Electronic Device and Storage Medium", the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a speech synthesis method, a speech synthesis system, an electronic device, and a storage medium. Background Art
[0003] Speech synthesis technology converts text into audio through mechanical and electronic methods, while personalized speech synthesis technology allows users to customize the timbre of the synthesized speech. As shown in Figure 1, the speaker representation (timbre / style representation) is first replicated to the same length as the content representation. The speaker and content representations are then directly concatenated to produce a content representation with global control information for timbre / style. In other words, personalized speech synthesis requires a model that can adapt to the timbre modeling and synthesis of speakers with different timbres and styles. Currently, most multi-timbre models / zero-shot learning timbre modeling models adopt a large model approach, stacking deep neural networks with larger parameters and more layers to increase model capacity and model diverse timbre variations. However, this comes at the cost of increased computing power and compromised inference efficiency and performance, significantly compromising the model's real-time performance. Summary of the Invention
[0004] This application aims to solve at least one of the technical problems existing in the prior art or related art.
[0005] To this end, the first aspect of the present application proposes a speech synthesis method.
[0006] A second aspect of the present application provides a speech synthesis system.
[0007] A third aspect of the present application provides an electronic device.
[0008] A fourth aspect of the present application provides a storage medium.
[0009] A fifth aspect of the present application proposes a computer program product.
[0010] In view of this, according to the first aspect of the present application, a speech synthesis method is proposed, including: obtaining a timbre representation of the timbre to be synthesized; determining a first convolutional neural network based on the timbre representation; obtaining a first representation based on the content representation of the text and the first convolutional neural network, wherein the text includes text content to be output; and obtaining speech information based on the first representation and the timbre representation.
[0011] The speech synthesis method provided in this application mainly includes: first obtaining a timbre representation of the timbre to be synthesized, which may include the timbre or style representation of a specific speaker or other specific timbre representation. The method of obtaining can be to obtain it directly from other places, or to obtain it by processing the acoustic features of the reference audio input by the speaker. After obtaining the timbre representation, a first convolutional neural network is constructed according to the determined timbre representation. That is, the first convolutional neural network is a conditional dynamic convolutional neural network, and its conditional dynamics are controlled by the timbre representation, so that the timbre representation is associated with the first convolutional neural network, and different timbre representations will result in different first convolutional neural networks. After obtaining the first convolutional neural network, a first representation is obtained based on the content representation of the text and the first convolutional neural network, that is, the content representation of the text is convolved using the first convolutional neural network. Since the first convolutional neural network is associated with the timbre representation, the speaker's timbre or style information is injected into the content representation of the text during the convolution operation of the first convolutional neural network on the content representation of the text, thereby obtaining the first representation. Among them, the content representation of the text refers to the text content that the model wants to output, which can be obtained by encoding the text. After obtaining the first representation, since the first representation is obtained by performing a convolution operation on the content representation of the text, it is also necessary to combine and process the first representation with the timbre representation to obtain the final voice information, wherein combining the first representation with the timbre representation can be to splice the first representation and the timbre representation along the time axis, and then process the spliced representation to obtain the final voice information. By combining and processing the first representation and the timbre representation, the timbre representation is injected into the content representation again. The present application constructs a conditional dynamic convolutional neural network based on the timbre representation of the speaker, and then uses the conditional dynamic convolutional neural network to perform a convolution operation on the content representation of the text, so that the timbre information of the speaker is injected into the content representation, thereby achieving the improvement of expressiveness without increasing the model parameters too much, and thus achieving more efficient and robust timbre modeling without increasing the model parameters on a large scale.
[0012] The speech synthesis method according to the present application may also have the following technical features:
[0013] In some technical solutions, optionally, the step of determining the first convolutional neural network based on the timbre representation includes: determining the first convolution kernel parameters based on the timbre representation and the multilayer perceptron; and determining the first convolutional neural network based on the first convolution kernel parameters.
[0014] In this technical solution, the step of determining the first convolutional neural network based on the timbre representation includes: first, it is necessary to construct a multilayer perceptron (MLP), wherein the multilayer perceptron can be composed of multiple stacked linear layer structures. Then, the first convolution kernel parameters are obtained based on the timbre representation and the multilayer perceptron, wherein the timbre representation is input into the multilayer perceptron, and the multilayer perceptron performs a nonlinear transformation on the input representation to obtain a higher-level representation. This process is achieved by learning weight parameters, which can be optimized and adjusted based on training data and target results. Then, in the multilayer perceptron, the first convolution kernel parameters are updated through the backpropagation algorithm and the gradient descent optimization algorithm. These parameters are adjusted based on the input timbre representation and the target result. That is, by inputting the timbre representation into the multilayer perceptron, the first convolution kernel parameters associated with the timbre representation are determined. Then, the first convolution neural network is constructed based on the obtained first convolution kernel parameters. This application determines the first convolution kernel parameters by utilizing timbre representation, thereby dynamically associating timbre representation with the first convolution kernel parameters, and then determining the first convolution neural network based on the first convolution kernel parameters, so that the first convolution neural network becomes a dynamic convolution neural network based on timbre representation, thereby achieving the technical effect of improving the expressiveness of the model without increasing the model size too much.
[0015] In some technical solutions, optionally, the first convolutional neural network is a conditional dynamic convolutional neural network.
[0016] In this technical solution, the first convolutional neural network can be a conditional dynamic convolutional neural network. Its conditional dynamics are controlled by the timbre representation, so that the timbre representation is associated with the first convolutional neural network. Different timbre representations will result in different first convolutional neural networks. By using a conditional dynamic convolutional neural network, the expressiveness of the model is improved without significantly increasing the model size.
[0017] In some technical solutions, optionally, the step of obtaining speech information based on the first representation and the timbre representation includes: splicing the first representation and the timbre representation to obtain a second representation; extending the second representation in time to obtain a third representation; and obtaining speech information based on the third representation and the timbre representation.
[0018] In this technical solution, the step of obtaining speech information based on a first representation and a timbre representation includes: first, concatenating the first representation and the timbre representation to obtain a second representation. That is, first, the length of the timbre representation is extended to the length of the first representation, and then concatenating the first representation and the timbre representation after the timbre extension, thereby achieving the technical effect of adding a global timbre representation along the time axis of the first representation to enhance timbre control. After obtaining the second representation, the second representation is time-extended to obtain a third representation. That is, the second representation is first used to predict the duration of the speech information. This can be achieved by training a duration prediction model. The duration prediction model can estimate the length of the speech information based on the second representation, i.e., the content of the speech information. In this way, an extended content representation, i.e., the third representation, is obtained, which has the same or similar length as the speech information. Finally, the third representation is processed based on the timbre representation to further enhance the timbre representation in the speech information, thereby obtaining the final speech information. By concatenating the first representation and the timbre representation, the timbre representation is reinjected into the content representation, and the second representation is time-extended so that the obtained speech information can maintain the preset duration.
[0019] In some technical solutions, optionally, the step of extending the duration of the second representation to obtain a third representation includes: determining the length of the voice information based on the content representation of the text; predicting the duration of the voice information based on the length of the voice information; and extending the duration of the second representation based on the duration of the voice information to obtain the third representation.
[0020] In this technical solution, the step of temporally extending the second representation to obtain the third representation includes: first, determining the length of the final voice message output based on the text content representation; then, determining the duration of the voice message based on the length of the voice message, i.e., the time required after the voice message is completely output; and finally, temporally extending the second representation based on the duration of the voice message, such that the second representation becomes an extended content representation, i.e., the third representation, having the same or a similar length as the voice message. By temporally extending the second representation to obtain the third representation, the preset duration of the voice message is guaranteed.
[0021] In some technical solutions, optionally, the step of obtaining speech information based on the third representation and the timbre representation includes: determining a second convolutional neural network based on the timbre representation; obtaining a fourth representation based on the third representation and the second convolutional neural network; and obtaining speech information based on the fourth representation and the timbre representation.
[0022] In this technical solution, the step of obtaining speech information based on the third representation and the timbre representation includes: first, determining a second convolutional neural network based on the timbre representation, wherein the second convolutional neural network can be the same as or different from the first convolutional neural network. That is, after obtaining the third representation, timbre feature injection is required again. During this injection process, the convolution kernel parameters obtained by the multilayer perceptron for the timbre feature can be the same as or different from the first convolution kernel parameters, and the determining factor is the parameters of the multilayer perceptron used. After obtaining the second convolutional neural network, the previously obtained third representation is input into the second convolutional neural network to obtain a fourth representation, and then the final speech information is obtained based on the fourth representation and the timbre representation. The fourth representation and the timbre representation can be concatenated in the same way as the first representation and the timbre representation. After concatenation, a content representation with enhanced style control is obtained, which is then decoded to obtain the final acoustic features, and then passed through the vocoder to obtain the final synthesized speech information. By injecting timbre features into the third representation again, the timbre features in the speech information are enhanced.
[0023] According to the second aspect of the present application, a speech synthesis system is proposed, including: a first acquisition module, the first acquisition module is used to obtain a timbre representation of the timbre to be synthesized; a first determination module, the first determination module is used to determine a first convolutional neural network based on the timbre representation; a first processing module, the first processing module is used to obtain a first representation based on the content representation of the text and the first convolutional neural network; and a second processing module, the second processing module is used to obtain speech information based on the first representation and the timbre representation.
[0024] The speech synthesis system provided by the present application specifically includes: a first acquisition module, a first determination module, a first processing module, and a second processing module. Among them, the first acquisition module first obtains the timbre representation of the timbre to be synthesized, and the timbre representation of the timbre to be synthesized may include the timbre or style representation of a specific speaker or other specific timbre representation. The acquisition method may be to obtain it directly from other places, or to obtain it after processing the reference audio acoustic features input by the speaker. After obtaining the timbre representation, the first determination module constructs a first convolutional neural network based on the determined timbre representation, that is, the first convolutional neural network is a conditional dynamic convolutional neural network, and its conditional dynamics is controlled by the timbre representation, so that the timbre representation is associated with the first convolutional neural network, and different timbre representations will obtain different first convolutional neural networks. After obtaining the first convolutional neural network, the first processing module obtains the first representation based on the content representation of the text and the first convolutional neural network, that is, the content representation of the text is convolved using the first convolutional neural network. Since the first convolutional neural network is associated with the timbre representation, the timbre or style information of the speaker will be injected into the content representation of the text during the convolution operation of the first convolutional neural network on the content representation of the text, thereby obtaining the first representation. Among them, the content representation of the text refers to the text content that the model wants to output, which can be obtained by encoding the text. After obtaining the first representation, since the first representation is obtained by performing a convolution operation on the content representation of the text, the second processing module also needs to combine and process the first representation and the timbre representation to obtain the final voice information, wherein the first representation and the timbre representation can be combined to splice the first representation and the timbre representation along the time axis, and then process the spliced representation to obtain the final voice information. By combining and processing the first representation and the timbre representation, the timbre representation is injected into the content representation again. The present application constructs a conditional dynamic convolutional neural network based on the timbre representation of the speaker, and then uses the conditional dynamic convolutional neural network to perform a convolution operation on the content representation of the text, so that the timbre information of the speaker is injected into the content representation, thereby achieving the improvement of expressiveness without increasing the model parameters too much, and then achieving more efficient and robust timbre modeling without increasing the model parameters on a large scale.
[0025] According to the third aspect of the present application, an electronic device is proposed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above-mentioned speech synthesis methods when executing the computer program.
[0026] The electronic device provided in this application implements the steps of the above-mentioned speech synthesis method when the processor executes the computer program, which can achieve the technical effects of any of the above-mentioned technical solutions and will not be repeated here.
[0027] According to a fourth aspect of the present application, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned speech synthesis methods are implemented.
[0028] The storage medium provided in this application implements the steps of the above-mentioned speech synthesis method when the computer program is executed by the processor, and can achieve the technical effects of any of the above-mentioned technical solutions, which will not be repeated here.
[0029] According to a fifth aspect of the present application, a computer program product is proposed, comprising a computer program, which, when executed by a processor, implements the steps of the speech synthesis method in any of the above technical solutions.
[0030] The computer program product provided by this technical solution implements the steps of the speech synthesis method of any technical solution of this application, and thus it has all the beneficial effects of the speech synthesis method of any technical solution of this application, which will not be repeated here.
[0031] Additional aspects and advantages of the present application will become apparent in the following description or may be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0033] FIG1 shows a schematic diagram of a speech synthesis method in the related art;
[0034] FIG2 shows a flow chart of a speech synthesis method according to an embodiment of the present invention;
[0035] FIG3 shows a second flow chart of a speech synthesis method according to an embodiment of the present application;
[0036] FIG4 shows a third flow chart of a speech synthesis method according to an embodiment of the present application;
[0037] FIG5 shows a fourth flow chart of a speech synthesis method according to an embodiment of the present application;
[0038] FIG6 shows a fifth flow chart of a speech synthesis method according to an embodiment of the present application;
[0039] FIG7 shows a sixth flow chart of a speech synthesis method according to an embodiment of the present application;
[0040] FIG8 shows a seventh flow chart of a speech synthesis method according to an embodiment of the present application;
[0041] FIG9 shows a schematic block diagram of a speech synthesis system according to an embodiment of the present application;
[0042] FIG10 shows a schematic block diagram of a first determination module according to an embodiment of the present application;
[0043] FIG11 shows a schematic block diagram of a second processing module according to an embodiment of the present application;
[0044] FIG12 shows a schematic block diagram of a third processing module according to an embodiment of the present application;
[0045] FIG13 shows a schematic block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0046] In order to more clearly understand the above-mentioned objects, features and advantages of the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features therein can be combined with each other in the absence of conflict.
[0047] In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present application is not limited to the specific embodiments disclosed below.
[0048] Figure 1 shows a schematic diagram of a speech synthesis method in the related art; as shown in Figure 1, in Figure 1, the speaker representation (timbre representation / style representation) is first copied into a dimension with the same length as the content representation. As shown in Figure 1, the content representation is "hello, world.", and its dimension length is 12, so the speaker representation also needs to be copied into a dimension of 12. The speaker representation and the content representation are then directly spliced together to obtain a content representation with timbre / style global control information. And the splicing process is generally located before the decoder. However, according to the inventor's research, it was found that the timbre modeling ability of this method is not robust enough and the generalization ability is not high.
[0049] FIG2 shows one of the flowcharts of the speech synthesis method according to an embodiment of the present application.
[0050] The method includes:
[0051] Step 202: Acquire a timbre representation of the timbre to be synthesized;
[0052] Step 204: determining a first convolutional neural network based on the timbre representation;
[0053] Step 206: Obtaining a first representation based on the content representation of the text and the first convolutional neural network, wherein the text includes text content to be output;
[0054] Step 208: Obtain voice information according to the first representation and the timbre representation.
[0055] The speech synthesis method provided in this application primarily comprises: first, obtaining a timbre representation of a timbre to be synthesized. The timbre representation of the timbre to be synthesized may include the timbre or style representation of a specific speaker, or other specific timbre representation. The timbre representation may be obtained directly from another source, or may be obtained by processing the acoustic features of reference audio input by the speaker.
[0056] After obtaining the timbre representation, a first convolutional neural network is constructed according to the determined timbre representation. That is to say, the first convolutional neural network is a conditional dynamic convolutional neural network, and its conditional dynamics is controlled by the timbre representation, so that the timbre representation is associated with the first convolutional neural network, and different timbre representations will obtain different first convolutional neural networks.
[0057] After obtaining the first convolutional neural network, a first representation is obtained based on the text content representation and the first convolutional neural network. That is, the first convolutional neural network is used to perform a convolution operation on the text content representation. Since the first convolutional neural network is associated with the timbre representation, the speaker's timbre or style information is injected into the text content representation during the convolution operation on the text content representation by the first convolutional neural network, thereby obtaining the first representation. The text content representation refers to the text content that the model wants to output, which can be obtained by encoding the text.
[0058] After obtaining the first representation, since the first representation is obtained by convolution of the text content representation, it is necessary to combine and process the first representation with the timbre representation to obtain the final voice information. Combining the first representation with the timbre representation can be done by splicing the first representation and the timbre representation along the time axis and then processing the spliced representation to obtain the final voice information. By combining and processing the first representation with the timbre representation, the timbre representation is reintegrated into the content representation.
[0059] This application constructs a conditional dynamic convolutional neural network based on the speaker's timbre representation, and then uses the conditional dynamic convolutional neural network to perform convolution operations on the content representation of the text, so that the speaker's timbre information is injected into the content representation, thereby improving the expressiveness without increasing the model parameters too much, and thus achieving more efficient and robust timbre modeling without increasing the model parameters on a large scale.
[0060] Figure 3 shows the second flow chart of the speech synthesis method of an embodiment of the present application; as shown in Figure 3, the speech synthesis method provided by the present application first uses the speaker representation to pass through one or more layers of linear layer networks to obtain convolution kernel parameters dynamically associated with the speaker representation, and then constructs a conditional dynamic convolutional neural network based on the convolution kernel parameters, wherein the conditional dynamic refers to the dynamic change of the convolution kernel parameters based on the speaker representation.
[0061] The content representation is then convolved using a conditional dynamic convolutional neural network to obtain a content representation infused with personalized timbre / style information. This content representation is then synthesized and decoded to obtain the final speech information. This demonstrates that compared to related technologies, this application can achieve more efficient and robust timbre modeling without significantly increasing the model size.
[0062] FIG4 shows a third flow chart of a speech synthesis method according to an embodiment of the present application; wherein the step of determining the first convolutional neural network based on the timbre representation includes:
[0063] Step 402: Determine the first convolution kernel parameters based on the timbre representation and the multi-layer perceptron;
[0064] Step 404: Determine a first convolutional neural network according to the first convolution kernel parameters.
[0065] In this embodiment, the step of determining the first convolutional neural network based on the timbre representation includes: first, it is necessary to construct a multilayer perceptron (MLP), wherein the multilayer perceptron can be composed of multiple stacked linear layer structures. Then, according to the timbre representation and the multilayer perceptron, the first convolution kernel parameters are obtained, wherein the timbre representation is input into the multilayer perceptron, and the multilayer perceptron performs a nonlinear transformation on the input representation to obtain a higher-level representation. This process is achieved by learning weight parameters, and the weight parameters can be optimized and adjusted according to the training data and the target result. Then, in the multilayer perceptron, the first convolution kernel parameters are updated by the back propagation algorithm and the gradient descent optimization algorithm. These parameters are adjusted according to the input timbre representation and the target result. That is, by inputting the timbre representation into the multilayer perceptron, the first convolution kernel parameters associated with the timbre representation are determined. Then, the first convolutional neural network is constructed based on the obtained first convolution kernel parameters. This application determines the first convolution kernel parameters by utilizing timbre representation, thereby dynamically associating timbre representation with the first convolution kernel parameters, and then determining the first convolution neural network based on the first convolution kernel parameters, so that the first convolution neural network becomes a dynamic convolution neural network based on timbre representation, thereby achieving the technical effect of improving the expressiveness of the model without increasing the model size too much.
[0066] In some embodiments, optionally, the first convolutional neural network is a conditional dynamic convolutional neural network.
[0067] In this embodiment, the first convolutional neural network can be a conditional dynamic convolutional neural network. Its conditional dynamics are controlled by the timbre representation, so that the timbre representation is associated with the first convolutional neural network. Different timbre representations will result in different first convolutional neural networks. By using a conditional dynamic convolutional neural network, the expressiveness of the model is improved without significantly increasing the model size.
[0068] FIG5 shows a fourth flow chart of a speech synthesis method according to an embodiment of the present application. The step of obtaining speech information according to the first representation and the timbre representation includes:
[0069] Step 502: obtaining a second representation by concatenating the first representation and the timbre representation;
[0070] Step 504: Extend the second representation to obtain a third representation;
[0071] Step 506: Obtain speech information according to the third representation and the timbre representation.
[0072] In this embodiment, the step of obtaining speech information based on the first representation and the timbre representation includes: first, concatenating the first representation and the timbre representation to obtain a second representation. That is, first, the length of the timbre representation is extended to the length of the first representation, and then concatenating the first representation and the timbre representation after the timbre extension, thereby achieving the technical effect of adding a global timbre representation along the time axis of the first representation to enhance timbre control. After obtaining the second representation, the second representation is time-extended to obtain a third representation. That is, the second representation is first used to predict the duration of the speech information. This can be achieved by training a duration prediction model. The duration prediction model can estimate the length of the speech information based on the second representation, i.e., the content of the speech information. The second representation is then time-extended based on the predicted duration, thereby obtaining an extended content representation, i.e., a third representation, having the same or similar duration as the speech information. Finally, the third representation is processed based on the timbre representation to further enhance the timbre representation in the speech information, thereby obtaining the final speech information. By splicing the first representation and the timbre representation, the timbre representation is injected into the content representation again, and at the same time, the second representation is time-extended so that the obtained voice information can ensure the preset duration.
[0073] FIG6 shows a fifth flow chart of a speech synthesis method according to an embodiment of the present application; wherein the step of performing time extension on the second representation to obtain a third representation includes:
[0074] Step 602: Determine the length of the voice information based on the content representation of the text;
[0075] Step 604: predicting the duration of the voice message based on the length of the voice message;
[0076] Step 606: Duration-extend the second representation according to the duration of the voice information to obtain a third representation.
[0077] In this embodiment, the step of temporally extending the second representation to obtain the third representation includes: first, determining the length of the final voice message output based on the text content representation; then, determining the duration of the voice message based on the length of the voice message, i.e., the time required after the voice message is completely output; and finally, temporally extending the second representation based on the duration of the voice message, such that the second representation becomes an extended content representation, i.e., the third representation, having the same or a similar length as the voice message. By temporally extending the second representation to obtain the third representation, the preset duration of the voice message is ensured.
[0078] FIG7 shows a sixth flow chart of a speech synthesis method according to an embodiment of the present application; wherein the step of obtaining speech information according to the third representation and the timbre representation includes:
[0079] Step 702: Determine a second convolutional neural network based on the timbre representation;
[0080] Step 704: Obtain a fourth representation based on the third representation and the second convolutional neural network;
[0081] Step 706: Obtain speech information according to the fourth representation and the timbre representation.
[0082] In this embodiment, the step of obtaining speech information based on the third representation and the timbre representation includes: first, determining a second convolutional neural network based on the timbre representation, wherein the second convolutional neural network can be the same as or different from the first convolutional neural network. That is, after obtaining the third representation, timbre feature injection is required again. During this injection process, the convolution kernel parameters obtained by the multilayer perceptron for the timbre feature can be the same as or different from the first convolution kernel parameters, and the determining factor is the parameters of the multilayer perceptron used. After obtaining the second convolutional neural network, the previously obtained third representation is input into the second convolutional neural network to obtain a fourth representation, and then the final speech information is obtained based on the fourth representation and the timbre representation. The fourth representation and the timbre representation can be concatenated in the same manner as the concatenation method for the first representation and the timbre representation. After concatenation, a content representation with enhanced style control is obtained, which is then decoded to obtain the final acoustic features, and then passed through a vocoder to obtain the final synthesized speech information. By injecting timbre features into the third representation again, the timbre features in the speech information are enhanced.
[0083] FIG8 shows the seventh flow chart of the speech synthesis method of an embodiment of the present application; wherein, the speech synthesis method is to first input the reference audio acoustic features into a reference audio encoder to obtain a speaker representation, which includes a timbre representation and a style representation. This speaker representation is then passed through a stacked linear layer structure to obtain a two-dimensional convolution kernel parameter about the dynamic change of the speaker representation, and this convolution kernel parameter is assigned to a two-dimensional convolution operation, that is, the convolution neural network is determined according to the convolution kernel parameter, and at the same time, the text of the speech content is input into the encoder for encoding to obtain a content representation, and then the content representation is convolved using the convolution neural network to obtain a content representation injected by the dynamic convolution of the speaker representation, and at the same time, the obtained content representation and the speaker representation are repeatedly expanded along the time axis. The vector is spliced in the feature dimension to obtain a content representation that strengthens the style control. Thus, the first speaker representation injection is completed. The content representation with enhanced style control is then used to predict the duration of the speech information. Based on the predicted duration, the duration is extended in the duration extension model to obtain a content representation with the same duration as the speech information. This content representation is then input into the decoder, where the speaker representation is repeatedly injected. This means that the same speaker representation is again injected through a linear layer structure to obtain convolution kernel parameters. The linear layer structure of this injection can be the same as or different from the linear layer structure of the first injection. If the linear layer structure is the same, the final convolution kernel parameters are also the same; if they are different, the final convolution kernel parameters are different. After obtaining the convolution kernel parameters, a convolutional neural network is determined based on the convolution kernel parameters. The convolutional neural network is then used to convolve the previously obtained content representation, i.e., the content representation with the same duration as the speech information, to obtain a new content representation. This new content representation is then concatenated with the timbre representation to obtain the final content representation. Finally, the final content representation is decoded to obtain the final acoustic features, which are then sent to the vocoder to obtain the synthesized audio, i.e., the final speech information.
[0084] FIG9 shows a schematic block diagram of a speech synthesis system according to an embodiment of the present application; wherein the speech synthesis system 90 includes:
[0085] A first acquisition module 902 is used to obtain a timbre representation of a timbre to be synthesized;
[0086] A first determining module 904 is configured to determine a first convolutional neural network based on the timbre representation after determining the timbre representation;
[0087] A first processing module 906 is configured to obtain a first representation based on the content representation of the text and the first convolutional neural network;
[0088] The second processing module 908 is configured to obtain speech information according to the first representation and the timbre representation.
[0089] The speech synthesis system 90 provided in the present application specifically includes: a first acquisition module 902, a first determination module 904, a first processing module 906, and a second processing module 908. Among them, first, the first acquisition module 902 obtains the timbre representation of the timbre to be synthesized, and the timbre representation of the timbre to be synthesized may include the timbre or style representation of a specific speaker or other specific timbre representation. The acquisition method may be to obtain it directly from other places, or to obtain it after processing the reference audio acoustic features input by the speaker. After obtaining the timbre representation, the first determination module 904 constructs a first convolutional neural network according to the determined timbre representation, that is, the first convolutional neural network is a conditional dynamic convolutional neural network, and its conditional dynamics are controlled by the timbre representation, so that the timbre representation is associated with the first convolutional neural network, and different timbre representations will obtain different first convolutional neural networks. After obtaining the first convolutional neural network, the first processing module 906 obtains the first representation based on the content representation of the text and the first convolutional neural network, that is, the content representation of the text is convolved using the first convolutional neural network. Since the first convolutional neural network is associated with the timbre representation, the speaker's timbre or style information is injected into the content representation of the text during the convolution operation of the first convolutional neural network on the content representation of the text, thereby obtaining the first representation. Among them, the content representation of the text refers to the text content that the model wants to output, which can be obtained by encoding the text. After obtaining the first representation, since the first representation is obtained by convolution operation of the content representation of the text, the second processing module 908 also needs to combine and process the first representation and the timbre representation to obtain the final voice information, wherein combining the first representation and the timbre representation can be splicing the first representation and the timbre representation along the time axis, and then processing the spliced representation to obtain the final voice information. By combining and processing the first representation and the timbre representation, the timbre representation is injected into the content representation again. This application constructs a conditional dynamic convolutional neural network based on the speaker's timbre representation, and then uses the conditional dynamic convolutional neural network to perform convolution operations on the content representation of the text, so that the speaker's timbre information is injected into the content representation, thereby improving the expressiveness without increasing the model parameters too much, and thus achieving more efficient and robust timbre modeling without increasing the model parameters on a large scale.
[0090] FIG10 shows a schematic block diagram of a first determination module according to an embodiment of the present application; wherein the first determination module 904 includes:
[0091] A second determination module 9042 is configured to determine first convolution kernel parameters based on the timbre representation and the multi-layer perceptron;
[0092] The third determination module 9044 is used to determine the first convolutional neural network according to the first convolution kernel parameters.
[0093] In this embodiment, the first determination module 904 includes: a second determination module 9042 and a third determination module 9044. First, a multilayer perceptron (MLP) needs to be constructed, wherein the MLP can be composed of multiple stacked linear layers. The second determination module 9042 then obtains first convolution kernel parameters based on the timbre representation and the MLP. The timbre representation is input into the MLP, which performs a nonlinear transformation on the input representation to obtain a higher-level representation. This process is achieved by learning weight parameters, which can be optimized and adjusted based on training data and target results. The first convolution kernel parameters are then updated in the MLP using a backpropagation algorithm and a gradient descent optimization algorithm. These parameters are adjusted based on the input timbre representation and the target result. In other words, by inputting the timbre representation into the MLP, the first convolution kernel parameters associated with the timbre representation are determined. Finally, the third determination module 9044 constructs a first convolutional neural network based on the obtained first convolution kernel parameters. This application determines the first convolution kernel parameters by utilizing timbre representation, thereby dynamically associating timbre representation with the first convolution kernel parameters, and then determining the first convolution neural network based on the first convolution kernel parameters, so that the first convolution neural network becomes a dynamic convolution neural network based on timbre representation, thereby achieving the technical effect of improving the expressiveness of the model without increasing the model size too much.
[0094] In some embodiments, optionally, the first convolutional neural network is a conditional dynamic convolutional neural network.
[0095] In this embodiment, the first convolutional neural network can be a conditional dynamic convolutional neural network. Its conditional dynamics are controlled by the timbre representation, so that the timbre representation is associated with the first convolutional neural network. Different timbre representations will result in different first convolutional neural networks. By using a conditional dynamic convolutional neural network, the expressiveness of the model is improved without significantly increasing the model size.
[0096] FIG11 shows a schematic block diagram of a second processing module according to an embodiment of the present application; wherein the second processing module 908 includes:
[0097] a concatenation module 9082 for concatenating the first representation and the timbre representation to obtain a second representation;
[0098] An expansion module 9084 is configured to perform time extension on the second representation to obtain a third representation;
[0099] The third processing module 9086 is configured to obtain speech information according to the third representation and the timbre representation.
[0100] In this embodiment, the second processing module 908 includes a splicing module 9082, an expansion module 9084, and a third processing module 9086. The splicing module 9082 splices the first representation and the timbre representation to obtain a second representation. Specifically, the splicing module 9082 first expands the length of the timbre representation to the length of the first representation, and then splices the first representation with the expanded timbre representation, thereby achieving the technical effect of adding a global timbre representation along the time axis of the first representation to enhance timbre control. After obtaining the second representation, the expansion module 9084 performs a time-length expansion on the second representation to obtain a third representation. Specifically, the second representation is first used to predict the duration of the speech information. This can be achieved by training a time-length prediction model that can estimate the length of the speech information based on the second representation, i.e., the content of the speech information. The second representation is then time-extended based on the predicted duration, thereby obtaining an expanded content representation, i.e., the third representation, having the same or similar duration as the speech information. Finally, the third processing module 9086 performs processing based on the third representation and the timbre representation to further enhance the timbre representation in the speech information, thereby obtaining the final speech information. By splicing the first representation and the timbre representation, the timbre representation is injected into the content representation again, and at the same time, the second representation is time-extended so that the obtained voice information can ensure the preset duration.
[0101] In some embodiments, the expansion module 9084 is specifically used to determine the length of the voice information based on the content representation of the text; predict the duration of the voice information based on the length of the voice information; and expand the second representation based on the duration of the voice information to obtain a third representation.
[0102] In this embodiment, expansion module 9084 first determines the length of the final output voice message based on the text content representation. It then determines the duration of the voice message based on the length of the voice message, i.e., the time required after the voice message is completely output. Finally, the second representation is time-extended based on the duration of the voice message, resulting in the second representation becoming an expanded content representation, i.e., a third representation, having the same or similar duration as the voice message. By time-extending the second representation to obtain the third representation, the preset duration of the voice message is ensured.
[0103] FIG12 shows a schematic block diagram of a third processing module according to an embodiment of the present application; wherein the third processing module 9086 includes:
[0104] A fourth determination module 9088 is configured to determine a second convolutional neural network based on the timbre representation;
[0105] A fourth processing module 9090 is configured to obtain a fourth representation based on the third representation and the second convolutional neural network;
[0106] The fifth processing module 9092 is configured to obtain speech information according to the fourth representation and the timbre representation.
[0107] In this embodiment, the third processing module 9086 includes: a fourth determination module 9088, a fourth processing module 9090, and a fifth processing module 9092. First, the fourth determination module 9088 determines a second convolutional neural network based on the timbre representation, wherein the second convolutional neural network can be the same as or different from the first convolutional neural network. That is, after obtaining the third representation, timbre feature injection is required again. During this injection process, the convolution kernel parameters obtained by the timbre feature through the multilayer perceptron can be the same as or different from the first convolution kernel parameters, and the determining factor is the parameters of the multilayer perceptron used. After obtaining the second convolutional neural network, the fourth processing module 9090 inputs the previously obtained third representation into the second convolutional neural network to obtain a fourth representation. Then, the fifth processing module 9092 obtains the final speech information based on the fourth representation and the timbre representation. The fourth representation and the timbre representation can be spliced together in the same way as the first representation and the timbre representation. After concatenation, a content representation with enhanced style control is obtained. This content representation is then decoded to obtain the final acoustic features, which are then passed through a vocoder to obtain the final synthesized speech information. By injecting timbre features into the third representation again, the timbre features in the speech information are strengthened.
[0108] Figure 13 shows a schematic block diagram of an electronic device of an embodiment of the present application; wherein, the electronic device 130 includes a memory 1302, a processor 1304, and a computer program stored in the memory 1302 and executable on the processor 1304, and when the processor 1304 executes the computer program, the steps of the speech synthesis method as described above are implemented.
[0109] The electronic device 130 provided in this application, in which the processor 1304 implements the steps of the above-mentioned speech synthesis method when executing the computer program, can achieve the technical effects of any of the above-mentioned embodiments and will not be repeated here.
[0110] One embodiment of the present application provides a storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of any of the above-mentioned speech synthesis methods are implemented.
[0111] The storage medium provided in this application implements the steps of the above-mentioned speech synthesis method when the computer program is executed by the processor, and can achieve the technical effects of any of the above-mentioned embodiments, which will not be repeated here.
[0112] A storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The storage medium can be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. A non-exhaustive list of more specific examples of storage media includes: a portable computer floppy disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory card, a floppy disk, an encoding mechanical device (such as a punched card or a groove with a raised structure on which instructions are recorded), and any suitable combination of the above. The storage medium used herein should not be understood as a transmission signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium, or electrical signals transmitted through wires.
[0113] One embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the speech synthesis method in any of the above embodiments.
[0114] The computer program product provided in this embodiment implements the steps of the speech synthesis method of any embodiment of the present application, and thus has all the beneficial effects of the speech synthesis method of any embodiment of the present application, which will not be repeated here.
[0115] In this specification, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance, unless otherwise expressly specified or limited. Terms such as "connect," "install," and "fix" should be interpreted broadly. For example, "connect" can refer to a fixed connection, a detachable connection, or an integral connection; it can refer to a direct connection or an indirect connection through an intermediary. Those skilled in the art will understand the specific meanings of these terms in this application based on specific circumstances.
[0116] Throughout this specification, terms such as "one embodiment," "some embodiments," and "specific embodiments" mean that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present application. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0117] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A speech synthesis method, wherein: include: Obtaining a timbre representation of a timbre to be synthesized; Determining a first convolutional neural network based on the timbre representation; Obtaining a first representation based on a content representation of a text and the first convolutional neural network, wherein the text includes text content to be output; Speech information is obtained according to the first representation and the timbre representation.
2. The speech synthesis method according to claim 1, wherein: The step of determining a first convolutional neural network according to the timbre representation comprises: Determining a first convolution kernel parameter according to the timbre representation and the multilayer perceptron; Determine the first convolutional neural network according to the first convolution kernel parameters.
3. The speech synthesis method according to claim 1 or 2, wherein: The first convolutional neural network is a conditional dynamic convolutional neural network.
4. The speech synthesis method according to claim 1 or 2, wherein: The step of obtaining voice information according to the first representation and the timbre representation includes: splicing the first representation and the timbre representation to obtain a second representation; performing time extension on the second representation to obtain a third representation; The speech information is obtained according to the third representation and the timbre representation.
5. The speech synthesis method according to claim 4, wherein: The step of performing time extension on the second representation to obtain a third representation includes: Determining the length of the voice information based on the content representation of the text; Predicting the duration of the voice information according to the length of the voice information; The second representation is time-extended according to the duration of the voice information to obtain the third representation.
6. The speech synthesis method according to claim 4, wherein: The step of obtaining the voice information according to the third representation and the timbre representation includes: Determining a second convolutional neural network based on the timbre representation; Obtaining a fourth representation based on the third representation and the second convolutional neural network; The speech information is obtained according to the fourth representation and the timbre representation.
7. A speech synthesis system, wherein: include: a first acquisition module, configured to acquire a timbre representation of a timbre to be synthesized; a first determining module, configured to determine a first convolutional neural network according to the timbre representation; a first processing module, configured to obtain a first representation based on a content representation of a text and the first convolutional neural network, wherein the text includes text content to be output; The second processing module is used to obtain voice information according to the first representation and the timbre representation.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the speech synthesis method according to any one of claims 1 to 6 are implemented.
9. A storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the steps of the speech synthesis method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising computer instructions, wherein: When the computer instructions are executed by a processor, the steps of the speech synthesis method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Speech synthesis method for generating new tone
CN110459201A
Speech synthesis method, device and equipment and storage medium
CN112382270A
Speech synthesis method and system for new tone generation
CN112802448A
Chinese speech synthesis method based on diffusion probability model
CN114023300A
Voice processing method and device, computer equipment and computer readable storage medium
CN114708849A