Speech synthesis method and apparatus, storage medium, and electronic device
By adjusting the parameters of the pre-trained decoder and constructing the target speech synthesis model, the stability and sound quality issues in the speech synthesis scheme were resolved, achieving high-quality speech synthesis results.
Patent Information
- Application Number
- CN202210179826.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-25
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2042-02-25
AI Technical Summary
Existing speech synthesis schemes rely on speaker feature extraction networks, resulting in poor stability and sound quality of synthesized speech. Furthermore, noise artifacts are easily embedded in the speech synthesis system, affecting the synthesis effect.
By acquiring the target speaker's speech, extracting feature vectors using the first encoder, and adjusting the parameters of the pre-trained decoder, a target speech synthesis model is constructed, avoiding adjustments to the second encoder and ensuring the stability and accuracy of speech synthesis.
It achieves stability and accuracy in speech synthesis under limited conditions, avoids solidification with noisy quality, and improves the sound quality of speech synthesis.
Smart Images

Figure CN114495901B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of audio processing, in particular, to a speech synthesis method and device, a storage medium and an electronic device. BACKGROUND
[0002] In the field of speech synthesis, in general application scenarios, a large amount of data (more than 5h) is needed to support the synthesis to have a relatively stable effect. For most users, it is unrealistic to record 5h of data according to strict specifications, and for regular users, when synthesizing their own voice, they pay more attention to the effect of the synthesized voice and their own voice in terms of timbre and tone. How to enhance the pronunciation stability of the speech synthesis system itself and improve the sound quality as much as possible while ensuring the timbre effect of the user is a problem that needs to be solved.
[0003] In existing speech synthesis solutions, it is usually necessary to absolutely rely on a speaker feature extraction network with extremely strong decoupling capability, that is, the synthesized speech and the target speaker speech authorized by the user for use absolutely rely on the capability of the speaker feature extraction network, but the capability of the speaker feature extraction network in the prior art cannot completely meet the needs in this scenario. In addition, some speech synthesis solutions first retrain a pre-trained speech synthesis system using the target speaker speech authorized by the user for use to achieve the effect of the synthesized timbre, but since the purpose of the speech synthesis system is to synthesize speech with sound quality information, if the target speaker speech authorized by the user for use is noisy, the speech synthesis system trained will also include the noisy sound quality information, resulting in the problem that the speech synthesized according to the text is also noisy. SUMMARY
[0004] This section is provided to summarize the concepts in a brief manner, which will be described in detail in the following detailed description section. This section is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0005] In a first aspect, the present disclosure provides a speech synthesis method, comprising:
[0006] obtaining a target speaker speech;
[0007] extracting a first feature vector of the target speaker speech through a first encoder, and extracting a target speaker voice feature of a target speaker in the target speaker speech through a speaker feature extraction network;
[0008] According to the first feature vector, the target speaker voice feature, and the target speaker voice, parameter adjustment is performed on a first decoder, wherein the first decoder is a decoder that has been pre-trained.
[0009] A target speech synthesis model is constructed by using the first decoder after the parameter adjustment and a second encoder, wherein the second encoder is pre-trained.
[0010] The target speech synthesis model is used to synthesize target speech corresponding to the target speaker by inputting a text to be synthesized and the target speaker voice feature into the target speech synthesis model.
[0011] In a second aspect, the present disclosure provides a speech synthesis device, and the device comprises:
[0012] An acquisition module is configured to acquire target speaker voice.
[0013] A first processing module is configured to extract a first feature vector of the target speaker voice by using a first encoder, and extract a target speaker voice feature of a target speaker from the target speaker voice by using a speaker feature extraction network.
[0014] A second processing module is configured to perform parameter adjustment on a first decoder according to the first feature vector, the target speaker voice feature, and the target speaker voice, wherein the first decoder is a decoder that has been pre-trained.
[0015] A third processing module is configured to construct a target speech synthesis model by using the first decoder after the parameter adjustment and a second encoder, wherein the second encoder is pre-trained.
[0016] A speech synthesis module is configured to synthesize target speech corresponding to the target speaker by inputting a text to be synthesized and the target speaker voice feature into the target speech synthesis model.
[0017] In a third aspect, the present disclosure provides a computer readable medium having a computer program stored thereon, wherein the program is executed by a processing device to implement the steps of the method according to the first aspect.
[0018] In a fourth aspect, the present disclosure provides an electronic device, and the device comprises:
[0019] A storage device having at least one computer program stored thereon.
[0020] At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method according to the first aspect.
[0021] By the technical solution, when a target speech corresponding to a target speaker who has obtained user authorization to use is to be generated according to the voice of the target speaker himself who has obtained user authorization to use, the first decoder in the speech synthesis model can be adjusted in parameters only by the speaker voice who has obtained user authorization to use and the first encoder, without adjusting the second encoder in the speech synthesis model, so that the noise in the voice of the speaker who has obtained user authorization to use used for parameter adjustment can be avoided from being fixed in the target speech synthesis model, and the problem that all the target speeches synthesized by the target speech synthesis model are noisy can be avoided, and since the first decoder can be adjusted in parameters by the speaker voice who has obtained user authorization to use before speech synthesis, the synthesis of the target speech does not need to completely rely on the ability of the speaker feature extraction network to extract the speaker voice features of the speaker who has obtained user authorization to use, and the stability and precision of speech synthesis according to the speaker voice features of the speaker who has obtained user authorization to use extracted under limited conditions are ensured.
[0022] Other features and advantages of the present disclosure will be described in detail in the following detailed description section. BRIEF DESCRIPTION OF DRAWINGS
[0023] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:
[0024] Figure 1 is a flowchart of a speech synthesis method according to an exemplary embodiment of the present disclosure.
[0025] Figure 2 is a flowchart of a speech synthesis method according to another exemplary embodiment of the present disclosure.
[0026] Figure 3 is a flowchart of a speech synthesis method according to another exemplary embodiment of the present disclosure.
[0027] Figure 4 is a model structure diagram in a speech synthesis method according to another exemplary embodiment of the present disclosure.
[0028] Figure 5 is a structural block diagram of a speech synthesis device according to an exemplary embodiment of the present disclosure.
[0029] Figure 6 is a structural block diagram of a speech synthesis device according to another exemplary embodiment of the present disclosure.
[0030] Figure 7 A structural diagram of an electronic device suitable for implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0031] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. While certain embodiments of the present disclosure will be shown and described, it is to be understood that the present disclosure is not limited to the embodiments to be shown and described, but can be implemented in various forms, and should not be construed as being limited to the embodiments set forth herein. It is to be understood that the drawings and embodiments of the present disclosure are merely for illustrative purposes and are not intended to limit the scope of the present disclosure.
[0032] It should be understood that each step recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.
[0033] The term "comprising" and variations thereof as used herein are used inclusively, i.e., "comprising, but not limited to." The term "based on" is "based, at least in part, on." The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments." Related definitions are given below in the description of the application.
[0034] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.
[0035] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not limiting, and those skilled in the art should understand that, unless otherwise explicitly stated in the context, it should be understood as "one or more".
[0036] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are merely for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0037] All actions of obtaining signals, information or data in the present disclosure are performed in compliance with the corresponding data protection regulations policy of the country of residence and with the authorization given by the owner. All speaker voice, target speaker voice, first feature vector of target speaker voice, target speaker voice feature of target speaker, etc. involved in the present disclosure are also performed in compliance with the corresponding data protection regulations policy of the country of residence and with the authorization given by the owner.
[0038] Figure 1 is a flowchart of a voice synthesis method according to an exemplary embodiment of the present disclosure. As shown in Figure 1 , the method comprises steps 101 to 105.
[0039] In step 101, a target speaker voice is obtained. The target speaker can be a user who needs to perform voice synthesis according to his own voice, or can be any speaker who needs to perform voice synthesis according to his own voice. No matter how the target speaker is determined, as long as the voice of the target speaker can be obtained, the voice of the target speaker is the voice that has been authorized by the user to use. When the target speaker is a user who needs to perform voice synthesis according to his own voice, the target speaker voice can be a piece of semantics that the user is required to input in real time, and if not, any piece of voice of the target speaker that has been authorized by the user to use can be used as the target speaker voice.
[0040] In step 102, a first feature vector of the target speaker voice is extracted by a first encoder, and a target speaker voice feature of the target speaker is extracted in the target speaker voice by a speaker feature extraction network.
[0041] The first encoder is an encoder in a deep learning model, and in the present disclosure, the specific structure of the first encoder is not limited, as long as the first encoder can achieve the functions required by the first encoder. The first feature vector is a feature vector of the voice extracted from the target speaker voice that has been authorized by the user to use. The first encoder can be an encoder that can extract the feature vector of the voice from the voice by being pre-trained in any way, and the specific training method of the encoder is not limited in the present application.
[0042] The speaker feature extraction network is also a network for extracting the target speaker voice feature that has been authorized by a user to use and is related to the voice of the target speaker, and is irrelevant to the text content in the target speaker voice and is only related to the voice tone, pitch, etc. of the target speaker. That is, in a possible implementation, the target speaker voice feature can be the voice tone feature and / or the pitch feature of the target speaker.
[0043] In step 103, a first decoder is parameter adjusted according to the first feature vector, the target speaker voice feature and the target speaker voice, wherein the first decoder is a decoder that has been pre-trained.
[0044] Since the decoder is used to decode the intermediate feature back to text or voice, when the first decoder is parameter adjusted according to the first feature vector, the target speaker voice feature that has been authorized by a user to use and the target speaker voice that has been authorized by a user to use, the first feature vector and the target speaker voice feature can be used as the input of the first decoder, and the target speaker voice can be used as the output of the first decoder, so as to train the first decoder, thereby achieving the effect of parameter adjusting the first decoder according to the target speaker voice.
[0045] If the first decoder before the parameter adjustment is directly used to synthesize the target voice corresponding to the target speaker according to the target speaker voice feature, the synthesized target voice is often quite different from the voice of the target speaker that has been authorized by a user to use, because the first decoder has not encountered the target speaker voice feature. Therefore, by parameter adjusting the first decoder according to the first feature vector, the target speaker voice feature and the target speaker voice, the first decoder can achieve the effect of voice synthesis when synthesizing voice according to the target speaker voice feature that has been authorized by a user to use.
[0046] In step 104, a target voice synthesis model is constructed by the first decoder after the parameter adjustment and a second encoder, wherein the second encoder is pre-trained.
[0047] Since the first encoder used in the parameter adjustment of the first decoder is an encoder for extracting feature vectors of speech, and the speech synthesis scenario to which the present scheme is applied is a scenario of synthesizing target speech that has not been said by a speaker directly from text and the voice of the target speaker, when constructing the target speech synthesis model, an encoder capable of extracting feature vectors from text needs to be additionally pre-trained as the second encoder. The specific structure and training method of the second encoder are not limited in the present disclosure, as long as the second encoder can extract feature vectors from text to achieve the purpose of speech synthesis.
[0048] In step 105, the text to be synthesized and the voice features of the target speaker are input into the target speech synthesis model to synthesize target speech corresponding to the target speaker.
[0049] In a possible implementation, the second encoder, the first encoder and the first decoder can be included in the target speech synthesis model. Since the first decoder is trained by taking the feature vectors extracted from speech by the first encoder as input, when the target speech synthesis model is used to synthesize target speech corresponding to the target speaker, the second encoder can be used to extract text-related feature vectors from the text to be synthesized, and the second decoder obtained by conventional training can be used to synthesize any type of speech corresponding to the text to be synthesized, then the first encoder can be used to extract relevant feature vectors from the speech corresponding to the text to be synthesized, and finally the first decoder can be input with the target speaker voice features that have been authorized by the user to obtain the target speech.
[0050] But the specific structure of the target speech synthesis model is not limited in the present disclosure, as long as the second encoder and the first decoder can be used to generate target speech corresponding to the target speaker from the text to be synthesized and the target speaker voice features that have been authorized by the user.
[0051] By the technical solution, when a target voice corresponding to a target speaker is to be generated according to a voice of the target speaker that has been authorized by a user to use, the first decoder in the voice synthesis model can be adjusted in parameters only by the voice of the speaker that has been authorized by the user to use and the first encoder, without adjusting the second encoder in the voice synthesis model, so that the noise in the voice of the speaker that has been authorized by the user to use used for the parameter adjustment can be prevented from being fixed in the target voice synthesis model, and thus the problem that all target voices synthesized by the target voice synthesis model are noisy can be avoided, and since the first decoder can be adjusted in parameters by the voice of the speaker that has been authorized by the user to use before voice synthesis, the synthesis of the target voice does not need to completely rely on the ability of the speaker feature extraction network to extract the voice feature of the speaker that has been authorized by the user to use, and the stability and precision of voice synthesis according to the voice feature of the speaker that has been authorized by the user to use extracted under limited conditions are ensured.
[0052] Figure 2 is a flowchart of a voice synthesis method according to yet another example embodiment of the present disclosure. As shown in Figure 2 The method further includes steps 201 to 203.
[0053] In step 201, a selection instruction input by a user is acquired, the selection instruction being used to represent a voice style that the user wants to synthesize.
[0054] In step 202, a target second encoder is determined in the at least one second encoder that has been pre-trained according to the selection instruction.
[0055] In step 203, a target voice synthesis model is constructed by the first decoder adjusted in parameters and the target second encoder.
[0056] That is, a plurality of second encoders respectively corresponding to different voice styles can be pre-trained, and before the user needs to synthesize a voice according to a voice of the user, the user can also select a desired voice style according to a demand, the voice style can be, for example, cheerful, sentimental, cute, or any other pre-defined voice style, and the second encoders corresponding to various voice styles can be obtained by training the second encoders according to the training data of the related styles of the pre-defined voice styles.
[0057] After the user selects a voice style of a voice to be synthesized, the corresponding second encoder is selected as the target second encoder to construct the target voice synthesis model with the first decoder adjusted in parameters, so that the effect of synthesizing the target voice of the voice style indicated by the selection instruction can be achieved.
[0058] In a possible implementation, the first encoder can be an encoder in a speech recognition model. The speech recognition model is pre-trained by first training data, the first training data including a plurality of sets of first speech training data and a plurality of sets of first text training data corresponding to the first speech training data one by one, the first speech training data being taken as an input of the speech recognition model, and the first text training data being taken as an output of the speech recognition model, so as to train the speech recognition model.
[0059] In a possible implementation, the first decoder can be pre-trained by determining second training data, the second training data being a plurality of second speech training data and including a plurality of speech styles, extracting, by the first encoder, a second feature vector of each second speech training data, and extracting, by the speaker feature extraction network, a training data speaker feature in each second speech training data, taking the second feature vector and the training data speaker feature as an input of the first decoder, and taking the second speech training data as an output of the first decoder, so as to pre-train the first decoder.
[0060] In the case where the first encoder is an encoder in the speech recognition model, the feature vector extracted from the speech by the first encoder, for example, the first feature vector in the target speaker speech that has obtained the authorization of the user to use, is a feature vector for recognizing text, the feature vector is weakly related to the speaker, is only related to the text to be recognized, and is clearly corresponding to the speech because it is extracted from the speech. Therefore, when the first decoder is pre-trained according to the second feature vector of the second speech training data extracted by the first encoder, the training difficulty of the first decoder can be reduced to a certain extent, that is, the training difficulty of the first decoder to restore the synthesized speech according to the second feature vector clearly corresponding to the speech is lower than that of synthesizing the speech according to the feature vector extracted from the text, and it is easier to train the first decoder to meet the accuracy condition.
[0061] In the case of training the first decoder according to the training method of the first decoder in the above-mentioned embodiment, the text to be synthesized can be first synthesized into speech irrelevant to the target speaker by the second encoder in combination with the average decoder pre-trained by a large amount of training data, as described in the above-mentioned embodiment, and then the feature vector in the speech irrelevant to the target speaker is extracted by the first encoder, and the target speech is synthesized by the first decoder with adjusted parameters according to the target speaker voice features authorized by the user. However, this scheme is too long in the link of speech synthesis, and needs to go through speech synthesis, then acquire the feature vector, and finally synthesize the speech again, which may have a slow time delay problem, and the stability of the overall speech synthesis model is relatively poor, because the accuracy of synthesizing speech according to the text to be synthesized in the first step cannot be guaranteed, and noise may still be contained in the synthesized speech irrelevant to the target speaker due to the training data of the second encoder and the average decoder during training.
[0062] To solve the problem in the above-mentioned scheme, the second encoder can be pre-trained in the following manner: determining third training data, wherein the third training data includes a plurality of sets of third speech training data and a plurality of sets of third text training data corresponding to the third speech training data one by one; extracting third feature vectors of the third speech training data by the first encoder; taking the third text training data as the input of the second encoder and taking the third feature vectors as the output of the second encoder to pre-train the second encoder. That is, the second encoder is trained by taking the feature vectors extracted from the speech by the first encoder as the supervision vectors, so that the second encoder can directly extract the feature vectors required by the first decoder from the text data, and the target speech can be obtained only by one synthesis of the first decoder, which shortens the link of speech synthesis and avoids the possibility of introducing secondary errors by multiple syntheses in the above-mentioned scheme, further ensuring the stability of the speech synthesis model in the present disclosure.
[0063] Figure 3 is a flowchart of a speech synthesis method according to another exemplary embodiment of the present disclosure. As shown in Figure 3 In the case of pre-training the second encoder by the method described in the above-mentioned embodiment, the method can further include steps 301 and 302.
[0064] In step 301, the text to be synthesized is input into the second encoder in the target speech synthesis model to obtain a fourth feature vector.
[0065] In step 302, the fourth feature vector and the target speaker voice feature are input into the first decoder in the target speech synthesis model after the parameter adjustment to obtain the target speech corresponding to the target speaker.
[0066] The fourth feature vector is equivalent to the feature vector extracted from the speech irrelevant to the target speaker obtained by the first encoder from the first speech synthesis according to the text to be synthesized, so as to be directly input into the first decoder to obtain the target speech corresponding to the target speaker.
[0067] Figure 4 is a model structure diagram in a speech synthesis method according to another exemplary embodiment of the present disclosure. As shown in Figure 4 The model structure shown in the dashed box is used to adjust the parameters of the first decoder 406 in the speech synthesis model 407 according to the target speaker voice 403, and the model structure shown in the dotted box is used to synthesize the target speech 411 according to the text to be synthesized 409. The first encoder 401 belongs to the speech recognition model 402, which can extract the first feature vector 404 from the target speaker voice 403. The first feature vector 404 is used to be input into the first decoder 406 together with the target speaker voice feature 405 extracted by a speaker feature extraction network (not shown), and the target speaker voice 403 is used as the output of the first decoder 406 to supervise the training of the first decoder 406, so as to realize the parameter adjustment of the first decoder 406 according to the target speaker voice 403. The first decoder 406 is the decoder in the speech synthesis model 407. After the training is completed, the fourth feature vector 410 is directly extracted from the text to be synthesized 409 by the second encoder 408 in the speech synthesis model 407. Since the second encoder 408 is trained by taking the output of the first encoder 401 as the supervision feature, the fourth feature vector 410 can be directly input into the first decoder 406. The first decoder 406 can directly synthesize the target speech 411 corresponding to the text to be synthesized and the target speaker according to the fourth feature vector 410 and the target speaker voice feature 405.
[0068] Figure 5 is a structural block diagram of a speech synthesis device according to an exemplary embodiment of the present disclosure. As shown in Figure 5As shown, the device comprises: an acquisition module 10, configured to acquire a target speaker voice; a first processing module 20, configured to extract a first feature vector of the target speaker voice through a first encoder, and extract a target speaker voice feature of a target speaker in the target speaker voice through a speaker feature extraction network; a second processing module 30, configured to perform parameter adjustment on a first decoder according to the first feature vector, the target speaker voice feature and the target speaker voice, wherein the first decoder is a decoder that has been pre-trained; a third processing module 40, configured to construct a target voice synthesis model through the first decoder after the parameter adjustment and a second encoder, wherein the second encoder is pre-trained; and a voice synthesis module 50, configured to input a text to be synthesized and the target speaker voice feature into the target voice synthesis model to synthesize a target voice corresponding to the target speaker.
[0069] Through the above technical solution, when it is required to generate a target voice corresponding to a target speaker according to the voice of the target speaker that has been authorized by a user for use, the first decoder in the voice synthesis model can be adjusted in parameters only through the speaker voice that has been authorized by the user for use and the first encoder, without adjusting the second encoder in the voice synthesis model, so that the noise in the voice of the target speaker that has been authorized by the user for use can be avoided from being fixed in the target voice synthesis model, and thus the problem that all the target voices synthesized through the target voice synthesis model are noisy can be avoided, and since the first decoder can be adjusted in parameters through the voice of the target speaker that has been authorized by the user for use before voice synthesis, the synthesis of the target voice does not need to completely rely on the ability of the speaker feature extraction network to extract the voice feature of the target speaker that has been authorized by the user for use, and the stability and precision of voice synthesis according to the voice feature of the target speaker that has been authorized by the user for use extracted under limited conditions are ensured.
[0070] Figure 6 is a structural block diagram of a voice synthesis device according to yet another exemplary embodiment of the present disclosure. As shown, Figure 6 The acquisition module is further configured to acquire a selection instruction input by a user, the selection instruction being used to represent a voice style that the user wants to synthesize; the device further comprises a fourth processing module 60, configured to determine a target second encoder in at least one second encoder pre-trained according to the selection instruction; and the third processing module 40 is further configured to construct a target voice synthesis model through the retrained first decoder and the target second encoder.
[0071] In a possible implementation, the first encoder is an encoder in a speech recognition model, and the speech recognition model is pre-trained by first training data, the first training data including a plurality of sets of first speech training data and a plurality of sets of first text training data corresponding to the first speech training data one by one, the first speech training data being input into the speech recognition model, and the first text training data being output from the speech recognition model, so as to train the speech recognition model.
[0072] In a possible implementation, the first decoder is pre-trained by determining second training data, the second training data being a plurality of second speech training data and including a plurality of speech styles, extracting a second feature vector of each second speech training data by the first encoder, and extracting a training data speaker feature in each second speech training data by the speaker feature extraction network, and taking the second feature vector and the training data speaker feature as input of the first decoder and taking the second speech training data as output of the first decoder, so as to pre-train the first decoder.
[0073] In a possible implementation, the second encoder is pre-trained by determining third training data, the third training data including a plurality of sets of third speech training data and a plurality of sets of third text training data corresponding to the third speech training data one by one, extracting a third feature vector of the third speech training data by the first encoder, and taking the third text training data as input of the second encoder and taking the third feature vector as output of the second encoder, so as to pre-train the second encoder.
[0074] In a possible implementation, the speech synthesis module 50 is further configured to input the to-be-synthesized text into the second encoder in the target speech synthesis model to obtain a fourth feature vector, and input the fourth feature vector and the target speaker voice feature into the first decoder in the target speech synthesis model after the parameter adjustment to obtain the target speech corresponding to the target speaker.
[0075] In a possible implementation, the target speaker voice feature is a timbre feature and / or a pitch feature of the target speaker.
[0076] Reference will be made to the following description of embodiments of the disclosure in conjunction with the accompanying drawings. Figure 7The diagram illustrates a structural schematic of an electronic device 700 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0077] like Figure 7 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0078] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0079] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.
[0080] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having at least one electrical conductor, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In this disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wire, cable, RF, infrared, or any suitable combination thereof.
[0081] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0082] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and can not be assembled into the electronic device.
[0083] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire a target speaker voice; extract a first feature vector of the target speaker voice through a first encoder, and extract a target speaker voice feature of a target speaker in the target speaker voice through a speaker feature extraction network; perform parameter adjustment on a first decoder according to the first feature vector, the target speaker voice feature, and the target speaker voice, wherein the first decoder is a decoder that has been pre-trained; construct a target speech synthesis model through the first decoder after the parameter adjustment and a second encoder, wherein the second encoder is pre-trained; and input a text to be synthesized and the target speaker voice feature into the target speech synthesis model to synthesize a target voice corresponding to the target speaker.
[0084] Computer program code for carrying out operations of the present disclosure can be written in any of one or more programming languages or combinations of languages including object or visual programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0085] The flow and block diagrams in the drawings show architectural, functional, and operational representations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0086] The modules described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the name of the module does not constitute a limitation on the module itself, for example, the acquisition module can also be described as a "module for acquiring target speaker voice".
[0087] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, and without limitation, example types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chips (SOCs), complex programmable logic devices (CPLDs), etc.
[0088] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage media can include, without limitation, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include one or more lines of electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical storage devices, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0089] According to one or more embodiments of the present disclosure, example 1 provides a speech synthesis method, the method comprising: acquiring a target speaker voice; extracting a first feature vector of the target speaker voice through a first encoder, and extracting a target speaker voice feature of a target speaker in the target speaker voice through a speaker feature extraction network; performing parameter adjustment on a first decoder according to the first feature vector, the target speaker voice feature, and the target speaker voice, wherein the first decoder is a decoder that has been pre-trained; constructing a target speech synthesis model through the first decoder after the parameter adjustment and a second encoder, wherein the second encoder is pre-trained; inputting a text to be synthesized and the target speaker voice feature into the target speech synthesis model to synthesize a target voice corresponding to the target speaker.
[0090] According to one or more embodiments of the present disclosure, example 2 provides the method of example 1, the method further comprising: obtaining a selection instruction input by a user, the selection instruction being used to represent a speech style that the user wants to synthesize; determining a target second encoder from the pre-trained at least one second encoder according to the selection instruction; and constructing a target speech synthesis model by the first decoder adjusted by the parameters and the target second encoder.
[0091] According to one or more embodiments of the present disclosure, example 3 provides the method of example 1, the first encoder being an encoder in a speech recognition model, the speech recognition model being pre-trained by first training data, the first training data including a plurality of sets of first speech training data and a plurality of sets of first text training data corresponding one-to-one to the first speech training data, the first speech training data being used as input of the speech recognition model, and the first text training data being used as output of the speech recognition model, so as to train the speech recognition model.
[0092] According to one or more embodiments of the present disclosure, example 4 provides the method of example 3, the first decoder being pre-trained by: determining second training data, the second training data being a plurality of second speech training data and including a plurality of speech styles; extracting a second feature vector of each second speech training data by the first encoder, and extracting a training data speaker feature in each second speech training data by the speaker feature extraction network; using the second feature vector and the training data speaker feature as input of the first decoder, and using the second speech training data as output of the first decoder, so as to pre-train the first decoder.
[0093] According to one or more embodiments of the present disclosure, example 5 provides the method of example 3, the second encoder being pre-trained by: determining third training data, the third training data including a plurality of sets of third speech training data and a plurality of sets of third text training data corresponding one-to-one to the third speech training data; extracting a third feature vector of the third speech training data by the first encoder; using the third text training data as input of the second encoder, and using the third feature vector as output of the second encoder, so as to pre-train the second encoder.
[0094] According to one or more embodiments of the present disclosure, example 6 provides the method of example 5, wherein the inputting the text to be synthesized and the target speaker voice feature into the target speech synthesis model to synthesize a target speech corresponding to the target speaker comprises: inputting the text to be synthesized into the second encoder in the target speech synthesis model to obtain a fourth feature vector; and inputting the fourth feature vector and the target speaker voice feature into the first decoder in the target speech synthesis model after the parameter adjustment to obtain the target speech corresponding to the target speaker.
[0095] According to one or more embodiments of the present disclosure, example 7 provides the method of example 1, wherein the target speaker voice feature is a timbre feature and / or a pitch feature of the target speaker.
[0096] According to one or more embodiments of the present disclosure, example 8 provides a speech synthesis apparatus, comprising: an acquisition module configured to acquire a target speaker speech; a first processing module configured to extract a first feature vector of the target speaker speech through a first encoder, and extract a target speaker voice feature of a target speaker in the target speaker speech through a speaker feature extraction network; a second processing module configured to perform parameter adjustment on a first decoder according to the first feature vector, the target speaker voice feature, and the target speaker speech, wherein the first decoder is a decoder that has been pre-trained; a third processing module configured to construct a target speech synthesis model through the first decoder after the parameter adjustment and a second encoder, wherein the second encoder is pre-trained; and a speech synthesis module configured to input text to be synthesized and the target speaker voice feature into the target speech synthesis model to synthesize a target speech corresponding to the target speaker.
[0097] According to one or more embodiments of the present disclosure, example 9 provides a computer readable medium having stored thereon a computer program, which, when executed by a processing apparatus, implements the steps of the method of any one of examples 1-7.
[0098] According to one or more embodiments of the present disclosure, example 10 provides an electronic device, comprising: a storage device having stored thereon at least one computer program; and at least one processing device configured to execute the at least one computer program in the storage device to implement the steps of the method of any one of examples 1-7.
[0099] The above description merely illustrates the preferred embodiment of the disclosure and a principle of applied technologies. It should be understood by those skilled in the art that the disclosed range of the disclosure is not limited to the technical solutions formed by the specific combinations of the technical features described above, and should also cover other technical solutions formed by the combinations of the technical features described above or their equivalent features without departing from the disclosed concept. For example, the technical solutions formed by the mutual replacement of the above-described features and the technical features with similar functions disclosed in the disclosure (but not limited to) can be formed.
[0100] Furthermore, although operations are depicted in a particular, sequential order, this should not be understood as requiring or implying that the operations are performed in the order illustrated or sequentially. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, although specific implementation details are contained in the above discussion, these should not be construed as limiting the scope of the disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0101] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely illustrative of specific forms of implementing the claims. With respect to the devices in the above-described embodiments, the specific manner in which the various modules perform operations has been described in detail in the embodiments related to the method, and will not be described here in detail.
Claims
1. A speech synthesis method characterized by, The method comprises: obtaining target speaker voice; extracting a first feature vector of the target speaker voice through a first encoder, and extracting target speaker voice features of a target speaker in the target speaker voice through a speaker feature extraction network; parameter adjusting a first decoder according to the first feature vector, the target speaker voice features and the target speaker voice, wherein the first decoder is a decoder that has been pre-trained; constructing a target speech synthesis model through the parameter-adjusted first decoder and a second encoder, wherein the second encoder is pre-trained; inputting a text to be synthesized and the target speaker voice features into the target speech synthesis model to synthesize target speech corresponding to the target speaker.
2. The method of claim 1, wherein, The method further comprises: obtaining a selection instruction input by a user, wherein the selection instruction represents a voice style that the user wants to synthesize; determining a target second encoder from at least one pre-trained second encoder according to the selection instruction; constructing the target speech synthesis model through the parameter-adjusted first decoder and the target second encoder.
3. The method of claim 1, wherein, The first encoder is an encoder in a speech recognition model, wherein the speech recognition model is pre-trained through first training data, the first training data comprises a plurality of sets of first speech training data and a plurality of sets of first text training data corresponding to the first speech training data one by one, the first speech training data is used as the input of the speech recognition model, and the first text training data is used as the output of the speech recognition model to train the speech recognition model.
4. The method of claim 3, wherein, The first decoder is pre-trained in the following manner: determining second training data, wherein the second training data comprises a plurality of second speech training data and a plurality of voice styles; extracting a second feature vector of each second speech training data through the first encoder, and extracting training data speaker features in each second speech training data through the speaker feature extraction network; using the second feature vector and the training data speaker features as the input of the first decoder, and using the second speech training data as the output of the first decoder to pre-train the first decoder.
5. The method of claim 3, wherein, The second encoder is pre-trained in the following manner: determining third training data, wherein the third training data comprises a plurality of sets of third speech training data and a plurality of sets of third text training data corresponding to the third speech training data one by one; extracting a third feature vector of the third speech training data through the first encoder; using the third text training data as the input of the second encoder, and using the third feature vector as the output of the second encoder to pre-train the second encoder.
6. The method of claim 5, wherein, The inputting of the text to be synthesized and the target speaker voice features into the target speech synthesis model to synthesize target speech corresponding to the target speaker comprises: input the to-be-synthesized text into the second encoder in the target voice synthesis model to obtain a fourth feature vector; input the fourth feature vector and the target speaker voice feature into the first decoder in the target voice synthesis model after the parameter adjustment to obtain the target voice corresponding to the target speaker.
7. The method of claim 1, wherein, The target speaker voice feature is a timbre feature and / or a pitch feature of the target speaker.
8. A speech synthesis apparatus characterized by comprising: The device comprises: an acquisition module configured to acquire a target speaker voice; a first processing module configured to extract a first feature vector of the target speaker voice by a first encoder and extract a target speaker voice feature of a target speaker in the target speaker voice by a speaker feature extraction network; a second processing module configured to perform parameter adjustment on a first decoder according to the first feature vector, the target speaker voice feature, and the target speaker voice, wherein the first decoder is a decoder that has been pre-trained; a third processing module configured to construct a target voice synthesis model by the first decoder after the parameter adjustment and a second encoder, wherein the second encoder is pre-trained; a voice synthesis module configured to input to-be-synthesized text and the target speaker voice feature into the target voice synthesis model to synthesize a target voice corresponding to the target speaker.
9. A computer readable medium having stored thereon a computer program, characterized in that The program, when executed by a processing device, implements the steps of the method of any one of claims 1-7.
10. An electronic device, comprising: comprise: a storage device having at least one computer program stored thereon; at least one processing device configured to execute the at least one computer program in the storage device to implement the steps of the method of any one of claims 1-7.
Citation Information
Patent Citations
Voice generation method, and device, equipment and computer readable medium
CN111785247A
Speech synthesis method and device, readable medium and electronic equipment
CN112489621A