Method, device, storage medium and electronic device for generating virtual voice

By using a multi-stream encoder and a generative adversarial network model, the problems of insufficient flexibility and reliability in virtual speech generation in existing technologies are solved, and effective training of cross-language data and high-quality virtual speech generation are achieved.

CN115985286BActive Publication Date: 2025-09-23BEIJING UNISOUND INFORMATION TECH CO LTD

Patent Information

Application Number
CN202211676955.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2025-09-23
Estimated Expiration
2042-12-26

AI Technical Summary

Technical Problem

The flexibility and reliability of virtual speech generation in existing technologies are low, especially in the generation of cross-language speakers, and the GMM model has poor robustness in nonlinear data modeling.

Method used

By adopting a multi-stream encoder and a generative adversarial network (GAN) model, the target acoustic model is trained by obtaining cross-language speech text samples and object speech attribute information. The multi-stream encoder is used to capture text features in different languages, and the robustness of the model is improved through GAN modeling, supporting cross-language data training and speaker generation.

Benefits of technology

It improves the flexibility and reliability of virtual speech generation, can effectively support cross-language data training and speaker generation, and improves the quality of virtual speech and the effect of making it difficult to distinguish from real speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115985286B_ABST
    Figure CN115985286B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, storage medium, and electronic device for generating virtual speech. The method comprises: obtaining multiple different speech text samples and speech attribute information, wherein each speech text sample in multiple different language speech text samples corresponds to a language and an object; inputting each speech text sample into a multi-stream encoder to obtain text features corresponding to each speech text sample; and training a preset speech acoustic model based on generative adversarial network modeling using the text features and speech features to obtain a target acoustic model for generating virtual speech. The present invention can support cross-language data training and the generation of cross-language speakers. The multi-stream encoder can better capture text features in different languages, improve the flexibility and reliability of virtual preset generation, and thus solve the technical problem of low flexibility and reliability in generating virtual speech in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field related to speech processing technology, and in particular to a method, device, storage medium and electronic device for generating virtual speech. Background Art

[0002] The Google team proposed TacoSpawn, a method for synthesizing speech from non-existent speakers. Based on Tacotron, TacoSpawn uses maximum likelihood estimation to learn the distribution of speaker embeddings. This is used to generate new speaker embeddings (for speakers not present in the training set), which are then synthesized using text-to-speech (TTS). This technology can be used for privacy protection because the generated speakers are not real.

[0003] In related solutions, the trained speaker embeddings serve as the training data model to learn the distribution of speaker embeddings. This distribution is parameterized using a Gaussian mixture model (GMM). During inference, new speakers are generated through distributed sampling. GMM parameter modeling cannot effectively model nonlinear or nearly linear data, resulting in poor robustness. TacoSpawn only supports English speaker generation and does not support cross-lingual speaker generation, resulting in limited flexibility in virtual speech generation.

[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0005] The embodiments of the present invention provide a method, device, storage medium and electronic device for generating virtual voice, so as to at least solve the technical problems of low flexibility and reliability in generating virtual voice in the prior art.

[0006] According to one aspect of an embodiment of the present invention, a method for generating virtual speech is provided, comprising: obtaining a plurality of different speech text samples and object speech attribute information corresponding to the plurality of different speech texts, wherein each speech text sample in the plurality of different language speech text samples corresponds to a language and an object, and each speech text sample includes international standard phonetic symbols; inputting each speech text sample into a multi-stream encoder to obtain text features corresponding to each speech text sample; inputting the object speech attribute information corresponding to the plurality of different speech texts into a speech feature module to obtain speech features corresponding to each speech text sample; and training a preset speech acoustic model based on generative adversarial network modeling through the text features and the speech features to obtain a target acoustic model, wherein the target acoustic model is used to generate virtual speech.

[0007] Optionally, the preset speech acoustic model based on generative adversarial network modeling is trained through the text features and the speech features to obtain the target acoustic model, including: obtaining the speech attribute information of the language corresponding to the text features; constructing a loss function based on the speech attribute information and the speech information predicted by the target acoustic model; and obtaining the target acoustic model when the loss function meets preset conditions.

[0008] Optionally, inputting each speech text sample into a multi-stream encoder to obtain text features corresponding to each speech text sample includes: inputting each speech text sample into a multi-stream encoder to obtain multiple types of calculated text features processed by multiple feature capture modules in the multi-stream encoder; and summing the multiple types of calculated text features to obtain the text features.

[0009] Optionally, after obtaining the target acoustic model by training the preset speech acoustic model based on generative adversarial network modeling through the text features and the speech features, the method also includes: obtaining the target speech text; inputting the target speech text into the target acoustic model to obtain the virtual speech corresponding to the target speech text.

[0010] According to the first aspect of an embodiment of the present application, a device for generating a virtual voice is provided, comprising: a first acquisition unit, configured to acquire a plurality of different voice text samples and object voice attribute information corresponding to the plurality of different voice texts, wherein each of the plurality of voice text samples in different languages ​​corresponds to a language and an object, and each voice text sample includes international standard phonetic symbols; a first feature processing unit, configured to input each voice text sample into a multi-stream encoder to obtain text features corresponding to each voice text sample; a second feature processing unit, configured to input the object voice attribute information corresponding to the plurality of different voice texts into a voice feature module to obtain voice features corresponding to each voice text sample; and a model training unit, configured to train a preset voice acoustic model based on generative adversarial network modeling through the text features and the voice features to obtain a target acoustic model, wherein the target acoustic model is used to generate a virtual voice.

[0011] Optionally, the model training unit includes: an acquisition module for acquiring speech attribute information of the language corresponding to the text feature; a construction module for constructing a loss function based on the speech attribute information and the speech information predicted by the target acoustic model; and a first determination module for obtaining the target acoustic model when the loss function meets preset conditions.

[0012] Optionally, the feature processing unit includes: a second determination module, used to input each speech text sample into the multi-stream encoder to obtain multiple types of calculated text features processed by multiple feature capture modules in the multi-stream encoder; and a third determination module, used to sum the multiple types of calculated text features to obtain the text feature.

[0013] Optionally, the device also includes: a second acquisition unit, used to train the preset speech acoustic model based on generative adversarial network modeling through the text features, and after obtaining the target acoustic model, obtain the target speech text; a determination unit, used to input the target speech text into the target acoustic model to obtain the virtual speech corresponding to the target speech text.

[0014] In an embodiment of the present invention, a plurality of different speech text samples and object speech attribute information corresponding to a plurality of different speech texts are obtained, wherein each speech text sample in a plurality of different language speech text samples corresponds to a language and an object, and each speech text sample includes international standard phonetic symbols; each speech text sample is input into a multi-stream encoder to obtain text features corresponding to each speech text sample; the object speech attribute information corresponding to the plurality of different speech texts is input into a speech feature module to obtain speech features corresponding to each speech text sample; a preset speech acoustic model based on generative adversarial network modeling is trained through text features and speech features to obtain a target acoustic model, wherein the target acoustic model is used to generate virtual speech, that is, the present invention increases the scalability of cross-language texts, can support cross-language data training and the generation of cross-language speakers, the multi-stream encoder can better capture text features in different languages, improve the flexibility and reliability of virtual preset generation, and thus solve the technical problem of low flexibility and reliability in generating virtual speech in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0016] Figure 1 1 is a hardware structure block diagram of a mobile terminal according to an optional method for generating virtual voice according to an embodiment of the present invention;

[0017] Figure 2 is a flow chart of an optional method for generating virtual voice according to an embodiment of the present invention;

[0018] Figure 3 is a schematic structural diagram of an optional multi-stream encoder according to an embodiment of the present invention;

[0019] Figure 4 is a schematic diagram of an optional speech model training structure according to an embodiment of the present invention;

[0020] Figure 5 2 is a diagram of an optional virtual voice generation device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0022] It should be noted that the terms "first," "second," and the like in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a sequence of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.

[0023] In order to better understand the contents of this plan, the relevant contents are explained as follows:

[0024] Generative Adversarial Networks (GAN) is a deep learning model that consists of two main parts: a generative model and a discriminative model.

[0025] Generator G: Generate data through generator G.

[0026] Discriminator D (Discriminator): Determines whether the image is real or machine-generated. The purpose is to determine whether the data is "fake data" made by the generator.

[0027] The generator and the discriminator compete with each other and constantly adjust parameters. The ultimate goal is to make it impossible for the discriminator network to determine whether the output of the generator network is real.

[0028] The embodiment of the method for generating virtual voice provided in the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 This is a hardware structure block diagram of a mobile terminal for a method of generating virtual voice according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal 10 may include one or more ( Figure 1 Only one is shown in the figure) processor 102 (processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data. Optionally, the mobile terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0029] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for generating virtual voice in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the mobile terminal 10 via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0030] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the telecommunications provider of the mobile terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0031] This embodiment also provides a method for generating a virtual voice. Figure 2 Flowchart of a method for generating a virtual voice according to an embodiment of the present invention. Figure 2 As shown, the virtual voice generation method includes the following steps:

[0032] Step S202 , obtaining a plurality of different speech text samples and object speech attribute information corresponding to the plurality of different speech texts, wherein each speech text sample in the plurality of different language speech text samples corresponds to a language and an object, and each speech text sample includes International Standard Phonetic Symbols.

[0033] Step S204: input each speech-text sample into a multi-stream encoder to obtain text features corresponding to each speech-text sample.

[0034] Step S206: inputting the object speech attribute information corresponding to the plurality of different speech texts into a speech feature module to obtain speech features corresponding to each speech text sample.

[0035] Step S208: training a preset speech acoustic model based on generative adversarial network modeling through the text features and speech features to obtain a target acoustic model, wherein the target acoustic model is used to generate virtual speech.

[0036] In this embodiment, the aforementioned virtual voice generation method may include, but is not limited to, converting user voices into virtual voices for various user privacy protection purposes, such as using voice-changing software. The virtual voices may include, but are not limited to, non-human voices, such as animal sounds, and human voices without corresponding user voices.

[0037] In this embodiment, an object can be understood as an individual producing speech, such as a person speaking. Acquiring multiple different speech text samples and the corresponding object speech attribute information can be understood as acquiring the speech attribute information and speech styles of different individuals producing the speech. For example, the speech attribute information of speaker A, including information such as the speaker's pitch, intensity, duration, and phonemes, as well as the speech style information of speaker A saying, "The weather is so nice today."

[0038] The aforementioned speech text samples may include, but are not limited to, different speech text data across different languages, i.e., speech data from different groups is collected and speech text is generated based on the speech data. The aforementioned speech attribute information includes, but is not limited to, the physical attributes (properties) of speech, including the four elements of pitch, intensity, duration, and timbre.

[0039] 1. Pitch: Pitch refers to the various pitches of sound, or the height of a sound, and is a fundamental characteristic of sound. The pitch of a sound is determined by the vibration frequency of the sound-producing body, and the two are directly proportional: higher frequencies indicate higher pitches, while lower frequencies indicate lower pitches.

[0040] The pitch of a sound is determined by the frequency of the sound waves. A higher frequency results in a high pitch, while a lower frequency results in a low pitch. Pitch is one of the elements that make up speech.

[0041] 2. Sound intensity: Also known as volume, it refers to the strength (loudness) of a sound. It is a fundamental characteristic of sound. The intensity of a sound is determined by the amplitude of the sound-producing organ's vibration (amplitude). The two are directly proportional: the larger the amplitude, the "louder" the sound, and vice versa.

[0042] 3. Sound Length Sound length refers to the length of a sound, which is determined by the duration of the vibration of the sound-producing body. The longer the vibration of the sound-producing body lasts, the longer the sound is, and vice versa.

[0043] 4. Timbre: Timbre refers to the sensory characteristics of sound. The pitch of a sound is determined by its frequency, while its loudness is determined by its amplitude. However, we can still distinguish the sounds produced by different objects by their timbre. Different materials and structures of different sound generators will produce different timbres.

[0044] In this embodiment, a multi-stream encoder is used. For cross-lingual data, texts in different languages ​​have distinct characteristics. This encoder can capture more text features. Different feature capture modules can be used within the encoder, each with its own strengths. For example, RNNs excel at capturing temporal features, while CNNs excel at capturing global features. Finally, these features are summed up to produce text features. Based on these text features and speech features, high-quality virtual speech can be generated. High quality means that it is difficult to distinguish the user's actual speech from the virtual speech.

[0045] Among them, using GAN for speaker embedding modeling can effectively model nonlinear or approximately linear data and improve the robustness of the model.

[0046] Through the embodiments provided by the present application, by obtaining multiple different speech text samples and object speech attribute information corresponding to the multiple different speech texts, wherein each speech text sample in the multiple different language speech text samples corresponds to a language and an object, and each speech text sample includes international standard phonetic symbols; each speech text sample is input into a multi-stream encoder to obtain text features corresponding to each speech text sample; the object speech attribute information corresponding to the multiple different speech texts is input into a speech feature module to obtain speech features corresponding to each speech text sample; a preset speech acoustic model based on generative adversarial network modeling is trained through text features and speech features to obtain a target acoustic model, wherein the target acoustic model is used to generate virtual speech, that is, the present invention increases the scalability of cross-language texts, can support cross-language data training and the generation of cross-language speakers, the multi-stream encoder can better capture text features in different languages, improve the flexibility and reliability of virtual preset generation, and thus solve the technical problem of low flexibility and reliability in generating virtual speech in the prior art.

[0047] Optionally, the training of a preset speech acoustic model based on generative adversarial network modeling through the text features and the speech features to obtain a target acoustic model may include: obtaining speech attribute information of the language corresponding to the text features; constructing a loss function based on the speech attribute information and the speech information predicted by the target acoustic model; and obtaining the target acoustic model when the loss function meets preset conditions.

[0048] Optionally, inputting each speech text sample into a multi-stream encoder to obtain text features corresponding to each speech text sample may include: inputting each speech text sample into a multi-stream encoder to obtain multiple types of calculated text features processed by multiple feature capture modules in the multi-stream encoder; and summing the multiple types of calculated text features to obtain the text features.

[0049] Optionally, after obtaining the target acoustic model by training the preset speech acoustic model based on generative adversarial network modeling through the text features and the speech features, the method also includes: obtaining the target speech text; inputting the target speech text into the target acoustic model to obtain the virtual speech corresponding to the target speech text.

[0050] As an optional embodiment, the present application also provides a method for generating a new cross-language speaker. The specific content of this solution is as follows.

[0051] The IPA standard (International Phonetic Alphabet, IPA for short) dictionary is used. The IPA dictionary is a standard representation of spoken pronunciation and can be used to represent all languages. It replaces the original English phoneme dictionary, increases language scalability, and can use cross-language data to train acoustic models.

[0052] IPA is a phonetic notation system used to represent all languages ​​in the world. It was first developed in 1888 by the International Phonetic Association.

[0053] The International Phonetic Alphabet (IPA) follows the strict standard of "one note, one symbol" and was originally used to mark the pronunciation of Western and African languages. Over the years, thanks to the efforts of Chinese linguists such as Zhao Yuanren, the IPA has been gradually refined and can now be used to mark the pronunciation of Eastern languages ​​such as Chinese.

[0054] Using a multi-stream encoder, for cross-language data, texts in different languages ​​have different characteristics. The multi-stream encoder can capture more text features, and different feature capture modules can be used in the multi-stream encoder. Different modules have different strengths in capturing features, such as RNN is good at capturing temporal features, and CNN is good at capturing global features, and finally summing them up. Figure 3 The structural diagram of the multi-stream encoder is shown in FIG.

[0055] Use the GAN model to learn the speaker embedding space. Figure 4 As shown in the figure, it is a schematic diagram of the speech model training structure.

[0056] Speech model training involves two parts of feature extraction. First, the speaker's speech attribute information (Speaker ID) is processed by the Speech Attribute Feature Recognition Module's Lookup Embedding feature to obtain speech recognition features. These features are then input into the Speaker Embedding Module to obtain speech features.

[0057] In the second part, the speaker's IPA-formatted style is passed through the multi-stream detection module multi-stream textencoder, and its output is input into the text feature module word embedding to obtain the style features.

[0058] A loss function is constructed based on text features and speech features to train the preset speech model. The trained speech model is obtained.

[0059] Among them, the noise is input into the pre-trained speech model through the generator, and then input into the discriminator from the preset speech model, thereby realizing adversarial training and adjusting the model parameters.

[0060] The solution provided in this embodiment uses IPA to represent text phonemes, increasing cross-lingual text scalability and supporting cross-lingual data training and cross-lingual speaker generation. Multi-stream encoders can better capture text features in different languages. Using GANs for speaker embedding modeling offers stronger modeling capabilities and more robust results than GMMs.

[0061] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0062] In this embodiment, a device for generating a virtual voice is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments, and will not be repeated hereafter. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and contemplated.

[0063] Figure 5 : is a structural block diagram of a device for generating virtual speech according to an embodiment of the present invention, such as Figure 5 As shown, the virtual voice generating device includes:

[0064] The first acquisition unit 51 is used to acquire multiple different speech text samples and object speech attribute information corresponding to the multiple different speech texts, wherein each speech text sample in the multiple different language speech text samples corresponds to a language and an object, and each speech text sample includes international standard phonetic symbols.

[0065] The first feature processing unit 53 is configured to input each speech text sample into the multi-stream encoder to obtain text features corresponding to each speech text sample.

[0066] The second feature processing unit 55 is configured to input the object speech attribute information corresponding to the plurality of different speech texts into a speech feature module to obtain speech features corresponding to each speech text sample.

[0067] The model training unit 57 is used to train a preset speech acoustic model based on generative adversarial network modeling through text features and speech features to obtain a target acoustic model, wherein the target acoustic model is used to generate virtual speech.

[0068] In the embodiment provided by the present application, a first acquisition unit 51 acquires multiple different speech text samples and object speech attribute information corresponding to the multiple different speech texts, wherein each speech text sample in the multiple different language speech text samples corresponds to a language and an object, and each speech text sample includes the International Standard Phonetic Alphabet. A first feature processing unit 53 inputs each speech text sample into a multi-stream encoder to obtain text features corresponding to each speech text sample. A second feature processing unit 55 inputs the object speech attribute information corresponding to the multiple different speech texts into a speech feature module to obtain speech features corresponding to each speech text sample. A model training unit 57 trains a preset speech acoustic model based on generative adversarial network modeling using text features and speech features to obtain a target acoustic model, wherein the target acoustic model is used to generate virtual speech. This invention increases the scalability of cross-language text, supports cross-language data training and the generation of cross-language speakers, and the multi-stream encoder can better capture text features in different languages, improving the flexibility and reliability of virtual preset generation, thereby solving the technical problem of low flexibility and reliability in generating virtual speech in the prior art.

[0069] Optionally, the model training unit may include: an acquisition module for acquiring speech attribute information of the language corresponding to the text feature; a construction module for constructing a loss function based on the speech attribute information and the speech information predicted by the target acoustic model; and a first determination module for obtaining the target acoustic model when the loss function meets preset conditions.

[0070] Optionally, the first feature processing unit includes: a second determination module, used to input each speech text sample into the multi-stream encoder to obtain multiple types of calculated text features processed by multiple feature capture modules in the multi-stream encoder; and a third determination module, used to sum the multiple types of calculated text features to obtain the text feature.

[0071] Optionally, the device may also include: a second acquisition unit, used to train a preset speech acoustic model based on generative adversarial network modeling through the text features and the speech features, and after obtaining the target acoustic model, obtain the target speech text; a determination unit, used to input the target speech text into the target acoustic model to obtain a virtual speech corresponding to the target speech text.

[0072] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.

[0073] An embodiment of the present invention further provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.

[0074] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:

[0075] S1, obtaining a plurality of different speech text samples and object speech attribute information corresponding to the plurality of different speech texts, wherein each speech text sample in the plurality of different language speech text samples corresponds to a language and an object, and each speech text sample includes International Standard Phonetic Symbols;

[0076] S2, inputting each of the speech and text samples into a multi-stream encoder to obtain text features corresponding to each of the speech and text samples;

[0077] S3, inputting object speech attribute information corresponding to a plurality of different speech texts into a speech feature module to obtain speech features corresponding to each speech text sample;

[0078] S4, training a preset speech acoustic model based on generative adversarial network modeling through the text features to obtain a target acoustic model, wherein the target acoustic model is used to generate virtual speech.

[0079] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.

[0080] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0081] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0082] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:

[0083] S1, obtaining a plurality of different speech text samples and object speech attribute information corresponding to the plurality of different speech texts, wherein each speech text sample in the plurality of different language speech text samples corresponds to a language and an object, and each speech text sample includes international standard phonetic symbols.

[0084] S2: Input each speech-text sample into a multi-stream encoder to obtain text features corresponding to each speech-text sample.

[0085] S3, inputting object speech attribute information corresponding to a plurality of different speech texts into a speech feature module to obtain speech features corresponding to each speech text sample;

[0086] S4, the text features and speech features are used to train a preset speech acoustic model based on generative adversarial network modeling to obtain a target acoustic model, wherein the target acoustic model is used to generate virtual speech.

[0087] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.

[0088] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, centralized on a single computing device, or distributed across a network of multiple computing devices. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0089] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for generating a virtual voice, characterized in that: include: Acquire a plurality of speech text samples in different languages ​​and object speech attribute information corresponding to the plurality of speech text samples in different languages, wherein each speech text sample in the plurality of speech text samples corresponds to a language and an object, where an object is an individual that makes a speech, and each speech text sample includes International Standard Phonetic Symbols; Inputting each of the speech and text samples into a multi-stream encoder to obtain text features corresponding to each of the speech and text samples, including: inputting each of the speech and text samples into the multi-stream encoder to obtain multiple types of calculated text features processed by multiple feature capture modules in the multi-stream encoder; summing the multiple types of calculated text features to obtain the text features; Inputting the object speech attribute information corresponding to the plurality of different speech texts into a speech feature module to obtain speech features corresponding to each speech text sample; A preset speech acoustic model based on generative adversarial network modeling is trained by the text features and the speech features to obtain a target acoustic model, including: obtaining speech attribute information of the language corresponding to the text features; constructing a loss function based on the speech attribute information and the speech information predicted by the target acoustic model; and obtaining the target acoustic model when the loss function meets preset conditions, wherein the target acoustic model is used to generate virtual speech.

2. The method according to claim 1, characterized in that After the preset speech acoustic model based on generative adversarial network modeling is trained by the text features and the speech features to obtain a target acoustic model, the method further includes: Get the target voice text; The target speech text is input into the target acoustic model to obtain a virtual speech corresponding to the target speech text.

3. A device for generating virtual speech, characterized in that: include: a first acquisition unit, configured to acquire a plurality of speech text samples in different languages ​​and object speech attribute information corresponding to the plurality of speech text samples in different languages, wherein each speech text sample in the plurality of speech text samples in different languages ​​corresponds to a language and an object, where an object is an individual uttering speech, and each speech text sample includes International Standard Phonetic Symbols; A first feature processing unit, configured to input each speech text sample into a multi-stream encoder to obtain text features corresponding to each speech text sample; A second feature processing unit is configured to input the object speech attribute information corresponding to the plurality of different speech texts into a speech feature module to obtain a speech feature corresponding to each speech text sample; A model training unit, configured to train a preset speech acoustic model based on generative adversarial network modeling using the text features and the speech features to obtain a target acoustic model, wherein the target acoustic model is used to generate virtual speech; The model training unit includes: an acquisition module for acquiring speech attribute information of the language corresponding to the text feature; a construction module for constructing a loss function based on the speech attribute information and the speech information predicted by the target acoustic model; and a first determination module for obtaining the target acoustic model when the loss function satisfies a preset condition; The feature processing unit includes: a second determination module, used to input each speech text sample into the multi-stream encoder to obtain multiple types of calculated text features processed by multiple feature capture modules in the multi-stream encoder; a third determination module, used to sum the multiple types of calculated text features to obtain the text feature.

4. The device according to claim 3, characterized in that The device further comprises: A second acquisition unit is configured to train a preset speech acoustic model based on a generative adversarial network modeling through the text features to obtain a target acoustic model and then acquire a target speech text; The determination unit is used to input the target speech text into the target acoustic model to obtain a virtual speech corresponding to the target speech text.

5. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 2 when executed.

6. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 2.

Citation Information

Patent Citations

  • Voice generation method and device, electronic equipment and readable storage medium

    CN113628608A

  • Speech synthesis model training method and system, electronic equipment and storage medium

    CN114267325A

Cited By

  • Virtual voice generation method

    CN116913242A

  • Method for generating a virtual voice

    CN116913242B