Voice conversion method, device, electronic device and storage medium

By extracting the prosodic parameters and unclassified phoneme features of the target speech, using an end-to-end speech recognition encoder to generate content vectors with consistent time resolution, and combining them with the target prosodic parameters for speech conversion, the problem of easy loss of detail information in speech conversion in existing technologies is solved, and the speech conversion effect is improved.

CN114333902BActive Publication Date: 2025-09-12CHENGDU IQIYI INTELLIGENT INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111671308.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-09-12
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

Existing speech conversion technologies are prone to losing detailed content information of the target speech, resulting in poor speech conversion effects.

Method used

By extracting the prosodic parameters and unclassified phoneme features of the target speech, an end-to-end speech recognition encoder is used to generate content vectors with consistent time resolution. The target prosodic parameters are then combined for speech conversion, avoiding dependence on phoneme classification and retaining speech details.

Benefits of technology

The speech conversion effect is improved, phoneme features are ensured to be consistent, target speech details are fully expressed, and the impact of inaccurate phoneme classification is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333902B_ABST
    Figure CN114333902B_ABST
Patent Text Reader

Abstract

The present application relates to a speech conversion method, device, electronic device and storage medium. The method includes: extracting the rhythm of a target speech to obtain target rhythm parameters, where the target speech is the speech to be speech converted; inputting the target speech into an end-to-end speech recognition encoder to obtain a content vector output by the end-to-end speech recognition encoder after processing the target speech; performing speech conversion on the target speech according to the target rhythm parameters and the content vector. Since the time resolution of the content vector is consistent with the time resolution of the target speech, the phoneme features represented by the content vector are consistent with the phoneme features contained in the target speech, thereby enabling the content vector to fully reflect the details of the target speech, thereby improving the effect of speech conversion. Moreover, since the phoneme features included in the content vector are not classified, the problem of inaccurate phoneme classification affecting speech conversion is avoided, thereby further improving the effect of speech conversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voice conversion, and in particular to a voice conversion method, device, electronic device and storage medium. Background Art

[0002] With the continuous development of deep learning technology, neural network-based voice conversion (VC) technology has become increasingly mature. Voice conversion involves modifying acoustic parameters related to the speaker's individual characteristics to make the timbre of a target speech sound like the target speaker's, while maintaining the semantic meaning. Currently, VC technology tends to lose detailed information about the target speech, resulting in poor voice conversion performance. Summary of the Invention

[0003] The present application provides a speech conversion method, device, electronic device and storage medium to solve the problem in related technologies that target speech detail content information is easily lost during speech conversion, resulting in poor speech conversion effect.

[0004] In a first aspect, the present application provides a speech conversion method, which includes: extracting the rhythm of a target speech to obtain target rhythm parameters, wherein the target speech is the speech that needs to be speech converted; inputting the target speech into an end-to-end speech recognition encoder to obtain a content vector output by the end-to-end speech recognition encoder after processing the target speech, wherein the content vector is used to represent unclassified phoneme features obtained from the target speech, and the time resolution of the content vector is consistent with the time resolution of the target speech, so that the phoneme features represented by the content vector are consistent with the phoneme features contained in the target speech; performing speech conversion on the target speech according to the target rhythm parameters and the content vector.

[0005] Optionally, extracting the rhythm of the target speech and obtaining target rhythm parameters includes: obtaining a speech vector of the target speech, the speech vector being used to characterize the audio features of the target speech, the audio features including: timbre features and rhythm features; extracting the rhythm features of the speech vector to obtain the target rhythm parameters.

[0006] Optionally, the target speech is input into an end-to-end speech recognition encoder to obtain a content vector output by the end-to-end speech recognition encoder after encoding the target speech, including: determining an energy vector corresponding to the target speech, the energy vector being used to characterize the emotional characteristics of the target speech; inputting the target speech and the energy vector into the end-to-end speech recognition encoder, so that the end-to-end speech recognition encoder encodes the target speech according to the energy vector, obtains the encoding result, and uses the encoding result as the content vector.

[0007] Optionally, performing speech conversion on the target speech according to the target prosodic parameters and the content vector includes: fusion encoding the content vector and the target prosodic parameters to obtain a coding vector; decoding the coding vector according to the target prosodic parameters to obtain a Melp feature; and inputting the Melp feature into a vocoder so that the vocoder parses the Melp feature to generate speech after the target speech is speech-converted.

[0008] Optionally, before decoding the encoding vector according to the target prosodic parameters to obtain the Melp feature, the method also includes: identifying the speaker of the target speech and confirming the target timbre based on the speaker, the target timbre being the timbre of the target speech that needs to be converted into speech, and the target timbre is used to decode the encoding vector to obtain the Melp feature.

[0009] Optionally, decoding the encoding vector according to the target prosodic parameter to obtain a melp feature includes: decoding the target vector at a target timbre according to the target prosodic parameter to obtain the melp feature.

[0010] In a second aspect, the present application provides a speech conversion device, which includes: a rhythm acquisition module, which is used to extract the rhythm of a target speech and obtain target rhythm parameters, and the target speech is the speech that needs to be speech converted; a content acquisition module, which is used to input the target speech into an end-to-end speech recognition encoder to obtain a content vector output by the end-to-end speech recognition encoder after processing the target speech, and the content vector is used to represent unclassified phoneme features obtained from the target speech, and the time resolution of the content vector is consistent with the time resolution of the target speech, so that the phoneme features represented by the content vector are consistent with the phoneme features contained in the target speech; a speech conversion module, which is used to perform speech conversion on the target speech according to the target rhythm parameters and the content vector.

[0011] Optionally, the content acquisition module is also used to: determine the energy vector corresponding to the target speech, the energy vector is used to characterize the emotional characteristics of the target speech; input the target speech and the energy vector into the end-to-end speech recognition encoder, so that the end-to-end speech recognition encoder encodes the target speech according to the energy vector, obtains the encoding result, and uses the encoding result as the content vector.

[0012] In a third aspect, an electronic device is provided, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0013] Memory for storing computer programs;

[0014] The processor is configured to implement the steps of the speech conversion method described in any one of the embodiments of the first aspect when executing the program stored in the memory.

[0015] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the speech conversion method as described in any embodiment of the first aspect are implemented.

[0016] The above technical solution provided by the embodiment of the present application has the following advantages compared with the prior art:

[0017] The method provided by an embodiment of the present application includes: extracting the prosody of a target speech to obtain target prosody parameters, wherein the target speech is the speech to be speech-converted; inputting the target speech into an end-to-end speech recognition encoder to obtain a content vector output by the end-to-end speech recognition encoder after processing the target speech, wherein the content vector is used to represent unclassified phoneme features obtained from the target speech, and the time resolution of the content vector is consistent with the time resolution of the target speech; performing speech conversion on the target speech according to the target prosody parameters and the content vector, and since the time resolution of the content vector is consistent with the time resolution of the target speech, the phoneme features represented by the content vector are consistent with the phoneme features contained in the target speech, thereby enabling the content vector to fully reflect the details of the target speech, thereby improving the effect of speech conversion. Moreover, since the phoneme features included in the content vector are not classified, the speech conversion method does not rely on the phoneme classification results, thereby avoiding the problem of inaccurate phoneme classification affecting speech conversion, further improving the effect of speech conversion, and thus solving the problem in the related art that the detailed content information of the target speech is easily lost during speech conversion, resulting in poor speech conversion effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0020] Figure 1 A flowchart of a voice conversion method provided in an embodiment of the present application;

[0021] Figure 2 A schematic diagram of the structure of a speech conversion system provided in an embodiment of the present application;

[0022] Figure 3 A schematic diagram of the structure of a speech conversion device provided in an embodiment of the present application;

[0023] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0024] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0025] Figure 1 This is a flow chart of a voice conversion method provided in an embodiment of the present application. Figure 1 As shown, the voice conversion method includes:

[0026] S101, extracting the prosody of the target speech and obtaining target prosody parameters;

[0027] It should be understood that the target voice is the voice that needs to be voice-converted; it should be understood that the voice conversion method provided in this embodiment can be applied to a terminal and / or a server, that is, each step in the voice conversion method can be performed by the terminal or the server alone, or by a combination of the terminal and the server; wherein the terminal can be implemented in various forms. For example, the terminal described in the present invention may include mobile terminals such as mobile phones, tablet computers, laptop computers, PDAs, PMPs, navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs and desktop computers. The subsequent description will be explained using mobile terminals as an example.

[0028] It should be understood that before extracting the rhythm of the target speech and obtaining the target rhythm parameters, the method also includes: obtaining the target speech, wherein the target speech can be obtained from a video file or from an audio file. Specifically, after obtaining the complete audio or a certain audio segment from the video file or audio file, the obtained audio is sentence-segmented through endpoint detection technology (VAD) to obtain each sentence audio, and the sentence audio is used as the target speech.

[0029] S102: Input the target speech into an end-to-end speech recognition encoder to obtain a content vector output by the end-to-end speech recognition encoder after processing the target speech;

[0030] It should be understood that the content vector is used to represent the unclassified phoneme features obtained from the target speech, and the phoneme features include but are not limited to: language content that changes rapidly in a short period of time, such as tone, intonation, and speaking speed. The phoneme features also include: language content such as text content, and the time resolution of the content vector is consistent with the time resolution of the target speech. For example, if the time resolution of the target speech is 10ms per frame, the time resolution of the content vector is also 10ms per frame. Each frame of the target speech contains several phoneme features. Since the time resolution of the content vector is the same as that of the target speech, the phoneme features contained in the content vector are the same as the phoneme features of the target speech, avoiding downsampling of several phoneme features contained in the target speech, so that the phoneme features represented by the content vector are consistent with the phoneme features contained in the target speech, thereby accurately expressing the tone, intonation, speaking speed and other phoneme features in the target speech.

[0031] S103: Perform voice conversion on the target speech according to the target prosody parameter and the content vector.

[0032] In some examples of this embodiment, extracting the prosody of a target speech and obtaining target prosody parameters includes: obtaining a speech vector of the target speech, the speech vector being used to characterize the audio features of the target speech, the audio features including timbre features and prosody features; extracting prosody features from the speech vector to remove the timbre features contained in the speech vector, and obtaining the target prosody parameters. Specifically, first, the target speech is converted into a speech vector that can be processed by a computer. In the process of converting the target speech into a speech vector that can be processed by a computer, redundant information in the target speech can be removed as needed, including but not limited to noise information, content, speaker information, etc.; then, prosody features are extracted from the speech vector to obtain target prosody parameters; for example, the speech vector is input into a prosody encoder, and prosody features are extracted from the speech vector, where prosody features include information such as pitch, loudness, speaking rate, and rhythm, to extract the prosody features carried in the speech vector to generate target prosody parameters corresponding to the target speech. At the same time, in the process of prosody feature extraction from the speech vector, redundant information carried in the target speech is further removed.

[0033] Continuing from the above example, it can be understood that in the speech conversion method provided in this embodiment, the target speech can be converted into a speech vector that can be processed by a computer through a pre-trained model. In the speech conversion method provided in this embodiment, the target prosodic parameters can be obtained by extracting prosodic features from the speech vector through a learning encoder.

[0034] In some examples of this embodiment, the target speech is input into an end-to-end speech recognition encoder to obtain a content vector output after the end-to-end speech recognition encoder encodes the target speech, including: inputting the target speech into an end-to-end speech recognition encoder so that the end-to-end speech recognition encoder encodes the target speech to obtain an encoding result, and using the encoding result as the content vector. In some examples, the end-to-end speech recognition encoder (E2E ASR Encoder Pretrained Model) is developed based on the ESPNET platform, and the end-to-end speech recognition encoder includes but is not limited to: a speech recognition encoder encoding module and a speech recognition decoder decoding module. The speech recognition encoder encoding module is used to encode the input data, and the speech recognition decoder module is used to decode the data obtained by the encoding process to obtain decoded data. Specifically, first, the target speech is framed by an end-to-end speech recognition encoder, and the data obtained by framing is preprocessed with Mel Frequency Cepstral Coefficents (MFCC) features. Then, the preprocessed data is input into the speech recognition encoder encoding module of the end-to-end speech recognition encoder to obtain the encoding result, which is used as the content vector.

[0035] Continuing with the above example, after the data is input into the end-to-end speech recognition encoder, the speech recognition encoder encoding module in the end-to-end speech recognition encoder encodes the input data to obtain an encoding result, and then directly uses the encoding result as the content vector. At this time, since the encoding result has not been decoded by the speech recognition decoder decoding module, the speech conversion method provided in this embodiment is not based on the speech recognition decoder decoding module, and the accuracy of decoding the encoding result, that is, it does not strongly depend on the training accuracy of the end-to-end speech recognition encoder; at the same time, since the obtained encoding result is not downsampled and decoded, the problem of losing phoneme features by downsampling the time resolution of the encoding result from 10ms per frame to 60ms per frame is avoided, that is, the time resolution of the encoding vector is kept consistent with the time resolution of the target speech, so that the phoneme features represented by the content vector are consistent with the phoneme features contained in the target speech, thereby retaining the details of the target speech and accurately expressing the phoneme features such as tone, intonation, and speaking speed of the target speech.

[0036] In some examples of this embodiment, the target speech is input into an end-to-end speech recognition encoder so that the end-to-end speech recognition encoder encodes the target speech to obtain an encoding result, and uses the encoding result as the content vector, including: determining an energy vector corresponding to the target speech, and the energy vector is used to characterize the emotional characteristics of the target speech; inputting the target speech and the energy vector into the end-to-end speech recognition encoder so that the end-to-end speech recognition encoder encodes the target speech according to the energy vector to obtain the encoding result, and uses the encoding result as the content vector. It is understandable that when the target speech is framed by an end-to-end speech recognition encoder and the data obtained by framing is preprocessed with MFCC features, if the emotion in the target speech is relatively weak, if only the data obtained by the target speech is encoded, the details of the target speech cannot be accurately obtained. Therefore, an energy vector representing the emotional characteristics of the target speech is added so that the end-to-end speech recognition encoder can encode the target speech based on the emotion of the target speech when encoding the content vector to obtain an encoding result, and then the final encoding result includes the emotion of the target speech. Using the encoding result as the content vector can more accurately reflect the details of the target speech.

[0037] In some examples of this embodiment, performing voice conversion on the target speech based on the target prosodic parameters and the content vector includes: fusion-encoding the content vector and the target prosodic parameters to obtain a coding vector; decoding the coding vector to obtain melp features; and inputting the melp features into a vocoder so that the vocoder parses the melp features to generate speech after the target speech is voice-converted. Specifically, the content vector is encoded using a voice conversion model to obtain a coding vector, and the coding vector is decoded based on the target prosodic parameters to obtain the melp features.

[0038] Continuing with the previous example, the speech conversion model has an encoder-decoder structure, consisting of a speech conversion encoder module and a speech conversion decoder module. The speech conversion encoder module aims to obtain a good representation of the input data. The speech conversion encoder module converts the input data into a one-hot vector, then passes it through a pre-net network structure. The output of the pre-net network is then input into the CBHG module. The final output from the CBHG module is a robust representation of the input data. The pre-net is a three-layer network structure whose main function is to perform a series of nonlinear transformations on the input, which helps the model converge and generalize. It has two hidden layers, and the connections between layers are fully connected. The number of hidden units in the first layer matches the number of input units, while the number of hidden units in the second layer is half of the first layer. Both hidden layers use the Reinforced Lu (Reinforced Lu) activation function with a dropout of 0.5 to improve the model's generalization ability. The CBHG module is used to further improve the model's generalization ability.

[0039] Continuing with the previous example, the speech conversion decoder module is mainly divided into three parts: pre-net, attention-RNN, and decoder-RNN. The structure of the pre-net is the same as that of the encoder, and it is used to perform some nonlinear transformations on the input data. The structure of the attention-RNN is a single-layer RNN containing 256 GRUs. It takes the output of the pre-net and the output of the attention as input, and outputs it to the decoder-RNN after passing through the GRU units. The decoder-RNN is a two-layer residual GRU, and its output is the sum of the input and the output after passing through the GRU units. Each layer also contains 256 GRU units. The input of the first decoder is a 0 matrix. After that, the output of step t is used as the input of step t+1, thereby converting the input data into another form of information.

[0040] It should be understood that the content vector and target prosodic parameters are input into the above-mentioned speech conversion model. The speech conversion encoder module of the speech conversion model fuses the content vector and target prosodic parameters and converts them into a one-hot vector. Then, it passes through a pre-net network structure. The output of the pre-net network is then input into the CBHG module. Finally, the output from the CBHG is a robust representation sequence of the input data, and this representation sequence is used as the encoding vector. The encoding vector is then decoded by the speech conversion decoder module, and the encoding vector is decoded in the form of Mel-spectrogram features, and then the encoding vector is decompressed into Mel-spectrogram features. The Mel-spectrogram features are then input into the vocoder so that the vocoder decodes the Mel-spectrogram features to generate speech. The speech generated by the vocoder decoding the Mel-spectrogram features is the speech obtained after speech conversion of the target speech.

[0041] In some examples of this embodiment, before decoding the encoding vector according to the target prosody parameters to obtain the Melp feature, the method further includes: identifying the speaker of the target speech and confirming the target timbre based on the speaker, the target timbre being the timbre of the target speech that needs to be converted into speech, that is, the target timbre is the timbre of the target speaker, and the target timbre is used to decode the encoding vector to obtain the Melp feature. Specifically, the speaker of the target speech is identified to confirm the speaker's identity, and the target timbre is confirmed based on the identity corresponding to the speaker. The target timbre is the timbre of the target speech that needs to be converted into speech, that is, the target timbre is the timbre of the target speaker, and the target timbre is used to decode the encoding vector to obtain the Melp feature. It should be understood that before identifying the speaker of the target speech, the speech conversion method also includes: setting the target timbre corresponding to the identity; that is, the speaker corresponds to the identity, that is, setting a corresponding target timbre for each speaker, and then when performing the identification of the speaker of the target speech, the target timbre corresponding to the target speech can be determined based on the identity determined by the determined speaker. It can be understood that in some examples, different speakers correspond to different target timbres, and in some examples, at least two different speakers correspond to the same target timbre.

[0042] Continuing with the above example, in some examples, if the speaker of the target voice cannot be identified, or the identified speaker of the target voice has not set the corresponding target timbre, then the preset default timbre will be used as the target timbre, and the timbre of the target voice will be converted.

[0043] In some examples of this embodiment, decoding the encoding vector according to the target prosody parameters to obtain the melp feature includes: decoding the target vector based on the target prosody parameters at the target timbre to obtain the melp feature. Specifically, a speech conversion encoder module of a speech conversion model performs fusion encoding on the target prosody parameters, the target timbre, and the target vector to obtain the encoding vector, and a speech conversion decoder module of the speech conversion model decodes the encoding vector to obtain the melp feature generated based on the target timbre.

[0044] It is understandable that the dimensions of the target timbre, target vector, and target prosodic parameter may be different. For example, the target timbre is 64-dimensional data, the target vector is 128-dimensional data, and the target prosodic parameter is 256-dimensional data. The speech conversion encoder module of the speech conversion model can fuse and encode data of different dimensions to obtain a coding vector. In some examples, the dimensions of the target timbre, target vector, and target prosodic parameter may be the same. For example, the target timbre, target vector, and target prosodic parameter are all 256-dimensional data. The speech conversion encoder module of the speech conversion model can fuse and encode data of the same dimension to obtain a coding vector.

[0045] The speech conversion method provided in this embodiment includes: extracting the rhythm of the target speech to obtain target rhythm parameters, wherein the target speech is the speech that needs to be speech converted; inputting the target speech into an end-to-end speech recognition encoder to obtain a content vector output by the end-to-end speech recognition encoder after processing the target speech, wherein the content vector is used to represent unclassified phoneme features obtained from the target speech, and the time resolution of the content vector is consistent with the time resolution of the target speech; performing speech conversion on the target speech according to the target rhythm parameters and the content vector, and since the phoneme features included in the content vector are not processed Row classification, since the time resolution of the content vector is consistent with the time resolution of the target speech, the phoneme features represented by the content vector are consistent with the phoneme features contained in the target speech, so that the content vector can fully reflect the details of the target speech, thereby improving the effect of speech conversion. Moreover, since the phoneme features included in the content vector are not classified, the speech conversion method does not rely on the phoneme classification results, avoiding the problem of inaccurate phoneme classification affecting speech conversion, further improving the effect of speech conversion, and thus solving the problem in related technologies that the target speech detail content information is easily lost during speech conversion, resulting in poor speech conversion effect.

[0046] Based on the same concept, this embodiment provides a voice conversion system, such as Figure 2As shown, the speech conversion system provided in this embodiment includes: a prosody module, a latent content extraction module, a speech conversion module, and a vocoder module; the specific functions of each module are as follows:

[0047] The prosody module includes: a pre-training module; the pre-training module uses a self-supervised learning algorithm to obtain quantized (discrete) representation features of the target speech by inputting the target speech into the pre-training module, and then converts the target speech into a speech vector; the prosody module also includes: a prosody encoding module; the prosody encoding module is a prosody encoder that is embedded in the speech vector to capture audio features that are independent of speech information and special speaker characteristics, such as stress, intonation, and speaking rate, and uses the captured audio features as target prosody parameters;

[0048] Among them, the implicit content extraction module includes: an end-to-end speech recognition encoder (E2EASR Encoder Pretrained Model), which includes an ASR encoder module and an ASR decoder module, and the above-mentioned end-to-end speech recognition encoder is trained on the ESPNET platform to obtain an implicit content extraction module based on the ASR encoder; the target speech is input into the implicit content extraction module, and then the EV extracted by the ASR encoder is output as the content vector, and used as the encoder input of the VC module.

[0049] Among them, the speech conversion module includes: a VC model, which uses an encoder-decoder structure to input a content vector and prosodic parameters, as well as a target timbre determined by determining the identity of the target voice, to encode a coding vector, and then decodes the coding vector in the form of a Melp feature to obtain a Melp feature.

[0050] Among them, the vocoder module includes: Parallel WaveGAN vocoder, which decodes the Melp features obtained by the VC module to obtain the converted generated speech, and then completes the speech conversion of the target speech.

[0051] like Figure 3 As shown, this embodiment provides a voice conversion device, which includes:

[0052] A prosody acquisition module 1 is used to extract the prosody of a target speech and obtain target prosody parameters, wherein the target speech is the speech to be converted into speech;

[0053] A content acquisition module 2, wherein the content acquisition module is configured to input the target speech into an end-to-end speech recognition encoder to obtain a content vector output by the end-to-end speech recognition encoder after processing the target speech, wherein the content vector is configured to represent unclassified phoneme features obtained from the target speech, and wherein the temporal resolution of the content vector is consistent with the temporal resolution of the target speech, so that the phoneme features represented by the content vector are consistent with the phoneme features contained in the target speech;

[0054] The speech conversion module 3 is used to perform speech conversion on the target speech according to the target prosodic parameters and the content vector.

[0055] Among them, the content acquisition module 2 is also used to: determine the energy vector corresponding to the target speech, and the energy vector is used to characterize the emotional characteristics of the target speech; input the target speech and the energy vector into the end-to-end speech recognition encoder, so that the end-to-end speech recognition encoder encodes the target speech according to the energy vector, obtains the encoding result, and uses the encoding result as the content vector.

[0056] It should be understood that the combination of the various modules of the voice conversion device provided in this embodiment can implement the various steps of the above-mentioned voice conversion method and achieve the same technical effects as the various steps of the above-mentioned voice conversion method, which will not be repeated here.

[0057] like Figure 4 As shown, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.

[0058] Memory 113, for storing computer programs;

[0059] In one embodiment of the present application, the processor 111 is configured to implement the steps of the speech conversion method provided in any one of the aforementioned method embodiments when executing the program stored in the memory 113 .

[0060] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the speech conversion method provided in any of the aforementioned method embodiments are implemented.

[0061] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0062] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A voice conversion method, characterized in that: The voice conversion method comprises: Extracting the prosody of a target speech and obtaining target prosody parameters includes: obtaining a speech vector of the target speech, the speech vector being used to characterize audio features of the target speech, the audio features including timbre features and prosody features; inputting the speech vector into a prosody encoder, removing the timbre features from the speech vector, extracting the prosody features from the speech vector, and obtaining the target prosody parameters, wherein the target speech is speech requiring speech conversion, and redundant information in the target speech is removed, the redundant information including noise information, content, and speaker information, and the prosody features including pitch, loudness, speaking rate, and rhythm; Identifying the speaker of the target speech and determining a target timbre based on the speaker, wherein the target timbre is the timbre of the target speech that needs to be converted into speech; Inputting the target speech into an end-to-end speech recognition encoder to obtain a content vector output by the end-to-end speech recognition encoder after processing the target speech, including: determining an energy vector corresponding to the target speech, wherein the energy vector is used to characterize the emotional characteristics of the target speech; inputting the target speech and the energy vector into the end-to-end speech recognition encoder, so that the end-to-end speech recognition encoder encodes the target speech according to the energy vector to obtain an encoding result, and uses the encoding result as the content vector, wherein the content vector is used to characterize unclassified phoneme features obtained from the target speech, and the time resolution of the content vector is consistent with the time resolution of the target speech, so that the phoneme features represented by the content vector are consistent with the phoneme features contained in the target speech; The target speech is converted into speech according to the target prosodic parameter, the content vector, and the target timbre.

2. The method according to claim 1, characterized in that Performing voice conversion on the target speech according to the target prosodic parameter and the content vector includes: Performing fusion encoding on the content vector and the target prosody parameter to obtain an encoding vector; Decoding the encoding vector to obtain a Melp feature; The melp features are input into a vocoder, so that the vocoder decodes the melp features and generates speech after the target speech is speech-converted.

3. The method according to claim 2, characterized in that The target timbre is used to decode the encoding vector to obtain the Melp feature.

4. The method according to claim 2, characterized in that The encoding vector is decoded according to the target prosody parameter to obtain a Melp feature, including: The target vector is decoded on the target timbre according to the target prosodic parameter to obtain the Melp feature.

5. A voice conversion device, characterized in that: The voice conversion device comprises: A prosody acquisition module is configured to extract the prosody of a target speech and obtain target prosody parameters, including: obtaining a speech vector of the target speech, wherein the speech vector is used to characterize audio features of the target speech, wherein the audio features include timbre features and prosody features; inputting the speech vector into a prosody encoder, removing the timbre features from the speech vector, extracting the prosody features from the speech vector, and obtaining the target prosody parameters; wherein the target speech is speech that requires speech conversion; and removing redundant information from the target speech, wherein the redundant information includes noise information, content, and speaker information; and the prosody features include pitch, loudness, speaking rate, and rhythm; a content acquisition module, the content acquisition module being configured to input the target speech into an end-to-end speech recognition encoder to obtain a content vector output by the end-to-end speech recognition encoder after processing the target speech, the content vector being configured to represent unclassified phoneme features obtained from the target speech, and the temporal resolution of the content vector being consistent with the temporal resolution of the target speech, so that the phoneme features represented by the content vector are consistent with the phoneme features contained in the target speech; a speech conversion module, configured to perform speech conversion on the target speech according to the target prosody parameter, the content vector, and the target timbre; The speech conversion device further includes a module for identifying a speaker of the target speech and determining a target timbre according to the speaker, the target timbre being the timbre of the target speech to be speech-converted; The content acquisition module is also used to: determine the energy vector corresponding to the target speech, where the energy vector is used to characterize the emotional characteristics of the target speech; input the target speech and the energy vector into the end-to-end speech recognition encoder, so that the end-to-end speech recognition encoder encodes the target speech according to the energy vector, obtains the encoding result, and uses the encoding result as the content vector.

6. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; The processor is configured to implement the steps of the speech conversion method according to any one of claims 1 to 4 when executing the program stored in the memory.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the speech conversion method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Cross-language timbre conversion system and method based on zero-order learning

    CN112767958A

  • Speech model training method and speech synthesis method based on front-end design

    CN113257221A