Artificial intelligence-based voice conversion method and device, computer device and medium

By performing feature extraction and mapping in the encoder and decoder, combined with random resampling training, the low accuracy of existing speech conversion methods in the fintech field is solved, enabling more natural and expressive robot customer service and improving customer experience and service quality.

CN116631422BActive Publication Date: 2026-05-01PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2023-05-26
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing voice conversion methods have low accuracy in the fintech field, resulting in unnatural voice service from chatbots, which affects customer experience and service quality.

Method used

By inputting the speech to be converted and the reference speech into the trained encoder for feature extraction, optimized semantic features and optimized prosodic features are obtained. These features are then input into the trained decoder for feature mapping. Combined with the reference speaker features of the target speaker, speech conversion is performed. At the same time, augmented speech is generated by random resampling for feature extraction and decoupling training of the encoder and decoder.

Benefits of technology

It improved the accuracy of voice conversion, provided more natural and expressive robot customer service, and enhanced customer experience and the quality of financial services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116631422B_ABST
    Figure CN116631422B_ABST
Patent Text Reader

Abstract

The application is suitable for the field of financial technology, and particularly relates to a voice conversion method and device based on artificial intelligence, computer equipment and medium.The application extracts optimized semantic features and optimized prosodic features of the voice to be converted and reference speaker features of the reference voice through an encoder, uses a decoder to obtain target converted voice, extracts first semantic features, first prosodic features and first speaker features of the voice to be converted and second semantic features and second prosodic features of the augmented voice through the encoder, inputs the first semantic features, the first prosodic features and the first speaker features into the decoder to obtain reconstructed voice, and calculates model loss to train the encoder and the decoder, thereby improving the accuracy of the encoder and the decoder, improving the accuracy of voice conversion, providing customers with robot customer service with higher naturalness, expressiveness and richness in the field of financial technology, and improving service quality and customer experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention is applicable to the field of financial technology, and in particular relates to a speech conversion method, device, computer equipment and medium based on artificial intelligence. Background Technology

[0002] Voice conversion, which makes one person's speech sound like another's without altering the content of the speech, has demonstrated significant application value in various fields such as driving navigation, video production, game development, and voice customer service. For example, in the voice customer service scenario of the fintech sector, due to the complexity and diversity of financial transactions, simple tasks such as numerous inquiries and after-sales services can severely consume the time and energy of sales personnel, reducing their efficiency and work quality. Using chatbots can save substantial labor costs. Furthermore, the naturalness and fluency of the chatbot's voice directly impacts the user experience. Therefore, in financial scenarios, the voice of chatbots can be converted to provide customers with more natural, expressive, and richer voice customer service, thereby improving customer experience and ultimately enhancing the service quality of financial transactions.

[0003] The voice conversion of robot customer service relies on speech conversion methods. Existing speech conversion methods typically extract the textual semantic information of the speech to be converted and the target speaker information, and then fuse and map the textual semantic information and speaker information to obtain the converted target speech. However, since speech synthesis is a highly upsampled process, the text-speech data pairs are mapped one-to-many, and the textual speech information does not contain prosodic information, the above-mentioned speech conversion methods do not decouple and extract the textual semantic information, prosodic information, and speaker information from the speech to be converted. This results in the converted target speech having shortcomings such as unclear high-frequency energy and a flat speech style, which greatly reduces the accuracy of the speech conversion results.

[0004] Therefore, in the context of voice customer service in the fintech field, improving the accuracy of voice conversion methods has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a speech conversion method, apparatus, computer device and medium based on artificial intelligence to solve the problem of low conversion accuracy of existing speech conversion methods.

[0006] In a first aspect, embodiments of the present invention provide an artificial intelligence-based speech conversion method, the speech conversion method comprising:

[0007] The process involves acquiring the speech to be converted and the reference speech of the target speaker, inputting the speech to be converted and the reference speech into a trained encoder for feature extraction, obtaining optimized semantic features and optimized prosodic features of the speech to be converted, and reference speaker features of the reference speech, and inputting the optimized semantic features, optimized prosodic features and reference speaker features into a trained decoder for feature mapping to obtain the target converted speech.

[0008] The training process for the encoder and the decoder includes:

[0009] The speech to be converted is randomly resampled to obtain the augmented speech of the speech to be converted;

[0010] The speech to be converted and the augmented speech are respectively input into the encoder for feature extraction to obtain the first semantic feature, the first prosodic feature and the first speaker feature of the speech to be converted, and the second semantic feature and the second prosodic feature of the augmented speech;

[0011] The first semantic feature, the first prosodic feature, and the first speaker feature are input into the decoder for feature mapping to obtain the reconstructed speech.

[0012] The loss of the calculation model is calculated based on the first semantic feature, the second semantic feature, the first prosodic feature, the second prosodic feature, the speech to be converted, and the reconstructed speech;

[0013] The encoder and decoder are trained based on the model loss to obtain a trained encoder and a trained decoder.

[0014] Secondly, embodiments of the present invention provide an artificial intelligence-based speech conversion device, the speech conversion device comprising:

[0015] The speech conversion module is used to acquire the speech to be converted and the reference speech of the target speaker. The speech to be converted and the reference speech are respectively input into a trained encoder for feature extraction to obtain optimized semantic features and optimized prosodic features of the speech to be converted, and reference speaker features of the reference speech. The optimized semantic features, optimized prosodic features and reference speaker features are input into a trained decoder for feature mapping to obtain the target converted speech.

[0016] The speech augmentation module is used to randomly resample the speech to be converted to obtain augmented speech of the speech to be converted;

[0017] The feature extraction module is used to input the speech to be converted and the augmented speech into the encoder for feature extraction, so as to obtain the first semantic feature, the first prosodic feature and the first speaker feature of the speech to be converted, and the second semantic feature and the second prosodic feature of the augmented speech.

[0018] The feature mapping module is used to input the first semantic feature, the first prosodic feature and the first speaker feature into the decoder for feature mapping to obtain reconstructed speech;

[0019] The loss calculation module is used to calculate the model loss based on the first semantic feature, the second semantic feature, the first prosodic feature, the second prosodic feature, the speech to be converted, and the reconstructed speech;

[0020] The model training module is used to train the encoder and the decoder based on the model loss to obtain a trained encoder and a trained decoder.

[0021] Thirdly, embodiments of the present invention provide a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the speech conversion method as described in the first aspect.

[0022] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the speech conversion method as described in the first aspect.

[0023] The beneficial effects of this invention compared to existing technologies are as follows: By inputting the speech to be converted and the reference speech into a trained encoder, optimized semantic features and optimized prosodic features of the speech to be converted, and reference speaker features of the reference speech are obtained. These optimized semantic features, prosodic features, and reference speaker features are then input into a trained decoder to obtain the target converted speech. The speaker information in the speech to be converted is replaced with the reference speaker information of the target speaker, and combined with the original speech and prosodic information for feature mapping, thus improving the accuracy of speech conversion. Furthermore, by inputting the speech to be converted and its augmented form into the encoder for feature extraction, the first semantic feature, first prosodic feature, and first speaker feature of the speech to be converted are obtained. The augmented speech's second semantic and prosodic features effectively extract and decouple features from the speech to be converted and augmented speech. The first semantic, prosodic, and first speaker features are input into the decoder for feature mapping to obtain the reconstructed speech. The model loss is calculated based on the first, second, first, and second semantic features, the first and second prosodic features, the speech to be converted, and the reconstructed speech, serving as the training basis for the encoder and decoder. This improves the accuracy of feature extraction and decoupling in the encoder, as well as the accuracy of feature mapping in the decoder, effectively enhancing the accuracy of speech conversion. In the voice customer service scenario of the fintech field, this provides customers with more natural, expressive, and richer robot customer service, improving service quality and customer experience. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a schematic diagram of an application environment for an artificial intelligence-based speech conversion method provided in Embodiment 1 of the present invention;

[0026] Figure 2 This is a flowchart illustrating an artificial intelligence-based speech conversion method provided in Embodiment 1 of the present invention.

[0027] Figure 3 This is a schematic diagram of the structure of an artificial intelligence-based speech conversion device provided in Embodiment 2 of the present invention;

[0028] Figure 4 This is a schematic diagram of the structure of a computer device provided in Embodiment 3 of the present invention. Detailed Implementation

[0029] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.

[0030] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0031] It should also be understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0032] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."

[0033] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0034] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0035] The embodiments of this invention can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0036] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0037] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0038] To illustrate the technical solution of the present invention, specific embodiments are described below.

[0039] The first embodiment of this invention provides an artificial intelligence-based speech conversion method, which can be applied to applications such as... Figure 1 In this application environment, the client communicates with the server. Clients include, but are not limited to, handheld computers, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, and personal digital assistants (PDAs). The server can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0040] This voice conversion method can be applied to various fields such as animation production, game development, voice navigation, and voice customer service. For example, in the voice customer service scenario of the fintech field, due to the complexity and diversity of financial business, a large number of simple tasks such as consultation and after-sales service will seriously occupy the energy and time of business personnel, reducing their work efficiency and quality. Using robot customer service can save a lot of labor costs. The naturalness and fluency of the robot customer service's voice service will directly affect the user experience. Therefore, this voice conversion method can convert the voice of robot customer service in financial scenarios to provide customers with more natural, expressive, and richer voice customer service, thereby improving the customer experience and improving the service quality of financial business.

[0041] See Figure 2 This is a flowchart illustrating an artificial intelligence-based speech conversion method provided in Embodiment 1 of the present invention. The speech conversion method described above can be applied to... Figure 1 In a client application, the speech conversion method may include the following steps:

[0042] Step S201: Obtain the speech to be converted and the reference speech of the target speaker. Input the speech to be converted and the reference speech into the trained encoder for feature extraction to obtain the optimized semantic features and optimized prosodic features of the speech to be converted, and the reference speaker features of the reference speech. Input the optimized semantic features, optimized prosodic features and reference speaker features into the trained decoder for feature mapping to obtain the target converted speech.

[0043] A given audio recording contains semantic information representing the content of the speech, speaker information representing the speaker's timbre and pronunciation habits, and prosodic information representing emotions, rhythm, and pauses. The speech-to-speech task aims to make one person's speech sound like another person's without altering the semantic and prosodic information; that is, replacing the original speaker information in the audio with the speaker information of the target speaker.

[0044] In this embodiment, in order to improve the accuracy of the speech conversion result when performing the speech conversion task, the speech to be converted is input into the trained encoder for feature extraction, and optimized semantic features and optimized prosodic features with high feature extraction and feature decoupling accuracy are obtained, which serve as the feature basis for speech conversion.

[0045] Then, the reference speech of the target speaker is obtained and input into the trained encoder for feature extraction. After decoupling the features of the reference speech, the reference speaker features are obtained. The reference speaker features of the reference speech are then combined with the optimized semantic features and optimized prosodic features of the speech to be converted and input into the trained decoder for feature mapping to obtain the target speech to be converted, thus completing the speech conversion task.

[0046] Correspondingly, in the voice customer service scenario of the fintech field, when converting the voice of a robot customer service representative, the voice to be converted can be the voice of the robot customer service representative interacting with the customer, and the reference voice can be the voice of the human customer service representative interacting with the customer. The reference speaker features of the human customer service representative can be combined with the optimized semantic features and optimized prosodic features of the voice to be converted, and then input into a trained decoder for feature mapping to obtain the target converted voice. The reference human customer service representative speaker features are used to synthesize a more natural, expressive, and richer voice for the robot customer service representative, thereby improving customer experience and the service quality of financial services.

[0047] The above-mentioned steps involve obtaining the reference speech of the target speaker, inputting the speech to be converted and the reference speech into a trained encoder for feature extraction, obtaining optimized semantic features and optimized prosodic features of the speech to be converted, and reference speaker features of the reference speech, and inputting the optimized semantic features, optimized prosodic features and reference speaker features into a trained decoder for feature mapping to obtain the target converted speech. Based on effective feature extraction and feature decoupling of the speech to be converted and the reference speech, the speaker information in the speech to be converted is replaced with the reference speaker information of the target speaker, and combined with the original speech information and prosodic information for feature mapping to obtain the target converted speech, thus improving the accuracy of speech conversion.

[0048] Step S202: Randomly resample the speech to be converted to obtain the augmented speech of the speech to be converted.

[0049] To improve the accuracy of feature extraction and decoupling of the encoder, as well as the accuracy of feature mapping of the decoder, this embodiment trains both the encoder and decoder to enhance the accuracy of the speech conversion results. This embodiment first processes the speech to be converted using methods such as random resampling, generating new augmented speech based on the original speech as the basis for analysis, thereby improving the accuracy of the speech conversion analysis.

[0050] Optionally, the speech to be converted is randomly resampled to obtain augmented speech, including:

[0051] The speech to be converted is randomly divided into N sub-speech segments. Each sub-speech segment is stretched or shortened according to a random multiplication factor to obtain N augmented sub-speech segments.

[0052] By concatenating N augmented speech segments, the augmented speech of the speech to be converted is obtained.

[0053] To improve the diversity of augmented speech categories, the speech to be converted is first randomly resampled and randomly divided into N sub-speech segments. Each sub-speech segment is then stretched or shortened at a random ratio to obtain N different augmented sub-speech segments. Then, the N augmented sub-speech segments are spliced ​​together according to their order in the speech to be converted to obtain the augmented speech of the speech to be converted.

[0054] Correspondingly, sub-speech, augmented sub-speech, and augmented speech all contain corresponding semantic information, speaker information, and prosodic information. Since augmented sub-speech is obtained by stretching or shortening augmented speech, compared with the corresponding sub-speech, augmented sub-speech changes prosodic information such as emotion, rhythm, and pauses, but does not change semantic information or speaker information such as timbre and pronunciation habits. Therefore, compared with the speech to be converted, augmented speech changes the prosodic information of the speech, but does not change the semantic information or speaker information of the speech.

[0055] The above steps of randomly resampling the speech to be converted to obtain augmented speech, and augmenting the speech to be converted to generate new augmented speech as the basis for the analysis of the speech to be converted, can effectively improve the accuracy of the analysis of the speech to be converted.

[0056] Step S203: Input the speech to be converted and the augmented speech into the encoder for feature extraction to obtain the first semantic feature, the first prosodic feature and the first speaker feature of the speech to be converted, as well as the second semantic feature and the second prosodic feature of the augmented speech.

[0057] Since a speech segment contains corresponding semantic information, speaker information, and prosodic information, and the speech conversion task requires replacing the original speaker information in the speech to be converted with the speaker information of the target speaker, feature extraction and feature decoupling are necessary. Specifically, the speech to be converted is input into an encoder for feature extraction, resulting in the first semantic feature, the first prosodic feature, and the first speaker feature of the speech to be converted.

[0058] In this embodiment, to improve the accuracy of feature extraction and feature decoupling of the encoder, augmented speech is used as a reference. The augmented speech is input into the encoder for feature extraction, resulting in the second semantic feature, second prosodic feature, and second speaker feature of the augmented speech. Since the augmented speech alters the prosodic information but not the semantic or speaker information compared to the speech to be converted, the encoder's feature extraction and feature decoupling accuracy can be improved by making the first and second semantic features increasingly similar and the first and second prosodic features increasingly dissimilar during encoder training. This, in turn, improves the accuracy of feature extraction and feature decoupling between the speech to be converted and the reference speech, ultimately enhancing the overall speech conversion accuracy.

[0059] Meanwhile, in order to improve the efficiency of speech conversion tasks and reduce the amount of data, in this embodiment, after the speech to be converted and the augmented speech are input into the encoder for feature extraction, the redundant feature information of the first semantic feature, first prosodic feature and first speaker feature of the speech to be converted, and the second semantic feature, second prosodic feature and second speaker feature of the augmented speech, which have been decoupled, is discarded. For example, the second speaker feature of the augmented speech is retained. Only the feature information that needs to be used is retained, such as the first semantic feature, first prosodic feature and first speaker feature of the speech to be converted, and the second semantic feature and second prosodic feature of the augmented speech.

[0060] The above steps involve inputting the speech to be converted and the augmented speech into the encoder for feature extraction, obtaining the first semantic feature, first prosodic feature, and first speaker feature of the speech to be converted, as well as the second semantic feature and second prosodic feature of the augmented speech. By performing feature extraction and feature decoupling on the speech to be converted and the augmented speech through the encoder, the feature basis of the speech conversion task is obtained. Using the features of the augmented speech and the speech to be converted as the training basis of the encoder can effectively improve the accuracy of feature extraction and feature decoupling of the encoder, thereby improving the accuracy of speech conversion.

[0061] Step S204: Input the first semantic feature, the first prosodic feature and the first speaker feature into the decoder for feature mapping to obtain the reconstructed speech.

[0062] The decoder can map the semantic features, prosodic features, and speaker features of the input to obtain the corresponding speech.

[0063] In this embodiment, in order to improve the accuracy of feature extraction and feature decoupling of the encoder, as well as the accuracy of feature mapping of the decoder, the first semantic features, first prosodic features, and first speaker features of the speech to be converted extracted by the encoder are input into the decoder for feature mapping to obtain reconstructed speech, which serves as a reference speech for the speech to be converted. When training the encoder and decoder, the accuracy of feature extraction and feature decoupling of the encoder, as well as the accuracy of feature mapping of the decoder, can be improved by making the reconstructed speech increasingly closer to the speech to be converted, thereby improving the accuracy of speech conversion.

[0064] The above steps of inputting the first semantic feature, the first prosodic feature, and the first speaker feature into the decoder for feature mapping to obtain the reconstructed speech, and using the reconstructed speech as a reference speech for the speech to be converted, as the training basis for the encoder and decoder, can effectively improve the feature extraction and feature decoupling accuracy of the encoder, as well as the feature mapping accuracy of the decoder, thereby improving the accuracy of speech conversion.

[0065] Step S205: Calculate the model loss based on the first semantic feature, the second semantic feature, the first prosodic feature, the second prosodic feature, the speech to be converted, and the reconstructed speech.

[0066] In this process, augmented speech alters the prosodic information of the speech but not the semantic information. Furthermore, the reconstructed speech is obtained by feature mapping the first semantic feature, first prosodic feature, and first speaker feature of the speech to be converted. Therefore, by training the encoder and decoder, the accuracy of feature extraction and feature decoupling of the encoder and the feature mapping accuracy of the decoder can be improved, thereby increasing the accuracy of speech conversion. This is achieved by making the first and second semantic features increasingly similar, the first and second prosodic features increasingly dissimilar, and the reconstructed speech increasingly similar to the speech to be converted.

[0067] Therefore, in this embodiment, the loss of the computational model based on the first semantic feature, the second semantic feature, the first prosodic feature, the second prosodic feature, the speech to be converted, and the reconstructed speech is used as the training basis for the encoder and the decoder, so as to improve the feature extraction and feature decoupling accuracy of the encoder and the feature mapping accuracy of the decoder, thereby improving the accuracy of speech conversion.

[0068] Optionally, the model loss is calculated based on the first semantic feature, the second semantic feature, the first prosodic feature, the second prosodic feature, the speech to be converted, and the reconstructed speech, including:

[0069] The first sub-loss is calculated based on the first semantic feature and the second semantic feature;

[0070] The second sub-loss is calculated based on the first and second prosodic features;

[0071] Calculate the third sub-loss based on the speech to be converted and the reconstructed speech;

[0072] The model loss is calculated based on the first sub-loss, the second sub-loss, and the third sub-loss.

[0073] Specifically, by training the encoder and decoder, the first semantic feature and the second semantic feature can become closer and closer, the first prosodic feature and the second prosodic feature can become less and less close, and the reconstructed speech can become closer and closer to the speech to be converted. This can improve the accuracy of feature extraction and feature decoupling of the encoder, as well as the accuracy of feature mapping of the decoder, thereby improving the accuracy of speech conversion.

[0074] Therefore, in this embodiment, a first sub-loss can be calculated based on the first semantic feature and the second semantic feature, a second sub-loss can be calculated based on the first prosodic feature and the second prosodic feature, and a third sub-loss can be calculated based on the speech to be converted and the reconstructed speech. Then, the model loss can be calculated based on the first sub-loss, the second sub-loss and the third sub-loss, which serve as the training basis for the encoder and the decoder.

[0075] Optionally, calculating the first sub-loss based on the first semantic feature and the second semantic feature includes:

[0076] Calculate the first similarity between the first semantic feature and the second semantic feature;

[0077] The first similarity score is normalized to obtain the normalized first similarity score;

[0078] The difference between the normalized first similarity and the preset value is used as the first sub-loss.

[0079] In this way, the accuracy of feature extraction and feature decoupling of the encoder and the feature mapping accuracy of the decoder can be improved by training the encoder and decoder so that the first semantic feature and the second semantic feature become closer and closer, thereby improving the accuracy of speech conversion.

[0080] Therefore, in this embodiment, a first similarity is calculated between the first semantic feature and the second semantic feature, and the first similarity is normalized to obtain a normalized first similarity. The larger the normalized first similarity, the closer the first semantic feature and the second semantic feature are, which indicates that the feature extraction and feature decoupling accuracy of the encoder and the feature mapping accuracy of the decoder are higher. Conversely, the smaller the normalized first similarity, the less close the first semantic feature and the second semantic feature are, which indicates that the feature extraction and feature decoupling accuracy of the encoder and the feature mapping accuracy of the decoder are lower.

[0081] The difference between the normalized first similarity and the preset value is used as the first sub-loss. The smaller the normalized first similarity, that is, the lower the feature extraction and feature decoupling accuracy of the encoder and the feature mapping accuracy of the decoder, the larger the corresponding first sub-loss.

[0082] The preset value can be set according to the actual situation. Since the first similarity is normalized in this embodiment, the preset value in this embodiment is set to 1.

[0083] Optionally, calculating the second sub-loss based on the first and second prosodic features includes:

[0084] Calculate the second similarity between the first prosodic feature and the second prosodic feature;

[0085] The second similarity is normalized to obtain the normalized second similarity.

[0086] The normalized second similarity is used as the second sub-loss.

[0087] One approach is to train the encoder and decoder so that the first prosodic feature and the second prosodic feature become increasingly dissimilar, thereby improving the accuracy of feature extraction and feature decoupling of the encoder, as well as the accuracy of feature mapping of the decoder, and thus improving the accuracy of speech conversion.

[0088] Therefore, in this embodiment, a second similarity between the first prosodic feature and the second prosodic feature is calculated, and the second similarity is normalized to obtain a normalized second similarity. The larger the normalized second similarity, the closer the first prosodic feature and the second prosodic feature are, indicating a lower feature extraction and feature decoupling accuracy of the encoder and a lower feature mapping accuracy of the decoder. Conversely, the smaller the normalized second similarity, the less close the first prosodic feature and the second prosodic feature are, indicating a higher feature extraction and feature decoupling accuracy of the encoder and a higher feature mapping accuracy of the decoder.

[0089] The normalized second similarity is then used as the second sub-loss. The larger the normalized second similarity, that is, the lower the accuracy of feature extraction and feature decoupling of the encoder and the accuracy of feature mapping of the decoder, the larger the corresponding second sub-loss.

[0090] Optionally, the calculation of the third sub-loss based on the speech to be converted and the reconstructed speech includes:

[0091] Obtain the first Mel spectrum matrix of the speech to be converted, and the second Mel spectrum matrix of the reconstructed speech;

[0092] Calculate the third similarity between the first Mel spectrum matrix and the second Mel spectrum matrix;

[0093] The third similarity is normalized to obtain the normalized third similarity.

[0094] The difference between the normalized third similarity and the preset value is used as the third sub-loss.

[0095] One approach is to train the encoder and decoder to make the speech to be converted and the reconstructed speech increasingly similar, thereby improving the accuracy of feature extraction and feature decoupling of the encoder and the accuracy of feature mapping of the decoder, and thus improving the accuracy of speech conversion.

[0096] In this embodiment, a third sub-loss needs to be calculated based on the speech to be converted and the reconstructed speech. In the field of speech processing, to improve the convenience of calculating the third sub-loss, we need to convert the speech signal into a corresponding spectrogram and use the data on the spectrogram as the corresponding speech information. Typically, spectrogram frequencies are linearly distributed, but the human ear's perception of frequency is logarithmic, meaning it is sensitive to changes in low frequencies and insensitive to changes in high frequencies. The non-linearly distributed Mel spectrum can effectively match the human ear's perception of frequency and is widely used in the field of speech processing.

[0097] Therefore, the first Mel spectrum matrix of the speech to be converted and the second Mel spectrum matrix of the reconstructed speech are obtained first. The third similarity between the first and second Mel spectrum matrices is then calculated as the similarity between the speech to be converted and the reconstructed speech. The third similarity is then normalized to obtain a normalized third similarity. A larger normalized third similarity indicates a closer similarity between the speech to be converted and the reconstructed speech, representing higher accuracy in feature extraction and decoupling by the encoder, and higher accuracy in feature mapping by the decoder. Conversely, a smaller normalized third similarity indicates a less similar similarity between the speech to be converted and the reconstructed speech, representing lower accuracy in feature extraction and decoupling by the encoder, and lower accuracy in feature mapping by the decoder.

[0098] The difference between the normalized third similarity and the preset value is used as the third sub-loss. The smaller the normalized third similarity, that is, the lower the accuracy of feature extraction and feature decoupling of the encoder and the accuracy of feature mapping of the decoder, the larger the corresponding third sub-loss.

[0099] Optionally, the model loss is calculated based on the first sub-loss, the second sub-loss, and the third sub-loss, including:

[0100] Obtain the first preset weight of the first sub-loss, the second preset weight of the second sub-loss, and the third preset weight of the third sub-loss;

[0101] Based on the first preset weight, the second preset weight, and the third preset weight, the first sub-loss, the second sub-loss, and the third sub-loss are weighted and summed, and the weighted summation result is used as the model loss.

[0102] Specifically, the first sub-loss is calculated based on the first semantic feature and the second semantic feature, the second sub-loss is calculated based on the first prosodic feature and the second prosodic feature, and the third sub-loss is calculated based on the speech to be converted and the reconstructed speech. The model loss can then be calculated based on the first sub-loss, the second sub-loss and the third sub-loss to characterize the accuracy of the encoder's feature extraction and feature decoupling, as well as the accuracy of the decoder's feature mapping.

[0103] To improve the accuracy of model loss calculation, in this embodiment, preset weights are set for the first sub-loss, the second sub-loss, and the third sub-loss according to the actual situation. The model loss with higher accuracy is obtained by weighted summation of the first sub-loss, the second sub-loss, and the third sub-loss.

[0104] Specifically, based on the actual situation, the first preset weight of the first sub-loss, the second preset weight of the second sub-loss, and the third preset weight of the third sub-loss are obtained. Then, the first sub-loss, the second sub-loss, and the third sub-loss can be weighted and summed according to the first, second, and third preset weights, and the weighted sum is used as the model loss. Correspondingly, the larger the model loss, the lower the accuracy of feature extraction and feature decoupling, as well as the feature mapping accuracy of the decoder; conversely, the smaller the model loss, the higher the accuracy of feature extraction and feature decoupling, as well as the feature mapping accuracy of the decoder.

[0105] The steps described above, which calculate the model loss based on the first semantic feature, the second semantic feature, the first prosodic feature, the second prosodic feature, the speech to be converted, and the reconstructed speech, use the similarity between the first and second semantic features, between the first and second prosodic features, and between the reconstructed speech and the speech to be converted to calculate the model loss. This loss is used to characterize the accuracy of the encoder's feature extraction and feature decoupling, as well as the accuracy of the decoder's feature mapping. As the training basis for the encoder and decoder, this can effectively improve the accuracy of speech conversion.

[0106] Step S206: Train the encoder and decoder according to the model loss to obtain the trained encoder and trained decoder.

[0107] The greater the model loss, the lower the accuracy of feature extraction and feature decoupling, as well as the accuracy of feature mapping in the decoder; conversely, the smaller the model loss, the higher the accuracy of feature extraction and feature decoupling, as well as the accuracy of feature mapping in the decoder.

[0108] Therefore, in this embodiment, the encoder and decoder are trained based on the model loss and the gradient descent method until the model loss converges, so as to improve the feature extraction and feature decoupling accuracy of the encoder and the feature mapping accuracy of the decoder, and obtain a trained encoder and a trained decoder, which are used to complete the speech conversion task and improve the speech conversion accuracy of the speech conversion task.

[0109] The above steps of training the encoder and decoder based on the model loss to obtain a trained encoder and a trained decoder improve the accuracy of feature extraction and feature decoupling of the encoder, as well as the accuracy of feature mapping of the decoder.

[0110] In this embodiment of the invention, the speech to be converted and the reference speech are respectively input into a trained encoder to obtain optimized semantic features and optimized prosodic features of the speech to be converted, and reference speaker features of the reference speech. The optimized semantic features, optimized prosodic features, and reference speaker features are then input into a trained decoder to obtain the target converted speech. The speaker information in the speech to be converted is replaced with the reference speaker information of the target speaker, and combined with the original speech information and prosodic information for feature mapping, thereby improving the accuracy of speech conversion. Furthermore, the speech to be converted and the augmented speech of the speech to be converted are respectively input into the encoder for feature extraction to obtain the first semantic features, the first prosodic features, and the first speaker features of the speech to be converted, and the second semantic features of the augmented speech. The system effectively extracts and decouples features from the first semantic feature and the second prosodic feature. The first semantic feature, the first prosodic feature, and the first speaker feature are input into the decoder for feature mapping to obtain the reconstructed speech. The model loss is calculated based on the first semantic feature, the second semantic feature, the first prosodic feature, the second prosodic feature, the speech to be converted, and the reconstructed speech, serving as the training basis for the encoder and decoder. This improves the accuracy of feature extraction and decoupling in the encoder, as well as the accuracy of feature mapping in the decoder, effectively enhancing the accuracy of speech conversion. In the voice customer service scenario of the fintech field, it provides customers with more natural, expressive, and richer robot customer service, improving service quality and customer experience.

[0111] Corresponding to the speech conversion method in the above embodiments, Figure 3 A structural block diagram of an artificial intelligence-based speech conversion device provided in Embodiment 2 of the present invention is given. For ease of explanation, only the parts related to the embodiments of the present invention are shown.

[0112] See Figure 3 The voice conversion device includes:

[0113] The speech conversion module 31 is used to acquire the speech to be converted and the reference speech of the target speaker. The speech to be converted and the reference speech are respectively input into the trained encoder for feature extraction to obtain the optimized semantic features and optimized prosodic features of the speech to be converted, as well as the reference speaker features of the reference speech. The optimized semantic features, optimized prosodic features and reference speaker features are input into the trained decoder for feature mapping to obtain the target converted speech.

[0114] The speech augmentation module 32 is used to randomly resample the speech to be converted to obtain augmented speech.

[0115] The feature extraction module 33 is used to input the speech to be converted and the augmented speech into the encoder for feature extraction, and obtain the first semantic feature, the first prosodic feature and the first speaker feature of the speech to be converted, as well as the second semantic feature and the second prosodic feature of the augmented speech.

[0116] Feature mapping module 34 is used to input the first semantic feature, the first prosodic feature and the first speaker feature into the decoder for feature mapping to obtain the reconstructed speech;

[0117] The loss calculation module 35 is used to calculate the model loss based on the first semantic feature, the second semantic feature, the first prosodic feature, the second prosodic feature, the speech to be converted, and the reconstructed speech;

[0118] The model training module 36 is used to train the encoder and decoder based on the model loss, so as to obtain the trained encoder and the trained decoder.

[0119] Optionally, the aforementioned voice augmentation module 32 includes:

[0120] The speech augmentation submodule is used to randomly divide the speech to be converted into N sub-speech segments, and stretch or shorten each sub-speech segment according to a random multiplication factor to obtain N augmented sub-speech segments;

[0121] The speech splicing submodule is used to splice N augmented speech segments to obtain the augmented speech of the speech to be converted.

[0122] Optionally, the loss calculation module 35 mentioned above includes:

[0123] The first sub-loss calculation submodule is used to calculate the first sub-loss based on the first semantic feature and the second semantic feature;

[0124] The second sub-loss calculation submodule is used to calculate the second sub-loss based on the first prosodic feature and the second prosodic feature;

[0125] The third sub-loss calculation submodule is used to calculate the third sub-loss based on the speech to be converted and the reconstructed speech.

[0126] The model loss calculation submodule is used to calculate the model loss based on the first sub-loss, the second sub-loss, and the third sub-loss.

[0127] Optionally, the aforementioned first sub-loss calculation submodule includes:

[0128] The first similarity calculation unit is used to calculate the first similarity between the first semantic feature and the second semantic feature;

[0129] The first similarity normalization unit is used to normalize the first similarity to obtain the normalized first similarity.

[0130] The first sub-loss calculation unit is used to take the difference between the normalized first similarity and the preset value as the first sub-loss.

[0131] Optionally, the aforementioned second sub-loss calculation submodule includes:

[0132] The second similarity calculation unit is used to calculate the second similarity between the first prosodic feature and the second prosodic feature;

[0133] The second similarity normalization unit is used to normalize the second similarity to obtain the normalized second similarity.

[0134] The second sub-loss determination unit is used to take the normalized second similarity as the second sub-loss.

[0135] Optionally, the aforementioned third sub-loss calculation submodule includes:

[0136] The Mel spectrum matrix acquisition unit is used to acquire the first Mel spectrum matrix of the speech to be converted and the second Mel spectrum matrix of the reconstructed speech;

[0137] The third similarity calculation unit is used to calculate the third similarity between the first Mel spectrum matrix and the second Mel spectrum matrix;

[0138] The third similarity normalization unit is used to normalize the third similarity to obtain the normalized third similarity.

[0139] The third sub-loss calculation unit is used to take the difference between the normalized third similarity and the preset value as the third sub-loss.

[0140] Optionally, the above model loss calculation submodule includes:

[0141] The preset weight acquisition unit is used to acquire the first preset weight of the first sub-loss, the second preset weight of the second sub-loss, and the third preset weight of the third sub-loss;

[0142] The model loss calculation unit is used to perform a weighted summation of the first sub-loss, the second sub-loss, and the third sub-loss based on the first preset weight, the second preset weight, and the third preset weight, and use the weighted summation result as the model loss.

[0143] It should be noted that the information interaction and execution process between the above modules are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.

[0144] Figure 4 This is a schematic diagram of the structure of a computer device provided in Embodiment 3 of the present invention. Figure 4As shown, the computer device of this embodiment includes: at least one processor ( Figure 4 Only one is shown in the diagram), a memory, and a computer program stored in the memory and executable on at least one processor, which, when executed by the processor, implements the steps in any of the above-described speech conversion method embodiments.

[0145] This computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 4 The examples of computer devices are merely examples and do not constitute a limitation on computer devices. Computer devices may include more or fewer components than shown in the illustration, or combinations of certain components, or different components, such as network interfaces, displays, and input devices.

[0146] The processor referred to can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0147] Memory includes readable storage media, internal memory, etc., wherein internal memory can be the RAM of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard drive of a computer device, or in other embodiments, it can be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, memory can include both internal storage units and external storage devices of a computer device. Memory is used to store the operating system, applications, bootloader, data, and other programs, such as program code for computer programs. Memory can also be used to temporarily store data that has been output or will be output.

[0148] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the functions described above can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the processes in the methods of the above embodiments by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code, a recording medium, a computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0149] The present invention can implement all or part of the processes in the methods of the above embodiments, or it can be accomplished by a computer program product. When the computer program product is run on a computer device, the computer device executes the steps in the above method embodiments.

[0150] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0151] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0152] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0153] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0154] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A speech conversion method based on artificial intelligence, characterized in that, The speech conversion method includes: The process involves acquiring the speech to be converted and the reference speech of the target speaker, inputting the speech to be converted and the reference speech into a trained encoder for feature extraction, obtaining optimized semantic features and optimized prosodic features of the speech to be converted, and reference speaker features of the reference speech, and inputting the optimized semantic features, optimized prosodic features and reference speaker features into a trained decoder for feature mapping to obtain the target converted speech. The training process for the encoder and the decoder includes: The speech to be converted is randomly resampled to obtain the augmented speech of the speech to be converted; The speech to be converted and the augmented speech are respectively input into the encoder for feature extraction to obtain the first semantic feature, the first prosodic feature and the first speaker feature of the speech to be converted, and the second semantic feature and the second prosodic feature of the augmented speech; The first semantic feature, the first prosodic feature, and the first speaker feature are input into the decoder for feature mapping to obtain the reconstructed speech. The loss of the calculation model is calculated based on the first semantic feature, the second semantic feature, the first prosodic feature, the second prosodic feature, the speech to be converted, and the reconstructed speech; The encoder and the decoder are trained based on the model loss to obtain a trained encoder and a trained decoder; The step of calculating the model loss based on the first semantic feature, the second semantic feature, the first prosodic feature, the second prosodic feature, the speech to be converted, and the reconstructed speech includes: Calculate the first sub-loss based on the first semantic feature and the second semantic feature; Calculate the second sub-loss based on the first prosodic feature and the second prosodic feature; Calculate the third sub-loss based on the speech to be converted and the reconstructed speech; Calculate the model loss based on the first sub-loss, the second sub-loss, and the third sub-loss; The calculation of the second sub-loss based on the first prosodic feature and the second prosodic feature includes: Calculate the second similarity between the first prosodic feature and the second prosodic feature; The second similarity is normalized to obtain the normalized second similarity. The normalized second similarity is used as the second sub-loss.

2. The speech conversion method according to claim 1, characterized in that, The step of randomly resampling the speech to be converted to obtain the augmented speech includes: The speech to be converted is randomly divided into N sub-speech segments. Each sub-speech segment is stretched or shortened according to a random multiplication factor to obtain N augmented sub-speech segments. The augmented speech segments are spliced ​​together to obtain the augmented speech of the speech to be converted.

3. The speech conversion method according to claim 1, characterized in that, The step of calculating the first sub-loss based on the first semantic feature and the second semantic feature includes: Calculate the first similarity between the first semantic feature and the second semantic feature; The first similarity is normalized to obtain the normalized first similarity. The difference between the normalized first similarity and the preset value is used as the first sub-loss.

4. The speech conversion method according to claim 1, characterized in that, The calculation of the third sub-loss based on the speech to be converted and the reconstructed speech includes: Obtain the first Mel spectrum matrix of the speech to be converted, and the second Mel spectrum matrix of the reconstructed speech; Calculate the third similarity between the first Mel spectrum matrix and the second Mel spectrum matrix; The third similarity is normalized to obtain the normalized third similarity. The difference between the normalized third similarity and the preset value is used as the third sub-loss.

5. The speech conversion method according to claim 1, characterized in that, The calculation of model loss based on the first sub-loss, the second sub-loss, and the third sub-loss includes: Obtain the first preset weight of the first sub-loss, the second preset weight of the second sub-loss, and the third preset weight of the third sub-loss; Based on the first preset weight, the second preset weight, and the third preset weight, the first sub-loss, the second sub-loss, and the third sub-loss are weighted and summed, and the weighted summation result is used as the model loss.

6. A speech conversion device based on artificial intelligence, characterized in that, The speech conversion device includes: The speech conversion module is used to acquire the speech to be converted and the reference speech of the target speaker. The speech to be converted and the reference speech are respectively input into a trained encoder for feature extraction to obtain optimized semantic features and optimized prosodic features of the speech to be converted, and reference speaker features of the reference speech. The optimized semantic features, optimized prosodic features and reference speaker features are input into a trained decoder for feature mapping to obtain the target converted speech. The speech augmentation module is used to randomly resample the speech to be converted to obtain augmented speech of the speech to be converted; The feature extraction module is used to input the speech to be converted and the augmented speech into the encoder for feature extraction, so as to obtain the first semantic feature, the first prosodic feature and the first speaker feature of the speech to be converted, and the second semantic feature and the second prosodic feature of the augmented speech. The feature mapping module is used to input the first semantic feature, the first prosodic feature and the first speaker feature into the decoder for feature mapping to obtain reconstructed speech; The loss calculation module is used to calculate the model loss based on the first semantic feature, the second semantic feature, the first prosodic feature, the second prosodic feature, the speech to be converted, and the reconstructed speech; The model training module is used to train the encoder and the decoder based on the model loss to obtain a trained encoder and a trained decoder. The loss calculation module includes: The first sub-loss calculation submodule is used to calculate the first sub-loss based on the first semantic feature and the second semantic feature; The second sub-loss calculation submodule is used to calculate the second sub-loss based on the first prosodic feature and the second prosodic feature; The third sub-loss calculation submodule is used to calculate the third sub-loss based on the speech to be converted and the reconstructed speech. The second sub-loss calculation submodule includes: The second similarity calculation unit is used to calculate the second similarity between the first prosodic feature and the second prosodic feature; The second similarity normalization unit is used to normalize the second similarity to obtain the normalized second similarity. The second sub-loss determination unit is used to take the normalized second similarity as the second sub-loss.

7. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the speech conversion method as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech conversion method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Speech synthesis model training method, speech synthesis method and speech synthesis device

    CN114141228A

  • Updating method and application method of sound conversion model

    CN115273777A