A voice conversion method, apparatus, device and medium
By extracting and integrating timbre and rhythm features through a trained voice conversion model, the problem of lack of personalization in voice conversion in traditional banking intelligent customer service systems is solved, and service quality is improved.
Patent Information
- Application Number
- CN202411760630.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2044-12-02
AI Technical Summary
Traditional banking intelligent customer service systems lack personalization in the voice conversion process, leading to customer aesthetic fatigue and difficulty in improving service quality.
Using the trained speech conversion model, the timbre, rhythm and content features of the speech are extracted through the timbre encoder, rhythm encoder and content encoder, and then decoded and reconstructed, and the rhythm features of the target converted speech are integrated to improve personalized services.
By introducing a prosody encoder to extract the prosody features of the target converted speech, the converted speech is prevented from being too mechanical and the personalized service quality of speech conversion is improved.
Smart Images

Figure CN119649833B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a voice conversion method, device, equipment and medium. Background Art
[0002] With the rapid development of financial technology, bank customer service is moving towards intelligent and personalized services. However, traditional bank intelligent customer service systems typically use voice conversion models to convert customer service voices into different timbres to provide voice answers. However, the lack of variation and mechanical nature of the spoken content can easily lead to customer fatigue and hinder service quality. Therefore, when using voice conversion models to convert customer service voices into different timbres, improving the personalized service of the converted voice has become an urgent problem. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a voice conversion method, apparatus, device, and medium to solve the problem of over-mechanization in the process of using a voice conversion model to convert customer service voice into voices with different timbres.
[0004] In a first aspect, an embodiment of the present invention provides a voice conversion method, the voice conversion method comprising:
[0005] Obtaining a content text of the speech to be converted and a target converted speech, as well as a trained speech conversion model, wherein the trained speech conversion model includes a trained timbre encoder, a trained temperament encoder, and a trained content encoder;
[0006] Inputting the content text into the trained content encoder and outputting content features to be converted;
[0007] Inputting the target converted speech into the trained temperament encoder and outputting the temperament features to be converted;
[0008] Inputting the target converted speech into the trained timbre encoder and outputting the timbre features to be converted;
[0009] The content features to be converted, the musicality features to be converted, and the timbre features to be converted are input into a decoder for decoding and reconstruction to obtain the converted speech.
[0010] In a second aspect, an embodiment of the present invention provides a speech conversion device, the speech conversion device comprising:
[0011] A second acquisition module is used to obtain the content text of the speech to be converted and the target converted speech, as well as a trained speech conversion model, wherein the trained speech conversion model includes a trained timbre encoder, a trained temperament encoder, and a trained content encoder;
[0012] A content feature extraction module, configured to input the content text into the trained content encoder and output content features to be converted;
[0013] A temperament feature extraction module, configured to input the target converted speech into the trained temperament encoder and output the temperament features to be converted;
[0014] A timbre feature extraction module, configured to input the target converted speech into the trained timbre encoder and output timbre features to be converted;
[0015] The reconstruction module is used to input the content features to be converted, the musicality features to be converted and the timbre features to be converted into a decoder for decoding and reconstruction to obtain the converted speech.
[0016] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the speech conversion method as described in the first aspect when executing the computer program.
[0017] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the speech conversion method as described in the first aspect is implemented.
[0018] Compared with the prior art, the present invention has the following beneficial effects:
[0019] In this application, the content text of the speech to be converted and the target conversion speech, as well as the trained speech conversion model, are obtained. The trained speech conversion model includes a trained timbre encoder, a trained pitch encoder and a trained content encoder. The content text is input into the trained content encoder, which outputs the content features to be converted. The target conversion speech is input into the trained pitch encoder, which outputs the pitch features to be converted. The target conversion speech is input into the trained pitch encoder, which outputs the pitch features to be converted. The content features to be converted, the pitch features to be converted and the pitch features to be converted are input into the decoder for decoding and reconstruction to obtain the converted speech. In this application, when using the trained speech conversion model for speech conversion, a pitch encoder is introduced to extract the pitch features of the target conversion speech, and the pitch features of the target conversion speech are integrated into the converted speech, so that the converted speech has the pitch features of the target conversion speech. In the process of using the speech conversion model to convert the customer service voice into a voice with different timbre, the problem of the converted speech being too mechanical is avoided, thereby improving the service quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 This is a schematic diagram of an application environment of a voice conversion method provided in Example 1 of the present application;
[0022] Figure 2 This is a flow chart of a voice conversion method provided in the second embodiment of the present invention;
[0023] Figure 3 This is a flow chart of a voice conversion method provided in Embodiment 3 of the present invention;
[0024] Figure 4 This is a structural block diagram of a speech conversion device provided by a fourth embodiment of the present invention;
[0025] Figure 5 This is a structural block diagram of a speech conversion device provided by Embodiment 5 of the present invention;
[0026] Figure 6 This is a structural diagram of a computer device provided in Example 6 of the present invention. DETAILED DESCRIPTION
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0028] In the following description, specific details such as particular system structures and techniques are provided for purposes of illustration, not limitation, to facilitate a thorough understanding of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present invention with unnecessary detail.
[0029] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0030] It will also be understood that the term "and / or" used in the present description and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0031] As used in the present specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0032] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0033] References to "one embodiment" or "some embodiments" in the present specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present invention. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0034] Embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.
[0035] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0036] It should be understood that the order of execution of the steps in the following embodiments does not necessarily mean the order in which they are executed. The order in which each process is executed should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0037] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.
[0038] The speech conversion method provided by the first embodiment of the present invention can be applied in the following situations: Figure 1In an application environment, clients communicate with servers. Clients include, but are not limited to, PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, personal digital assistants (PDAs), and other computer devices. Servers can be standalone servers or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0039] See also Figure 2 , is a flow chart of a voice conversion method provided by the second embodiment of the present invention, the above-mentioned voice conversion method can be applied to Figure 1 The server in Figure 2 As shown, the voice conversion method may include the following steps.
[0040] S201: Obtain content text of the speech to be converted, target conversion speech, and a trained speech conversion model. The trained speech conversion model includes a trained timbre encoder, a trained temperament encoder, and a trained content encoder.
[0041] S202: Input the content text into the trained content encoder and output the content features to be converted;
[0042] S203: Input the target converted speech into the trained temperament encoder and output the temperament features to be converted;
[0043] S204: Input the target converted speech into the trained timbre encoder and output the timbre features to be converted;
[0044] S205: Input the content features to be converted, the musicality features to be converted, and the timbre features to be converted into a decoder for decoding and reconstruction to obtain the converted speech.
[0045] In this embodiment, the content text of the speech to be converted, the target conversion speech, and a trained speech conversion model are obtained. The trained speech conversion model includes a trained timbre encoder, a trained prosody encoder, and a trained content encoder. The content text of the speech to be converted is the content corresponding to the converted speech, and the timbre features and prosody features of the target conversion speech are the timbre features and prosody features of the converted speech. The content text is input into the trained content encoder, which outputs the content features to be converted. The target conversion speech is input into the trained prosody encoder, which outputs the prosody features to be converted. The target conversion speech is input into the trained timbre encoder, which outputs the timbre features to be converted.
[0046] It should be noted that when the trained timbre encoder, the trained temperament encoder, and the trained content encoder are used for feature extraction, it is not necessary to segment the target converted speech.
[0047] The content features to be converted, the musical characteristics to be converted, and the timbre features to be converted are input into the decoder for decoding and reconstruction to obtain the converted speech.
[0048] See also Figure 3 , is a flowchart of a training method for a trained speech conversion model provided by the third embodiment of the present invention. The training method for the trained speech conversion model can be applied to Figure 1 The server in Figure 3 As shown, the training method may include the following steps.
[0049] S301: Obtain source speech and an initial speech conversion model for financial services, segment the source speech, and obtain segmented first and second speech. The initial speech conversion model includes an initial prosody encoder, an initial content encoder, an initial timbre encoder, and an adversarial module.
[0050] In step S301, the source speech is a training sample for training the initial speech conversion model for financial services. The source speech is segmented to obtain a segmented first speech and a second speech, that is, the source speech is divided into two parts. The initial speech conversion model for financial services is a deep learning model. The initial speech conversion model includes an initial prosody encoder, an initial content encoder, an initial timbre encoder and an adversarial module, wherein the initial prosody encoder is used to extract the prosodic features of the speech, the initial content encoder is used to extract the content features of the speech, the initial timbre encoder is used to extract the timbre features of the speech, and the adversarial module is used for adversarial training.
[0051] In this embodiment, a source voice is obtained, where the source voice may be voice collected for financial services, and the source voice is segmented to obtain a segmented first voice and a segmented second voice, that is, the source voice is segmented into two voice sequences.
[0052] It should be noted that the source speech is divided into the first speech and the second speech, and the initial speech conversion model is trained by combining the features of the first speech and the second speech, as well as the features of the source speech. This makes full use of limited speech data, learns the speaker's speech representation as fully as possible, and reduces the cost of manual voice recording for customer service.
[0053] It should be noted that, when the source speech is segmented, the source speech may also be segmented into speech sequences of three or more parts, which is not limited in this embodiment.
[0054] Obtain an initial speech conversion model for financial services. This model is used to convert speech into a different timbre or prosody. The initial speech conversion model includes an initial prosody encoder, an initial content encoder, an initial timbre encoder, and an adversarial module. The initial speech conversion model is trained using source speech as training samples.
[0055] S302: Input the source speech into the initial prosody encoder, output the prosody features of the source speech, input the prosody features of the source speech into the adversarial module, and output the predicted content features generated based on the prosody features of the source speech.
[0056] In step S302, the prosodic features of the source speech are extracted using the initial prosodic encoder to obtain the prosodic features of the source speech, the adversarial module is used to generate predicted content features, and the adversarial module is used to perform adversarial training on the initial prosodic encoder so that the prosodic features output by the trained prosodic encoder do not contain the content information of the speech.
[0057] In this embodiment, the initial prosody encoder is connected to the adversarial module, so that the initial prosody encoder is adversarially trained according to the adversarial module to expand the difference between the predicted content features generated based on the prosodic features of the source speech and the source speech content features output by the initial content encoder, so that the prosodic features output by the trained prosody encoder do not contain the content information of the speech.
[0058] Optionally, the source speech prosodic features are input into the adversarial module, and the output of the predicted content features generated based on the source speech prosodic features includes:
[0059] The parameters of the initial prosody encoder are updated inversely through the gradient reversal layer, so that the prosody features of the source speech output by the initial prosody encoder do not contain the content information of the source speech;
[0060] The content predictor is used to perform content prediction on the prosodic features of the source speech to obtain predicted content features.
[0061] In this embodiment, the adversarial module includes a gradient reversal layer and a content predictor. The initial prosody encoder is connected to the gradient reversal layer, which performs a reverse update on the parameters of the initial prosody encoder. During the reverse update, the gradient is multiplied by a negative number to perform the reverse update. The gradient reversal layer is connected to the content predictor and generates predicted content features based on the content predictor, where the content predictor can be a generator.
[0062] The gradient reversal layer is connected to the initial prosody encoder and the content predictor respectively. It can perform adversarial training on the initial prosody encoder by enlarging the difference between the predicted content features and the meta-speech content features output by the initial content encoder, so that the prosody features output by the trained prosody encoder do not contain the content information of the speech.
[0063] S303: Input the first speech and the second speech into the initial timbre encoder, output the first timbre feature of the first speech and the second timbre feature of the second speech, input the first speech and the second speech into the initial prosody encoder, output the first prosody feature of the first speech and the second prosody feature of the second speech, input the source speech into the initial content encoder, and output the source speech content feature.
[0064] In step S303, the initial timbre encoder is used to extract timbre features of the first speech and the second speech after segmentation, so as to determine whether the first timbre features of the first speech are similar to the second timbre features of the second speech. The initial prosody encoder is used to extract prosody features of the first speech and the second speech after segmentation, so as to determine whether the prosody features of the segmented speech are similar to the prosody features of the unsegmented source speech. The initial content encoder is used to extract content features of the source speech, so as to determine whether the difference between the content features of the source speech and the predicted content features increases.
[0065] In this embodiment, the first speech and the second speech are input into the initial timbre encoder, and the first timbre feature of the first speech and the second timbre feature of the second speech are output; the first speech and the second speech are input into the initial prosody encoder, and the first prosody feature of the first speech and the second prosody feature of the second speech are output; the source speech is input into the initial content encoder, and the source speech content feature is output.
[0066] It should be noted that the timbre features of the same source speech should be similar in different parts. Therefore, the training loss corresponding to the timbre encoder can be calculated based on the similarity between the first timbre feature and the second timbre feature. The prosodic features of the same source speech should be similar to the prosodic features of the segmented source speech. Therefore, the training loss corresponding to the prosodic encoder can be calculated based on whether the prosodic features of the source speech are similar to the first prosodic features and the second prosodic features. The training loss of the content encoder is calculated by predicting the differences between speech features based on the content features of the source speech.
[0067] S304: Calculate the timbre loss based on the first timbre feature and the second timbre feature, calculate the prosodic loss based on the first prosodic feature, the second prosodic feature and the prosodic feature of the source speech, and calculate the adversarial loss based on the predicted content feature and the content feature of the source speech.
[0068] In step S304, the timbre loss is the difference between the timbre features of different parts of the source speech, the prosody loss is the difference between the prosody features of the speech after segmentation and the prosody features of the source speech before segmentation, and the adversarial loss is the difference between the predicted content features generated by the adversarial module and the content features of the source speech.
[0069] In this embodiment, when the timbre loss is calculated based on the first timbre feature and the second timbre feature, the cosine similarity between the first timbre feature and the second timbre feature can be calculated, and the cosine similarity is used as the timbre loss. The calculation formula is as follows:
[0070]
[0071] in, For the loss of timbre, is the first timbre characteristic, is the second timbre characteristic, Other methods may also be used for calculation, such as calculating the difference between the first timbre feature and the second timbre feature to obtain the timbre feature, which is not limited in this embodiment.
[0072] According to the first prosodic feature, the second prosodic feature and the prosodic feature of the source speech, the prosodic feature of the source speech can be subtracted from the prosodic feature obtained by splicing the first prosodic feature and the second prosodic feature according to time to obtain the prosodic loss. The calculation formula is as follows:
[0073]
[0074] in, is the rhythm loss, is the prosodic feature of the source speech, is the first rhythmic feature, is the second rhythmic feature, The first rhythmic feature and the second rhythmic feature are spliced according to time.
[0075] In this embodiment, the first prosodic features of the segmented first speech and the second prosodic features of the segmented second speech are compared with the prosodic features of the source speech to calculate the prosodic loss. This allows the initial prosodic encoder to learn the prosodic features of speech of varying durations. Furthermore, the accuracy of the prosodic features extracted by the initial prosodic encoder for long speech is similar to that for short speech. This allows the initial prosodic encoder to learn the prosodic features of speech of any length, thereby improving the accuracy of prosodic feature extraction.
[0076] When the adversarial loss is calculated based on the predicted content features and the source speech content features, the difference between the predicted content features and the source speech content features can be calculated using the following formula:
[0077]
[0078] in, To combat losses, is the source speech content feature, is the predicted content feature generated based on the prosodic features of the source speech. Other methods can also be used to calculate the adversarial loss, such as using cosine similarity for calculation, which is not limited in this embodiment.
[0079] The prosodic loss is calculated based on the first prosodic feature, the second prosodic feature, and the prosodic feature of the source speech, including:
[0080] splicing the first rhythmic feature and the second rhythmic feature to obtain a spliced rhythmic feature;
[0081] The similarity between the prosodic features of the concatenated speech and the prosodic features of the source speech is calculated to obtain the prosodic loss.
[0082] In this embodiment, based on the first prosodic feature, the second prosodic feature, and the prosodic feature of the source speech, the cosine similarity between the first prosodic feature of the segmented first speech and the second prosodic feature of the second speech after time concatenation and the prosodic feature of the source speech of the source language can be calculated. The calculation formula is as follows:
[0083]
[0084] in, is the rhythm loss, is the first rhythmic feature, is the second rhythmic feature, The first rhythmic feature and the second rhythmic feature are spliced according to time, is the transposition of the prosodic features of the source speech, Other methods can also be used for calculation, such as subtracting the prosodic feature of the source speech from the prosodic feature obtained by splicing the first prosodic feature and the second prosodic feature according to time to obtain the prosodic loss, which is not limited in this embodiment.
[0085] S305: The initial speech conversion model is trained according to the timbre loss, the prosody loss, and the adversarial loss to obtain a trained speech conversion model. The trained speech conversion model includes a trained timbre encoder, a trained prosody encoder, and a trained content encoder.
[0086] In step S305, the total loss is calculated based on the timbre loss, prosody loss and adversarial loss, and the initial speech conversion model is trained based on the total loss. When the total loss converges, a trained speech conversion model is obtained. The trained speech conversion model includes a trained timbre encoder, a trained prosody encoder and a trained content encoder.
[0087] In this embodiment, the total loss is calculated based on the timbre loss, prosody loss, and adversarial loss, where the total loss includes the losses during the training of the initial prosody encoder, the initial content encoder, and the initial timbre encoder. The initial speech conversion model is trained so that the total loss during the training process converges, or when the total loss is minimized, the training is stopped to obtain a trained speech conversion model.
[0088] It should be noted that when calculating the total loss based on timbre loss, rhythm loss and adversarial loss, different weight values can be set for each loss, and the sum of the different weight values is 1.
[0089] It should be noted that when calculating the adversarial loss, the difference between the predicted content features and the source speech content features can be calculated. The larger the difference, the better. Therefore, when calculating the total loss, subtract the adversarial loss to make the total loss as small as possible.
[0090] The total loss is calculated as follows:
[0091]
[0092] in, For the total loss, For the loss of timbre, is the rhythm loss, To combat losses, 、 and is the weight value, 、 and The sum of is 1.
[0093] Optionally, the initial speech conversion model is trained according to timbre loss, prosody loss, and adversarial loss to obtain a trained speech conversion model, including:
[0094] Add the timbre loss and rhythm loss to calculate the sum loss;
[0095] Subtract the sum loss from the adversarial loss to calculate the total loss;
[0096] According to the total loss, the initial speech conversion model is trained to obtain a trained speech conversion model.
[0097] In this embodiment, when calculating the total loss, the timbre loss and the rhythm loss can also be added to obtain the sum loss, and the sum loss can be subtracted from the adversarial loss to calculate the total loss. Based on the total loss, the initial speech conversion model is trained. When the total loss is minimized or the total loss converges, the training is stopped to obtain a trained speech conversion model.
[0098] Optionally, the total loss is calculated, further comprising:
[0099] Input the source speech into the initial timbre encoder and output the timbre features of the source speech;
[0100] Input the source speech timbre features, source speech prosody features and source speech content features into the decoder for decoding and reconstruction, and output the reconstructed speech;
[0101] Calculate the reconstruction loss based on the reconstructed speech and the source speech;
[0102] The total loss is calculated based on reconstruction loss, timbre loss, temperament loss and adversarial loss.
[0103] In this embodiment, the initial speech conversion model also includes a decoder, which is used to decode and reconstruct the source speech content features, source speech timbre features and source speech prosody features to obtain reconstructed speech. The source speech is input into the initial timbre encoder, and the source speech timbre features are output. The source speech timbre features, source speech prosody features and source speech content features are input into the decoder for decoding and reconstruction, and the reconstructed speech is output. Based on the reconstructed speech and the source speech, the reconstruction loss is calculated, wherein the reconstruction loss can be calculated based on the difference between the spectral features of the source speech and the spectral features of the reconstructed speech. Or other methods are used for calculation, which are not limited in this embodiment. Based on the reconstruction loss, timbre loss, prosody loss and adversarial loss, the total loss is calculated. The calculation formula for the total loss is as follows:
[0104]
[0105] in, For the total loss, is the reconstruction loss, For the loss of timbre, is the rhythm loss, To combat the loss, different weight values may be set for each loss, and the sum of the different weight values is 1. This embodiment does not limit this.
[0106] In this application, the content text of the speech to be converted and the target conversion speech, as well as the trained speech conversion model, are obtained. The trained speech conversion model includes a trained timbre encoder, a trained pitch encoder and a trained content encoder. The content text is input into the trained content encoder, which outputs the content features to be converted. The target conversion speech is input into the trained pitch encoder, which outputs the pitch features to be converted. The target conversion speech is input into the trained pitch encoder, which outputs the pitch features to be converted. The content features to be converted, the pitch features to be converted and the pitch features to be converted are input into the decoder for decoding and reconstruction to obtain the converted speech. In this application, when using the trained speech conversion model for speech conversion, a pitch encoder is introduced to extract the pitch features of the target conversion speech, and the pitch features of the target conversion speech are integrated into the converted speech, so that the converted speech has the pitch features of the target conversion speech. In the process of using the speech conversion model to convert the customer service voice into a voice with different timbre, the problem of the converted speech being too mechanical is avoided, thereby improving the service quality.
[0107] See also Figure 4 , Figure 4 This is a structural block diagram of a voice conversion device provided by the fourth embodiment of the present invention, which is applied to the above-mentioned server. For ease of explanation, only the parts related to the embodiment of this application are shown. Figure 4 The speech conversion device 40 includes a second acquisition module 41 , a content feature extraction module 42 , a musicality feature extraction module 43 , a timbre feature extraction module 44 , and a reconstruction module 45 .
[0108] The second acquisition module 41 is used to obtain the content text of the speech to be converted and the target converted speech, as well as a trained speech conversion model. The trained speech conversion model includes a trained timbre encoder, a trained temperament encoder and a trained content encoder.
[0109] The content feature extraction module 42 is used to input the content text into the trained content encoder and output the content features to be converted.
[0110] The music feature extraction module 43 is used to input the target converted speech into the trained music coder and output the music feature to be converted.
[0111] The timbre feature extraction module 44 is used to input the target converted speech into a trained timbre encoder and output the timbre features to be converted.
[0112] The reconstruction module 45 is used to input the content features to be converted, the musicality features to be converted and the timbre features to be converted into the decoder for decoding and reconstruction to obtain the converted speech.
[0113] See also Figure 5 , Figure 5 This is a structural block diagram of a speech conversion device provided by the fifth embodiment of the present invention. The above training device is applied to the above server. For the sake of convenience, only the part related to the embodiment of this application is shown. Figure 5 The training device 50 includes a first acquisition module 51 , a prediction module 52 , a feature extraction module 53 , a calculation module 54 , and a training module 55 .
[0114] The first acquisition module 51 is used to obtain the source speech and the initial speech conversion model for financial services, segment the source speech to obtain the segmented first speech and second speech, and the initial speech conversion model includes an initial prosody encoder, an initial content encoder, an initial timbre encoder and an adversarial module.
[0115] The prediction module 52 is used to input the source speech into the initial prosody encoder, output the prosody features of the source speech, input the prosody features of the source speech into the adversarial module, and output the predicted content features generated based on the prosody features of the source speech.
[0116] The feature extraction module 53 is used to input the first speech and the second speech into the initial content encoder, output the first timbre feature corresponding to the first speech and the second timbre feature of the second speech, input the first speech and the second speech into the initial prosody encoder, output the first prosody feature of the first speech and the second prosody feature of the second speech, input the source speech into the initial content encoder, and output the source speech content feature.
[0117] The calculation module 54 is used to calculate the timbre loss based on the first timbre feature and the second timbre feature, calculate the prosodic loss based on the first prosodic feature, the second prosodic feature and the prosodic feature of the source speech, and calculate the adversarial loss based on the predicted content feature and the content feature of the source speech.
[0118] The training module 55 is used to train the initial speech conversion model according to the timbre loss, prosody loss and adversarial loss to obtain a trained speech conversion model. The trained speech conversion model includes a trained timbre encoder, a trained prosody encoder and a trained content encoder.
[0119] Optionally, the prediction module 52 includes:
[0120] The updating unit is used to reversely update the parameters of the initial prosody encoder through a gradient reversal layer so that the prosody features of the source speech output by the initial prosody encoder do not contain the content information of the source speech.
[0121] The prediction unit is used to perform content prediction on the prosodic features of the source speech through a content predictor to obtain predicted content features.
[0122] Optionally, the calculation module 54 includes:
[0123] The splicing unit is used to splice the first rhythmic feature and the second rhythmic feature to obtain a spliced rhythmic feature.
[0124] The first calculation unit is used to calculate the similarity between the prosodic features after concatenation and the prosodic features of the source speech, and calculate the prosodic loss.
[0125] Optionally, the training module 55 includes:
[0126] The second calculation unit is used to add the timbre loss and the rhythm loss to calculate the sum loss.
[0127] The third calculation unit is used to subtract the sum loss from the adversarial loss to calculate the total loss.
[0128] The training unit is used to train the initial speech conversion model according to the total loss to obtain a trained speech conversion model.
[0129] Optionally, the training device 50 further includes:
[0130] The output module is used to input the source speech into the initial timbre encoder and output the timbre characteristics of the source speech.
[0131] The reconstruction module is used to input the source speech timbre features, source speech prosody features and source speech content features into the decoder for decoding and reconstruction, and output the reconstructed speech.
[0132] The reconstruction loss calculation module is used to calculate the reconstruction loss based on the reconstructed speech and the source speech.
[0133] The total loss calculation module is used to calculate the total loss based on reconstruction loss, timbre loss, temperament loss and adversarial loss.
[0134] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0135] Figure 6 This is a schematic diagram of the structure of a computer device provided by Example 6 of the present invention. Figure 6 As shown, the computer device of this embodiment includes: at least one processor ( Figure 6 Only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, the steps of any of the above-mentioned speech conversion method embodiments are implemented.
[0136] The computer device may include, but is not limited to, a processor and a memory. It will be understood by those skilled in the art that Figure 6 The above is merely an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include a network interface, a display screen, and an input device.
[0137] The processor may be a CPU, other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0138] Memory includes readable storage media, internal memory, and the like. Internal memory can be the internal memory of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage medium. The readable storage medium can be the computer device's hard drive. In other embodiments, it can also be an external storage device, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, or a flash memory card. Furthermore, memory can include both the computer device's internal storage unit and external storage devices. Memory is used to store the operating system, application programs, boot loaders, data, and other programs, such as the program code of computer programs. Memory can also be used to temporarily store data that has been output or is about to be output.
[0139] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process steps in the above-described method embodiments by instructing the relevant hardware through a computer program. The computer program may be stored in a computer-readable storage medium. When executed by a processor, the computer program implements the steps of the above-described method embodiments. The computer program includes computer program code, which may be in source code form, object code form, executable file, or some intermediate form. Computer-readable media may include at least: any entity or device capable of carrying computer program code, recording media, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunications signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunications signals.
[0140] The present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiment when executing it.
[0141] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0142] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0143] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely schematic. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of the apparatus or unit, which can be electrical, mechanical or other forms.
[0144] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0145] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A voice conversion method, characterized in that: The voice conversion method comprises: Obtaining a content text of the speech to be converted and a target converted speech, as well as a trained speech conversion model, wherein the trained speech conversion model includes a trained timbre encoder, a trained temperament encoder, and a trained content encoder; Inputting the content text into the trained content encoder and outputting content features to be converted; Inputting the target converted speech into the trained temperament encoder and outputting the temperament features to be converted; Inputting the target converted speech into the trained timbre encoder and outputting the timbre features to be converted; Inputting the content feature to be converted, the musicality feature to be converted, and the timbre feature to be converted into a decoder for decoding and reconstruction to obtain the converted speech; The training method of the trained speech conversion model includes: Obtaining a source speech and an initial speech conversion model for financial services, segmenting the source speech to obtain a first speech and a second speech after segmentation, wherein the initial speech conversion model includes an initial prosody encoder, an initial content encoder, an initial timbre encoder, and an adversarial module; Inputting the source speech into the initial prosody encoder to output the prosody features of the source speech, inputting the prosody features of the source speech into the adversarial module to output the predicted content features generated based on the prosody features of the source speech; Inputting the first speech and the second speech into the initial timbre encoder, outputting a first timbre feature of the first speech and a second timbre feature of the second speech; inputting the first speech and the second speech into the initial prosody encoder, outputting a first prosody feature of the first speech and a second prosody feature of the second speech; inputting the source speech into the initial content encoder, outputting a source speech content feature; Calculating a timbre loss based on the first timbre feature and the second timbre feature, calculating a prosodic loss based on the first prosodic feature, the second prosodic feature, and the prosodic feature of the source speech, and calculating an adversarial loss based on the predicted content feature and the content feature of the source speech; training the initial speech conversion model according to the timbre loss, the prosody loss, and the adversarial loss to obtain a trained speech conversion model, wherein the trained speech conversion model includes a trained timbre encoder, a trained prosody encoder, and a trained content encoder; The training of the initial speech conversion model according to the timbre loss, the prosody loss, and the adversarial loss to obtain a trained speech conversion model includes: Adding the timbre loss and the rhythm loss to calculate a sum loss; Subtract the sum loss from the adversarial loss to calculate the total loss; The initial speech conversion model is trained according to the total loss to obtain a trained speech conversion model.
2. The voice conversion method according to claim 1, wherein: The adversarial module includes a gradient reversal layer and a content predictor; The step of inputting the source speech prosodic features into the adversarial module and outputting predicted content features generated based on the source speech prosodic features comprises: Reversely updating the parameters of the initial prosody encoder through the gradient reversal layer so that the prosody features of the source speech output by the initial prosody encoder do not contain content information of the source speech; The content predictor performs content prediction on the source speech prosodic features to obtain predicted content features.
3. The voice conversion method according to claim 1, wherein: The calculating the prosodic loss according to the first prosodic feature, the second prosodic feature, and the prosodic feature of the source speech includes: splicing the first rhythmic feature and the second rhythmic feature to obtain a spliced rhythmic feature; The similarity between the prosodic features after concatenation and the prosodic features of the source speech is calculated to obtain the prosodic loss.
4. The voice conversion method according to claim 1, wherein: The initial speech conversion model also includes a decoder; The calculation to obtain the total loss also includes: Inputting the source speech into the initial timbre encoder and outputting the timbre characteristics of the source speech; Inputting the source speech timbre feature, the source speech prosody feature and the source speech content feature into the decoder for decoding and reconstruction, and outputting the reconstructed speech; Calculating a reconstruction loss based on the reconstructed speech and the source speech; A total loss is calculated according to the reconstruction loss, the timbre loss, the rhythm loss and the adversarial loss.
5. A voice conversion device, characterized in that: The voice conversion device comprises: A second acquisition module is used to obtain the content text of the speech to be converted and the target converted speech, as well as a trained speech conversion model, wherein the trained speech conversion model includes a trained timbre encoder, a trained temperament encoder, and a trained content encoder; A content feature extraction module, configured to input the content text into the trained content encoder and output content features to be converted; A temperament feature extraction module, configured to input the target converted speech into the trained temperament encoder and output the temperament features to be converted; A timbre feature extraction module, configured to input the target converted speech into the trained timbre encoder and output timbre features to be converted; A reconstruction module, configured to input the content features to be converted, the musicality features to be converted, and the timbre features to be converted into a decoder for decoding and reconstruction to obtain converted speech; The voice conversion device also includes: A first acquisition module is configured to acquire a source speech and an initial speech conversion model for financial services, segment the source speech to obtain a segmented first speech and a second speech, wherein the initial speech conversion model includes an initial prosody encoder, an initial content encoder, an initial timbre encoder, and an adversarial module; A prediction module, configured to input the source speech into the initial prosody encoder, output source speech prosody features, input the source speech prosody features into the adversarial module, and output predicted content features generated based on the source speech prosody features; a feature extraction module configured to input the first speech and the second speech into the initial content encoder and output a first timbre feature corresponding to the first speech and a second timbre feature corresponding to the second speech; input the first speech and the second speech into the initial prosody encoder and output a first prosody feature of the first speech and a second prosody feature of the second speech; and input the source speech into the initial content encoder and output a source speech content feature; a calculation module, configured to calculate a timbre loss based on the first timbre feature and the second timbre feature, calculate a prosodic loss based on the first prosodic feature, the second prosodic feature, and the prosodic feature of the source speech, and calculate an adversarial loss based on a predicted content feature and the content feature of the source speech; a training module, configured to train the initial speech conversion model according to the timbre loss, the prosody loss, and the adversarial loss to obtain a trained speech conversion model, wherein the trained speech conversion model includes a trained timbre encoder, a trained prosody encoder, and a trained content encoder; The training of the initial speech conversion model according to the timbre loss, the prosody loss, and the adversarial loss to obtain a trained speech conversion model includes: Adding the timbre loss and the rhythm loss to calculate a sum loss; Subtract the sum loss from the adversarial loss to calculate the total loss; The initial speech conversion model is trained according to the total loss to obtain a trained speech conversion model.
6. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the speech conversion method according to any one of claims 1 to 4 is implemented.
7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the speech conversion method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Voice conversion method based on automatic encoder framework
CN113921025A
Tone conversion method, electronic equipment and computer readable storage medium
CN115273890A