Voice conversion method and device, computer equipment and storage medium
By using content encoder and conditional stream matching decoder in the speech conversion model, extracting and separating speech content and sound characteristics, the problem of insufficient generalization capabilities in the prior art is solved, and better speech conversion effect and robustness are achieved.
Patent Information
- Application Number
- CN202510454951.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-24
AI Technical Summary
The existing voice conversion technology has insufficient generalization ability and robustness in zero-sample scenarios, resulting in poor voice conversion effect.
Using a speech conversion model containing a content encoder and a conditional stream matching decoder, the speech content features are extracted and fused through multiple self-supervised learning intermediate layers and adapters, and the speech sound features are separated through vector quantization layers to improve the feature separation quality.
It improves the voice conversion effect, enhances the generalization ability and robustness of the model in zero-sample scenarios, and ensures the quality and user experience of voice conversion.
Smart Images

Figure CN120199260A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence technology and medical and health fields, and particularly relates to a voice conversion method, device, computer device, and storage medium. Background Art
[0002] Voice conversion technology is a research branch of speech signal processing, covering areas such as speaker recognition, speech recognition, and speech synthesis. The conversion goal is to convert the speech of one speaker into the voice of another speaker, making the speech of the source speaker sound like that of the target speaker while preserving the speech content of the source speaker.
[0003] Voice conversion technology can be applied in the field of medical and health. For example, in intelligent medical consultations and remote consultations, if a user (patient or doctor) likes a certain type of voice, the user can convert the current voice (the speech of the patient or doctor) into the preferred voice and retain the speech content, thereby improving the fun and comfort of the consultation. Voice conversion technology can also be applied in the financial field. For example, if a user (customer or insurance agent) has a more preferred voice, they can also use voice conversion technology to convert the current voice (the speech of the customer or insurance agent) into the preferred voice, thereby improving the insurance marketing effect.
[0004] The core challenge of voice conversion lies in how to extract independent speech content features from the source speech and speech sound features from the target speech. Although existing methods have tried various techniques to achieve this separation, the generalization ability and robustness of the model in zero-shot scenarios are still a difficult problem, resulting in poor voice conversion effects. Summary of the Invention
[0005] Embodiments of this application provide a voice conversion method, device, computer device, and storage medium, which can improve the voice conversion effect.
[0006] In a first aspect, embodiments of this application provide a voice conversion method. The method is applied to a voice conversion system, and a voice conversion model is preset in the voice conversion system. The voice conversion model includes a content encoder and a conditional flow matching decoder. The method includes:
[0007] Obtain source speech and target speech;
[0008] Extract and fuse the speech content features of the source speech through multiple first self-supervised learning intermediate layers in the content encoder and a first adapter in the content encoder to obtain initial speech content features;
[0009] Performing speaker voice feature separation processing on the initial speech content features through a vector quantization layer in the content encoder to obtain target speech content features;
[0010] Obtaining the target speaker voice feature corresponding to the target speech;
[0011] Inputting the target speech content features and the target speaker voice features into the conditional flow matching decoder for speech conversion processing to obtain a mel spectrogram;
[0012] Generating the converted speech of the source speech according to the mel spectrogram.
[0013] In a second aspect, an embodiment of the present application further provides a speech conversion system. A speech conversion model is preset in the speech conversion system. The speech conversion model includes a content encoder and a conditional flow matching decoder. The speech conversion system includes a transceiver unit and a processing unit, wherein:
[0014] The transceiver unit is configured to obtain a source speech and a target speech;
[0015] The processing unit is configured to perform speech content feature extraction and fusion processing on the source speech through multiple first self-supervised learning intermediate layers in the content encoder and a first adapter in the content encoder to obtain initial speech content features; performing speaker voice feature separation processing on the initial speech content features through a vector quantization layer in the content encoder to obtain target speech content features; obtaining the target speaker voice feature corresponding to the target speech; inputting the target speech content features and the target speaker voice features into the conditional flow matching decoder for speech conversion processing to obtain a mel spectrogram; generating the converted speech of the source speech according to the mel spectrogram.
[0016] In a third aspect, an embodiment of the present application further provides a computer device, which includes a memory and a processor. A computer program is stored on the memory. When the processor executes the computer program, the above method is implemented.
[0017] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium. The storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the above method can be implemented.
[0018] Embodiments of the present application provide a voice conversion method, apparatus, computer device, and storage medium. Among them, the method is applied to a voice conversion system, and a voice conversion model is preset in the voice conversion system. The voice conversion model includes a content encoder and a conditional flow matching decoder. The method includes: obtaining a source voice and a target voice; extracting and fusing the speaking content features of the source voice through a plurality of first self-supervised learning intermediate layers in the content encoder and a first adapter in the content encoder to obtain initial speaking content features; separating the speaking voice features from the initial speaking content features through a vector quantization layer in the content encoder to obtain target speaking content features; obtaining target speaking voice features corresponding to the target voice; inputting the target speaking content features and the target speaking voice features into the conditional flow matching decoder for voice conversion processing to obtain a Mel spectrogram; generating a converted voice of the source voice according to the Mel spectrogram. The voice conversion model provided in the embodiments of the present application can automatically fuse the outputs of the self-supervised learning intermediate layers through an adapter, and further separate the speaking voice features in the speaking content features through a vector quantization layer to improve the separation quality of the speaking content features, thereby improving the voice conversion effect. Description of the Drawings
[0019] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a schematic diagram of the application scenario of the voice conversion method provided by the embodiment of the present application;
[0021] Figure 2 It is a schematic diagram of the structure of a voice conversion model provided by the embodiment of the present application;
[0022] Figure 3 It is a schematic flow chart of the voice conversion method provided by the embodiment of the present application;
[0023] Figure 4 It is a schematic sub-flow chart of the voice conversion method provided by the embodiment of the present application;
[0024] Figure 5 It is a schematic sub-flow chart of the voice conversion method provided by the embodiment of the present application;
[0025] Figure 6 It is a schematic block diagram of the voice conversion system provided by the embodiment of the present application;
[0026] Figure 7 Schematic block diagram of the computer device provided by the embodiment of the present application. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0028] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0029] It should also be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0030] It should be further understood that the term "and / or" used in this specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0031] The embodiments of the present application provide a voice conversion method, device, computer device, and storage medium.
[0032] The execution subject of the voice conversion method may be the voice conversion system provided by the embodiment of the present application, or a computer device integrated with the voice conversion system. Among them, the voice conversion system may be implemented in a hardware or software manner, the computer device may be a terminal or a server, and the terminal may be a smart phone, a tablet computer, a personal digital assistant, or a notebook computer, etc.
[0033] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the application scenario of the voice conversion method provided by the embodiment of the present application. The voice conversion method is applied to Figure 1In the computer device, a voice conversion system is deployed in the computer device. A voice conversion model is preset in the voice conversion system. The voice conversion model includes a content encoder and a conditional flow matching decoder. When performing voice conversion, the computer device acquires a source voice and a target voice; the source voice is subjected to speech content feature extraction and fusion processing through multiple first self-supervised learning intermediate layers and a first adapter in the content encoder to obtain an initial speech content feature; the initial speech content feature is subjected to speech voice feature separation processing through a vector quantization layer in the content encoder to obtain a target speech content feature; a target speech voice feature corresponding to the target voice is acquired; the target speech content feature and the target speech voice feature are input into the conditional flow matching decoder for voice conversion processing to obtain a Mel spectrogram; and a converted voice of the source voice is generated according to the Mel spectrogram.
[0034] Please refer to Figure 2 , in some embodiments, the voice conversion model provided by the embodiments of the present application includes a content encoder, a speaker encoder, and a conditional flow matching decoder. The content encoder includes multiple first self-supervised learning intermediate layers, a first adapter, and a vector quantization layer; the speaker encoder includes multiple second self-supervised learning intermediate layers and a second adapter; the conditional flow matching decoder includes multiple cross-attention modules and multiple decoder modules (for example, including 4 cross-attention modules and decoder modules, Figure 2 only one example is shown in
[0035] Specifically, in this embodiment, the first self-supervised learning intermediate layer and the second self-supervised learning intermediate layer may specifically be the Hidden-unitBERT (HuBERT) model in the Self-Supervised Learning (SSL) model. The conditional flow matching decoder may be designed as a Transformer-based U-Net architecture, and the self-attention layer in the Transformer is replaced by a cross-attention layer.
[0036] Figure 3 is a schematic flowchart of the voice conversion method provided by the embodiments of the present application. As Figure 3 shown, the method includes the following steps S110-S160.
[0037] S110. Acquire a source voice and a target voice.
[0038] In this embodiment, the source voice is the original voice input by the user, and the target voice is the voice corresponding to the voice characteristics that the user wants to become. For example, the source voice is the voice of the first user, and the target voice is the voice of the second user. At this time, the content of the first user's speech is retained in the finally obtained converted voice, and the voice characteristics of the speech are the voice characteristics of the second user. Among them, the voice characteristics include timbre characteristics, volume characteristics, pitch characteristics, etc.
[0039] For example, in the field of medical and health, a remote consultation application with a voice conversion function (that is, the voice conversion system provided in this application is deployed) is provided. The remote consultation application is divided into a patient-side application and a doctor-side application. In the patient-side application, a voice conversion function for the inquirer and a voice conversion function for the doctor are provided. Among them, the voice conversion function for the inquirer is used to convert the voice of the patient into the voice of another person, and the voice conversion function for the doctor is used to convert the voice of the doctor into the voice of another person; in the doctor-side application, a voice conversion function for the inquirer and a voice conversion function for the doctor are also provided. In addition, according to the actual application, the voice conversion function can also be only opened for the patient-side application or the doctor-side application, or only the voice conversion function for the inquirer or the doctor is opened in the patient-side application and / or the doctor-side application.
[0040] In the voice conversion function for the inquirer, the source voice is the patient voice input by the patient, and the target voice is the voice that the patient wants to become. In the voice conversion function for the doctor, the source voice is the doctor voice input by the doctor, and the target voice is the voice that the doctor wants to become.
[0041] Among them, the target voice can be the voice selected by the user (patient or doctor) who actively triggers the voice conversion function from a variety of characteristic voices (such as lolita voice, queen voice, uncle voice) preset in the voice conversion system, or the voice uploaded by the user.
[0042] For example, in a specific example in the field of healthcare, patient A uses a remote consultation application and enables the doctor voice conversion function. The patient selects Feature Voice 1 from multiple preset feature voices. At this time, the voice conversion system uses this Feature Voice 1 as the target voice and the voice of the doctor's reply as the source voice. For example, the patient inputs "What are the precautions after tooth extraction?" in text or voice in the remote consultation application. After the doctor sees / hears the information, the doctor replies in voice "For the first 24 hours after tooth extraction, mainly have warm and cool liquid food, and brushing teeth and rinsing the mouth are prohibited." Then, "For the first 24 hours after tooth extraction, mainly have warm and cool liquid food, and brushing teeth and rinsing the mouth are prohibited." is the source voice. The voice conversion system needs to convert the voice characteristics of this segment of voice into the voice characteristics corresponding to the target voice, retain the voice content, and play the converted voice. At this time, patient A hears the voice he wants to hear, thus improving the user experience. Among them, during the current remote consultation, the patient only needs to select the target voice before the consultation starts. In subsequent multiple Q&A sessions of this consultation, the system will automatically convert the voice of the doctor's reply into the voice corresponding to the target voice.
[0043] In addition, in the remote consultation application, in addition to being able to conduct doctor consultations by uploading voice or text, patients can also upload electronic personal health records including medical records, electrocardiograms, and / or medical images, etc. After the doctor views the personal health records uploaded by the patient, the doctor then replies to the patient by voice.
[0044] Among them, if the target voice is the voice uploaded by the user, then after the voice conversion system obtains this target voice, it also needs to perform a qualification check on this target voice. For example, check the voice length and noise situation of the voice. Only when the voice length exceeds the preset length and the noise detection value is lower than the preset noise threshold, it is determined that the target voice is qualified; otherwise, it is determined that the target voice is unqualified, and the customer is reminded to upload again.
[0045] S120: Extract and fuse the speech content features of the source speech through multiple first self-supervised learning intermediate layers in the content encoder and the first adapter in the content encoder to obtain initial speech content features.
[0046] In this embodiment, after the voice conversion system obtains the source voice and the target voice, it is necessary to extract the speech content features in the source voice and the speech sound features in the target voice. For example, in the field of healthcare, after the doctor inputs the source voice "For the first 24 hours after tooth extraction, mainly have warm and cool liquid food, and brushing teeth and rinsing the mouth are prohibited." of the reply into the remote consultation application, the remote consultation application inputs this source voice into the content encoder in the voice conversion model.
[0047] Specifically, in some embodiments, please refer to Figure 4 , step S120 includes:
[0048] S1201. Extract the speech content features of the source speech through multiple said first self-supervised learning intermediate layers to obtain multiple initial encoded speech content sub-features;
[0049] S1202. Perform weighted summation processing on the multiple initial encoded speech content sub-features through multiple learnable weights preset in the first adapter to obtain the initial speech content feature.
[0050] In this embodiment, when extracting the speech content features of the source speech, first input the source speech into the content encoder in the speech conversion model, then perform speech content feature extraction processing on the source speech through multiple first self-supervised learning intermediate layers in the content encoder, and output the initial encoded speech content sub-features corresponding to each first self-supervised learning intermediate layer respectively. Then, perform weighted summation processing on the initial encoded speech content sub-features output by each first self-supervised learning intermediate layer through multiple learnable weights preset in the first adapter to obtain the initial speech content feature.
[0051] S130. Perform speech voice feature separation processing on the initial speech content feature through the vector quantization layer in the content encoder to obtain the target speech content feature.
[0052] In this embodiment, since the training objective of non-parallel speech conversion is to reconstruct the original speech, the content encoder naturally tends to generate features that contain both speech content and speech voice features. To further guide the model to separate the speech voice features, this application applies a vector quantization (VQ) layer after the first adapter. Quantifying the latent features can produce a discrete and compact representation. From the perspective of speech coding, the output of the adapter is guided to map the similar content information from different speakers to the closest embeddings, and finally generate accurate language information (speech content features) independent of the speaker (speech voice features).
[0053] Specifically, in some embodiments, please refer to Figure 5 , step S130 includes:
[0054] S1301. Map the initial speech content feature to different codebook vectors in the vector quantization layer according to the vector similarity;
[0055] S1302. Determine the target speech content feature according to the token corresponding to the mapped codebook vector.
[0056] In this embodiment, the initial speech content feature is a sequence of speech content feature vectors. For each feature vector in the sequence of speech content feature vectors, calculate its similarity or distance from all codebook vectors in the codebook. Then select the codebook vector with the highest similarity (or the smallest distance), record its corresponding token (index), and use these codebook vectors to replace the original features to obtain the target speech content feature; alternatively, regenerate the target speech content feature according to the token sequence, for example, generate the target speech content feature by looking up the codebook.
[0057] It can be seen that the goal of the content encoder provided by the embodiments of this application is to extract the speech content features in the speech while minimizing the influence of speaker-specific attributes (speech sound features).
[0058] S140. Obtain the target speech sound feature corresponding to the target speech.
[0059] In some embodiments, when the target speech is a preset characteristic speech in the system, its corresponding speech sound feature can be preset in advance to improve the speech conversion efficiency. At this time, step S140 includes: obtaining the speech identifier corresponding to the target speech; determining the candidate speech sound feature corresponding to the speech identifier among the preset multiple candidate speech sound features as the target speech sound feature. For example, the lolita voice corresponds to candidate speech sound feature 1, the mature female voice corresponds to candidate speech sound feature 2, and the uncle voice corresponds to candidate speech sound feature 3.
[0060] In other embodiments, when the target speech is a newly input speech by the user or a speech preset in the system but for which the speech sound feature has not been extracted, at this time, it is necessary to perform speech sound feature extraction processing on the target speech through multiple second self-supervised learning intermediate layers to obtain multiple initial encoded speech sound sub-features; then perform weighted summation processing on the multiple initial encoded speech sound sub-features through multiple learnable weights preset in the second adapter to obtain the target speech sound feature.
[0061] For example, when a patient uses the doctor voice conversion function in a remote consultation application and cannot find the voice they want to hear, they can upload a piece of speech by themselves and use the uploaded speech as the target speech. When the remote consultation application obtains the target speech, it recognizes that the target speech is a newly uploaded speech and extracts the speech sound feature of this speech as the voice feature to be converted when the doctor answers.
[0062] S150. Input the target speech content feature and the target speech sound feature into the conditional flow matching decoder for speech conversion processing to obtain a mel spectrogram.
[0063] In this embodiment, after obtaining the target speech content feature and the target speech sound feature, the target speech content feature and the target speech sound feature are input into a conditional flow matching decoder for speech conversion processing to obtain a Mel spectrogram.
[0064] For example, in the field of medical and health, patient A uses a remote consultation application and uses the doctor speech conversion function. At this time, the speech conversion system in the remote consultation application obtains the target speech content feature according to the doctor's response speech (source speech), and obtains the target speech sound feature according to the target speech specified by patient A. Then, the target speech content feature and the target speech sound feature are input into a conditional flow matching decoder for speech conversion processing to obtain a Mel spectrogram.
[0065] Among them, in order to improve the generation efficiency and speech quality, the conditional flow matching decoder in this embodiment adopts an optimal transport conditional flow matching objective, and matches the mapping between data and the target distribution through a regression transformation. The optimal transport conditional flow matching provides higher efficiency and robustness than the diffusion mechanism, which simulates the random transformation of data. Specifically, the conditional flow matching decoder provided in this embodiment is designed as a Transformer-based U-Net architecture, and provides speaker conditions in multiple ways. The self-attention layer in the Transformer block is replaced by a cross-attention layer, where the encoded speaker features are used as keys and values. By providing multiple conditions through cross-attention, the conditional flow matching decoder can faithfully simulate the acoustic details of various speakers. Moreover, the first adapter in the speaker encoder in this embodiment optimizes the generation of rich speaker information by combining multiple outputs of HuBERT (i.e., multiple first self-supervised learning intermediate layers). Thereby improving the speech conversion quality.
[0066] S160. Generate the converted speech of the source speech according to the Mel spectrogram.
[0067] In this embodiment, after obtaining the Mel spectrogram through the speech conversion model in the speech conversion system, the converted speech corresponding to the Mel spectrogram is generated through a preset neural vocoder.
[0068] For example, in the field of medical and health, patient A uses a remote consultation application and uses the doctor speech conversion function. The doctor's response heard by patient A is the converted speech of the doctor.
[0069] In addition, the present application can also be applied in the financial field. For example, through big data analysis, it is found that a certain customer prefers a certain type of voice. At this time, when an insurance agent (which can be a robot) conducts marketing or return visits to the customer, etc., the voice conversion system provided by the present application can convert the voice into the voice liked by the customer, so as to win more favor from the customer.
[0070] The voice conversion model provided in this embodiment is trained based on a target loss function, and the target loss function includes a commitment loss function of the content encoder, a prior loss function, and an optimal transport conditional flow matching loss function of the conditional flow matching decoder, where:
[0071] The commitment loss function aims to make the input of the VQ layer consistent with the codebook vectors, and the commitment loss function is:
[0072]
[0073] where, is the name of the commitment loss function, E x~p(x) is the data expectation, x is the source speech input by the content encoder, p(x) is the probability distribution of x in the feature space, [E(x)] is the output of the content encoder, e is the codebook vector, and sg[·] represents the stop gradient operation;
[0074] The prior loss function is used to ensure the uniform use of the codebook vectors, avoiding some codebook vectors being overused or underused, and the prior loss function is:
[0075]
[0076] where, is the name of the prior loss function;
[0077] The optimal transport conditional flow matching loss function is used to optimize the conditional flow matching decoder so that the mapping between the generated Mel spectrogram and the target Mel spectrogram is as matched as possible. This loss function is based on the optimal transport theory and aims to minimize the transportation cost between the source distribution and the target distribution. At this time, when the optimal transport conditional flow matching loss function is the first optimal transport conditional flow matching loss function, the first optimal transport conditional flow matching loss function is:
[0078]
[0079] where, is the name of the optimal transport conditional flow matching loss function, E (x,y)~π is the data expectation, π is the joint distribution between x and y, Φ is the set of all transport mappings, c(·,·) is the cost function, φ is the mapping function from x to y, φ(x) is the predicted converted speech, and y is the true converted speech;
[0080] The loss function method of the conditional flow matching decoder utilizes optimal transport to estimate a vector field with a linear trajectory. Specifically, the conditional flow matching objective aims to optimize the mel spectrogram generated by the decoder by minimizing the transportation cost between the source data distribution and the target data distribution. The core idea is to find an optimal mapping that can effectively transform the source data into the target data while preserving the data structure. At this time, when the optimal transport conditional flow matching loss function is the second optimal transport conditional flow matching loss function, the second optimal transport conditional flow matching loss function is:
[0081] The second optimal transport conditional flow matching loss function is:
[0082]
[0083] where, represents the square of the Euclidean distance;
[0084] The target loss function is:
[0085]
[0086] where, is the name of the target loss function, is the name of the optimal transport conditional flow matching loss function, and λ1, λ2, and λ3 are weight hyperparameters that balance the various loss functions.
[0087] In summary, the voice conversion method provided in this application is applied to a voice conversion system. A voice conversion model is preset in the voice conversion system. The voice conversion model includes a content encoder and a conditional flow matching decoder. The method includes: obtaining a source voice and a target voice; extracting and fusing the speaking content features of the source voice through multiple first self-supervised learning intermediate layers and a first adapter in the content encoder to obtain initial speaking content features; performing speaking voice feature separation processing on the initial speaking content features through a vector quantization layer in the content encoder to obtain target speaking content features; obtaining target speaking voice features corresponding to the target voice; inputting the target speaking content features and the target speaking voice features into the conditional flow matching decoder for voice conversion processing to obtain a mel spectrogram; generating a converted voice of the source voice according to the mel spectrogram. The voice conversion model provided in the embodiments of this application can automatically fuse the outputs of the self-supervised learning intermediate layers through the adapter, and further perform separation processing on the speaking voice features in the speaking content features through the vector quantization layer to improve the separation quality of the speaking content features, thereby improving the voice conversion effect.
[0088] Figure 6It is a schematic block diagram of a voice conversion system 600 provided by an embodiment of the present application. As Figure 6 shown, corresponding to the above voice conversion method, the present application also provides a voice conversion system 600. The voice conversion system 600 includes units for performing the above voice conversion method, and the voice conversion system 600 can be configured in terminals such as desktop computers, tablet computers, laptop computers, etc. Specifically, a voice conversion model is preset in the voice conversion system 600, and the voice conversion model includes a content encoder and a conditional flow matching decoder. Please refer to Figure 6 , the voice conversion system 600 includes a transceiver unit 601 and a processing unit 602, where:
[0089] The transceiver unit 601 is used to obtain a source voice and a target voice;
[0090] Among them, the source voice is the original voice input by the user, and the target voice is the voice corresponding to the speaking voice characteristics that the user wants to become. For example, the source voice is the voice of the first user, and the target voice is the voice of the second user. At this time, the final converted voice retains the speaking content of the first user, and the speaking voice characteristics of the voice are the speaking voice characteristics of the second user. Among them, the speaking voice characteristics include timbre characteristics, volume characteristics, pitch characteristics, etc.
[0091] In the voice conversion function of the inquirer, the source voice is the patient voice input by the patient, and the target voice is the voice that the patient wants to become. In the voice conversion function of the doctor, the source voice is the doctor voice input by the doctor, and the target voice is the voice that the doctor wants to become.
[0092] Among them, the target voice can be a voice selected by the user (patient or doctor) who actively triggers the voice conversion function from multiple featured voices (such as lolita voice, mature female voice, uncle voice) preset in the voice conversion system 600, or a voice uploaded by the user.
[0093] For example, in a specific example in the field of healthcare, patient A uses a remote consultation application and uses the doctor voice conversion function to select feature voice 1 from multiple preset feature voices. At this time, the voice conversion system 600 uses this feature voice 1 as the target voice and the doctor's input response voice as the source voice. For example, the patient inputs "What are the precautions after tooth extraction?" in text or voice in the remote consultation application. After the doctor sees / hears the information, the doctor's voice response is "Mainly have warm and cool liquid food within 24 hours after tooth extraction, and brushing teeth and rinsing the mouth are prohibited." Then, "Mainly have warm and cool liquid food within 24 hours after tooth extraction, and brushing teeth and rinsing the mouth are prohibited." is the source voice. The voice conversion system 600 needs to convert the voice characteristics of this voice into the voice characteristics corresponding to the target voice, retain the voice content, and play the converted voice. At this time, patient A hears the voice he wants to hear, thus improving the user experience. Among them, during the current remote consultation, the patient only needs to select the target voice before the consultation starts. In the subsequent multiple Q&A sessions of this consultation, the system will automatically convert the doctor's response voice into the voice corresponding to the target voice.
[0094] Among them, if the target voice is a voice uploaded by the user, then after the voice conversion system 600 obtains this target voice, it also needs to perform a qualification check on this target voice. For example, check the voice length and noise situation of the voice. Only when the voice length exceeds the preset length and the noise detection value is lower than the preset noise threshold, it is determined that the target voice is qualified; otherwise, it is determined that the target voice is unqualified, and the customer is reminded to upload again.
[0095] The processing unit 602 is configured to extract and fuse the speaking content features of the source voice through multiple first self-supervised learning intermediate layers in the content encoder and the first adapter in the content encoder to obtain initial speaking content features; perform speaking voice feature separation processing on the initial speaking content features through the vector quantization layer in the content encoder to obtain target speaking content features; obtain the target speaking voice features corresponding to the target voice; input the target speaking content features and the target speaking voice features into the conditional flow matching decoder for voice conversion processing to obtain a mel spectrogram; generate the converted voice of the source voice according to the mel spectrogram.
[0096] Among them, after the voice conversion system 600 obtains the source voice and the target voice, it is necessary to extract the speaking content features in the source voice and the speaking voice features in the target voice. For example, in the field of healthcare, after the doctor inputs the source voice "Mainly have warm and cool liquid food within 24 hours after tooth extraction, and brushing teeth and rinsing the mouth are prohibited." of the response into the remote consultation application, the remote consultation application inputs this source voice into the content encoder in the voice conversion model.
[0097] Since the training objective of non - parallel voice conversion is to reconstruct the original voice, the content encoder naturally tends to generate features that contain both the speaking content and the speaking voice features. To further guide the model to separate the speaking voice features, this application applies a VQ layer after the first adapter. Quantifying the latent features can produce discrete and compact representations.
[0098] Among them, when the target voice is a preset characteristic voice in the system, its corresponding speaking voice features can be preset in advance, thereby improving the voice conversion efficiency; when the target voice is a newly input voice by the user or a voice preset in the system but not subjected to speaking voice feature extraction, at this time, the target speaking voice features of this target voice need to be extracted.
[0099] After obtaining the target speaking content features and the target speaking voice features, the target speaking content features and the target speaking voice features are input into the conditional flow matching decoder for voice conversion processing to obtain a mel - spectrogram.
[0100] For example, in the field of medical and health, patient A uses a remote consultation application and uses the doctor voice conversion function. At this time, the voice conversion system 600 in the remote consultation application obtains the target speaking content features according to the doctor's reply voice (source voice), and obtains the target speaking voice features according to the target voice specified by patient A. Then, the target speaking content features and the target speaking voice features are input into the conditional flow matching decoder for voice conversion processing to obtain a mel - spectrogram.
[0101] In some embodiments, when the processing unit 602 executes the step of performing speaking content feature extraction and fusion processing on the source voice through multiple first self - supervised learning intermediate layers in the content encoder and the first adapter in the content encoder to obtain initial speaking content features, it is specifically configured to:
[0102] Perform speaking content feature extraction processing on the source voice through multiple first self - supervised learning intermediate layers to obtain multiple initial encoded speaking content sub - features; perform weighted summation processing on the multiple initial encoded speaking content sub - features through multiple learnable weights preset in the first adapter to obtain the initial speaking content features.
[0103] In some embodiments, when the processing unit 602 executes the step of performing speaking voice feature separation processing on the initial speaking content features through the vector quantization layer in the content encoder to obtain target speaking content features, it is specifically configured to:
[0104] Map the initial speaking content features to different codebook vectors in the vector quantization layer according to vector similarity; determine the target speaking content features according to the tokens corresponding to the mapped codebook vectors.
[0105] In some embodiments, the voice conversion system 600 further includes a speaker encoder, and the speaker encoder includes a plurality of second self-supervised learning intermediate layers and a second adapter; when the processing unit 602 executes the step of obtaining the target speaker voice feature corresponding to the target voice, it specifically is used for:
[0106] Performing speaker voice feature extraction processing on the target voice through the plurality of second self-supervised learning intermediate layers to obtain a plurality of initial encoded speaker voice sub-features; performing weighted summation processing on the plurality of initial encoded speaker voice sub-features through a plurality of learnable weights preset in the second adapter to obtain the target speaker voice feature.
[0107] In some embodiments, when the processing unit 602 executes the step of obtaining the target speaker voice feature corresponding to the target voice, it specifically is used for:
[0108] Obtaining the voice identifier corresponding to the target voice;
[0109] Determining the candidate speaker voice feature corresponding to the voice identifier among the plurality of preset candidate speaker voice features as the target speaker voice feature.
[0110] In some embodiments, the voice conversion model is trained based on a target loss function, and the target loss function includes a commitment loss function of the content encoder, a prior loss function, and an optimal transport conditional flow matching loss function of the conditional flow matching decoder, where:
[0111] The commitment loss function is:
[0112]
[0113] Wherein, is the name of the commitment loss function, E x~p(x) is the data expectation, x is the source voice input by the content encoder, p(x) is the probability distribution of x in the feature space, [E(x)] is the output of the content encoder, e is the codebook vector, and sg[·] represents the stop gradient operation;
[0114] The prior loss function is:
[0115]
[0116] Wherein, is the name of the prior loss function;
[0117] The optimal transport conditional flow matching loss function is the first optimal transport conditional flow matching loss function or the second optimal transport conditional flow matching loss function;
[0118] The target loss function is as follows:
[0119]
[0120] where, is the name of the target loss function, is the name of the optimal transport conditional flow matching loss function, and λ1, λ2, and λ3 are weight hyperparameters for balancing each loss function.
[0121] In some embodiments, the first optimal transport conditional flow matching loss function is:
[0122]
[0123] where, is the name of the optimal transport conditional flow matching loss function, E (x,y)~π is the data expectation, π is the joint distribution between x and y, Φ is the set of all transport maps, c(·, ·) is the cost function, φ is the mapping function from x to y, φ(x) is the predicted converted speech, and y is the true converted speech;
[0124] The second optimal transport conditional flow matching loss function is:
[0125]
[0126] where, represents the square of the Euclidean distance.
[0127] In summary, the speech conversion model provided in the speech conversion system 600 of the present application can automatically fuse the output of the self-supervised learning intermediate layer through the adapter, and further separate the speech voice features in the speech content features through the vector quantization layer to improve the separation quality of the speech content features, thereby improving the speech conversion effect.
[0128] It should be noted that those skilled in the art can clearly understand that the specific implementation processes of the above speech conversion system 600 and each unit can refer to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and brevity of description, they are not elaborated herein.
[0129] The above speech conversion system can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 7 shown.
[0130] Please refer to Figure 7 , Figure 7It is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 700 can be a terminal or a server. A voice conversion system is deployed in the computer device. The voice conversion model includes a content encoder and a conditional flow matching decoder. Among them, the terminal can be an electronic device with communication functions such as a smart phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device. The server can be an independent server or a server cluster composed of multiple servers.
[0131] Refer to Figure 7 , the computer device 700 includes a processor 702, a memory, and a network interface 705 connected through a system bus 701. Among them, the memory can include a non-volatile storage medium 703 and an internal memory 704.
[0132] The non-volatile storage medium 703 can store an operating system 7031 and a computer program 7032. The computer program 7032 includes program instructions. When the program instructions are executed, the processor 702 can be made to execute a voice conversion method.
[0133] The processor 702 is used to provide computing and control capabilities to support the operation of the entire computer device 700.
[0134] The internal memory 704 provides an environment for the operation of the computer program 7032 in the non-volatile storage medium 703. When the computer program 7032 is executed by the processor 702, the processor 702 can be made to execute a voice conversion method.
[0135] The network interface 705 is used for network communication with other devices. Those skilled in the art can understand that Figure 7 the structure shown in
[0136] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device 700 to which the solution of the present application is applied. The specific computer device 700 may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.
[0136] Among them, the processor 702 is used to run the computer program 7032 stored in the memory to implement the following steps:
[0137] Obtain a source voice and a target voice;
[0138] Extract and fuse the speech content features of the source voice through multiple first self-supervised learning intermediate layers in the content encoder and a first adapter in the content encoder to obtain initial speech content features;
[0139] Perform speaker voice feature separation processing on the initial speaker content features through the vector quantization layer in the content encoder to obtain target speaker content features;
[0140] Obtain the target speaker voice feature corresponding to the target speech;
[0141] Input the target speaker content features and the target speaker voice features into the conditional flow matching decoder for speech conversion processing to obtain a Mel spectrogram;
[0142] Generate the converted speech of the source speech according to the Mel spectrogram.
[0143] It should be understood that in the embodiments of the present application, the processor 702 may be a central processing unit (CPU), and this processor 702 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0144] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program includes program instructions, and the computer program can be stored in a storage medium, and this storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0145] Therefore, the present application also provides a storage medium. This storage medium may be a computer-readable storage medium. The storage medium stores a computer program, where the computer program includes program instructions. When the program instructions are executed by the processor, the processor is caused to execute the following steps:
[0146] Obtain the source speech and the target speech;
[0147] Perform speaker content feature extraction and fusion processing on the source speech through multiple first self-supervised learning intermediate layers in the content encoder and the first adapter in the content encoder to obtain initial speaker content features;
[0148] The initial speech content features are subjected to speech voice feature separation processing through a vector quantization layer in the content encoder to obtain target speech content features;
[0149] Obtain the target speech voice features corresponding to the target speech;
[0150] Input the target speech content features and the target speech voice features into the conditional flow matching decoder for speech conversion processing to obtain a mel spectrogram;
[0151] Generate the converted speech of the source speech according to the mel spectrogram.
[0152] The storage medium may be various computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes.
[0153] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0154] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0155] The steps in the method embodiments of this application can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of this application can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0156] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application.
[0157] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A speech conversion method, characterized in that: The method is applied to a speech conversion system, wherein a speech conversion model is preset in the speech conversion system, and the speech conversion model includes a content encoder and a conditional stream matching decoder. The method includes: Obtain source speech and target speech; Extracting and fusing the speech content features of the source speech through the multiple first self-supervised learning intermediate layers in the content encoder and the first adapter in the content encoder to obtain initial speech content features; Perform speech sound feature separation processing on the initial speech content feature through a vector quantization layer in the content encoder to obtain a target speech content feature; Acquire target speech sound features corresponding to the target speech; Inputting the target speech content feature and the target speech sound feature into the conditional stream matching decoder for speech conversion processing to obtain a Mel-spectrogram; Generate converted speech of the source speech according to the mel-spectrogram.
2. The method according to claim 1, characterized in that: The extracting and fusing the speech content features of the source speech through the multiple first self-supervised learning intermediate layers in the content encoder and the first adapter in the content encoder to obtain the initial speech content features includes: Perform speech content feature extraction processing on the source speech through the plurality of the first self-supervised learning intermediate layers to obtain a plurality of initial encoded speech content sub-features; The initial speech content feature is obtained by performing weighted summation processing on the multiple initial encoded speech content sub-features through the multiple learnable weights preset in the first adapter.
3. The method according to claim 1, characterized in that The step of performing speech sound feature separation processing on the initial speech content feature through a vector quantization layer in the content encoder to obtain a target speech content feature includes: Mapping the initial speech content features to different codebook vectors in the vector quantization layer according to vector similarity; The target speech content feature is determined according to the token corresponding to the mapped codebook vector.
4. The method according to claim 1, characterized in that The speech conversion system further includes a speaker encoder, wherein the speaker encoder includes a plurality of second self-supervised learning intermediate layers and a second adapter; the step of obtaining the target speech sound feature corresponding to the target speech includes: Performing speech sound feature extraction processing on the target speech through the plurality of the second self-supervised learning intermediate layers to obtain a plurality of initial encoded speech sound sub-features; The target speaking voice feature is obtained by performing weighted summation processing on the multiple initially encoded speaking voice sub-features through the multiple learnable weights preset in the second adapter.
5. The method according to claim 1, characterized in that The step of obtaining a target speech sound feature corresponding to the target speech includes: Obtaining a voice identifier corresponding to the target voice; A candidate speaking voice feature corresponding to the voice identifier among a plurality of preset candidate speaking voice features is determined as the target speaking voice feature.
6. The method according to any one of claims 1 to 5, characterized in that The speech conversion model is trained based on a target loss function, wherein the target loss function includes a commitment loss function of the content encoder, a priori loss function, and an optimal transmission conditional stream matching loss function of the conditional stream matching decoder, wherein: The commitment loss function is: in, is the name of the commitment loss function, E x~p(x) is the data expectation, x is the source speech input by the content encoder, p(x) is the probability distribution of x in the feature space, [E(x)] is the output of the content encoder, e is the codebook vector, and sg[·] indicates stopping the gradient operation; The prior loss function is: in, is the name of the prior loss function; The optimal transmission condition flow matching loss function is a first optimal transmission condition flow matching loss function or a second optimal transmission condition flow matching loss function; The objective loss function is: in, is the name of the objective loss function, is the name of the optimal transmission condition flow matching loss function, and λ1, λ2 and λ3 are weight hyperparameters for balancing various loss functions.
7. The method according to claim 6, characterized in that The first optimal transmission condition flow matching loss function is: in, is the name of the optimal transmission condition flow matching loss function, E (x,y)~π is the data expectation, π is the joint distribution between x and y, Φ is the set of all transfer mappings, c(·, ·) is the cost function, φ is the mapping function from x to y, φ(x) is the predicted converted speech, and y is the actual converted speech; The second optimal transmission condition flow matching loss function is: in, Represents the square of the Euclidean distance.
8. A speech conversion system, characterized in that: The speech conversion system is preset with a speech conversion model, the speech conversion model includes a content encoder and a conditional stream matching decoder, and the speech conversion system includes a transceiver unit and a processing unit, wherein: The transceiver unit is used to obtain the source voice and the target voice; The processing unit is used to extract and fuse the speech content features of the source speech through the multiple first self-supervised learning intermediate layers in the content encoder and the first adapter in the content encoder to obtain initial speech content features; perform speech sound feature separation processing on the initial speech content features through the vector quantization layer in the content encoder to obtain target speech content features; obtain target speech sound features corresponding to the target speech; input the target speech content features and the target speech sound features into the conditional stream matching decoder for speech conversion processing to obtain a Mel-spectrogram; and generate the converted speech of the source speech according to the Mel-spectrogram.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: A speech conversion system is deployed in the computer device, the speech conversion model includes a content encoder and a conditional stream matching decoder, and the processor implements the speech conversion method according to any one of claims 1 to 7 when executing the computer program.
10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the speech conversion method according to any one of claims 1 to 7.