Voice conversion method and device based on transposition double attention, equipment and medium

By using technical means such as transposed dual attention mechanism and HuBERT model in speech conversion, the problem of accurate conversion of speech styles in the existing technology is solved, and a more natural and high-quality speech conversion effect is achieved.

CN120108375APending Publication Date: 2025-06-06PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510340215.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to accurately convert source speech into a specific style of the target speaker, and cannot reproduce the speaking style of the target speaker.

Method used

Using a transpose dual attention-based speech conversion method, combining HuBERT model, content encoder and score-based diffusion model, the initial model is constructed and trained through the target loss function to generate a style library containing multiple style embeddings to meet complex conversion needs.

Benefits of technology

The naturalness and conversion quality of speech conversion are improved, so that the converted speech is more natural and accurate in style to meet the characteristics of the target speaker.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108375A_ABST
    Figure CN120108375A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, finance and medical health, and provides a voice conversion method and device based on transposition double attention, equipment and a medium, on one hand, a transposition double attention mechanism is adopted, multiple voice style representations can be extracted in a content-related mode, and text transcription or speaker labels are not needed; on one hand, self-supervised learning features are extracted based on a HuBERT model and a content encoder to obtain content information, and the self-supervised learning features are discretized into content embedding by using a vector quantization method, so that efficient representation of data can be realized on the premise of not losing key semantics, and subsequent processing and calculation are facilitated; on the other hand, the voice of the target speaker is used for generating a style library containing multiple style embedding, each style embedding corresponds to pronunciation of different voice contents, more complex conversion requirements can be met, and the naturalness and conversion quality of voice conversion are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence, finance and medical health technology, and in particular to a speech conversion method, device, equipment and medium based on transposed dual attention. Background Art

[0002] Speech conversion refers to the task of converting the speech of one person (source) to the speech of another person (target) while preserving the linguistic content (e.g., phonemes, words) expressed by the source speaker.

[0003] Specifically, it is expected that the content of the converted speech is given by the content of the source speech, while the speaking style (e.g., speaker identity, accent) is similar to that of the target speaker. The goal of any-to-any speech conversion is to process the speech of any unseen speaker without any prior knowledge about the speaker.

[0004] Many existing any-to-any speech conversion models successfully transfer some of the target speech's style information to the converted speech, but still lack the ability to reproduce the target speaker's speaking style. Most works still rely on a single global vector (usually speaker embeddings extracted from a pre-trained speaker verification model) to condition all frames, and thus fail to faithfully transfer the target speaker's style to the source speech content. Summary of the invention

[0005] In view of the above, it is necessary to provide a speech conversion method, device, equipment and medium based on transposed dual attention, aiming to solve the problem that the source speech cannot be accurately converted into a certain style.

[0006] A speech conversion method based on transposition dual attention, the speech conversion method based on transposition dual attention comprising:

[0007] Building an initial model based on the transposed dual attention mechanism, the HuBERT model, the content encoder, and the score-based diffusion model, and building a target loss function for the initial model;

[0008] Collect speech data of different speaking styles to build a target speech library, collect multiple source speech segments, and randomly initialize the learnable query set;

[0009] Based on the target loss function, the initial model is trained according to the target speech library, the multiple source speech segments, and the learnable query set to obtain a speech conversion model;

[0010] In response to a speech conversion instruction for a target source speech, the target source speech is processed using the speech conversion model to obtain a target converted speech;

[0011] The target converted voice is played by using a designated voice playing device.

[0012] A speech conversion device based on transposition dual attention, the speech conversion device based on transposition dual attention comprising:

[0013] A construction unit, used to construct an initial model based on a transposed dual attention mechanism, a HuBERT model, a content encoder, and a score-based diffusion model, and to construct a target loss function of the initial model;

[0014] The construction unit is further used to collect speech data of different speaking styles to construct a target speech library, collect multiple source speech segments, and randomly initialize a learnable query set;

[0015] A training unit, configured to train the initial model based on the target loss function, the target speech library, the multiple source speech segments, and the learnable query set to obtain a speech conversion model;

[0016] A processing unit, configured to, in response to a speech conversion instruction for a target source speech, process the target source speech using the speech conversion model to obtain a target converted speech;

[0017] The playing unit is used to play the target converted voice by using a designated voice playing device.

[0018] A computer device, comprising:

[0019] a memory storing at least one instruction; and

[0020] A processor executes the instructions stored in the memory to implement the speech conversion method based on transposed dual attention.

[0021] A computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the speech conversion method based on transposed dual attention.

[0022] It can be seen from the above technical solutions that, on the one hand, the present invention adopts a transposed dual attention mechanism, which can extract multiple speech style representations in a content-related manner without text transcription or speaker labels; on the one hand, the present invention extracts self-supervised learning features based on the HuBERT model and content encoder to obtain content information, and uses a vector quantization method to discretize the self-supervised learning features into content embeddings, which can achieve efficient data representation without losing key semantics, and facilitate subsequent processing and calculation; on the other hand, the speech of the target speaker is used to generate a style library containing multiple style embeddings, each style embedding corresponds to the pronunciation of different speech content, can adapt to more complex conversion requirements, and improve the naturalness and conversion quality of speech conversion. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a flow chart of a preferred embodiment of the speech conversion method based on transposed dual attention of the present invention.

[0024] Figure 2 It is the overall architecture diagram of voice conversion performed by the present invention.

[0025] Figure 3 It is a functional module diagram of a preferred embodiment of the speech conversion device based on transposed dual attention of the present invention.

[0026] Figure 4 It is a structural schematic diagram of a computer device of a preferred embodiment of the present invention for implementing a speech conversion method based on transposed dual attention. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0028] like Figure 1 As shown, it is a flow chart of a preferred embodiment of the speech conversion method based on transposition dual attention of the present invention. According to different requirements, the order of the steps in the flow chart can be changed, and some steps can be omitted.

[0029] The speech conversion method based on transposed dual attention is applied to one or more computer devices, wherein the computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASIC), programmable gate arrays (FPGA), digital signal processors (DSP), embedded devices, etc.

[0030] The computer device may be any electronic product that can perform human-computer interaction with a user, such as a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an interactive network television (IPTV), a smart wearable device, etc.

[0031] The computer device may also include a network device and / or a user device, wherein the network device includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud consisting of a large number of hosts or network servers based on cloud computing.

[0032] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), as well as big data and artificial intelligence platforms.

[0033] Among them, Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0034] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0035] The network where the computer device is located includes but is not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.

[0036] S10, constructing an initial model based on the transposed dual attention mechanism, the HuBERT model, the content encoder, and the score-based diffusion model, and constructing a target loss function of the initial model.

[0037] In this embodiment, the transposed dual attention mechanism includes a first multi-head attention layer and a second multi-head attention layer that is in a transposed relationship with the first multi-head attention layer.

[0038] In this embodiment, the HuBERT model is a model that performs self-supervised speech representation learning through hidden unit mask prediction. The core idea is to use an offline clustering step to generate noisy labels and apply prediction loss only to the masked area, forcing the model to learn a high-level representation of the unmasked input, thereby inferring the target of the masked input. In this way, HuBERT significantly improves the quality of speech representation and has advanced performance in multiple speech recognition tasks.

[0039] In this embodiment, the content encoder is a device that compiles and converts signals or data into a signal form that can be used for communication, transmission and storage. The content encoder is mainly used to compress and encode the original data to reduce the amount of data and improve the transmission efficiency. The content encoder plays an important role in the field of digital signal processing and multimedia, and is usually used in conjunction with a decoder to achieve data compression, transmission and recovery.

[0040] In this embodiment, the backbone of the score-based diffusion model adopts a U-Net (U-shaped Network) structure for speech synthesis.

[0041] In this embodiment, the objective loss function of constructing the initial model includes:

[0042] Construct diffusion model loss and content encoder loss;

[0043] The sum of the diffusion model loss and the content encoder loss is calculated to obtain the target loss function.

[0044] Through the above embodiments, the diffusion model loss and the content encoder loss can be taken into account during the training process to improve the accuracy of the model.

[0045] S11, collect speech data of different speaking styles to build a target speech library, collect multiple source speech segments, and randomly initialize the learnable query set.

[0046] In this embodiment, the target speech database contains a speech data set with a specific speaking style, which is used to provide a target style reference for speech conversion.

[0047] For example, the target voice library can store the content of a designated host, the voice of a designated artist, etc., to enrich the voice style. Specifically, if a Mandarin voice is to be converted into a voice with the style of a well-known host, multiple audio clips of the host hosting the program can be collected to form the target voice library. These audio clips will be used to extract style features in the voice conversion system and guide the source voice to transform into the host style.

[0048] In this embodiment, the multiple source speech segments may be speech to be converted.

[0049] For example, the source speech may be a normal person's Mandarin speech, such as a recording of Zhang San reading a news report. The speech will retain the content information in the speech conversion system while changing the style so that it has the characteristics of the speech in the target speech library (such as the host style).

[0050] During the training process, the content of the source speech is taken as the final output, and the model is trained to faithfully reconstruct the input speech within the autoencoder framework.

[0051] In this embodiment, the learnable query set is a learnable embedding set randomly initialized during the model training process, which can be used as a query vector in the attention mechanism to explore the relationship between different features to assist in style feature extraction and fusion.

[0052] S12, based on the target loss function, the initial model is trained according to the target speech library, the multiple source speech segments, and the learnable query set to obtain a speech conversion model.

[0053] In this embodiment, when the value of the target loss function converges, the training is stopped and the speech conversion model is obtained.

[0054] S13, in response to a speech conversion instruction for a target source speech, the target source speech is processed using the speech conversion model to obtain a target converted speech.

[0055] In this embodiment, the process of processing the target source speech by using the speech conversion model to obtain the target converted speech includes:

[0056] Processing the target speech library using the HuBERT model and the content encoder to obtain content embedding;

[0057] Performing style modeling according to the transposed dual attention mechanism, the learnable query set, and the content embedding to obtain a style library;

[0058] The target source speech is processed according to the style library and the diffusion model to obtain the target converted speech.

[0059] In this embodiment, the process of processing the target speech library by using the HuBERT model and the content encoder to obtain content embedding includes:

[0060] Using a first HuBERT model to extract features from the speech data in the target speech library to obtain a first self-supervised learning feature;

[0061] Performing vector quantization processing on the first self-supervised learning feature to obtain a first vector quantization feature;

[0062] The first vector quantization feature is content encoded by using a first content encoder to obtain the content embedding.

[0063] In the above embodiment, self-supervised learning features (SSL) can be extracted as content information based on the HuBERT model, and the SSL features can be discretized into content embedding using a quantization method.

[0064] SSL features are usually continuous high-dimensional vectors that contain a lot of detailed information, but there may be redundancy. Therefore, the content embedding obtained after discretization can represent the key information of the speech content in a more concise way, remove unnecessary details, and achieve efficient data representation without losing key semantics, which is convenient for subsequent processing and calculation.

[0065] Continuous SSL features may be sensitive to slight changes or noise in the data, while discretized content embedding can smooth out some local noise and subtle changes in the data to a certain extent by mapping the features to a limited discrete space, making the model more tolerant to input changes, improving the robustness and stability of the model, and reducing performance degradation caused by data fluctuations.

[0066] Discrete content embedding can transform the representation of speech content into a more semantically and structurally meaningful form, making it easier for the model to learn the relationship and pattern between speech content and other related information (such as speaking style, etc.). This discretized representation method helps the model to better generalize on different speech data, more accurately capture the commonalities and regularities in speech, and improve the performance of the model in various scenarios.

[0067] In subsequent model calculations, processing discrete content embeddings is usually less computationally intensive than processing continuous high-dimensional SSL features. Discretization reduces the dimensionality and complexity of the data, reduces the computational cost during model training and inference, improves the model's operating efficiency, and facilitates the realization of real-time or efficient speech processing tasks.

[0068] In this embodiment, the style library obtained by performing style modeling according to the transposed dual attention mechanism, the learnable query set and the content embedding includes:

[0069] Using a Mel filter bank to perform Mel filtering on the speech data in the target speech library to obtain first data;

[0070] Mel-encoding the first data using a Mel encoder to obtain second data;

[0071] Performing style encoding on the second data using a style encoder to obtain a style encoding output;

[0072] Determine the content embedding as the key of the first multi-head attention layer, determine the style encoding output as the value of the first multi-head attention layer, and determine the learnable embedding in the learnable query set as the query of the first multi-head attention layer for processing, and obtain the output of the first multi-head attention layer to construct the style library.

[0073] In the above embodiment, establishing the style library based on the attention mechanism mainly has the following advantages:

[0074] (1) Focusing on key style information: The attention mechanism allows the model to assign attention weights based on the similarity between content embedding and query set entries, thereby focusing on key style information related to the current speech content, filtering out irrelevant or minor style features, and extracting more representative style representations, so that the style library can more accurately characterize the style characteristics of the target speaker.

[0075] (2) Dynamically adapt to content changes: In speech conversion, different speech content may require different speaking styles. The attention mechanism can dynamically adjust the degree of attention and selection method of style information according to the input speech content, so that the style library can flexibly adapt to various content changes, provide the most appropriate style embedding for different speech content, and improve the naturalness and accuracy of speech conversion.

[0076] (3) Fixed style library size: By predetermining the query set size, the style library output by the attention layer is fixed in terms of the number of vectors. This helps to keep the size of the style library stable during model training and inference, reduces computational complexity and storage requirements, and also facilitates the management and operation of the style library.

[0077] (4) Enhanced style modeling capabilities: The use of multi-head attention layers can interact and model the content embedding and style encoder output from multiple different angles and subspaces, capturing richer and more complex style information and the relationship between content and style, thereby enhancing the ability to model the target speaker's style, improving the quality and expressiveness of the style library, and providing better quality style features for subsequent speech synthesis.

[0078] In this embodiment, the processing of the target source speech according to the style library and the diffusion model to obtain the target converted speech includes:

[0079] Using a second HuBERT model to extract features of the target source speech to obtain a second self-supervised learning feature;

[0080] Performing vector quantization processing on the second self-supervised learning feature to obtain a second vector quantization feature;

[0081] Using a second content encoder to perform content encoding on the second vector quantization feature to obtain target content embedding;

[0082] Performing linear mapping on the target content embedding to obtain target features;

[0083] Determine the learnable embedding in the learnable query set as the key of the second multi-head attention layer, determine the style library as the value of the second multi-head attention layer, and determine the target content embedding as the query of the second multi-head attention layer for processing to obtain a style embedding;

[0084] adding random noise and time embedding to the diffusion model;

[0085] Using the diffusion model, under the random noise, different style embeddings are selected for processing in each frame of the target content embedding according to the time embedding to obtain a mel-spectrogram;

[0086] The Mel spectrum is processed by a vocoder to obtain the target converted speech.

[0087] In the above embodiment, under the transposed dual attention mechanism, the first multi-head attention layer uses the randomly initialized learnable embedding (query set) as the query, the content embedding as the key, and the style encoder output as the value, and can filter out relevant style information from the style encoder output according to the content embedding, focusing on the style features related to the current speech content to establish a style library. For example: in speech with different semantic contents, this mechanism can help capture the unique intonation, rhythm and other style characteristics of the target speaker under the corresponding content.

[0088] The second multi-head attention layer transposes the roles of query and key, embedding the content extracted from the source speech as the query, using the same query set generated by the style library as the key, and the value is the style library. In this way, the target speaker's speaking style on the specific content can be accurately found from the style library based on the source speech content, achieving the adaptation of content and style, so that the converted speech is more natural and accurately matches the characteristics of the target speaker in style. For example, during speech conversion, the converted speech can have the corresponding style of the target speaker while maintaining the content of the source speech.

[0089] In the above embodiment, based on time embedding, different style embeddings are used in each frame, which means that in the process of speech synthesis, different style representations (i.e., style embeddings) in the style library are input into the diffusion model frame by frame. Since speech is a time-series signal, different frames carry different speech information. During training, the diffusion model learns the mapping relationship between speech content and style, and understands which style features correspond to different speech content segments. During reasoning, based on this learning result, the model selects appropriate style embeddings from the style library according to the content of each frame of the source speech and inputs them into the diffusion model. When generating the Mel-spectrogram, the model can fully consider the style changes of the speech at different time points, generate more natural speech that meets the expected style, avoid the single style of the entire speech, and realize the dynamic and precise conversion of speech style in the time dimension.

[0090] In the above embodiments, random noise can be used to simulate the diversity of data in the actual environment. In speech synthesis, real-world speech data has certain variability and noise. By randomly adding noise to the model, the natural changes of speech data in the actual environment can be simulated, so that the model can learn a more generalized speech feature representation. For example, different recording environments, changes in the speaker's state and other factors will bring various subtle changes to the speech, and noise addition helps the model capture these changes. Let the model learn how to recover the original speech (i.e., Mel spectrogram) from the noise, which can enhance the robustness of the model to noise. This means that in practical applications, even if some unseen noise interference is encountered, the model can generate high-quality speech more stably. For example, when processing source speech recorded in a noisy environment, the model can still effectively perform style conversion and speech synthesis. The diffusion model is based on the principle of gradually adding noise to convert data into noise distribution, and then reverse denoising to generate target data. Randomly adding noise is the basis for building a forward diffusion process, which provides a starting point for the subsequent reverse denoising process, so that the model can learn the mapping relationship from noise to Mel spectrogram, thereby realizing speech synthesis.

[0091] Specifically, the above processing process can refer to Figure 2 , Figure 2The output of the first multi-head attention layer is the weighted sum of the outputs of the style encoder; the attention weight of the first multi-head attention layer is determined according to the similarity between the content embedding and the learnable query set.

[0092] S14, using a designated voice playing device to play the target converted voice.

[0093] In this embodiment, the designated voice playback device may be a speaker of a smart terminal used by the source voice user, etc.

[0094] In this embodiment, when this embodiment is applied to the intelligent agent in the financial scenario, personalized voice assistant and voice customer service can be provided through voice style conversion. When this embodiment is applied to the medical and health scenario, the patient's privacy can be protected through voice style conversion, and a more relaxed Q&A atmosphere can be provided for the patient.

[0095] It can be seen from the above technical solutions that, on the one hand, the present invention adopts a transposed dual attention mechanism, which can extract multiple speech style representations in a content-related manner without text transcription or speaker labels; on the one hand, the present invention extracts self-supervised learning features based on the HuBERT model and content encoder to obtain content information, and uses a vector quantization method to discretize the self-supervised learning features into content embeddings, which can achieve efficient data representation without losing key semantics, and facilitate subsequent processing and calculation; on the other hand, the speech of the target speaker is used to generate a style library containing multiple style embeddings, each style embedding corresponds to the pronunciation of different speech content, can adapt to more complex conversion requirements, and improve the naturalness and conversion quality of speech conversion.

[0096] like Figure 3 As shown, it is a functional module diagram of a preferred embodiment of the speech conversion device based on transposition dual attention of the present invention. The speech conversion device based on transposition dual attention 11 includes a construction unit 110, a training unit 111, a processing unit 112, and a playback unit 113. The module / unit referred to in the present invention refers to a series of computer program segments that can be executed by a processor and can perform fixed functions, which are stored in a memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0097] The construction unit 110 is used to construct an initial model based on the transposed dual attention mechanism, the HuBERT model, the content encoder and the score-based diffusion model, and to construct a target loss function of the initial model.

[0098] In this embodiment, the transposed dual attention mechanism includes a first multi-head attention layer and a second multi-head attention layer that is in a transposed relationship with the first multi-head attention layer.

[0099] In this embodiment, the HuBERT model is a model that performs self-supervised speech representation learning through hidden unit mask prediction. The core idea is to use an offline clustering step to generate noisy labels and apply prediction loss only to the masked area, forcing the model to learn a high-level representation of the unmasked input, thereby inferring the target of the masked input. In this way, HuBERT significantly improves the quality of speech representation and has advanced performance in multiple speech recognition tasks.

[0100] In this embodiment, the content encoder is a device that compiles and converts signals or data into a signal form that can be used for communication, transmission and storage. The content encoder is mainly used to compress and encode the original data to reduce the amount of data and improve the transmission efficiency. The content encoder plays an important role in the field of digital signal processing and multimedia, and is usually used in conjunction with a decoder to achieve data compression, transmission and recovery.

[0101] In this embodiment, the backbone of the score-based diffusion model adopts a U-Net (U-shaped Network) structure for speech synthesis.

[0102] In this embodiment, the construction unit 110 constructs the target loss function of the initial model including:

[0103] Construct diffusion model loss and content encoder loss;

[0104] The sum of the diffusion model loss and the content encoder loss is calculated to obtain the target loss function.

[0105] Through the above embodiments, the diffusion model loss and the content encoder loss can be taken into account during the training process to improve the accuracy of the model.

[0106] The construction unit 110 is further used to collect speech data of different speaking styles to construct a target speech library, collect multiple source speech segments, and randomly initialize a learnable query set.

[0107] In this embodiment, the target speech database contains a speech data set with a specific speaking style, which is used to provide a target style reference for speech conversion.

[0108] For example, the target voice library can store the content of a designated host, the voice of a designated artist, etc., to enrich the voice style. Specifically, if a Mandarin voice is to be converted into a voice with the style of a well-known host, multiple audio clips of the host hosting the program can be collected to form the target voice library. These audio clips will be used to extract style features in the voice conversion system and guide the source voice to transform into the host style.

[0109] In this embodiment, the multiple source speech segments may be speech to be converted.

[0110] For example, the source speech may be a normal person's Mandarin speech, such as a recording of Zhang San reading a news report. The speech will retain the content information in the speech conversion system while changing the style so that it has the characteristics of the speech in the target speech library (such as the host style).

[0111] During the training process, the content of the source speech is taken as the final output, and the model is trained to faithfully reconstruct the input speech within the autoencoder framework.

[0112] In this embodiment, the learnable query set is a learnable embedding set randomly initialized during the model training process, which can be used as a query vector in the attention mechanism to explore the relationship between different features to assist in style feature extraction and fusion.

[0113] The training unit 111 is used to train the initial model based on the target loss function, according to the target speech library, the multiple source speech segments, and the learnable query set to obtain a speech conversion model.

[0114] In this embodiment, when the value of the target loss function converges, the training is stopped and the speech conversion model is obtained.

[0115] The processing unit 112 is used for processing the target source speech by using the speech conversion model in response to the speech conversion instruction of the target source speech to obtain the target converted speech.

[0116] In this embodiment, the processing unit 112 processes the target source speech using the speech conversion model to obtain the target converted speech including:

[0117] Processing the target speech library using the HuBERT model and the content encoder to obtain content embedding;

[0118] Performing style modeling according to the transposed dual attention mechanism, the learnable query set, and the content embedding to obtain a style library;

[0119] The target source speech is processed according to the style library and the diffusion model to obtain the target converted speech.

[0120] In this embodiment, the process of processing the target speech library by using the HuBERT model and the content encoder to obtain content embedding includes:

[0121] Using a first HuBERT model to extract features from the speech data in the target speech library to obtain a first self-supervised learning feature;

[0122] Performing vector quantization processing on the first self-supervised learning feature to obtain a first vector quantization feature;

[0123] The first vector quantization feature is content encoded by using a first content encoder to obtain the content embedding.

[0124] In the above embodiment, self-supervised learning features (SSL) can be extracted as content information based on the HuBERT model, and the SSL features can be discretized into content embedding using a quantization method.

[0125] SSL features are usually continuous high-dimensional vectors that contain a lot of detailed information, but there may be redundancy. Therefore, the content embedding obtained after discretization can represent the key information of the speech content in a more concise way, remove unnecessary details, and achieve efficient data representation without losing key semantics, which is convenient for subsequent processing and calculation.

[0126] Continuous SSL features may be sensitive to slight changes or noise in the data, while discretized content embedding can smooth out some local noise and subtle changes in the data to a certain extent by mapping the features to a limited discrete space, making the model more tolerant to input changes, improving the robustness and stability of the model, and reducing performance degradation caused by data fluctuations.

[0127] Discrete content embedding can transform the representation of speech content into a more semantically and structurally meaningful form, making it easier for the model to learn the relationship and pattern between speech content and other related information (such as speaking style, etc.). This discretized representation method helps the model to better generalize on different speech data, more accurately capture the commonalities and regularities in speech, and improve the performance of the model in various scenarios.

[0128] In subsequent model calculations, processing discrete content embeddings is usually less computationally intensive than processing continuous high-dimensional SSL features. Discretization reduces the dimensionality and complexity of the data, reduces the computational cost during model training and inference, improves the model's operating efficiency, and facilitates the realization of real-time or efficient speech processing tasks.

[0129] In this embodiment, the style library obtained by performing style modeling according to the transposed dual attention mechanism, the learnable query set and the content embedding includes:

[0130] Using a Mel filter bank to perform Mel filtering on the speech data in the target speech library to obtain first data;

[0131] Mel-encoding the first data using a Mel encoder to obtain second data;

[0132] Performing style encoding on the second data using a style encoder to obtain a style encoding output;

[0133] Determine the content embedding as the key of the first multi-head attention layer, determine the style encoding output as the value of the first multi-head attention layer, and determine the learnable embedding in the learnable query set as the query of the first multi-head attention layer for processing, and obtain the output of the first multi-head attention layer to construct the style library.

[0134] In the above embodiment, establishing the style library based on the attention mechanism mainly has the following advantages:

[0135] (1) Focusing on key style information: The attention mechanism allows the model to assign attention weights based on the similarity between content embedding and query set entries, thereby focusing on key style information related to the current speech content, filtering out irrelevant or minor style features, and extracting more representative style representations, so that the style library can more accurately characterize the style characteristics of the target speaker.

[0136] (2) Dynamically adapt to content changes: In speech conversion, different speech content may require different speaking styles. The attention mechanism can dynamically adjust the degree of attention and selection method of style information according to the input speech content, so that the style library can flexibly adapt to various content changes, provide the most appropriate style embedding for different speech content, and improve the naturalness and accuracy of speech conversion.

[0137] (3) Fixed style library size: By predetermining the query set size, the style library output by the attention layer is fixed in terms of the number of vectors. This helps to keep the size of the style library stable during model training and inference, reduces computational complexity and storage requirements, and also facilitates the management and operation of the style library.

[0138] (4) Enhanced style modeling capabilities: The use of multi-head attention layers can interact and model the content embedding and style encoder output from multiple different angles and subspaces, capturing richer and more complex style information and the relationship between content and style, thereby enhancing the ability to model the target speaker's style, improving the quality and expressiveness of the style library, and providing better quality style features for subsequent speech synthesis.

[0139] In this embodiment, the processing of the target source speech according to the style library and the diffusion model to obtain the target converted speech includes:

[0140] Using a second HuBERT model to extract features of the target source speech to obtain a second self-supervised learning feature;

[0141] Performing vector quantization processing on the second self-supervised learning feature to obtain a second vector quantization feature;

[0142] Using a second content encoder to perform content encoding on the second vector quantization feature to obtain target content embedding;

[0143] Performing linear mapping on the target content embedding to obtain target features;

[0144] Determine the learnable embedding in the learnable query set as the key of the second multi-head attention layer, determine the style library as the value of the second multi-head attention layer, and determine the target content embedding as the query of the second multi-head attention layer for processing to obtain a style embedding;

[0145] adding random noise and time embedding to the diffusion model;

[0146] Using the diffusion model, under the random noise, different style embeddings are selected for processing in each frame of the target content embedding according to the time embedding to obtain a mel-spectrogram;

[0147] The Mel spectrum is processed by a vocoder to obtain the target converted speech.

[0148] In the above embodiment, under the transposed dual attention mechanism, the first multi-head attention layer uses the randomly initialized learnable embedding (query set) as the query, the content embedding as the key, and the style encoder output as the value, and can filter out relevant style information from the style encoder output according to the content embedding, focusing on the style features related to the current speech content to establish a style library. For example: in speech with different semantic contents, this mechanism can help capture the unique intonation, rhythm and other style characteristics of the target speaker under the corresponding content.

[0149] The second multi-head attention layer transposes the roles of query and key, embedding the content extracted from the source speech as the query, using the same query set generated by the style library as the key, and the value is the style library. In this way, the target speaker's speaking style on the specific content can be accurately found from the style library based on the source speech content, achieving the adaptation of content and style, so that the converted speech is more natural and accurately matches the characteristics of the target speaker in style. For example, during speech conversion, the converted speech can have the corresponding style of the target speaker while maintaining the content of the source speech.

[0150] In the above embodiment, based on time embedding, different style embeddings are used in each frame, which means that in the process of speech synthesis, different style representations (i.e., style embeddings) in the style library are input into the diffusion model frame by frame. Since speech is a time-series signal, different frames carry different speech information. During training, the diffusion model learns the mapping relationship between speech content and style, and understands which style features correspond to different speech content segments. During reasoning, based on this learning result, the model selects appropriate style embeddings from the style library according to the content of each frame of the source speech and inputs them into the diffusion model. When generating the Mel-spectrogram, the model can fully consider the style changes of the speech at different time points, generate more natural speech that meets the expected style, avoid the single style of the entire speech, and realize the dynamic and precise conversion of speech style in the time dimension.

[0151] In the above embodiments, random noise can be used to simulate the diversity of data in the actual environment. In speech synthesis, real-world speech data has certain variability and noise. By randomly adding noise to the model, the natural changes of speech data in the actual environment can be simulated, so that the model can learn a more generalized speech feature representation. For example, different recording environments, changes in the speaker's state and other factors will bring various subtle changes to the speech, and noise addition helps the model capture these changes. Let the model learn how to recover the original speech (i.e., Mel spectrogram) from the noise, which can enhance the robustness of the model to noise. This means that in practical applications, even if some unseen noise interference is encountered, the model can generate high-quality speech more stably. For example, when processing source speech recorded in a noisy environment, the model can still effectively perform style conversion and speech synthesis. The diffusion model is based on the principle of gradually adding noise to convert data into noise distribution, and then reverse denoising to generate target data. Randomly adding noise is the basis for building a forward diffusion process, which provides a starting point for the subsequent reverse denoising process, so that the model can learn the mapping relationship from noise to Mel spectrogram, thereby realizing speech synthesis.

[0152] Specifically, the above processing process can refer to Figure 2 , Figure 2The output of the first multi-head attention layer is the weighted sum of the outputs of the style encoder; the attention weight of the first multi-head attention layer is determined according to the similarity between the content embedding and the learnable query set.

[0153] The playing unit 113 is used to play the target converted voice using a designated voice playing device.

[0154] In this embodiment, the designated voice playback device may be a speaker of a smart terminal used by the source voice user, etc.

[0155] In this embodiment, when this embodiment is applied to the intelligent agent in the financial scenario, personalized voice assistant and voice customer service can be provided through voice style conversion. When this embodiment is applied to the medical and health scenario, the patient's privacy can be protected through voice style conversion, and a more relaxed Q&A atmosphere can be provided for the patient.

[0156] It can be seen from the above technical solutions that, on the one hand, the present invention adopts a transposed dual attention mechanism, which can extract multiple speech style representations in a content-related manner without text transcription or speaker labels; on the one hand, the present invention extracts self-supervised learning features based on the HuBERT model and content encoder to obtain content information, and uses a vector quantization method to discretize the self-supervised learning features into content embeddings, which can achieve efficient data representation without losing key semantics, and facilitate subsequent processing and calculation; on the other hand, the speech of the target speaker is used to generate a style library containing multiple style embeddings, each style embedding corresponds to the pronunciation of different speech content, can adapt to more complex conversion requirements, and improve the naturalness and conversion quality of speech conversion.

[0157] like Figure 4 , which is a schematic diagram of the structure of a computer device of a preferred embodiment of the present invention for implementing a speech conversion method based on transposed dual attention.

[0158] The computer device 1 may include a memory 12, a processor 13 and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a speech conversion program based on transposed dual attention.

[0159] Those skilled in the art will appreciate that the schematic diagram is merely an example of the computer device 1 and does not constitute a limitation on the computer device 1. The computer device 1 may be a bus-type structure or a star-type structure. The computer device 1 may also include more or less other hardware or software than shown in the diagram, or a different arrangement of components. For example, the computer device 1 may also include input and output devices, network access devices, etc.

[0160] It should be noted that the computer device 1 is only an example, and other existing or future electronic products that are suitable for the present invention should also be included in the protection scope of the present invention and included here by reference.

[0161] Among them, the memory 12 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (for example: SD or DX memory, etc.), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as a mobile hard disk of the computer device 1. In other embodiments, the memory 12 can also be an external storage device of the computer device 1, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the computer device 1. Further, the memory 12 can also include both an internal storage unit of the computer device 1 and an external storage device. The memory 12 can not only be used to store application software and various types of data installed in the computer device 1, such as the code of the speech conversion program based on transposition dual attention, but also can be used to temporarily store data that has been output or is to be output.

[0162] In some embodiments, the processor 13 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips. The processor 13 is the control core (Control Unit) of the computer device 1, and uses various interfaces and lines to connect various components of the entire computer device 1, and executes or executes programs or modules stored in the memory 12 (for example, executing a speech conversion program based on transposition dual attention, etc.), and calls data stored in the memory 12 to execute various functions of the computer device 1 and process data.

[0163] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the above-mentioned various embodiments of the speech conversion method based on transposition dual attention, for example Figure 1 Steps shown.

[0164] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to implement the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of implementing specific functions, which are used to describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a construction unit 110, a training unit 111, a processing unit 112, and a playback unit 113.

[0165] The above-mentioned integrated unit implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, and includes a number of instructions for enabling a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute the part of the speech conversion method based on transposition dual attention described in various embodiments of the present invention.

[0166] If the module / unit integrated in the computer device 1 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware devices through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of each of the above-mentioned method embodiments can be implemented.

[0167] The computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. The computer readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.

[0168] Furthermore, the computer-readable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.

[0169] The blockchain referred to in this invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.

[0170] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 The bus is represented by only one straight line, but it does not mean that there is only one bus or one type of bus. The bus is configured to realize the connection and communication between the memory 12 and at least one processor 13, etc.

[0171] Although not shown, the computer device 1 may also include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 13 through a power management device, so that the power management device can realize functions such as charging management, discharging management, and power consumption management. The power source may also include any components such as one or more DC or AC power sources, recharging devices, power failure detection circuits, power converters or inverters, power status indicators, etc. The computer device 1 may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.

[0172] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the computer device 1 and other computer devices.

[0173] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), or a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touch device. The display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the computer device 1 and to display a visual user interface.

[0174] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.

[0175] It can be understood by those skilled in the art that Figure 4 The structure shown does not constitute a limitation on the computer device 1, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.

[0176] Combination Figure 1 , the memory 12 in the computer device 1 stores a plurality of instructions to implement a speech conversion method based on transposed dual attention, and the processor 13 can execute the plurality of instructions to implement:

[0177] Building an initial model based on the transposed dual attention mechanism, the HuBERT model, the content encoder, and the score-based diffusion model, and building a target loss function for the initial model;

[0178] Collect speech data of different speaking styles to build a target speech library, collect multiple source speech segments, and randomly initialize the learnable query set;

[0179] Based on the target loss function, the initial model is trained according to the target speech library, the multiple source speech segments, and the learnable query set to obtain a speech conversion model;

[0180] In response to a speech conversion instruction for a target source speech, the target source speech is processed using the speech conversion model to obtain a target converted speech;

[0181] The target converted voice is played by using a designated voice playing device.

[0182] Specifically, the specific implementation method of the processor 13 for the above instructions can refer to Figure 1 The description of the relevant steps in the corresponding embodiments will not be repeated here.

[0183] It should be noted that the data involved in this case were all obtained legally. The software tools or components not produced by our company that appear in the embodiments of this application are only examples and do not represent actual use.

[0184] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0185] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0186] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0187] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.

[0188] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0189] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.

[0190] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the present invention can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.

[0191] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.

Claims

1. A speech conversion method based on transposed dual attention, characterized in that: The speech conversion method based on transposition dual attention includes: Building an initial model based on the transposed dual attention mechanism, the HuBERT model, the content encoder, and the score-based diffusion model, and building a target loss function for the initial model; Collect speech data of different speaking styles to build a target speech library, collect multiple source speech segments, and randomly initialize the learnable query set; Based on the target loss function, the initial model is trained according to the target speech library, the multiple source speech segments, and the learnable query set to obtain a speech conversion model; In response to a speech conversion instruction for a target source speech, the target source speech is processed using the speech conversion model to obtain a target converted speech; The target converted voice is played by using a designated voice playing device.

2. The speech conversion method based on transposed dual attention as claimed in claim 1, characterized in that: The objective loss function of constructing the initial model includes: Construct diffusion model loss and content encoder loss; The sum of the diffusion model loss and the content encoder loss is calculated to obtain the target loss function.

3. The speech conversion method based on transposed dual attention as claimed in claim 1, characterized in that: The using the speech conversion model to process the target source speech to obtain the target converted speech comprises: Processing the target speech library using the HuBERT model and the content encoder to obtain content embedding; Performing style modeling according to the transposed dual attention mechanism, the learnable query set, and the content embedding to obtain a style library; The target source speech is processed according to the style library and the diffusion model to obtain the target converted speech.

4. The speech conversion method based on transposed dual attention as claimed in claim 3, characterized in that: The using the HuBERT model and the content encoder to process the target speech library to obtain content embedding includes: Using a first HuBERT model to extract features from the speech data in the target speech library to obtain a first self-supervised learning feature; Performing vector quantization processing on the first self-supervised learning feature to obtain a first vector quantization feature; The first vector quantization feature is content encoded by using a first content encoder to obtain the content embedding.

5. The speech conversion method based on transposed dual attention as claimed in claim 4, characterized in that: The transposed dual attention mechanism includes a first multi-head attention layer and a second multi-head attention layer in a transposed relationship with the first multi-head attention layer; The style library obtained by performing style modeling according to the transposed dual attention mechanism, the learnable query set and the content embedding includes: Using a Mel filter bank to perform Mel filtering on the speech data in the target speech library to obtain first data; Mel-encoding the first data using a Mel encoder to obtain second data; Performing style encoding on the second data using a style encoder to obtain a style encoding output; Determine the content embedding as the key of the first multi-head attention layer, determine the style encoding output as the value of the first multi-head attention layer, and determine the learnable embedding in the learnable query set as the query of the first multi-head attention layer for processing, and obtain the output of the first multi-head attention layer to construct the style library.

6. The speech conversion method based on transposed dual attention as claimed in claim 5, characterized in that: The processing of the target source speech according to the style library and the diffusion model to obtain the target converted speech includes: Using a second HuBERT model to extract features of the target source speech to obtain a second self-supervised learning feature; Performing vector quantization processing on the second self-supervised learning feature to obtain a second vector quantization feature; Using a second content encoder to perform content encoding on the second vector quantization feature to obtain target content embedding; Performing linear mapping on the target content embedding to obtain target features; Determine the learnable embedding in the learnable query set as the key of the second multi-head attention layer, determine the style library as the value of the second multi-head attention layer, and determine the target content embedding as the query of the second multi-head attention layer for processing to obtain a style embedding; adding random noise and time embedding to the diffusion model; Using the diffusion model, under the random noise, different style embeddings are selected for processing in each frame of the target content embedding according to the time embedding to obtain a mel-spectrogram; The Mel spectrum is processed by a vocoder to obtain the target converted speech.

7. In the speech conversion method based on transposed dual attention as described in claim 5, the output of the first multi-head attention layer is the weighted sum of the outputs of the style encoder; the attention weight of the first multi-head attention layer is determined according to the similarity between the content embedding and the learnable query set.

8. A speech conversion device based on transposed dual attention, characterized in that: The speech conversion device based on transposition dual attention comprises: A construction unit, used to construct an initial model based on a transposed dual attention mechanism, a HuBERT model, a content encoder, and a score-based diffusion model, and to construct a target loss function of the initial model; The construction unit is further used to collect speech data of different speaking styles to construct a target speech library, collect multiple source speech segments, and randomly initialize a learnable query set; A training unit, configured to train the initial model based on the target loss function, the target speech library, the multiple source speech segments, and the learnable query set to obtain a speech conversion model; A processing unit, configured to, in response to a speech conversion instruction for a target source speech, process the target source speech using the speech conversion model to obtain a target converted speech; The playing unit is used to play the target converted voice by using a designated voice playing device.

9. A computer device, characterized in that: The computer device comprises: a memory storing at least one instruction; and A processor executes instructions stored in the memory to implement the speech conversion method based on transposed dual attention as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the speech conversion method based on transposed dual attention as described in any one of claims 1 to 7.