Voice conversion method and device, computer equipment and storage medium

Through the combination of the multi-head attention module and the graph encoder, the problems of long computing cycles and large computing volume in speech conversion are solved, and an efficient speech conversion process is realized, reducing dependence on the target speaker's data is improved, and computing efficiency is improved.

CN120236597APending Publication Date: 2025-07-01MOBILITY ASIA SMART TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311706804.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing voice conversion technology has the problem of long computing cycles and large computing volumes, especially when adding new target speaker voices or performing custom replicas of user voices, a large amount of data is required for model training, resulting in high costs and long cycles.

Method used

The multi-head attention module is used to align the text features and audio features of the source speaker, combine the graph encoder and graph convolution network to extract features, and perform voice conversion through an end-to-end method to reduce dependence on the target speaker data, and realize multi-feature parallel extraction.

Benefits of technology

It improves the computing efficiency of speech conversion, reduces calculation costs and model training time, and realizes a fast speech conversion process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236597A_ABST
    Figure CN120236597A_ABST
Patent Text Reader

Abstract

The invention relates to a voice conversion method and device, computer equipment and a storage medium. The method comprises the following steps: obtaining a source speaker text feature and a source speaker audio feature according to source speaker voice data; aligning the source speaker text features and the source speaker audio features by adopting a multi-head attention module to obtain audio mapping data for representing an alignment relationship; matrix multiplication processing is carried out on the audio mapping data and the source speaker text features to obtain source speaker language features; according to the audio features of the source speaker, determining the fundamental frequency features of the source speaker; obtaining voice data of a target speaker, and extracting features of the target speaker from the voice data of the target speaker; and inputting the source speaker language features, the source speaker fundamental frequency features and the target speaker features into a target speaker decoder, and then processing and converting the features into target voices through a vocoder. By adopting the method, the operation efficiency of voice conversion can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of speech processing, and in particular, to a speech conversion method, apparatus, computer device, and storage medium. Background Art

[0002] With the development of speech processing technology, there has emerged the technology of voice conversion. While retaining the speech content, voice conversion converts the timbre of the original speaker into the timbre of a specified speaker. It has important applications in aspects such as movie dubbing, character imitation, and replicating the timbre of a person. There are also rich application scenarios in the vehicle field, such as navigation voice conversion, IP customization, and personalized user timbre customization.

[0003] Currently, great progress has been made in voice conversion to a specific target speaker based on deep learning. For example, voice conversion methods based on CycleGAN (Cycle Generative Adversarial Network, a machine learning algorithm that can achieve data conversion by transforming input samples), VAE (Variant Autoencoder), and ASR (Automatic Speech Recognition) can all convert speech into the timbre of speakers within the training set. However, if you want to add a new target speaker timbre or perform custom replication of user timbre, usually a large amount of speaker data is required to retrain a voice conversion model with the timbre of this speaker as the target timbre, or adaptively train the existing model with a small amount of data. By the ways of retraining the target speaker model and adaptively training the existing model, there are problems of long database recording cycle and high cost.

[0004] In addition, by the way of sharing the attention network between the text-assisted attention model and the TTS model, although it can also achieve one-to-many voice conversion, there are also disadvantages such as long training cycle and large computational amount.

[0005] Therefore, the current voice conversion technology has problems of long operation cycle and large computational amount. Summary of the Invention

[0006] Based on this, it is necessary to provide a voice conversion method, apparatus, computer device, and storage medium that can improve the operation efficiency of voice conversion for the above technical problems.

[0007] In a first aspect, an embodiment of the present disclosure provides a voice conversion method, and the method includes:

[0008] Obtain source speaker text features and source speaker audio features according to the source speaker speech data;

[0009] The multi - head attention module is used to align the source speaker's text features and audio features to obtain audio mapping data for representing the alignment relationship;

[0010] Based on matrix multiplication processing of the audio mapping data and the source speaker's text features, the source speaker's language features are obtained;

[0011] According to the source speaker's audio features, the source speaker's fundamental frequency features are determined;

[0012] The target speaker's speech data is obtained, and the target speaker's features are extracted from the target speaker's speech data;

[0013] The source speaker's language features, the source speaker's fundamental frequency features, and the target speaker's features are input into the target speaker decoder for decoding to obtain target spectral data;

[0014] The target spectral data is processed through a vocoder to be converted into target speech.

[0015] In some embodiments, according to the source speaker's speech data, obtaining the source speaker's text features and audio features includes:

[0016] According to the source speaker's speech data, the source speaker's text data and voice data are obtained. According to the source speaker's text data, the source speaker's phoneme data is obtained;

[0017] The graph encoder is used to perform syntactic relationship encoding on the source speaker's text data according to the syntax graph of the source speaker's text data to obtain relationship encoding features, and based on the graph attention mechanism, the relationship encoding features are used to perform text encoding on the source speaker's phoneme data to obtain the source speaker's text features;

[0018] The source speaker's voice data is pre - processed to obtain the source speaker's audio features.

[0019] In some embodiments, using the graph encoder to perform syntactic relationship encoding on the source speaker's text data according to the syntax graph of the source speaker's text data to obtain relationship encoding features, and based on the graph attention mechanism, using the relationship encoding features to perform text encoding on phoneme data to obtain the source speaker's text features includes:

[0020] Generating a syntax tree corresponding to the source speaker's text data according to the source speaker's text data;

[0021] Parsing the syntactic relationships between words in the syntax tree, using phonemes as nodes and the syntactic relationships between phonemes as edges to generate a syntax graph;

[0022] Based on a bidirectional gated recurrent unit network, the syntactic relationship between two phonemes in a syntactic graph is bidirectionally encoded to obtain relationship encoding features; wherein, the relationship encoding features include a forward relationship encoding vector and a backward relationship encoding vector;

[0023] Based on a graph attention network of a graph encoder and the relationship encoding features, attention scores are calculated;

[0024] According to the attention scores, the source speaker phoneme data is encoded to obtain source speaker text features.

[0025] In some embodiments, based on a bidirectional gated recurrent unit network, the syntactic relationship between two phonemes in a syntactic graph is bidirectionally encoded to obtain relationship encoding features, including:

[0026] When two phonemes belong to the same word, based on the bidirectional gated recurrent unit network and using a self-loop edge encoding algorithm, the syntactic relationship between the two phonemes is bidirectionally encoded;

[0027] When two phonemes belong to different words, based on the bidirectional gated recurrent unit network, the syntactic relationship between the words to which the two phonemes respectively belong is bidirectionally encoded.

[0028] In some embodiments, the source speaker language features, the source speaker fundamental frequency features, and the target speaker features are input into a target speaker decoder for decoding to obtain target spectral data, including:

[0029] The source speaker language features and the source speaker fundamental frequency features are concatenated to obtain a source speaker concatenated feature;

[0030] The source speaker concatenated feature and the target speaker features are input into the target speaker decoder for decoding to obtain target spectral data.

[0031] In some embodiments, target speaker speech data is obtained, and target speaker features are extracted from the target speaker speech data, including:

[0032] A speaker embedding vector is extracted from the target speaker speech data by using a speaker embedding extraction network;

[0033] An affinity graph is constructed based on the speaker embedding vector and the embedding vectors of the training samples of the speaker embedding extraction network; wherein, in the affinity graph, the embedding vectors are used as nodes, and at least one nearest neighbor embedding vector determined by cosine similarity to each embedding vector is used as an edge;

[0034] Based on a graph convolutional network, clustering processing is performed on the nodes in the affinity graph to generate target speaker clusters;

[0035] Based on the target speaker clusters, target speaker features are obtained.

[0036] In some embodiments, clustering the nodes in the affinity graph based on a graph convolutional network to generate target speaker clusters includes:

[0037] Clustering the nodes in the affinity graph to generate a plurality of target speaker cluster proposals;

[0038] Extracting the clustering features of each target speaker cluster proposal by a clustering detection unit of the graph convolutional network, and identifying candidate target speaker cluster proposals from the generated plurality of target speaker cluster proposals according to the clustering features;

[0039] Predicting the probability values of the nodes in the candidate target speaker cluster proposal by a clustering segmentation unit based on graph convolution, and removing the nodes with probability values less than a preset threshold as abnormal nodes from the nodes of the candidate target speaker cluster proposal;

[0040] Sorting the nodes according to the magnitudes of the probability values of the nodes in the candidate target speaker cluster proposal, and selecting at least one node with a probability value greater than the preset threshold to generate a target speaker cluster.

[0041] In some embodiments, the voice conversion method further includes:

[0042] Assigning pseudo-labels to the target speaker clusters;

[0043] Optimizing the network parameters of the speaker embedding extraction network through a noise reduction loss function to denoise the pseudo-labels generated in the clustering process, and re-inputting the denoised pseudo-labels into the speaker embedding extraction network for iterative training.

[0044] In a second aspect, an embodiment of the present disclosure provides a voice conversion device, and the device includes:

[0045] A source speaker data acquisition module, configured to obtain source speaker text features and source speaker audio features according to source speaker voice data;

[0046] An audio feature alignment module, configured to align the source speaker text features and the source speaker audio features by using a multi-head attention module to obtain audio mapping data for representing the alignment relationship;

[0047] A source speaker fundamental frequency feature acquisition module, configured to determine source speaker fundamental frequency features according to the source speaker audio features;

[0048] A target speaker feature acquisition module, configured to acquire target speaker voice data and extract target speaker features from the target speaker voice data;

[0049] A decoding module, configured to input the source speaker's language features, the source speaker's fundamental frequency features, and the target speaker's features into a target speaker decoder for decoding to obtain target spectral data;

[0050] A conversion module, configured to convert the target spectral data into target speech through vocoder processing.

[0051] In a third aspect, embodiments of the present disclosure provide a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, corresponding steps in the voice conversion method in any embodiment of the first aspect of the present disclosure are implemented.

[0052] In a fourth aspect, embodiments of the present disclosure provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, corresponding steps in the voice conversion method in any embodiment of the first aspect of the present disclosure are implemented.

[0053] The above voice conversion method, device, computer device, and storage medium utilize a multi-head attention module to output audio mapping data representing the alignment relationship between the source speaker's text features and the source speaker's audio features, multiply the audio mapping data with the source speaker's text feature matrix to obtain the source speaker's language features, extract the target speaker's features from the target speaker's speech data, and input the source speaker's text features, the source speaker's fundamental frequency features, and the target speaker's features into a target speaker decoder for decoding, and after further processing, convert them into target speech. By introducing an end-to-end graph attention parallel method for voice conversion, it is not necessary to use a large amount of target speaker speech data to train the target speaker model, saving computational costs and model training time. In addition, while using a graph speaker encoder to extract the target speaker's features, a source speaker feature determination module determines the source speaker's language features in parallel, and a fundamental frequency extraction module also synchronously extracts the source speaker's fundamental frequency features, realizing parallel extraction of multiple features in the voice conversion process and improving the operation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a schematic flowchart of a voice conversion method in some embodiments;

[0055] Figure 2 It is a schematic data flow diagram for obtaining the source speaker's language features based on a source speaker feature determination module in some embodiments;

[0056] Figure 3 It is a schematic diagram of relevant network units and data flow in a voice conversion method in some embodiments;

[0057] Figure 4 It is a schematic data flow diagram for extracting the target speaker's features based on a graph speaker encoder in some embodiments;

[0058] Figure 5 Schematic diagram of data flow for obtaining target spectral data based on a target speaker decoder in some embodiments;

[0059] Figure 6 Block diagram of the structure of a voice conversion device in some embodiments;

[0060] Figure 7 Internal structure diagram of a computer device in some embodiments. Detailed implementation manners

[0061] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0062] The voice conversion method provided by the present application can be applied to a computer device, which can be a server or a terminal. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, portable wearable devices, and intelligent terminal devices such as vehicle-mounted terminals, smart speakers, and intelligent robots. The server can be implemented by an independent server or a server cluster composed of multiple servers.

[0063] In a first aspect, an embodiment of the present disclosure provides a voice conversion method, as Figure 1 shown. Taking the application of this method to an intelligent terminal device as an example, the method includes the following steps S101 to S107.

[0064] Step S101: Obtain source speaker text features and source speaker audio features according to source speaker voice data.

[0065] Among them, the source speaker voice data refers to the voice data of the original speaker in voice conversion technology. The source speaker text features can refer to the information representing the text features corresponding to the original speaker's voice, such as the information representing the sentence or phrase structure and content features corresponding to the original speaker's voice. The source speaker audio features can refer to the information representing the audio features corresponding to the original speaker's voice, such as the Mel spectrum data corresponding to the original speaker's voice.

[0066] By obtaining the source speaker voice data and performing operations such as speech recognition and audio extraction on it, the source speaker text features and source speaker audio features can be obtained.

[0067] Specifically, the source speaker text features can be obtained through traditional speech recognition technology methods, by converting the speech signal into a text signal for acquisition. For example, methods based on feature parameter extraction can be used, or methods for extracting text features of other existing models can be adopted.

[0068] For the source speaker audio features, audio feature extraction technologies can be used, such as methods for extracting features such as Mel cepstral distance and frequency variance spectral entropy.

[0069] In some embodiments, the source speaker text features and the source speaker audio features can be represented in the form of vectors. For example, the source speaker text features can be represented in forms such as word vectors, sentence vectors, or phoneme vectors, and the source speaker audio features can be represented in forms such as Mel spectrograms and cepstral coefficients to represent audio features. In some embodiments, the source speaker text features and the source speaker audio features can also be represented in the form of matrices.

[0070] In some embodiments, the source speaker audio features can be Mel spectrogram data. By preprocessing the speech signal in the source speaker speech data and then performing Fourier transform, its spectral data can be obtained. Further, through the Mel filter bank, the Mel spectrogram can be obtained.

[0071] For the specific ways to obtain the source speaker text features and the source speaker audio features according to the source speaker speech data, those skilled in the art can select appropriate ways according to actual needs, and no special limitations are imposed here.

[0072] Step S102: Use the multi-head attention module to align the source speaker text features and the source speaker audio features to obtain audio mapping data for representing the alignment relationship.

[0073] Among them, the multi-head attention module includes three matrix operation groups and an attention mechanism. Through the attention mechanism, different weights can be used to perform attention operations on different parts of the input sequence. Each attention operation is grouped and feature information is refined from multiple dimensions. In NLP (Natural Language Processing) tasks, self-attention can reconstruct the representation of the target word according to the context words.

[0074] The audio mapping data can be the audio features after the alignment operation, and the audio features can include the alignment relationship between the text features.

[0075] Specifically, use the multi-head attention module to align the source speaker text features and the source speaker audio features to obtain the aligned source speaker audio features, that is, obtain the audio mapping data.

[0076] In some embodiments, the alignment process using the multi-head attention module may include concatenating the source speaker's text features and audio features, taking the source speaker's text features as the query (Q), and the source speaker's audio features as the key (K) and value (V), calculating the weights of the query and the key through the attention mechanism to obtain attention weights, and then multiplying the value by the attention weights and outputting. Further, it may also include independent mapping through multiple heads, concatenating the outputs of each head to obtain audio mapping data.

[0077] Step S103: Based on matrix multiplication processing of the audio mapping data and the source speaker's text features, obtain the source speaker's language features.

[0078] Among them, matrix multiplication processing refers to matrix multiplication operation, which is a basic operation method in machine learning.

[0079] Specifically, the text features can be used as the rows of the matrix, the audio mapping data as the columns of the matrix, and then matrix multiplication operation is performed to obtain the source speaker's language features.

[0080] Step S104: Determine the fundamental frequency features of the source speaker according to the source speaker's audio features.

[0081] Specifically, the CWT (Continuous Wavelet Transform) algorithm can be used to extract the fundamental frequency features of the source speaker, or the source speaker's audio features (such as Mel spectrum data) can be input into the end-to-end CREPE (convolutional representation for pitch estimation) model to extract the fundamental frequency features of the source speaker.

[0082] Those skilled in the art can select a suitable algorithm or model to extract the fundamental frequency features of the source speaker according to actual needs, which is not limited here.

[0083] Step S105: Obtain the target speaker's speech data, and extract the target speaker's features from the target speaker's speech data.

[0084] Among them, the target speaker's speech data refers to the speech data of the specified speaker converted in the speech conversion technology, and the target speaker's features can be feature vectors that can characterize the speech characteristics of the target speaker.

[0085] Specifically, the target speaker's features can be extracted from the input target speaker's speech data through a graph speaker encoder. Exemplarily, the target speaker's features can be extracted through an encoder in any traditional method, or the target speaker's feature extraction method based on the graph speaker encoder described in the embodiments below of this application can also be used.

[0086] Step S106: Input the source speaker's language features, the source speaker's fundamental frequency features, and the target speaker's features into the target speaker decoder for decoding to obtain target spectral data.

[0087] In some embodiments, step S106 includes: concatenating the source speaker's language features and the source speaker's fundamental frequency features to obtain the source speaker's concatenated features, and inputting the source speaker's concatenated features and the target speaker's features into the target speaker decoder for decoding to obtain target spectral data.

[0088] The purpose of concatenation is to converge the source speaker's language features and the source speaker's fundamental frequency features. For example, a concatenation matrix is used to record the source speaker's language features and the source speaker's fundamental frequency features. There is no special limitation on the concatenation method here, and those skilled in the art can select a suitable concatenation method according to actual needs.

[0089] In some embodiments, the target spectral data can be a spectrogram. In some cases, the spectrogram can be a Mel spectrogram, and in some other cases, the spectrogram can be a linear spectrogram.

[0090] Step S107: Process the target spectral data through a vocoder to convert it into target speech.

[0091] Specifically, input the target spectrum into the vocoder to convert it into target speech. Among them, the vocoder can be a griffin-lim vocoder (a vocoder based on speech signals that generates an unknown phase spectrum from a known magnitude spectrum through an iterative process and reconstructs the speech waveform using the generated phase spectrum), or a neural network vocoder. Those skilled in the art can also select other vocoders for the conversion of synthesized audio according to actual needs, such as WaveNet vocoder, WaveRNN vocoder, etc.

[0092] By performing the voice conversion method of steps S101 to S107, the multi-head attention module is used to output audio mapping data representing the alignment relationship between the source speaker's text features and audio features. The audio mapping data is matrix-multiplied with the source speaker's text feature matrix to obtain the source speaker's language features. Then, the target speaker's features are extracted from the target speaker's voice data. The source speaker's language features, the source speaker's fundamental frequency features, and the target speaker's features are input into the target speaker decoder for decoding and are further processed to be converted into the target voice. By introducing the method of end-to-end graph attention parallelism for voice conversion, there is no need to use a large amount of target speaker voice data to train the target speaker model, saving computational costs and model training time. In addition, in this method, while using the graph speaker encoder to extract the target speaker's features, the source speaker feature determination module simultaneously determines the source speaker's language features in parallel, and the fundamental frequency extraction module also synchronously extracts the source speaker's fundamental frequency features, realizing the parallel extraction of multiple features in the voice conversion process and improving the operation efficiency.

[0093] In some embodiments, step S101 may include the following steps: obtaining the source speaker's text data and source speaker's voice data according to the source speaker's voice data; obtaining the source speaker's phoneme data according to the source speaker's text data; using a graph encoder to perform syntactic relationship encoding on the source speaker's text data according to the syntactic graph of the source speaker's text data to obtain relationship encoding features, and performing text encoding on the source speaker's phoneme data using the relationship encoding features based on the graph attention mechanism to obtain the source speaker's text features; preprocessing the source speaker's voice data to obtain the source speaker's audio features.

[0094] Among them, the source speaker's phoneme data may be a phoneme sequence carrying position encoding information.

[0095] In some embodiments, to obtain the source speaker's text data and source speaker's voice data according to the source speaker's voice data, the source speaker's text data can be obtained by speech recognition, and the source speaker's voice data can be obtained by using audio extraction technology.

[0096] In some embodiments, to obtain the source speaker's phoneme data, it can be obtained through the front-end processing module. Specifically, the source speaker's text data can be preprocessed by the front-end processing module. For example, it can go through text regularization processing (Text Normalization), grapheme-to-phoneme conversion processing (Grapheme-to-Phoneme). For Chinese text data, it can also go through polyphone classification processing (Polyphone Classification), prosody prediction processing (ProsodyPrediction), etc. for preprocessing and be converted into a phoneme sequence.

[0097] In some embodiments, the source speaker text data can be first converted into a phoneme sequence, and then position encoding can be further inserted into the phoneme sequence. Inserting position encoding can be understood as associating or concatenating the phoneme sequence with the position encoding. For a language, the position and order of words in a sentence are very important. Therefore, by inserting position encoding, the phonemes in the phoneme sequence can carry position information, enabling the subsequent model network to determine the phonemes to be processed based on the position information.

[0098] In some embodiments, the source speaker text features can be obtained through a graph encoder. Specifically, the graph encoder is used to perform syntactic relationship encoding on the source speaker text data according to the syntactic graph of the source speaker text data to obtain relationship encoding features, and based on the graph attention mechanism, the relationship encoding features are used to perform text encoding on the source speaker phoneme data to obtain the source speaker text features.

[0099] Among them, the graph encoder can include a relationship encoding unit and a text encoding unit. The relationship encoding unit is used to perform syntactic relationship encoding on the source speaker text data according to the syntactic graph of the source speaker text data to obtain relationship encoding features. The text encoding unit can include a graph attention network. By performing text encoding on the source speaker phoneme data through the graph attention mechanism combined with the relationship feature encoding, text encoding features can be obtained.

[0100] The syntactic graph is an extension of the syntactic tree. The syntactic tree is used to represent the linguistic dependency relationship between words. In the tree structure, the words in the sentence that have an association relationship are directly connected. Usually, in order to mine the syntactic relationship between two words in a sentence, the topological structure of the syntactic tree can be extended to establish a fully connected communication.

[0101] The text encoding features refer to the text features used to represent the syntactic relationship of the source speaker text data.

[0102] Specifically, the phoneme data corresponding to the preprocessed source speaker text data and the source speaker text data can be input into the graph encoder together. In the graph encoder, the relationship encoding features are processed by the graph attention mechanism to indicate the character relationship. The purpose of the graph encoder is to convert the input phoneme data and relationship encoding features into the corresponding syntax-driven character embedding sequence (text encoding features).

[0103] Preprocess the source speaker voice data to obtain the source speaker audio features.

[0104] In some embodiments, the source speaker audio features can be mel spectrum data. The spectrum data can be obtained by preprocessing the speech signal in the source speaker speech data and then performing Fourier transform, and further, the mel spectrum can be obtained through the mel filter bank.

[0105] In some embodiments, the source speaker audio features can be preprocessed and then input into the masked multi-head attention module for attention mechanism processing.

[0106] In some embodiments, the source speaker audio features can be Mel spectrum data carrying positional encoding information.

[0107] In some implementations, the source speaker text features can also be marked with identity recognition by the speaker label.

[0108] In some embodiments, a source speaker feature determination module can be used to obtain the source speaker language features. The source speaker feature determination module can include a graph encoder and a source speaker decoder. The graph encoder includes a relational encoding unit and a text encoding unit, and the source speaker decoder includes a masked multi-head attention module and a multi-head attention module, etc. Exemplarily, the source speaker feature determination module can refer to Figure 2 as shown Figure 2 Fig. shows a schematic data flow diagram of obtaining the source speaker language features based on the source speaker feature determination module 200 in some embodiments. The arrow direction in the figure indicates the data flow.

[0109] As Figure 2 shown, the method of obtaining the source speaker language features through the source speaker feature determination module can include the following steps from the first step to the twelfth step.

[0110] First step: Input the source speaker text data (such as a text sequence).

[0111] Second step: Input the source speaker text data into the front-end processing module. After text regularization (TextNormalization), grapheme-to-phoneme conversion, and for Chinese, polyphone classification (PolyphoneClassification) and prosody prediction (Prosody Prediction) are also performed, and finally it is converted into a phoneme sequence.

[0112] Third step: Input the source speaker text data into the relational encoding unit. Through the syntax tree, the syntax graph is input into the graph attention module through a bidirectional GRU (bidirectional gated recurrent unit).

[0113] Fourth step: The phoneme sequence output in the second step is input into the encoder prenet. After being processed by the prenet, it is concatenated with the positional encoding.

[0114] Fifth step: Input the phoneme data output in the fourth step into the graph attention module, and output after processing such as residual and normalization, forward feedback layer, and linear layer.

[0115] Step 6: Concatenate the features output in Step 5 with the speaker label features to obtain the source speaker text features.

[0116] Step 7: Input the source speaker text features output in Step 6 into the multi-head attention module of the source speaker decoder.

[0117] Step 8: Input the source speaker audio features (such as Mel spectrogram) into the decoder prenet, concatenate them with the positional encoding features, and then input them into the masked multi-head attention module of the source speaker decoder.

[0118] Step 9: Input the features output in Step 8 into the residual and normalization layer for processing, and then input them into the multi-head attention module of the source speaker decoder.

[0119] Step 10: Output after the alignment operation of the multi-head attention module.

[0120] Step 11: Output the output of Step 10 through the residual and normalization layer, the forward feedback layer residual and normalization layer, and the linear layer.

[0121] Step 12: Perform matrix multiplication on the output of Step 11 and the output of Step 5 to output the source speaker language features.

[0122] In some embodiments, the network units and data flow related to the voice conversion method may be as Figure 3 shown, including steps S310 to S343.

[0123] Step S310: Input the source speaker text data (such as text sequence) into the source speaker feature determination module.

[0124] Step S311: The source speaker feature determination module outputs the source speaker language features.

[0125] Step S320: Input the source speaker audio features (such as Mel spectrogram) into the source speaker fundamental frequency extraction module.

[0126] Step S322: The source speaker fundamental frequency extraction module outputs the source speaker fundamental frequency features.

[0127] Step S330: Input the target speaker voice data into the graph speaker editor.

[0128] Step S331: The graph speaker editor outputs the target speaker features.

[0129] Step S340: Concatenate the source speaker language features and the source speaker fundamental frequency features and input them into the target speaker decoder.

[0130] Step S341: Input the target speaker features into the target speaker decoder.

[0131] Step S342: The target speaker decoder outputs the converted Mel spectrogram and inputs it to the vocoder.

[0132] Step S343: The vocoder outputs the converted target speech.

[0133] In some embodiments, obtaining the source speaker text feature by using a graph encoder may include: generating a syntax tree corresponding to the source speaker text data according to the source speaker text data; parsing the syntactic relationships between words in the syntax tree, using phonemes as nodes and the syntactic relationships between phonemes as edges to generate a syntax graph; based on a bidirectional gated recurrent unit network, bidirectionally encoding the syntactic relationships between two phonemes in the syntax graph to obtain relationship encoding features; where the relationship encoding features include a forward relationship encoding vector and a backward relationship encoding vector; calculating attention scores based on the graph attention network of the graph encoder and the relationship encoding features; encoding the source speaker phoneme data according to the attention scores to obtain the source speaker text feature.

[0134] Specifically, the graph encoder may include a text encoding unit and a relationship encoding unit.

[0135] The relationship encoding unit converts the syntax tree of the input source speaker text data into a syntax graph that describes the global relationships between the involved input data.

[0136] By performing dependency syntactic analysis on the source speaker text data, a syntax tree is obtained. Dependency syntactic analysis is a natural language processing technique whose purpose is to identify the dependency relationships between individual words in a sentence. In natural language processing, dependency syntactic analysis can help understand the semantic structure of a sentence and better perform tasks such as text analysis, information extraction, and speech recognition.

[0137] Specifically, the dependency analysis module identifies the dependency relationships between words in the source speaker text data and constructs a dependency tree. The dependency relationship between words refers to the dependency relationship of one word on another word, and this dependency relationship can be a lexical, syntactic, or semantic relationship (such as subject-predicate relationship, verb-object relationship, parallel relationship, etc.). Constructing a dependency tree means organizing all words into a tree-structured data, with each word as a node and the dependency relationships between words as tree edges.

[0138] The syntax graph is an extension of the syntax tree, turning the unidirectional connection into a bidirectional connection by adding reverse connections. In addition, self-loop edges are introduced, and each word has a specific label. In this way, the words (or phonemes) in a sentence can be represented by nodes, and their connection relationships are represented by edges. Through the bidirectional connection, a word can directly receive and send information to any other word, regardless of whether they are directly related.

[0139] To model the relationship between two nodes (phonemes), the relationship between node pairs is described as the shortest relationship path between them. A recurrent neural network with GRU can be used to convert the relationship sequence into a distributed representation.

[0140] For example, the shortest relationship path sp between node i and node j i→j is represented as [sp1,..., sp t ,..., sp n+1 = [e(i, k1), e(k1, k2),..., e(k n , j)], where e(·, ·) represents the edge label, and k1:n are relay nodes. A bidirectional GRU is used for path relationship encoding:

[0141]

[0142]

[0143] where represents the forward relationship encoding vector, represents the backward relationship encoding vector, GRU f represents the forward GRU network, GRU b represents the backward GRU network, and the last hidden states of the forward and backward GRU networks are concatenated to form the final relationship encoding r ij = [s n+1 ; s0]. The final relationship encoding represents the linguistic relationship between two words (or phonemes). Constructing the relationship encoding provides a global view of how to collect and distribute information for the graph encoder model in speech conversion.

[0144] The graph encoder in the embodiments of this application incorporates the syntactic relationship encoding into the self-attention mechanism to indicate the relationship between characters (or phonemes). The purpose of the graph encoder is to convert the input character embedding sequence (or phoneme sequence) and the relationship encoding into the corresponding syntax-driven character embedding sequence (or phoneme sequence). By incorporating the explicit relationship representation between two nodes in the syntax graph into the graph attention calculation, a syntax-aware graph attention mechanism is formed, which can be abbreviated as syntax-aware graph attention. Further, the final encoding representation can also be calculated by stacking multiple syntax-aware graph attention and feed-forward layer modules. In each module, the character embedding sequence can be updated based on all other character embedding sequences and the corresponding relationship encoding.

[0145] In some embodiments, based on a bidirectional GRU network (bidirectional gated recurrent unit network), the syntactic relationship between two phonemes is bidirectionally encoded to obtain relationship encoding data, including: when the two phonemes belong to the same word, the syntactic relationship between the two phonemes is bidirectionally encoded based on the bidirectional GRU network and using the self-loop edge encoding algorithm; when the two phonemes belong to different words, the syntactic relationship between the words to which the two phonemes belong respectively is bidirectionally encoded based on the bidirectional GRU network.

[0146] In this embodiment, the basic unit of the sentences in the source speaker text data can be phoneme markers. The relationship encoding between words in NLP (Natural Language Processing) can be extended to the relationship encoding between phonemes. When the two phonemes belong to the same word, the self-loop edge encoding algorithm can be used to bidirectionally encode the syntactic relationship between the two phonemes; when the two phonemes belong to different words, the syntactic relationship between the words to which the two phonemes belong respectively is bidirectionally encoded based on the bidirectional GRU network. By extending the relationship encoding between words to the relationship encoding between phonemes, the encoding accuracy can be improved.

[0147] In some embodiments, bidirectionally encoding the syntactic relationship between two phonemes can be based on the shortest path between the two phonemes for relationship encoding.

[0148] In some embodiments, calculating the attention score based on the graph attention mechanism of the graph encoder and the relationship encoding vector includes: capturing the addressing relationship based on the basic syntactic content in the relationship encoding by the graph attention mechanism of the graph encoder; calculating the forward relationship deviation between phonemes according to the forward relationship encoding vector; controlling the backward relationship deviation between phonemes according to the backward relationship encoding vector; calculating the comprehensive deviation based on the forward relationship encoding vector and the backward relationship encoding vector; obtaining the attention score according to the addressing relationship, the forward relationship deviation, the backward relationship deviation, and the comprehensive deviation.

[0149] Specifically, in order to also encode the connection direction between phonemes when calculating the attention, first divide the relationship encoding vector r ij into a forward relationship encoding vector r i→j and a backward relationship encoding vector r j→i ; [r i→j ; r j→i = W rrij .

[0150] Then, use the syntax-aware graph attention mechanism to calculate the attention score, and the score is based on the phoneme representation and its bidirectional relationship representation as follows:

[0151]

[0152] Among them, (a) represents the addressing relationship capturing the basic syntax content. (b) represents calculating the forward relationship deviation between phonemes based on the forward relationship encoding vector. (c) represents controlling the backward relationship deviation between phonemes according to the backward relationship encoding vector; (d) represents the comprehensive deviation (encoding the general relationship deviation) obtained based on the forward relationship encoding vector and the backward relationship encoding vector.

[0153] In some embodiments, step S105 may include: extracting a speaker embedding vector from the target speaker's speech data using a speaker embedding extraction network; constructing an affinity graph based on the speaker embedding vector and the embedding vectors of the training samples of the speaker embedding extraction network; wherein, in the affinity graph, the embedding vectors are used as nodes and at least one nearest neighbor embedding vector determined by cosine similarity to each embedding vector is used as an edge; performing clustering processing on the nodes in the affinity graph based on a graph convolutional network to generate a target speaker cluster; and obtaining target speaker features based on the target speaker cluster.

[0154] Specifically, the speaker embedding extraction network can be trained in a supervised manner using historical speaker speech data as samples.

[0155] Among them, the affinity graph can be denoted as G. The affinity graph G=(V, E) is a relational graph, where V is the set of nodes in the graph and E is the set of edges in the graph. Here, the nodes represent the input sample data, and the edges represent the similarity between the input data. Based on the embedding vectors extracted from a pre-trained speaker embedding extraction network, each sample (or the embedding vector corresponding to the sample) is regarded as a node, and K nearest neighbor samples are found for each sample using cosine similarity. By connecting the neighbor nodes, an affinity graph containing all samples and the relationship of closeness and distance between samples is constructed.

[0156] By constructing the affinity graph, a relationship network of closeness and distance between the currently extracted speaker embedding vector and the sample embedding vectors of the training samples can be constructed. Based on the constructed affinity graph and performing clustering processing using a graph convolutional network, the accuracy of speaker feature extraction can be further improved.

[0157] In some embodiments, performing clustering processing on the nodes in the affinity graph based on a graph convolutional network to generate a target speaker cluster includes the following steps 1 to 4.

[0158] Step 1: Generate clustering proposals.

[0159] Specifically, by clustering the nodes in the affinity graph, multiple target speaker clustering proposals are generated. Among them, the target speaker clustering proposal is a subgraph of the affinity graph, which is generated based on supernodes, and the supernodes contain a small number of nodes that are closely related to each other.

[0160] In some embodiments, clustering the nodes in the affinity graph to generate multiple target speaker clustering proposals, including: obtaining multiple supernodes by setting the edge weights between nodes in the affinity graph; performing a clustering operation based on each supernode, using the centroid of each supernode as a node and the relationship between centroids as an edge to generate multiple target speaker clustering proposals.

[0161] Specifically, a set of supernodes can be generated by setting various thresholds (preliminary clustering) for the edge weights of the affinity graph. Although the samples in the same supernode may be from the same speaker, each speaker may correspond to multiple supernodes. Therefore, further clustering of the supernodes is required. In this way, a higher-level graph based on supernodes is constructed, with the centroid of the supernodes as nodes and the relationship between centroids as edges.

[0162] Step two: Clustering detection based on graph convolution.

[0163] Specifically, the clustering feature extraction unit based on the graph convolutional network extracts the clustering features of each target speaker clustering proposal, and identifies candidate target speaker clustering proposals from the multiple generated target speaker clustering proposals. Through the detection of the clustering detection unit, the possibility of whether the proposal is correct can be determined, and high-quality target speaker clustering proposals can be identified as candidate target speaker clustering proposals.

[0164] Exemplarily, the graph convolutional network is used to extract the features of each proposal. The calculation formula of each GCN (Graph Convolutional Network) layer can be expressed as:

[0165]

[0166] where, represents the pair-wise angle matrix. P i represents the target speaker clustering proposal, H (k) (P i ) represents the feature vector of the k-th layer. θ (k) represents the matrix for transforming vectors, and σ is the ReLU non-linear activation function. During the training process, GCN optimizes by minimizing the mean square error (MSE) objective function between the true score and the predicted score. Then, high-quality clusters are selected from the multiple generated target speaker clustering proposals as candidate target speaker clustering proposals.

[0167] Step three: Clustering segmentation based on graph convolution.

[0168] Specifically, the clustering and segmentation unit based on graph convolution predicts the probability values of each node in the candidate target speaker clustering proposal, and removes the nodes with probability values less than the preset threshold as abnormal nodes from the nodes of the candidate target speaker clustering proposal.

[0169] More specifically, even if the GCNs identify high-quality clusters, they may contain some outliers (abnormal nodes) that need to be removed. The outlier can be excluded from the clustering proposal by using the clustering and segmentation unit. During the clustering detection process, each sample in the candidate clustering proposal consists of a set of feature vectors (representing a node), an affinity matrix, and a binary vector indicating whether the node is positive. Then, the node binary cross-entropy is used as the loss function to train the clustering and segmentation unit. The clustering and segmentation unit outputs a predicted probability value for each node to indicate the likelihood of it being a true member rather than an abnormal node. In the clustering and segmentation unit, the abnormal nodes are removed from the clustering proposal.

[0170] Step 4: Overlap removal processing.

[0171] Specifically, each node is sorted according to the magnitude of the probability value of each node in the candidate target speaker clustering proposal, and at least one node with a probability value greater than the preset threshold is selected to generate the target speaker clustering.

[0172] More specifically, there should be no overlap between different clusters. In the "overlap removal" process, the scores predicted by the GCN are used and the nodes are sorted in descending order, and the nodes with the highest GCN scores are selected to form the clustering proposal, and the unlabeled data set is divided into appropriate clusters.

[0173] In the above embodiments, through the generated multiple target speaker clustering proposals, graph convolution-based clustering detection, clustering segmentation, and overlap removal processing are performed, which can improve the speaker clustering effect, improve the similarity of target speakers, and obtain high-quality target speaker clustering, thereby improving the output accuracy of the graph speaker encoder.

[0174] In some embodiments, the voice conversion method further includes assigning pseudo-labels to the target speaker clustering; optimizing the network parameters of the speaker embedding extraction network through a denoising loss function to denoise the pseudo-labels generated during the clustering process, and re-inputting the denoised pseudo-labels into the speaker embedding extraction network for iterative training.

[0175] By assigning pseudo-labels to the target speaker clustering, the clustering effect in a noisy environment can be improved, and the robustness of the speech to be converted in a noisy environment can be enhanced.

[0176] To avoid the speaker embedding extraction network from fitting noise data, the parameters of the speaker embedding extraction network can be updated through an improved loss function, and its formula is:

[0177]

[0178] Among them, B is the minimum batch size, indicating that the sample x i is classified as the posterior probability of the true label y i . is the posterior probability that the sample x i is classified as the predicted label. Indicates the predicted label of x i . α t ∈[0,1] is the and confidence weight at the t-th training iteration between them, which determines whether the loss function depends more on the true label or the predicted label. α t dynamically increases and its formula can be expressed as:

[0179] α t = α T ·(t / T) λ

[0180] Among them, α T ∈[0,1] represents the value of α t at the final iteration. The dynamically increasing confidence weight means that the loss tends to depend more on the predicted label because the prediction becomes more and more accurate. By dynamically learning the labels of the samples, the network can be prevented from overfitting incorrect samples.

[0181] To better understand the operation process of the graph speaker encoder in the above embodiments, reference can be made to Figure 4 as shown in Figure 4 which shows a schematic data flow diagram of extracting target speaker features based on the graph speaker encoder 400 in some embodiments.

[0182] In some embodiments, the relationship encoding unit in the graph encoder in the source speaker feature determination module obtains the syntax tree, and Stanza (pre-trained dependency parsing module) can be used. Relationship encoding is performed in two directions based on bidirectional GRU. N = 6 encoding blocks are used in the text encoding unit of the graph encoder. The multi-head attention in the graph attention unit in the graph encoder can be 4 heads or 8 heads. The input source speaker text data is first converted into a 256-dimensional character embedding sequence and then used as the input of the encoder.

[0183] In some embodiments, the source speaker concatenated features and the target speaker features are input into the target speaker decoder for decoding, which can pass through multiple processing layers, including an upsampling layer, a one-dimensional convolutional layer, a normalization layer, etc.

[0184] Specifically, the structure of the target speaker decoder can be referred to Figure 5 as shown Figure 5 Fig. 5 shows a schematic structural diagram of the data flow of the target speaker decoder 500 for obtaining target spectral data in some embodiments.

[0185] It should be understood that although Figure 1 each step in the flowchart of Figure 1 is shown sequentially according to the indication of the arrow, these steps are not necessarily executed sequentially in the order indicated by the arrow. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,

[0186] In a second aspect, an embodiment of the present disclosure provides a voice conversion device, as Figure 6 shown. The voice conversion device 600 includes: a source speaker data acquisition module 601, an audio feature alignment module 602, a source speaker fundamental frequency feature acquisition module 603, a target speaker feature acquisition module 604, a decoding module 605, and a conversion module 606.

[0187] The source speaker data acquisition module 601 is configured to obtain source speaker text features and source speaker audio features according to the source speaker voice data.

[0188] The audio feature alignment module 602 is configured to align the source speaker text features and the source speaker audio features by using a multi-head attention module to obtain audio mapping data representing the alignment relationship.

[0189] The source speaker fundamental frequency feature acquisition module 603 is configured to determine the source speaker fundamental frequency feature according to the source speaker audio feature.

[0190] The target speaker feature acquisition module 604 is configured to obtain target speaker voice data and extract target speaker features from the target speaker voice data.

[0191] The decoding module 605 is configured to input the source speaker language features, the source speaker fundamental frequency features, and the target speaker features into the target speaker decoder for decoding to obtain target spectral data.

[0192] The conversion module 606 is configured to process the target spectral data through a vocoder to convert it into target speech.

[0193] In some embodiments, the source speaker data acquisition module 601 may include a source speaker phoneme data acquisition unit, a source speaker text feature acquisition unit, and a source speaker audio feature acquisition unit.

[0194] The source speaker phoneme data acquisition unit is configured to obtain source speaker phoneme data according to the source speaker text data.

[0195] The source speaker text feature acquisition unit is configured to use a graph encoder to perform syntactic relationship encoding on the source speaker text data according to the syntax graph of the source speaker text data, obtain relationship encoding features, and perform text encoding on the source speaker phoneme data by using the relationship encoding features based on a graph attention mechanism to obtain source speaker text features.

[0196] The source speaker audio feature acquisition unit is configured to preprocess the source speaker voice data to obtain source speaker audio features.

[0197] In some embodiments, the source speaker text feature acquisition unit may include a syntax graph generation subunit, a relationship encoding feature acquisition subunit, and an encoding subunit.

[0198] The syntax graph generation subunit is configured to generate a syntax tree corresponding to the source speaker text data according to the source speaker text data; parse the syntactic relationships between words in the syntax tree, use phonemes as nodes, and the syntactic relationships between phonemes as edges to generate a syntax graph.

[0199] The bidirectional encoding subunit is configured to perform bidirectional encoding on the syntactic relationships between two phonemes in the syntax graph based on a bidirectional gated recurrent unit network to obtain relationship encoding features. Among them, the relationship encoding features include a forward relationship encoding vector and a backward relationship encoding vector.

[0200] The encoding subunit is configured to calculate attention scores based on the graph attention network of the graph encoder and the relationship encoding features; encode the source speaker phoneme data according to the attention scores to obtain source speaker text features.

[0201] In some embodiments, the bidirectional encoding subunit may perform bidirectional encoding in different cases. When two phonemes belong to the same word, the bidirectional encoding unit is configured to perform bidirectional encoding on the syntactic relationships between the two phonemes based on a bidirectional gated recurrent unit network and use a self-loop edge encoding algorithm; when two phonemes belong to different words, the bidirectional encoding unit is configured to perform bidirectional encoding on the syntactic relationships between the words to which the two phonemes belong respectively based on a bidirectional gated recurrent unit network.

[0202] In some embodiments, the decoding module 605 may include a feature splicing unit and a decoding unit.

[0203] The feature splicing unit is used to splice the source speaker's language features and the source speaker's fundamental frequency features to obtain the source speaker's spliced features.

[0204] The decoding unit is used to input the source speaker's spliced features and the target speaker's features into the target speaker decoder for decoding to obtain the target spectral data.

[0205] In some embodiments, the target speaker feature acquisition module 604 may include a speaker embedding vector extraction unit, an affinity graph construction unit, a target speaker clustering unit, and a speaker feature acquisition unit.

[0206] The speaker embedding vector extraction unit is used to extract the speaker embedding vector from the target speaker's speech data by using the speaker embedding extraction network.

[0207] The affinity graph construction unit is used to construct an affinity graph based on the speaker embedding vector and the embedding vectors of the training samples of the speaker embedding extraction network. Among them, the embedding vectors are used as nodes in the affinity graph, and at least one nearest neighbor embedding vector determined by the cosine similarity to each embedding vector is used as an edge.

[0208] The target speaker clustering unit is used to perform clustering processing on the nodes in the affinity graph based on the graph convolutional network to generate the target speaker clustering.

[0209] The speaker feature acquisition unit is used to obtain the target speaker's features based on the target speaker clustering.

[0210] In some embodiments, the target speaker clustering unit may include a clustering proposal sub-unit, a clustering detection sub-unit, a clustering segmentation sub-unit, and an overlapping removal processing sub-unit.

[0211] The clustering proposal sub-unit is used to cluster the nodes in the affinity graph to generate multiple target speaker clustering proposals.

[0212] The clustering detection sub-unit is used to extract the clustering features of each target speaker clustering proposal based on the clustering detection unit of the graph convolutional network, and identify the candidate target speaker clustering proposals from the multiple generated target speaker clustering proposals according to the clustering features.

[0213] The clustering segmentation sub-unit is used to predict the probability values of the nodes in the candidate target speaker clustering proposal based on the clustering segmentation unit of the graph convolution, and remove the nodes with probability values less than the preset threshold as abnormal nodes from the nodes of the candidate target speaker clustering proposal.

[0214] The overlapping removal processing sub-unit is used to sort the nodes according to the magnitude of the probability values of the nodes in the candidate target speaker clustering proposal, and select at least one node with a probability value greater than the preset threshold to generate the target speaker clustering.

[0215] In some embodiments, the voice conversion device 600 may further include a label denoising module. The label denoising module is used to optimize the network parameters of the speaker embedding extraction network through a denoising loss function, so as to denoise the pseudo-labels generated during the clustering process, and re-enter the denoised pseudo-labels into the speaker embedding extraction network for iterative training.

[0216] For the specific limitations of the voice conversion device 600, reference may be made to the limitations on the voice conversion method in the foregoing text, which will not be elaborated here. Each module in the above voice conversion device 600 can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0217] In a third aspect, an embodiment of the present disclosure provides a computer device, which may be a terminal, and its internal structure diagram may be as Figure 7 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the corresponding steps in the voice conversion method in any embodiment of the first aspect of the present disclosure. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0218] Those skilled in the art can understand that Figure 7 the structure shown in

[0219] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0220] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0221] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0222] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A voice conversion method, characterized in that, The method includes: Obtaining source speaker text features and source speaker audio features according to source speaker speech data; Using a multi-head attention module to align the source speaker text features and source speaker audio features to obtain audio mapping data for representing the alignment relationship; Obtaining source speaker language features based on matrix multiplication processing of the audio mapping data and the source speaker text features; Determining source speaker fundamental frequency features according to the source speaker audio features; Obtaining target speaker speech data and extracting target speaker features from the target speaker speech data; Inputting the source speaker language features, the source speaker fundamental frequency features, and the target speaker features into a target speaker decoder for decoding to obtain target spectral data; Converting the target spectral data into target speech through vocoder processing.

2. The method according to claim 1, wherein The obtaining source speaker text features and source speaker audio features according to source speaker speech data includes: Obtaining source speaker text data and source speaker voice data according to source speaker speech data; Obtaining source speaker phoneme data according to the source speaker text data; Using a graph encoder to perform grammatical relationship encoding on the source speaker text data according to the grammar graph of the source speaker text data to obtain relationship encoding features, and based on a graph attention mechanism, using the relationship encoding features to perform text encoding on the source speaker phoneme data to obtain the source speaker text features; Performing preprocessing on the source speaker voice data to obtain the source speaker audio features.

3. The method according to claim 2, wherein The using a graph encoder to perform grammatical relationship encoding on the source speaker text data according to the grammar graph of the source speaker text data to obtain relationship encoding features, and based on a graph attention mechanism, using the relationship encoding features to perform text encoding on the phoneme data to obtain source speaker text features includes: Generating a grammar tree corresponding to the source speaker text data according to the source speaker text data; Parsing the grammatical relationships between words in the grammar tree, using phonemes as nodes and the grammatical relationships between phonemes as edges to generate a grammar graph; Based on a bidirectional gated recurrent unit network, performing bidirectional encoding on the grammatical relationships between two phonemes in the grammar graph to obtain the relationship encoding features; wherein, the relationship encoding features include a forward relationship encoding vector and a backward relationship encoding vector; Calculating attention scores based on the graph attention network of the graph encoder and the relationship encoding features; Encoding the source speaker phoneme data according to the attention scores to obtain the source speaker text features.

4. The method according to claim 3, characterized in that, The based on a bidirectional gated recurrent unit network, performing bidirectional encoding on the grammatical relationships between two phonemes in the grammar graph to obtain relationship encoding features includes: When two phonemes belong to the same word, performing bidirectional encoding on the grammatical relationships between the two phonemes based on a bidirectional gated recurrent unit network and using a self-loop edge encoding algorithm; When two phonemes belong to different words, performing bidirectional encoding on the grammatical relationships between the words to which the two phonemes respectively belong based on a bidirectional gated recurrent unit network.

5. The method according to claim 1, wherein Inputting the source speech language feature, the source speaker fundamental frequency feature, and the target speaker feature into a target speaker decoder for decoding to obtain target spectral data includes: Concatenating the source speaker language feature and the source speaker fundamental frequency feature to obtain a source speaker concatenated feature; Inputting the source speaker concatenated feature and the target speaker feature into the target speaker decoder for decoding to obtain the target spectral data.

6. The method according to claim 1, characterized in that Obtaining target speaker speech data and extracting target speaker features from the target speaker speech data includes: Extracting a speaker embedding vector from the target speaker speech data by using a speaker embedding extraction network; Constructing an affinity graph based on the speaker embedding vector and the embedding vectors of the training samples of the speaker embedding extraction network; wherein, in the affinity graph, the embedding vectors are used as nodes and at least one nearest neighbor embedding vector determined by cosine similarity for each embedding vector is used as an edge; Performing clustering processing on the nodes in the affinity graph based on a graph convolutional network to generate target speaker clusters; Obtaining the target speaker features based on the target speaker clusters.

7. The method according to claim 6, wherein Performing clustering processing on the nodes in the affinity graph based on a graph convolutional network to generate target speaker clusters includes: Performing clustering on the nodes in the affinity graph to generate a plurality of target speaker cluster proposals; Extracting the clustering features of each of the target speaker cluster proposals based on the clustering detection unit of the graph convolutional network, and identifying candidate target speaker cluster proposals from the generated plurality of target speaker cluster proposals according to the clustering features; Predicting the probability values of the nodes in the candidate target speaker cluster proposals based on the clustering segmentation unit based on graph convolution, and removing the nodes with probability values less than a preset threshold as abnormal nodes from the nodes of the candidate target speaker cluster proposals; Sorting the nodes according to the magnitudes of the probability values of the nodes in the candidate target speaker cluster proposals, and selecting at least one node with a probability value greater than the preset threshold to generate the target speaker clusters.

8. The method according to claim 6, wherein The method further includes: Assigning pseudo-labels to the target speaker clusters; Optimizing the network parameters of the speaker embedding extraction network through a noise reduction loss function to perform noise reduction on the pseudo-labels generated during the clustering process, and re-inputting the denoised pseudo-labels into the speaker embedding extraction network for iterative training.

9. A voice conversion device, characterized in that, The apparatus includes: A source speaker data acquisition module, configured to obtain a source speaker text feature and a source speaker audio feature according to source speaker speech data; An audio feature alignment module, configured to align the source speaker text feature and the source speaker audio feature by using a multi-head attention module to obtain audio mapping data for representing the alignment relationship; A source speaker fundamental frequency feature acquisition module, configured to determine a source speaker fundamental frequency feature according to the source speaker audio feature; A target speaker feature acquisition module, configured to obtain target speaker speech data and extract target speaker features from the target speaker speech data; A decoding module, configured to input the source speaker language feature, the source speaker fundamental frequency feature, and the target speaker feature into a target speaker decoder for decoding to obtain target spectral data; A conversion module, configured to process and convert the target spectral data through a vocoder into target speech.

10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.