Voice data processing method and device, electronic equipment and readable medium

CN119559957BActive Publication Date: 2026-07-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2024-11-29
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In existing speech conversion technologies, the training data for training models includes information such as semantics, timbre, and rhythm. This results in the converted audio retaining the timbre of the original speaker, affecting the similarity between the converted audio and the target speaker's voice.

Method used

By acquiring source speech data from the source speaker and speaker information from the target speaker, feature extraction is performed. Using content representations and posterior probability vectors from a semantic dictionary, the content representation of the speech frame is determined and speech conversion is performed to reduce timbre leakage and improve the similarity of the target speaker's voice.

Benefits of technology

By using a weighted combination of semantic dictionaries, timbre leakage is reduced and the similarity between the converted audio and the target speaker's voice is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559957B_ABST
    Figure CN119559957B_ABST
Patent Text Reader

Abstract

The application provides a voice data processing method and device, electronic equipment and readable medium, comprising: obtaining source voice data of a source speaker and speaker information of a target speaker; performing feature extraction on the source voice data to obtain a posterior probability vector containing the K speech units to which the speech frames in the source voice data belong; determining a content re-expression of the speech frames in the source voice data according to the content expression corresponding to the K speech units in a semantic dictionary and the posterior probability vector, the semantic dictionary containing the content expression corresponding to the K speech units, and the content expression being obtained by statistical calculation according to semantic expressions in voice data from at least two speakers and posterior probabilities; and performing voice conversion according to the content re-expression of the speech frames in the source voice data and the speaker information to obtain target voice data of the target speaker. The method can reduce timbre leakage in the converted audio and improve the sound similarity of the converted audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, electronic device, and readable medium for processing voice data. Background Technology

[0002] With the development of computer technology, speech conversion technology has developed rapidly. The task of speech conversion (VC) is to convert the semantic representation of source speech into the speech of the target speaker while preserving the language content.

[0003] In related technologies, speech conversion services extract semantic content representations from trained speech from a trained model, and then convert the speech audio of the target speaker based on these content representations.

[0004] However, in such schemes, the training data of the training model contains information such as semantics, timbre, and rhythm. The semantic content representation also contains timbre information. As a result, the timbre information in the semantic content representation causes the converted audio to retain some of the timbre of the source speaker, thus affecting the similarity between the converted audio and the target speaker's voice. Summary of the Invention

[0005] To address the aforementioned technical issues, this application provides a method, apparatus, electronic device, and readable medium for processing voice data, thereby reducing the timbre leaked into the final generated target voice data and improving the similarity between the converted audio and the target speaker's voice.

[0006] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0007] According to one aspect of the embodiments of this application, a method for processing voice data is provided, including:

[0008] Obtain the source speech data of the source speaker and the speaker information of the target speaker;

[0009] Feature extraction is performed on the source speech data to obtain a posterior probability vector containing the speech frames in the source speech data belonging to K speech units, wherein the speech unit includes at least one of phonemes and soft speech units, and K is an integer greater than 1;

[0010] Based on the content representations corresponding to the K speech units in the semantic dictionary and the posterior probability vector, the content re-representation of the speech frames in the source speech data is determined. The semantic dictionary contains the content representations corresponding to the K speech units. The content representation of each speech unit in the semantic dictionary is obtained by statistical calculation based on the semantic representations and posterior probabilities in the speech data from at least two speakers.

[0011] Based on the content of the speech frames in the source speech data and the speaker information, speech conversion is performed to obtain the target speech data of the target speaker.

[0012] According to one aspect of the embodiments of this application, a voice data processing apparatus is provided, comprising:

[0013] The information acquisition module is configured to acquire the source speech data of the source speaker and the speaker information of the target speaker;

[0014] The feature extraction module is configured to extract features from the source speech data to obtain a posterior probability vector containing the speech frames in the source speech data belonging to K speech units, wherein the speech unit includes at least one of phonemes and soft speech units, and K is an integer greater than 1.

[0015] The expression conversion module is configured to determine the content re-expression of the speech frame in the source speech data based on the content expression corresponding to the K speech units in the semantic dictionary and the posterior probability vector. The semantic dictionary contains the content expression corresponding to the K speech units, and the content expression of each speech unit in the semantic dictionary is obtained by statistical calculation based on the semantic expression and posterior probability in the speech data from at least two speakers.

[0016] The speech conversion module is configured to perform speech conversion based on the content of the speech frames in the source speech data and the speaker information to obtain the target speech data of the target speaker.

[0017] In some embodiments of this application, based on the above technical solutions, the semantic dictionary includes a global dictionary, and the speech conversion module is further configured to: acquire the semantic expression of each speech frame in each dictionary speech data of the dictionary audio set and the posterior probability of each speech frame belonging to the Kth speech unit, wherein the dictionary audio set contains speech data from at least two speakers; statistically analyze the posterior probability of the extracted speech frame belonging to the Kth speech unit to obtain a first statistic of the Kth speech unit; obtain the semantic expression of the Kth speech unit based on the statistical result of the product of the semantic expression of each speech frame and the posterior probability of each speech frame belonging to the Kth speech unit and the first statistic; and determine the global dictionary based on the semantic expression of the Kth speech unit and the semantic expressions of the other K-1 speech units among the K speech units.

[0018] In some embodiments of this application, based on the above technical solutions, the semantic dictionary includes a global dictionary, and the speech conversion module is specifically configured to: input the source speech data into a pre-trained first feature extraction model, the first feature extraction model being used to calculate the posterior probability that each speech frame in the source speech data belongs to the K phonemes respectively; and obtain the output result of the bottleneck layer in the first feature extraction model as the posterior probability vector of the source speech data.

[0019] In some embodiments of this application, based on the above technical solutions, the speech conversion module is specifically configured to: input the source speech data into a pre-trained second feature extraction model, wherein the second feature extraction model is used to calculate the posterior probability that each speech frame in the source speech data belongs to the K soft speech units; and use the output result of the specified converter layer in the second feature extraction model as the posterior probability vector of the source speech data.

[0020] In some embodiments of this application, based on the above technical solutions, the content re-expression of the source speech data is generated based on the global semantic dictionary, and the speaker information includes the prompt speech data of the target speaker; the speech conversion module is specifically configured to: obtain the content re-expression of the prompt speech data based on the global semantic dictionary; input the content re-expression of the speech frame in the source speech data, the content re-expression of the prompt speech data, and the encoded tokens of the source speech data and the prompt speech into the first speech generation model to generate a target decoder token; and perform audio decoding based on the target decoder token to obtain the target speech data of the target speaker.

[0021] In some embodiments of this application, based on the above technical solutions, the semantic dictionary further includes a speaker dictionary, and the speech conversion module is further configured to: statistically analyze the posterior probability of each speaker's speech frame belonging to the Kth speech unit, and obtain a second statistic corresponding to each speaker for the Kth speech unit; for each speaker, based on the statistical result of the product of the semantic expression of each speech frame and the posterior probability of each speech frame belonging to the Kth speech unit and the second statistic, obtain the semantic expression corresponding to each speaker for the Kth speech unit; and determine the speaker dictionary corresponding to each speaker based on the semantic expression corresponding to each speaker for the Kth speech unit and the semantic expressions corresponding to each speaker for the other K-1 speech units among the K speech units.

[0022] In some embodiments of this application, based on the above technical solutions, the expression conversion module is specifically configured to: multiply the content representations corresponding to the K speech units in the global semantic dictionary and the speaker dictionary by the posterior probability vector to obtain the global content re-representation and the speaker content re-representation; calculate the weighted sum of the global content re-representation and the speaker content re-representation to obtain the content re-representation of the speech frames in the source speech data;

[0023] The speech conversion module is specifically configured to: input the content of the speech frame in the source speech data and the speaker identifier of the target speaker in the speaker information into the second speech generation model for speech conversion, so as to obtain the target speech data of the target speaker.

[0024] In some embodiments of this application, based on the above technical solutions, the expression conversion module is specifically configured to: obtain the source content expression of the source speech data; calculate the weighted sum of the source content expression, the global content re-expression, and the speaker content re-expression to obtain the content re-expression of the speech frames in the source speech data.

[0025] In some embodiments of this application, based on the above technical solutions, the speech conversion module is specifically configured to: take the content re-expression of the speech frames in the source speech data and the embedded expression of the speaker information as conditions, and compare them with a specified spectrum containing noise. Figure 1 The third speech generation model is used to predict the spectrum of the input, and the resulting spectrogram is obtained. The third speech generation model contains embedded information of time step and speaker information. The resulting spectrogram is then converted into speech by a pre-trained vocoder to obtain the target speech data of the target speaker.

[0026] In some embodiments of this application, based on the above technical solutions, the speech conversion module is further configured to: extract features from the training speech data based on the semantic dictionary to obtain a content representation of the training speech data; concatenate the content representation with a specified spectrogram containing noise to form training data, wherein the specified spectrogram is noise-added according to the time step; train the third model to be trained using the training data to obtain the third speech generation model, wherein the third model to be trained includes a feature linear modulation layer, a convolutional layer, and a diffusion transform layer, the feature linear modulation layer includes information about the time step, and the convolutional layer and the diffusion transform layer include speaker embedding information of the speaker of the training speech data.

[0027] According to one aspect of the embodiments of this application, an electronic device is provided, the electronic device comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform a voice data processing method as described above by executing the executable instructions.

[0028] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the voice data processing method as described in the above technical solutions.

[0029] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the voice data processing method provided in the various optional implementations described above.

[0030] In embodiments of this application, a speech conversion device acquires source speech data from a source speaker and speaker information from a target speaker. Then, it performs feature extraction on the source speech data to obtain a posterior probability vector containing the speech frames in the source speech data belonging to K speech units, where each speech unit includes at least one of phonemes and soft speech units, and K is an integer greater than 1. Based on the content representations corresponding to the K speech units in a semantic dictionary and the posterior probability vector, it determines the content re-representation of the speech frames in the source speech data. The semantic dictionary contains the content representations corresponding to the K speech units, and the content representation of each speech unit in the semantic dictionary is obtained by statistical calculation based on semantic representations and posterior probabilities from speech data from at least two speakers. Finally, it performs speech conversion based on the content re-representation of the speech frames in the source speech data and the speaker information to obtain the target speech data of the target speaker. In this way, the source speech data is converted using a semantic dictionary, which is constructed by weighted combination of the semantic expressions of different speakers. This results in the re-expression of content unrelated to the speaker, thereby reducing the timbre leaked into the final generated target speech data and improving the similarity between the converted audio and the target speaker's voice.

[0031] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0032] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0033] Figure 1 The voice data processing method described in this application is applied to the system architecture of a cloud host platform.

[0034] Figure 2 This is a flowchart of a voice data processing method according to an embodiment of this application.

[0035] Figure 3 This is a flowchart of a voice data processing method according to an embodiment of this application.

[0036] Figure 4 This is a schematic structural diagram of the PPG extractor in the embodiments of this application.

[0037] Figure 5 This is a schematic structural diagram of the Hubert extractor in the embodiments of this application.

[0038] Figure 6 This is a schematic diagram of the conversion process based on a large language model in the embodiments of this application.

[0039] Figure 7 This is a schematic flowchart illustrating the speech conversion using a second speech generation model in the embodiments of this application.

[0040] Figure 8 This is a schematic diagram of the diffusion model in an embodiment of this application.

[0041] Figure 9 A schematic block diagram of the speech data processing apparatus in an embodiment of this application is shown.

[0042] Figure 10 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation

[0043] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0044] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0045] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve the predetermined function, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement at least one module or unit. Furthermore, each module or unit can be part of an overall module or unit that includes the functions of that module or unit.

[0046] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in at least one hardware module or integrated circuit, or in different network and / or processor devices and / or microcontroller devices.

[0047] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0048] It should be understood that the solution of this application can be applied to speech conversion systems, specifically in the speech conversion services provided by such systems. Specifically, the task of speech conversion includes converting the semantic representation of source speech into the speech of the target speaker while preserving the language content. Specifically, it typically involves converting source speech into the speech of the target speaker, or directly generating the target speaker's speech based on text information. The solution of this application uses semantic representations from different speakers to construct an offline global semantic dictionary. Each entry in the dictionary is a stable and timbre-neutral representation of a certain phoneme class. In the speech conversion task, each content feature from the source speaker can be re-represented as a weighted combination of entries in the semantic dictionary. The resulting new content representation contains stable semantic information and preserves contextual information. Furthermore, for one-to-many or many-to-many speech conversion tasks, it is feasible to construct a speaker-dependent semantic dictionary for a single target speaker, and the resulting new content representation will contain target timbre information beneficial to the conversion task. With the development of computer technology, speech conversion technology has rapidly advanced. The task of speech conversion (VC) includes converting the semantic representation of source speech into the speech of the target speaker while preserving the language content. In related technologies, speech conversion services extract semantic content representations from trained models to convert the target speaker's audio. However, in such schemes, because the training data of the model contains semantic, timbre, and rhythm information, and the semantic content representation also contains timbre information, the timbre information in the semantic content representation causes the converted audio to retain some of the source speaker's timbre, thus affecting the similarity between the converted audio and the target speaker's voice.

[0049] Based on this, the technical solution of this application proposes a voice data processing scheme. Specifically, please refer to... Figure 1The voice data processing method according to the embodiments of this application, applied to a voice conversion system, can mainly include a terminal device 110, a network 120, and a server 130. The terminal device 110 can include smartphones, tablets, laptops, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. The server 130 can be a server providing various services; it can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The network 120 can be a communication medium of various connection types capable of providing a communication link between the terminal device 110 and the server 130, such as a wired communication link or a wireless communication link.

[0050] Depending on the implementation requirements, the system architecture in this application embodiment can have any number of playback terminals, networks, and servers. For example, server 130 can be a server group composed of multiple server devices. In addition, the technical solutions provided in this application embodiment can be applied to terminal device 110, or to server 130, or can be implemented jointly by terminal device 110 and server 130. This application does not impose any special limitations on this.

[0051] like Figure 1 As shown, the speech conversion system in this application can be deployed in server 130. Server 130 obtains source speech data of the source speaker and speaker information of the target speaker from terminal device 110. Then, server 130 performs feature extraction on the source speech data to obtain a posterior probability vector containing the speech frames in the source speech data belonging to K speech units, wherein the speech unit includes at least one of phonemes and soft speech units, and K is an integer greater than 1. Server 130 then determines the content re-expression of the speech frames in the source speech data based on the content expressions corresponding to the K speech units in the semantic dictionary and the posterior probability vector, wherein the semantic dictionary contains the content expressions corresponding to the K speech units, and the content expression of each speech unit in the semantic dictionary is obtained by statistical calculation based on the semantic expressions and posterior probabilities from the speech data of at least two speakers. Finally, server 130 performs speech conversion based on the content re-expression of the speech frames in the source speech data and the speaker information to obtain the target speech data of the target speaker.

[0052] The implementation details of the technical solutions in the embodiments of this application are described in detail below: Figure 2A flowchart illustrating a voice data processing method according to an embodiment of this application is shown. This voice data processing method can be executed by a device with computing processing capabilities, such as a server or a terminal device. The voice data processing method of this application will be described below from the perspective of a speech conversion system on a server. (Refer to...) Figure 2 As shown, the method for processing this voice data includes at least steps S210 to S240, which are described in detail below:

[0053] Step S210: Obtain the source speech data of the source speaker and the speaker information of the target speaker.

[0054] The speech-to-speech device acquires source speech data from the source speaker and speaker information from the target speaker. Speaker information is typically the speaker identifier of the target speaker. Specifically, the speech-to-speech device encodes each speaker it can convert and uses the speaker identifier to determine which speaker's voice to convert to. Source speech data is the audio data of the source speaker, such as the source speaker's speech or a song.

[0055] Step S220: Perform feature extraction on the source speech data to obtain a posterior probability vector containing the speech frames in the source speech data belonging to K speech units, wherein the speech unit includes at least one of phonemes and soft speech units, and K is an integer greater than 1.

[0056] The speech conversion device extracts features from the source speech data. This is done using a dedicated feature extractor. The feature extractor can be a pre-trained machine learning model, such as a supervised-trained semantic posterior graph (PPG) model, or a self-supervised machine learning model, such as the Hubert model. The speech conversion device uses these models to obtain the corresponding content representation from the source speech data. Content representation is typically obtained on a per-speech-frame basis. The source speech data contains multiple speech frames, and each speech frame has its corresponding content representation extracted by the feature extractor. The speech conversion device also obtains a posterior probability vector indicating which speech frame belongs to which K speech units. It is important to emphasize that although this application uses a posterior probability vector as an example, other forms of probability data can be used in specific implementations, such as antecedent probability, likelihood probability, conditional probability, total probability, or prediction probability. Furthermore, depending on the specific implementation, data forms such as probability matrices or tensors can also be used. Specifically, the posterior probability vector contains the posterior probability of the speech frame belonging to each speech unit. The number of speech units varies depending on the specific type of speech unit. A speech unit is typically one of a phoneme or a soft speech unit. Both phonemes and soft speech units are the smallest processing units in audio processing. Phonemes are usually determined by summarizing the pronunciation rules of different languages, while soft speech units are obtained by segmenting specific speech data. The speech conversion device predetermines the number of speech units, for example, during the generation of a semantic dictionary. Assuming that the source speech data contains Q speech frames, the speech conversion device determines the probability that each of the Q speech frames belongs to a K speech unit, thus obtaining a posterior probability vector. Depending on the specific implementation, the speech conversion device can determine the probability of each speech frame belonging to a K speech unit separately, or it can determine only the semantic unit to which each speech frame is most likely to belong and the posterior probability of belonging to that speech unit.

[0057] Step S230: Determine the content re-expression of the speech frame in the source speech data based on the content expression corresponding to the K speech units in the semantic dictionary and the posterior probability vector. The semantic dictionary contains the content expression corresponding to the K speech units, and the content expression of each speech unit in the semantic dictionary is obtained by statistical calculation based on the semantic expression and posterior probability in the speech data from at least two speakers.

[0058] Specifically, the speech conversion device obtains the content representation of speech frames in the source speech data based on the mapping relationship between the content representations of K speech units in the semantic dictionary and the posterior probability vector. This mapping relationship can be a suitable data computation relationship, such as addition, multiplication, or weighted calculation. The semantic dictionary contains the content representations corresponding to the K speech units. The content representations in the semantic dictionary are pre-obtained by the speech conversion device through weighted merging of speech data from at least two speakers. Therefore, this content representation is independent of the timbre of a specific speaker and only related to semantic content. It can be understood that the semantic dictionary can be a row of data containing K columns, each corresponding to the content representation of a semantic unit. The posterior probability vector is a column of data containing K rows, each corresponding to the posterior probability between a speech frame and a semantic unit. The product of the two yields the content representation of a speech frame in the source speech data based on the content representation statistically derived from the semantic dictionary. Combining the content representations of each speech frame in the source speech data yields the content representation of the source speech data. It can be understood that since the representations in the semantic dictionary are independent of timbre, the content representation of the source speech data is also independent of timbre and only related to semantics. In some embodiments, phonemes and soft speech units can be used simultaneously, for example, extracted and calculated separately to obtain semantic re-expressions, and finally the results are fused, or the posterior probabilities of the two can be combined and calculated with a semantic dictionary.

[0059] Step S240: Based on the content re-expression of the speech frames in the source speech data and the speaker information, perform speech conversion to obtain the target speech data of the target speaker.

[0060] Speech-to-speech devices can convert content re-expressed using a pre-trained transducer. The transducer receives speaker information as input to determine the timbre of the target speaker to be converted. The transducer can employ specialized machine learning models, such as variational inference and adversarial learning models, vocoders in large language models, or diffusion models.

[0061] In embodiments of this application, a speech conversion device acquires source speech data from a source speaker and speaker information from a target speaker. Then, it performs feature extraction on the source speech data to obtain a posterior probability vector containing the speech frames in the source speech data belonging to K speech units, where each speech unit includes at least one of phonemes and soft speech units, and K is an integer greater than 1. Based on the content representations corresponding to the K speech units in a semantic dictionary and the posterior probability vector, it determines the content re-representation of the speech frames in the source speech data. The semantic dictionary contains the content representations corresponding to the K speech units, and the content representation of each speech unit in the semantic dictionary is obtained by statistical calculation based on semantic representations and posterior probabilities from speech data from at least two speakers. Finally, it performs speech conversion based on the content re-representation of the speech frames in the source speech data and the speaker information to obtain the target speech data of the target speaker. In this way, the source speech data is converted using a semantic dictionary, which is constructed by weighted combination of the semantic expressions of different speakers. This results in the re-expression of content unrelated to the speaker, thereby reducing the timbre leaked into the final generated target speech data and improving the similarity between the converted audio and the target speaker's voice.

[0062] In embodiments of this application, a method for... Figure 2 Other detailed embodiments of the technical solution shown in the example are as follows: Figure 3 As shown, the semantic dictionary includes a global dictionary. In one embodiment of the speech data processing method of this application, the following steps may be included:

[0063] Step S310: Obtain the semantic representation of each speech frame in each dictionary speech data of the dictionary audio set and the posterior probability of each speech frame belonging to the Kth speech unit. The dictionary audio set contains speech data from at least two speakers.

[0064] Step S320: Statistically calculate the posterior probability of the extracted speech frame belonging to the Kth speech unit to obtain the first statistic of the Kth speech unit;

[0065] Step S330: Based on the statistical result of the product of the semantic expression of each speech frame and the posterior probability of each speech frame belonging to the Kth speech unit, and the first statistic, the semantic expression of the Kth speech unit is obtained.

[0066] Step S340: Determine the global dictionary based on the semantic expression of the Kth speech unit and the semantic expressions of the other K-1 speech units among the K speech units.

[0067] In this embodiment, the dictionary audio set is a pre-selected training speech set containing speech data from multiple speakers. Depending on the conversion target, this speech data can come from people of different genders or from different languages. The speech conversion device first acquires the semantic representation of each speech frame in each dictionary speech data in the dictionary audio set, as well as the posterior probability that each speech frame belongs to the Kth speech unit. It is understood that the semantic representation of each speech frame can be extracted during previous processing of other speech units, or it can be obtained from the dictionary audio set by an extractor. The method of extracting the semantic representation and posterior probability is the same as that used in the speech conversion process in the above embodiment, and the speech units are also the same. That is, if the extracted speech unit is a phoneme, then the speech units in the semantic dictionary are also phonemes. Taking a phoneme as an example, let X... i,j,t ∈R d×1 Where d is the dimension of the semantic representation, i = 1…S, j = 1…N i , t=1…T ij Let be the semantic representation of the j-th speech data of the i-th speaker in the t-th frame. Let represent the posterior probability that the current frame belongs to the k-th phoneme class. Then the first statistic n... k The following method can be used to calculate:

[0068]

[0069] The semantic representation m of the Kth speech unit k The calculation method is as follows:

[0070]

[0071] Finally, the global dictionary M g It can then be represented as follows:

[0072] M g = [m1,m2,...m K ]

[0073] Where K is the total number of phoneme classes or discrete speech units, and m1 to m k-1 These are the semantic expressions of the other K-1 speech units, which can be calculated by the speech conversion device in the previous process. Through the above method, the semantic expressions of speech frames from different speakers corresponding to a speech unit are comprehensively calculated during the calculation process to obtain the corresponding semantic expression, thus combining them into a global dictionary. This reduces the influence of individual speaker timbre on the semantic expressions in the global dictionary, which is beneficial to improving the generalization ability of the semantic expressions in the global dictionary.

[0074] Step S350: Obtain the source speech data of the source speaker and the speaker information of the target speaker.

[0075] Optionally, the implementation details of step S350 are the same as... Figure 2 The steps S210 shown are the same and will not be repeated here.

[0076] Step S360: Perform feature extraction on the source speech data to obtain a posterior probability vector containing the speech frames in the source speech data belonging to K speech units, wherein the speech unit includes at least one of phonemes and soft speech units, and K is an integer greater than 1.

[0077] Optionally, the implementation details of step S360 are the same as... Figure 2 The steps S220 shown are the same and will not be repeated here.

[0078] Step S370: Determine the content re-expression of the speech frame in the source speech data based on the content expression corresponding to the K speech units in the semantic dictionary and the posterior probability vector. The semantic dictionary contains the content expression corresponding to the K speech units, and the content expression of each speech unit in the semantic dictionary is obtained by statistical calculation based on the semantic expression and posterior probability in the speech data from at least two speakers.

[0079] Optionally, the implementation details of step S370 are the same as... Figure 2 The steps S230 shown are the same and will not be repeated here.

[0080] Step S380: Based on the content re-expression of the speech frames in the source speech data and the speaker information, perform speech conversion to obtain the target speech data of the target speaker.

[0081] Optionally, the implementation details of step S380 are the same as... Figure 2 The steps S240 shown are the same and will not be repeated here.

[0082] In embodiments of this application, a speech conversion device acquires source speech data of a source speaker and speaker information of a target speaker. Then, it performs feature extraction on the source speech data to obtain a posterior probability vector containing the speech frames in the source speech data belonging to K speech units, where each speech unit includes at least one of phonemes and soft speech units, and K is an integer greater than 1. Based on the content representations corresponding to the K speech units in a semantic dictionary and the posterior probability vector, it determines the content re-representation of the speech frames in the source speech data. The semantic dictionary contains the content representations corresponding to the K speech units, and the content representation of each speech unit in the semantic dictionary is obtained by statistical calculation based on semantic representations and posterior probabilities from speech data of at least two speakers. Finally, it performs speech conversion based on the content re-representation of the speech frames in the source speech data and the speaker information to obtain the target speech data of the target speaker. In this way, the source speech data is converted using a semantic dictionary, which is constructed by weighted combination of the semantic expressions of different speakers. This results in the re-expression of content unrelated to the speaker, thereby reducing the timbre leaked into the final generated target speech data and improving the similarity between the converted audio and the target speaker's voice.

[0083] In some embodiments of this application, based on the technical solutions of this application, during the process of acquiring the semantic representation of each speech frame in each dictionary speech data of the dictionary audio set and the posterior probability of each speech frame belonging to the Kth speech unit, the speech conversion device inputs the source speech data into a pre-trained first feature extraction model. The first feature extraction model is used to calculate the posterior probability of each speech frame in the source speech data belonging to the K phonemes, and then obtains the output result of the bottleneck layer in the first feature extraction model as the posterior probability vector of the source speech data. In this embodiment, the first feature extraction model can be implemented using PPG. Specifically, please refer to... Figure 4 , Figure 4 This is a schematic structural diagram of the PPG extractor in an embodiment of this application. Figure 4 As shown, the PPG extractor is trained using cross-entropy and center loss. The speech conversion device obtains the corresponding semantic representation from the output of the bottleneck layer of the PPG structure and the corresponding posterior probability from the output of the PPG.

[0084] In some embodiments of this application, based on the technical solutions of this application, during the process of acquiring the semantic representation of each speech frame in each dictionary speech data of the dictionary audio set and the posterior probability of each speech frame belonging to the Kth speech unit, the speech conversion device inputs the source speech data into a pre-trained second feature extraction model. The second feature extraction model is used to calculate the posterior probability of each speech frame in the source speech data belonging to the K soft speech units, and then uses the output result of the specified converter layer in the second feature extraction model as the posterior probability vector of the source speech data. In this embodiment, the first feature extraction model can be implemented using the Hubert model. Specifically, please refer to [link to specific documentation]. Figure 5 , Figure 5 This is a schematic structural diagram of the Hubert extractor in an embodiment of this application. Figure 5 As shown, in this embodiment, two fully connected layers are added after the backbone of the Hubert model. The speech conversion device uses a converter with a specified number of layers in the Hubert model to extract semantic features; for example, a seventh-layer converter is used to obtain semantic features. For the obtained semantic features, this application uses clustering to extract the semantic representations corresponding to soft speech units from the speech of different speakers.

[0085] In some embodiments of this application, based on the technical solutions of this application, the content re-expression of the source speech data is generated based on the global semantic dictionary, and the speaker information includes the prompt speech data of the target speaker; in the process of obtaining the target speech data of the target speaker by performing speech conversion based on the content re-expression of the speech frames in the source speech data and the speaker information, the speech conversion device obtains the content re-expression of the prompt speech data based on the global semantic dictionary, and then inputs the content re-expression of the speech frames in the source speech data, the content re-expression of the prompt speech data, and the encoded tokens of the source speech data and the prompt speech into the first speech generation model to generate a target decoder token, and finally performs audio decoding based on the target decoder token to obtain the target speech data of the target speaker. In the embodiments of this application, Large Language Models (LLM) are used as the model for performing speech conversion in the speech conversion device. Specifically, the first speech generation model is implemented using an LLM model. This model receives source speech data and a prompt speech data as input. The prompt speech data is usually data related to the target speaker. In some embodiments, the prompting voice data is the voice information of the target speaker on a portion of the source voice data, for example, the audio of the first three seconds of the source voice data. In this embodiment, the voice conversion device converts both the source voice information and the prompting information into a content representation based on a global dictionary. Specifically, please refer to... Figure 6 , Figure 6 This is a schematic diagram of the conversion process based on a large language model in an embodiment of this application. For example... Figure 6 As shown, the semantic representation TICR based on the global dictionary, along with the encoded tokens of the prompt and source speech data, is input into the large language model to obtain the output decoded token. The large language model in the figure adopts the Neural Codec Language Model. Then, the decoded token is transcoded by the decoder to obtain the output audio content.

[0086] In some embodiments of this application, based on the technical solutions of this application, the semantic dictionary further includes a speaker dictionary. The speech conversion device statistically analyzes the posterior probability of each speaker's speech frame belonging to the Kth speech unit, obtaining a second statistic corresponding to each speaker for the Kth speech unit. Then, for each speaker, based on the statistical result of the product of the semantic expression of each speech frame and the posterior probability of each speech frame belonging to the Kth speech unit, and the second statistic, the semantic expression corresponding to each speaker for the Kth speech unit is obtained. Finally, based on the semantic expression corresponding to each speaker for the Kth speech unit and the semantic expressions corresponding to each speaker for the other K-1 speech units in the Kth speech units, the speaker dictionary corresponding to each speaker is determined. In this embodiment, the speech conversion device also generates a speaker dictionary corresponding to each speaker. Specifically, the generation method of the speaker dictionary is similar to that of the global dictionary, the difference being that the statistical and calculation process is organized according to the speaker. Specifically, the second statistic n... ik The following method can be used to calculate:

[0087]

[0088] The Kth speech unit corresponds to the semantic expression m of each speaker. ik The calculation method is as follows:

[0089]

[0090] Finally, the Pronunciation Dictionary M S It can then be represented as follows:

[0091] M s =[m i1 m i2 ,...m iK ]

[0092] In some embodiments of this application, based on the technical solutions in this application, during the process of determining the content re-expression of the speech frame in the source speech data according to the content expression corresponding to the K speech units in the semantic dictionary and the posterior probability vector, the speech conversion device will multiply the content expression corresponding to the K speech units in the global semantic dictionary and the speaker dictionary and the posterior probability vector respectively to obtain the global content re-expression and the speaker content re-expression, and calculate the weighted sum of the global content re-expression and the speaker content re-expression to obtain the content re-expression of the speech frame in the source speech data.

[0093] Specifically, the restatement of the overall content and the restatement of the speaker's content can be determined in the following ways:

[0094]

[0095] in, To further express the overall content, The speaker reiterates the content, M g For the global dictionary, M S For the speaker dictionary, P i,j,t ∈R K×1 Let be the posterior probability vector.

[0096] In the process of obtaining target speech data of the target speaker by re-representing the content of speech frames in the source speech data and the speaker information, the speech conversion device inputs the re-representation of the content of speech frames in the source speech data and the speaker identifier of the target speaker in the speaker information into a second speech generation model for speech conversion to obtain the target speech data of the target speaker. In this embodiment, a variational inference with adversarial learning for end-to-end text-to-speech (VITS) model is used for speech conversion. For details, please refer to [link to relevant documentation]. Figure 7 , Figure 7 This is a schematic flowchart illustrating the speech conversion using a second speech generation model in an embodiment of this application. Figure 7 As shown, after the input source speech data is extracted, it is calculated with the global dictionary and the speaker dictionary respectively to obtain a semantically independent global content re-expression and a speaker-related speaker content re-expression. Then, the global content re-expression and the speaker content re-expression are weighted and summed, and the result is input into the VITS model for speech conversion to obtain the target speech data of the target speaker.

[0097] In some embodiments of this application, based on the technical solutions of this application, during the process of calculating the weighted sum of the global content re-expression and the speaker content re-expression to obtain the content re-expression of the speech frame in the source speech data, the speech conversion device acquires the source content expression of the source speech data. Then, it calculates the weighted sum of the source content expression, the global content re-expression, and the speaker content re-expression to obtain the content re-expression of the speech frame in the source speech data. In this embodiment, the source content expression of the source speech data is further added to the input data of the VITS model. For example... Figure 7 As shown by the dotted line, the speech conversion device performs a weighted summation of the source content representation, the global content re-representation, and the speaker content re-representation, and inputs the result into the VITS model for speech conversion to obtain the target speech data of the target speaker.

[0098] In some embodiments of this application, based on the technical solutions in this application, during the process of obtaining target speech data of the target speaker by performing speech conversion based on the content re-expression of speech frames in the source speech data and the speaker information, the speech conversion device will use the content re-expression of speech frames in the source speech data and the embedded expression of the speaker information as conditions, and compare them with a specified spectrum containing noise. Figure 1 The input is a third speech generation model to perform spectrum prediction, and the resulting spectrogram is obtained. The third speech generation model contains embedded information of time step and speaker information. Then, the resulting spectrogram is converted into speech by a pre-trained vocoder to obtain the target speech data of the target speaker.

[0099] In some embodiments of this application, based on the technical solutions in this application, the speech conversion device further performs feature extraction on the training speech data based on the semantic dictionary to obtain the content representation of the training speech data. Then, the content representation is concatenated with a specified spectrogram containing noise to form training data. The specified spectrogram is noise-added according to the time step. The training data is then used to train the third model to be trained to obtain the third speech generation model. The third model to be trained includes a feature linear modulation layer, a convolutional layer, and a diffusion transform layer. The feature linear modulation layer contains information about the time step, and the convolutional layer and the diffusion transform layer contain the speaker embedding information of the speaker in the training speech data.

[0100] In this embodiment, a diffusion model is used for the speech conversion process. For details, please refer to [link to relevant documentation]. Figure 8 , Figure 8 This is a schematic diagram of the diffusion model in an embodiment of this application. Figure 8As shown, the speech conversion device takes content re-expression and speaker information embedding as conditions, and compares them with a specified spectrum containing noise. Figure 1 The input to the third speech generation model is used for spectral prediction. Specifically, the specified spectrogram can be a Mel spectrogram, and the content re-expression and the Mel spectrogram are combined and input into the third speech generation model. The third speech generation model contains a feature linear modulation layer, a convolutional layer, and a diffusion transformer layer. The feature linear modulation layer embeds the time step used in the diffusion process, while the convolutional and diffusion transformer layers embed the speaker's identifier. During training, the third speech generation model outputs the predicted Mel spectrogram, while in the actual conversion process, it outputs the resulting spectrogram. The resulting spectrogram is then converted into target speech data of the target speaker by a pre-trained vocoder.

[0101] It should be noted that although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0102] The following describes the implementation of the apparatus of this application, which can be used to execute the voice data processing method in the above embodiments of this application. Figure 9 A schematic block diagram illustrating the composition of a voice data processing apparatus according to an embodiment of this application is shown. Figure 9 As shown, the voice data processing device 900 mainly includes:

[0103] Information acquisition module 910 is configured to acquire source speech data of the source speaker and speaker information of the target speaker;

[0104] The feature extraction module 920 is configured to extract features from the source speech data to obtain a posterior probability vector containing the speech frames in the source speech data belonging to K speech units, wherein the speech unit includes at least one of phonemes and soft speech units, and K is an integer greater than 1.

[0105] The expression conversion module 930 is configured to determine the content re-expression of the speech frame in the source speech data based on the content expression corresponding to the K speech units in the semantic dictionary and the posterior probability vector. The semantic dictionary contains the content expression corresponding to the K speech units, and the content expression of each speech unit in the semantic dictionary is obtained by statistical calculation based on the semantic expression and posterior probability in the speech data from at least two speakers.

[0106] The speech conversion module 940 is configured to perform speech conversion based on the content of the speech frames in the source speech data and the speaker information to obtain the target speech data of the target speaker.

[0107] In some embodiments of this application, based on the above technical solutions, the semantic dictionary includes a global dictionary, and the speech conversion module 940 is further configured to: acquire the semantic expression of each speech frame in each dictionary speech data of the dictionary audio set and the posterior probability of each speech frame belonging to the Kth speech unit, wherein the dictionary audio set contains speech data from at least two speakers; statistically analyze the posterior probability of the extracted speech frame belonging to the Kth speech unit to obtain a first statistic of the Kth speech unit; obtain the semantic expression of the Kth speech unit based on the statistical result of the product of the semantic expression of each speech frame and the posterior probability of each speech frame belonging to the Kth speech unit and the first statistic; and determine the global dictionary based on the semantic expression of the Kth speech unit and the semantic expressions of the other K-1 speech units among the K speech units.

[0108] In some embodiments of this application, based on the above technical solutions, the semantic dictionary includes a global dictionary, and the speech conversion module 940 is specifically configured to: input the source speech data into a pre-trained first feature extraction model, the first feature extraction model being used to calculate the posterior probability that each speech frame in the source speech data belongs to the K phonemes respectively; and obtain the output result of the bottleneck layer in the first feature extraction model as the posterior probability vector of the source speech data.

[0109] In some embodiments of this application, based on the above technical solutions, the speech conversion module 940 is specifically configured to: input the source speech data into a pre-trained second feature extraction model, wherein the second feature extraction model is used to calculate the posterior probability that each speech frame in the source speech data belongs to the K soft speech units; and use the output result of the specified converter layer in the second feature extraction model as the posterior probability vector of the source speech data.

[0110] In some embodiments of this application, based on the above technical solutions, the content re-expression of the source speech data is generated based on the global semantic dictionary, and the speaker information includes the prompt speech data of the target speaker; the speech conversion module 940 is specifically configured to: obtain the content re-expression of the prompt speech data based on the global semantic dictionary; input the content re-expression of the speech frame in the source speech data, the content re-expression of the prompt speech data, and the encoded tokens of the source speech data and the prompt speech into the first speech generation model to generate a target decoder token; and perform audio decoding based on the target decoder token to obtain the target speech data of the target speaker.

[0111] In some embodiments of this application, based on the above technical solutions, the semantic dictionary further includes a speaker dictionary, and the speech conversion module 940 is further configured to: statistically analyze the posterior probability of each speaker's speech frame belonging to the Kth speech unit, and obtain a second statistic corresponding to each speaker for the Kth speech unit; for each speaker, based on the statistical result of the product of the semantic expression of each speech frame and the posterior probability of each speech frame belonging to the Kth speech unit and the second statistic, obtain the semantic expression corresponding to each speaker for the Kth speech unit; and determine the speaker dictionary corresponding to each speaker based on the semantic expression corresponding to each speaker for the Kth speech unit and the semantic expressions corresponding to each speaker for the other K-1 speech units among the K speech units.

[0112] In some embodiments of this application, based on the above technical solutions, the expression conversion module 930 is specifically configured to: multiply the content representations corresponding to the K speech units in the global semantic dictionary and the speaker dictionary by the posterior probability vectors respectively to obtain the global content re-representation and the speaker content re-representation; calculate the weighted sum of the global content re-representation and the speaker content re-representation to obtain the content re-representation of the speech frames in the source speech data;

[0113] The speech conversion module 940 is specifically configured to: input the content re-expression of the speech frame in the source speech data and the speaker identifier of the target speaker in the speaker information into the second speech generation model for speech conversion, so as to obtain the target speech data of the target speaker.

[0114] In some embodiments of this application, based on the above technical solutions, the expression conversion module 930 is specifically configured to: obtain the source content expression of the source speech data; calculate the weighted sum of the source content expression, the global content re-expression, and the speaker content re-expression to obtain the content re-expression of the speech frames in the source speech data.

[0115] In some embodiments of this application, based on the above technical solutions, the speech conversion module 940 is specifically configured to: take the content re-expression of the speech frames in the source speech data and the embedded expression of the speaker information as conditions, and combine them with a specified spectrum containing noise. Figure 1 The third speech generation model is used to predict the spectrum of the input, and the resulting spectrogram is obtained. The third speech generation model contains embedded information of time step and speaker information. The resulting spectrogram is then converted into speech by a pre-trained vocoder to obtain the target speech data of the target speaker.

[0116] In some embodiments of this application, based on the above technical solutions, the speech conversion module 940 is further configured to: extract features from the training speech data based on the semantic dictionary to obtain a content representation of the training speech data; concatenate the content representation with a specified spectrogram containing noise to form training data, wherein the specified spectrogram is noise-added according to the time step; train the third model to be trained using the training data to obtain the third speech generation model, wherein the third model to be trained includes a feature linear modulation layer, a convolutional layer, and a diffusion transform layer, the feature linear modulation layer includes information about the time step, and the convolutional layer and the diffusion transform layer include speaker embedding information of the speaker of the training speech data.

[0117] It should be noted that the apparatus provided in the above embodiments and the method provided in the above embodiments belong to the same concept, and the specific way in which each module performs the operation has been described in detail in the method embodiments, and will not be repeated here.

[0118] Figure 10 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.

[0119] It should be noted that, Figure 10 The computer system 1000 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0120] like Figure 10 As shown, the computer system 1000 includes a Central Processing Unit (CPU) 1001, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1002 or programs loaded from storage section 1008 into Random Access Memory (RAM) 1003. The RAM 1003 also stores various programs and data required for system operation. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An Input / Output (I / O) interface 1005 is also connected to the bus 1004.

[0121] The following components are connected to I / O interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to I / O interface 1005 as needed. Removable media 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1010 as needed so that computer programs read from them can be installed into storage section 1008 as needed.

[0122] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1009, and / or installed from removable medium 1011. When the computer program is executed by central processing unit (CPU) 1001, it performs various functions defined in the system of this application.

[0123] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0125] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0126] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0127] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0128] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A method for processing voice data, characterized in that, include: Obtain the source speech data of the source speaker and the speaker information of the target speaker; Feature extraction is performed on the source speech data to obtain a posterior probability vector containing the speech frames in the source speech data belonging to K speech units, wherein the speech unit includes at least one of phonemes and soft speech units, and K is an integer greater than 1; The content representations corresponding to the K speech units in the global dictionary and the speaker dictionary are multiplied by the posterior probability vector to obtain the global content re-representation and the speaker content re-representation. The global dictionary contains the content representations corresponding to the K speech units. The content representation of each speech unit in the global dictionary is obtained by statistical calculation based on the semantic representations and posterior probabilities from speech data from at least two speakers. The speaker dictionary contains the semantic representations of the K speech units corresponding to each speaker. Calculate the weighted sum of the global content re-expression and the speaker content re-expression to obtain the content re-expression of the speech frame in the source speech data; The content of the speech frames in the source speech data and the speaker identifier of the target speaker in the speaker information are input into the second speech generation model for speech conversion to obtain the target speech data of the target speaker.

2. The processing method according to claim 1, characterized in that, The method further includes: Obtain the semantic representation of each speech frame in each dictionary speech data of the dictionary audio set and the posterior probability of each speech frame belonging to the Kth speech unit. The dictionary audio set contains speech data from at least two speakers. The posterior probability of the extracted speech frame belonging to the Kth speech unit is statistically analyzed to obtain the first statistic of the Kth speech unit; The semantic expression of the Kth speech unit is obtained by combining the statistical result of the product of the semantic expression of each speech frame and the posterior probability of each speech frame belonging to the Kth speech unit with the first statistic. The global dictionary is determined based on the semantic expression of the Kth speech unit and the semantic expressions of the other K-1 speech units among the K speech units.

3. The processing method according to claim 2, characterized in that, The step of obtaining the semantic representation of each speech frame in each dictionary speech data of the dictionary audio set and the posterior probability of each speech frame belonging to the Kth speech unit includes: The source speech data is input into a pre-trained first feature extraction model, which is used to calculate the posterior probability that each speech frame in the source speech data belongs to the K phonemes. The output of the bottleneck layer in the first feature extraction model is obtained and used as the posterior probability vector of the source speech data.

4. The processing method according to claim 2, characterized in that, The step of obtaining the semantic representation of each speech frame in each dictionary speech data of the dictionary audio set and the posterior probability of each speech frame belonging to the Kth speech unit includes: The source speech data is input into a pre-trained second feature extraction model, which is used to calculate the posterior probability that each speech frame in the source speech data belongs to the K soft speech units. The output of the specified converter layer in the second feature extraction model is used as the posterior probability vector of the source speech data.

5. The processing method according to claim 2, characterized in that, The method further includes: The posterior probability of a speech frame belonging to the Kth speech unit in the speech data of each speaker is statistically analyzed to obtain the second statistic corresponding to each speaker for the Kth speech unit; For each speaker, the semantic expression corresponding to the Kth speech unit for each speaker is obtained based on the statistical result of the product of the semantic expression of each speech frame and the posterior probability of each speech frame belonging to the Kth speech unit, and the second statistic. Based on the semantic expression of the Kth speech unit corresponding to each speaker and the semantic expression of the other K-1 speech units among the K speech units corresponding to each speaker, determine the speaker dictionary corresponding to each speaker.

6. The processing method according to claim 2, characterized in that, Calculating the weighted sum of the global content re-expression and the speaker content re-expression to obtain the content re-expression of the speech frames in the source speech data includes: Obtain the source content representation of the source speech data; The weighted sum of the source content representation, the global content re-representation, and the speaker content re-representation is calculated to obtain the content re-representation of the speech frame in the source speech data.

7. A voice data processing apparatus, characterized in that, include: The information acquisition module is configured to acquire the source speech data of the source speaker and the speaker information of the target speaker; The feature extraction module is configured to extract features from the source speech data to obtain a posterior probability vector containing the speech frames in the source speech data belonging to K speech units, wherein the speech unit includes at least one of phonemes and soft speech units, and K is an integer greater than 1. The expression conversion module is configured to multiply the content expressions corresponding to the K speech units in the global dictionary and the speaker dictionary with the posterior probability vector to obtain the global content re-expression and the speaker content re-expression, respectively. The global dictionary contains the content expressions corresponding to the K speech units. The content expression of each speech unit in the global dictionary is obtained by statistical calculation based on the semantic expression and posterior probability from speech data from at least two speakers. The speaker dictionary contains the semantic expression of each speaker corresponding to the K speech units. The expression conversion module is also configured to: calculate the weighted sum of the global content re-expression and the speaker content re-expression to obtain the content re-expression of the speech frame in the source speech data; The speech conversion module is configured to input the content of the speech frames in the source speech data and the speaker identifier of the target speaker in the speaker information into the second speech generation model for speech conversion, so as to obtain the target speech data of the target speaker.

8. An electronic device, characterized in that, include: processor; Memory for storing the executable instructions of the processor; The processor is configured to perform the voice data processing method according to any one of claims 1 to 6 by executing the executable instructions.

9. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for processing voice data as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium, a processor of a computer device reading the computer instructions from the computer-readable storage medium, and the processor executing the computer instructions to cause the computer device to perform a method for processing voice data as described in any one of claims 1 to 6.