A speech conversion model training method, a speech conversion method and device

CN115762543BActive Publication Date: 2026-09-25SHANGHAI ZHENGDA XIMALAYA NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211415018.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-11
Publication Date
2026-09-25
Estimated Expiration
2042-11-11

AI Technical Summary

Technical Problem

[0003]当前语音转换的建模方法是基于音素后验图(Phone Posteriorgram,PPG)与说话人向量结合的方式,达到基础的音色转换的效果,然而现有的语音转换方式中,语音转换效果欠佳

Benefits of technology

[0053]本申请通过确定训练语料的音素编码向量、音色表征向量以及风格表征向量,计算音素编码向量、音色表征向量以及风格表征向量的互信息,将音素编码向量、音色表征向量以及风格表征向量级联得到目标向量,确定目标向量的预测梅尔谱,计算预测梅尔谱与真实梅尔谱的损失信息,基于互信息与损失信息对待训练的语音转换模型的参数进行调整,得到更新后的语音转换模型,返回执行计算音素编码向量、音色表征向量以及风格表征向量的互信息至基于互信息与损失信息对待训练的语音转换模型的参数进行调整,得到更新后的语音转换模型的步骤,直至达到训练次数,基于上述训练过程,实现音频中不同特征的解耦,进一步提升语音转换的效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115762543B_ABST
    Figure CN115762543B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a speech conversion model training method, a speech conversion method and device, relating to the technical field of speech conversion, the method comprising: determining a phoneme encoding vector, a timbre representation vector and a style representation vector of a training corpus, calculating mutual information of the phoneme encoding vector, the timbre representation vector and the style representation vector, concatenating the phoneme encoding vector, the timbre representation vector and the style representation vector to obtain a target vector, determining a predicted mel spectrum of the target vector, calculating loss information of the predicted mel spectrum and a real mel spectrum, adjusting parameters of a speech conversion model to be trained based on the mutual information and the loss information, returning to the step of adjusting the parameters of the speech conversion model to be trained based on the mutual information and the loss information to calculate mutual information of the phoneme encoding vector, the timbre representation vector and the style representation vector until a training number is reached, thereby improving the effect of speech conversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech conversion technology, and more specifically, to a speech conversion model training method, a speech conversion method, and an apparatus. Background Technology

[0002] With the maturation of the voice technology industry, voice conversion (VC), as a key component of voice technology, is widely used in intelligent voice creation and AI voice changing. VC is a technology that converts the timbre of a speech to that of a target speaker while preserving the content of the speech.

[0003] Current speech conversion modeling methods are based on combining phone posterior maps (PPG) with speaker vectors to achieve basic timbre conversion. However, the speech conversion effect of existing methods is not ideal. Summary of the Invention

[0004] The purpose of this application is to provide a speech conversion model training method, device, electronic device and storage medium that can decouple different features in audio and improve the speech conversion effect.

[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows:

[0006] In a first aspect, embodiments of this application provide a method for training a speech conversion model, the method comprising:

[0007] Determine the phoneme encoding vector, timbre representation vector, and style representation vector of the training corpus;

[0008] Calculate the mutual information of phoneme encoding vector, timbre representation vector, and style representation vector;

[0009] The target vector is obtained by concatenating the phoneme encoding vector, timbre representation vector, and style representation vector.

[0010] Determine the predicted Mel spectrum of the target vector;

[0011] Calculate the loss information between the predicted Mel spectrum and the true Mel spectrum;

[0012] The parameters of the speech conversion model to be trained are adjusted based on the mutual information and the loss information to obtain the updated speech conversion model.

[0013] The process of returning to the step of calculating the mutual information of the phoneme encoding vector, timbre representation vector, and style representation vector, and then adjusting the parameters of the speech conversion model to be trained based on the mutual information and the loss information to obtain the updated speech conversion model, continues until the training iterations are reached.

[0014] In an optional implementation, the step of calculating the mutual information of the phoneme encoding vector, timbre representation vector, and style representation vector includes:

[0015] The training corpus is sampled to obtain first training data, wherein the first training data contains multiple data pairs, each data pair including a phoneme encoding vector, a timbre representation vector, a style representation vector, and a mel corresponding to the phoneme encoding vector, the timbre representation vector, and the style representation vector;

[0016] Based on the first training data, determine the maximum likelihood function of the first training data;

[0017] Based on the maximum likelihood function, the joint distribution of the phoneme encoding vector, timbre representation vector, style representation vector, and Mel corresponding to the phoneme encoding vector, timbre representation vector, and style representation vector in the training corpus is updated.

[0018] Based on the updated joint distribution, determine the upper bound of the variational contrastive logarithm scale for each data pair in the first training data.

[0019] The mutual information of the first training data is calculated based on the upper bound of the logarithmic scale of all variational contrasts.

[0020] In an optional implementation, the step of determining the upper bound of the variational contrastive logarithmic scale for each data pair in the first training data based on the updated joint distribution includes:

[0021] Determine the uniform sampling value of the first training data;

[0022] Based on the updated joint distribution and the uniform sampled values, the upper limit of the variational ratio logarithmic scale for each data pair in the first training data is determined.

[0023] In an optional implementation, the step of determining the maximum likelihood function of the first training data based on the first training data includes:

[0024] Determine the likelihood function for each data pair in the first training data;

[0025] The largest likelihood function is determined from all the likelihood functions and used as the maximum likelihood function of the first training data.

[0026] In an optional implementation, the likelihood function satisfies the following formula:

[0027]

[0028] Where N is the number of data pairs in the first training data, L(θ) is the likelihood function, and q θ (y i |x i ) is the joint variational release of x and y, where x is the phoneme encoding vector, timbre representation vector, and style representation vector, and y is the mel corresponding to the phoneme encoding vector, timbre representation vector, and style representation vector.

[0029] In an optional implementation, the upper limit of the logarithmic scale of the variational pair satisfies the following formula:

[0030] Where Ui is the upper bound of the variational contrast logarithm scale, xi is the phoneme encoding vector, timbre representation vector, and style representation vector, and y is the upper bound of the variational contrast logarithm scale. i Let si` be the melm value corresponding to the phoneme encoding vector, timbre representation vector, and style representation vector, and q be the uniform sampling value. θ (y i |x i ) is a joint variational release of x and y.

[0031] In an optional implementation, the upper limit of the logarithmic scale of the variational pair satisfies the following formula:

[0032]

[0033] Ui is the upper bound of the variational contrast logarithm scale, xi is the phoneme encoding vector, timbre representation vector, and style representation vector, and y is the upper bound of the variational contrast logarithm scale. i Let q be the mel for the phoneme encoding vector, timbre representation vector, and style representation vector, where N is the number of data pairs in the first training data, and q is the number of data pairs in the first training data. θ (y i |x i ) is a joint variational release of x and y.

[0034] In an optional implementation, the mutual information of the first training data is calculated based on the following formula:

[0035]

[0036] Among them, I vclub For mutual information, Ui is the upper bound of the variational contrastive logarithmic scale, and N is the number of data pairs in the first training data.

[0037] Secondly, embodiments of this application provide a speech conversion method, the method comprising: determining the target speaker ID and target voiceprint features corresponding to the target synthesized speech;

[0038] Extract the target speaker timbre representation vector based on the target speaker ID and the target voiceprint features;

[0039] Determine the source phoneme encoding vector and source style representation vector of the source audio;

[0040] The target speaker timbre representation vector, the source phoneme encoding vector, and the source style representation vector are concatenated to obtain a concatenated vector;

[0041] The concatenated vectors are used as input to the speech conversion model, and the target audio is output.

[0042] Thirdly, embodiments of this application provide a speech conversion model training device, the device comprising:

[0043] The first determining module is used to determine the phoneme encoding vector, timbre representation vector, and style representation vector of the training corpus;

[0044] The first calculation module is used to calculate the mutual information of the phoneme encoding vector, the timbre representation vector, and the style representation vector;

[0045] The cascading module is used to cascade the phoneme encoding vector, timbre representation vector, and style representation vector to obtain the target vector.

[0046] The second determining module is used to determine the predicted Mel spectrum of the target vector;

[0047] The second calculation module is used to calculate the loss information between the predicted Mel spectrum and the true Mel spectrum;

[0048] The adjustment module is used to adjust the parameters of the speech conversion model to be trained based on the mutual information and the loss information, so as to obtain the updated speech conversion model.

[0049] The training module is used to return the mutual information of the calculation of the phoneme encoding vector, timbre representation vector, and style representation vector to the parameters of the speech conversion model to be trained based on the mutual information and the loss information, so as to obtain the updated speech conversion model, until the training number is reached.

[0050] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the speech conversion model training method and the speech conversion method.

[0051] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the speech conversion model training method and the speech conversion method.

[0052] This application has the following beneficial effects:

[0053] This application determines the phoneme encoding vector, timbre representation vector, and style representation vector of the training corpus, calculates the mutual information of these vectors, concatenates them to obtain the target vector, determines the predicted Mel spectrum of the target vector, calculates the loss information between the predicted and true Mel spectra, adjusts the parameters of the speech conversion model to be trained based on the mutual information and loss information, and obtains an updated speech conversion model. The process of calculating the mutual information of the phoneme encoding vector, timbre representation vector, and style representation vector, and adjusting the parameters of the speech conversion model to be trained based on the mutual information and loss information, is repeated until the required number of training iterations is reached. Based on this training process, the decoupling of different features in the audio is achieved, further improving the speech conversion effect. Attached Figure Description

[0054] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1 A block diagram illustrating an electronic device provided in an embodiment of this application;

[0056] Figure 2 This is one of the flowcharts illustrating a speech conversion model training method provided in an embodiment of this application;

[0057] Figure 3 A second schematic flowchart illustrating a speech conversion model training method provided in this application embodiment;

[0058] Figure 4 The third schematic flowchart of a speech conversion model training method provided in this application embodiment;

[0059] Figure 5 The fourth flowchart illustrates a speech conversion model training method provided in this application embodiment;

[0060] Figure 6 A flowchart illustrating a speech conversion method provided in an embodiment of this application;

[0061] Figure 7 This is a structural block diagram of a speech conversion model training device provided in an embodiment of this application. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0063] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0064] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0065] In the description of this application, it should be noted that if terms such as "upper," "lower," "inner," or "outer" are used to indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of the invention is usually placed during use, they are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.

[0066] Furthermore, the terms "first" and "second" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.

[0067] In the description of this application, it should also be noted that, unless otherwise expressly specified and limited, the terms "set up," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0068] Through extensive research, the inventors discovered that, with the maturation of the voice technology industry, Voice Conversion (VC), as a key component of voice technology, is widely used in intelligent voice creation and AI voice changing. VC is a technology that converts the timbre of a speech to that of a target speaker while preserving the content of the speech.

[0069] Current speech conversion modeling methods are based on combining phone posterior maps (PPG) with speaker vectors to achieve basic timbre conversion. However, the speech conversion effect of existing methods is not ideal.

[0070] In view of the above-mentioned problems, this embodiment provides a speech conversion model training method, speech conversion method, and apparatus. The method involves determining the phoneme encoding vector, timbre representation vector, and style representation vector of the training corpus; calculating the mutual information of these vectors; concatenating them to obtain a target vector; determining the predicted Mel spectrum of the target vector; calculating the loss information between the predicted and true Mel spectra; adjusting the parameters of the speech conversion model to be trained based on the mutual information and loss information; obtaining an updated speech conversion model; and returning to the steps of calculating the mutual information of the phoneme encoding vector, timbre representation vector, and style representation vector and adjusting the parameters of the speech conversion model to be trained based on the mutual information and loss information to obtain an updated speech conversion model. This process is repeated until the required number of training iterations is reached. Based on the above training process, the decoupling of different features in the audio is achieved, further improving the speech conversion effect. The solution provided in this embodiment is described in detail below.

[0071] This embodiment provides an electronic device capable of training a speech conversion model. In one possible implementation, the electronic device can be a user terminal, such as, but not limited to, a server, smartphone, personal computer (PC), tablet computer, personal digital assistant (PDA), mobile internet device (MID), etc.

[0072] Please refer to Figure 1 , Figure 1 This is a schematic diagram of the structure of the electronic device 100 provided in the embodiments of this application. The electronic device 100 may further include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown. Figure 1 The components shown can be implemented using hardware, software, or a combination thereof.

[0073] The electronic device 100 includes a voice conversion model device 110, a memory 120, and a processor 130.

[0074] The components of the memory 120 and processor 130 are electrically connected directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The voice conversion model device 110 includes at least one software function module that can be stored in the memory 120 in the form of software or firmware or embedded in the operating system (OS) of the electronic device 100. The processor 130 is used to execute the executable modules stored in the memory 120, such as the software function modules and computer programs included in the voice conversion model device 110.

[0075] The memory 120 may be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory 120 is used to store programs, and the processor 130 executes the programs after receiving execution instructions.

[0076] Please refer to Figure 2 , Figure 2 For application Figure 1 The flowchart of a speech conversion model method for an electronic device 100 is shown below, and the method includes each step in detail.

[0077] Step 201: Determine the phoneme encoding vector, timbre representation vector, and style representation vector of the training corpus.

[0078] Step 202: Calculate the mutual information of the phoneme encoding vector, timbre representation vector, and style representation vector.

[0079] Step 203: Concatenate the phoneme encoding vector, timbre representation vector, and style representation vector to obtain the target vector.

[0080] Step 204: Determine the predicted Mel spectrum of the target vector.

[0081] Step 205: Calculate the loss information between the predicted Mel spectrum and the true Mel spectrum.

[0082] Step 206: Adjust the parameters of the speech conversion model to be trained based on mutual information and loss information to obtain the updated speech conversion model.

[0083] Step 207: Return to the step of calculating the mutual information of the phoneme encoding vector, timbre representation vector, and style representation vector, and then adjust the parameters of the speech conversion model to be trained based on the mutual information and loss information to obtain the updated speech conversion model, until the training iterations are reached.

[0084] An ASR model trained on audio data from different domains determines the PPG of phoneme representations in the training corpus, and learns the phoneme encoding vector of the training corpus using the Mean Square Error (MSE) as the loss function through a phoneme encoder.

[0085] Based on the VP model, the timbre representation of the training corpus is determined. The speaker phoneme encoding vector of the training corpus is determined by the speaker encoder. The speaker vector and the phoneme encoding vector are concatenated to form the timbre representation vector.

[0086] Determine the phoneme-level fundamental frequency envelope and energy envelope in the training corpus, and concatenate the fundamental frequency envelope and energy envelope to obtain the style representation vector of the training corpus.

[0087] To ensure that the various vectors, namely the phoneme encoding vector, timbre representation vector, and style representation vector, are as independent as possible, and that the target vector obtained by cascading the phoneme encoding vector, timbre representation vector, and style representation vector does not affect each other, the mutual information of the phoneme encoding vector, timbre representation vector, and style representation vector is calculated.

[0088] The target vector is obtained by concatenating the phoneme encoding vector, timbre representation vector, and style representation vector of the training corpus. Based on the encoder, the predicted Mel spectrum of the target vector, which is the concatenation of the phoneme encoding vector, timbre representation vector, and style representation vector, is determined.

[0089] Based on the mutual information of phoneme encoding vectors, timbre representation vectors, and style representation vectors, as well as the loss information between the predicted Mel spectrum of the target vector and the true Mel spectrum, the parameters of the speech conversion model are adjusted. In particular, the parameters of the speech conversion model are adjusted based on mutual information to achieve decoupling between phoneme encoding vectors, timbre representation vectors, and style representation vectors.

[0090] The process of repeatedly calculating the mutual information of the phoneme encoding vector, timbre representation vector, and style representation vector is used to adjust the parameters of the speech conversion model to be trained based on the mutual information and loss information, so as to obtain an updated speech conversion model. This process continues until the loss information converges or the number of training iterations is reached, thus completing the training of the speech conversion model.

[0091] There are multiple ways to calculate the mutual information of phoneme encoding vectors, timbre representation vectors, and style representation vectors. In one implementation, such as... Figure 3 As shown, it may include the following steps:

[0092] Step 202-1: Sample the training corpus to obtain the first training data.

[0093] The first training data contains multiple data pairs, each of which includes a phoneme encoding vector, a timbre representation vector, a style representation vector, and a mel corresponding to the phoneme encoding vector, timbre representation vector, and style representation vector.

[0094] Step 202-2: Determine the maximum likelihood function of the first training data based on the first training data.

[0095] Step 202-3: Based on the maximum likelihood function, update the phoneme encoding vector, timbre representation vector, style representation vector, and the joint distribution of Mel with the corresponding phoneme encoding vector, timbre representation vector, and style representation vector in the training corpus.

[0096] Step 202-4: Based on the updated joint distribution, determine the upper bound of the variational contrastive logarithmic scale for each data pair in the first training data.

[0097] Step 202-5: Calculate the mutual information of the first training data based on the upper bound of the logarithmic scale of all variational contrasts.

[0098] It should be noted that the first training data can be one of the data in the training corpus, and the first training data can include a preset number of data pairs.

[0099] After obtaining mutual information based on the first training data and tuning the speech conversion model based on the mutual information, the second training data is obtained from the training corpus. The number of data pairs in the second training data is the same as the number of data pairs in the first training data.

[0100] After obtaining mutual information based on the second training data, the steps of obtaining new training data from the training corpus and calculating mutual information based on the newly collected training data are repeated until all corpora in the training corpus have been collected.

[0101] The following explanation uses the example of collecting the first training data from the training corpus:

[0102] The likelihood function of each data pair in the first training data is calculated. Based on the likelihood function of each data pair, the maximum likelihood function in the first training data is determined. Based on the maximum likelihood function, the joint distribution of the phoneme encoding vector, timbre representation vector, style representation vector, and their corresponding Mel values ​​in the training corpus is updated. Based on the updated joint distribution, the upper limit of the variational contrast logarithm scale for each data pair in the first training data is calculated using the formula for calculating the upper limit of the variational contrast logarithm scale. Finally, based on the calculated upper limits of the variational contrast logarithm scale for all data pairs in the first training data, the mutual information of the first training data is obtained, i.e., the mutual information between the phoneme encoding vector, timbre representation vector, and style representation vector in the first training data. The parameters of the speech conversion model are updated for the first time based on this mutual information.

[0103] For example, when the first training data contains 10 data pairs, the likelihood function corresponding to the 10 data pairs is calculated, and the largest likelihood function among the 10 likelihood functions is determined as the maximum likelihood function of the first training data. The joint distribution is updated based on the maximum likelihood function. The upper limit of the variational contrast logarithm scale of the 10 data pairs is calculated based on the formula for calculating the upper limit of the variational contrast logarithm scale. The mutual information of the first training data is calculated based on the upper limit of the variational contrast logarithm scale of the 10 data pairs.

[0104] Specifically, the mutual information of the first training data is calculated based on the following formula:

[0105]

[0106] Among them, I vclub For mutual information, Ui is the upper bound of the variational contrast logarithmic scale, and N is the number of data pairs in the first training data.

[0107] There are multiple ways to determine the upper bound of the variational contrast logarithmic scale for each data pair. In one implementation, such as... Figure 4 As shown, it may include the following steps:

[0108] Step 202-3-1: Determine the uniform sampling value of the first training data.

[0109] Step 202-3-2: Based on the updated joint distribution and uniform sampling values, determine the upper limit of the variational ratio logarithmic scale for each data pair in the first training data.

[0110] If the first training data is sampled, the upper limit of the variational ratio logarithmic scale is calculated based on the following formula.

[0111] The upper bound of the logarithmic scale of variational pairs satisfies the following formula:

[0112] Where Ui is the upper bound of the variational contrast logarithm scale, xi is the phoneme encoding vector, timbre representation vector, and style representation vector, and y is the upper bound of the variational contrast logarithm scale. i Let si` be the melm value corresponding to the phoneme encoding vector, timbre representation vector, and style representation vector, and q be the uniform sampling value. θ (y i |x i ) is a joint variational release of x and y.

[0113] In another example, when the first training data is not uniformly sampled, the variational pair on the logarithmic scale is calculated using the following formula:

[0114] Where Ui is the upper bound of the variational contrast logarithm scale, xi is the phoneme encoding vector, timbre representation vector, and style representation vector, and y is the upper bound of the variational contrast logarithm scale. i Let si` be the melm value corresponding to the phoneme encoding vector, timbre representation vector, and style representation vector, and q be the uniform sampling value. θ (y i |x i ) is a joint variational release of x and y.

[0115] There are multiple ways to determine the maximum likelihood function for the first training data. In one implementation, such as... Figure 5 As shown, it may include the following steps:

[0116] Step 202-2-1: Determine the likelihood function for each data pair in the first training data.

[0117] Step 202-2-2: Determine the largest likelihood function from all likelihood functions, and use it as the largest likelihood function for the first training data.

[0118] The likelihood function satisfies the following formula:

[0119]

[0120] Where N is the number of data pairs in the first training data, L(θ) is the likelihood function, and q θ (y i |x i ) is the joint variational release of x and y, where x is the phoneme encoding vector, timbre representation vector, and style representation vector, and y is the mel corresponding to the phoneme encoding vector, timbre representation vector, and style representation vector.

[0121] For example, when the first training data contains 10 data pairs, the likelihood function of each of the 10 data pairs is calculated, and the largest likelihood function is selected from the 10 likelihood functions. Based on the largest likelihood function, the joint distribution of the phoneme encoding vector, timbre representation vector, style representation vector, and Mel corresponding to the phoneme encoding vector, timbre representation vector, and style representation vector in the training corpus is updated.

[0122] Please refer to Figure 6 , Figure 6 For application Figure 1 The flowchart below shows a voice conversion method for an electronic device 100, and the method includes a detailed description of each step.

[0123] Step 301: Determine the target speaker ID and target voiceprint features corresponding to the target synthesized speech.

[0124] Step 302: Extract the target speaker timbre representation vector based on the target speaker ID and target voiceprint features.

[0125] Step 303: Determine the source phoneme encoding vector and source style representation vector of the source audio.

[0126] Step 304: Concatenate the target speaker timbre representation vector, the source phoneme encoding vector, and the source style representation vector to obtain a concatenated vector.

[0127] Step 305: Use the concatenated vector as input to the trained speech conversion model and output the target audio. It should be noted that the target audio corresponds to the timbre of the target speaker in the target synthesized speech, and corresponds to the phonemes and styles in the source audio.

[0128] For example, when user A needs to synthesize target audio with their own timbre, the speaker ID corresponding to user A and user A's voiceprint features are determined to obtain user A's timbre representation vector. The source phoneme encoding vector and source style representation vector of the source audio are determined. For example, if the source audio corresponds to speaker B, then A's timbre representation vector, B's source audio's source phoneme encoding vector, and source style representation vector are input into the trained speech conversion model to obtain the target audio with A's timbre and B's prosody and style.

[0129] Please refer to Figure 7 This application embodiment also provides an application for Figure 1 The speech conversion model training device 110 of the electronic device 100 includes:

[0130] The first determining module 111 is used to determine the phoneme encoding vector, timbre representation vector and style representation vector of the training corpus;

[0131] The first calculation module 112 is used to calculate the mutual information of the phoneme encoding vector, the timbre representation vector, and the style representation vector;

[0132] The cascade module 113 is used to cascade the phoneme encoding vector, timbre representation vector and style representation vector to obtain the target vector.

[0133] The second determining module 114 is used to determine the predicted Mel spectrum of the target vector;

[0134] The second calculation module 115 is used to calculate the loss information between the predicted Mel spectrum and the true Mel spectrum;

[0135] The adjustment module 116 is used to adjust the parameters of the speech conversion model to be trained based on the mutual information and the loss information to obtain an updated speech conversion model.

[0136] The training module 117 is used to return the mutual information of the calculation of the phoneme encoding vector, timbre representation vector and style representation vector to the step of adjusting the parameters of the speech conversion model to be trained based on the mutual information and the loss information to obtain the updated speech conversion model, until the training number is reached.

[0137] It should be noted that the speech conversion model training device provided in this embodiment has the same basic principle and technical effect as the speech conversion model training method embodiment described above. For the sake of brevity, any parts not mentioned in this embodiment can be referred to the corresponding content in the above method embodiment.

[0138] This application also provides an electronic device 100, which includes a processor 130 and a memory 120. The memory 120 stores computer-executable instructions, which, when executed by the processor 130, implement the speech conversion model training and speech conversion method.

[0139] This application embodiment also provides a computer-readable storage medium storing a computer program. When the computer program is executed by the processor 130, it implements the speech conversion model training and speech conversion method.

[0140] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0141] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part. If the function is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.

[0142] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0143] The above descriptions are merely various embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for training a speech conversion model, characterized in that, The method includes: Determine the phoneme encoding vector, timbre representation vector, and style representation vector of the training corpus; Calculate the mutual information of phoneme encoding vector, timbre representation vector, and style representation vector; The target vector is obtained by concatenating the phoneme encoding vector, timbre representation vector, and style representation vector. Determine the predicted Mel spectrum of the target vector; Calculate the loss information between the predicted Mel spectrum and the true Mel spectrum; The parameters of the speech conversion model to be trained are adjusted based on the mutual information and the loss information to obtain the updated speech conversion model. The step of returning to the step of calculating the mutual information of the phoneme encoding vector, timbre representation vector, and style representation vector and adjusting the parameters of the speech conversion model to be trained based on the mutual information and the loss information to obtain the updated speech conversion model, continues until the training iterations are reached; the step of calculating the mutual information of the phoneme encoding vector, timbre representation vector, and style representation vector includes: The training corpus is sampled to obtain first training data, wherein the first training data contains multiple data pairs, each data pair including a phoneme encoding vector, a timbre representation vector, a style representation vector, and a mel corresponding to the phoneme encoding vector, the timbre representation vector, and the style representation vector; Based on the first training data, determine the maximum likelihood function of the first training data; Based on the maximum likelihood function, the joint distribution of the phoneme encoding vector, timbre representation vector, style representation vector, and Mel corresponding to the phoneme encoding vector, timbre representation vector, and style representation vector in the training corpus is updated. Based on the updated joint distribution, determine the upper bound of the variational contrastive logarithm scale for each data pair in the first training data. The mutual information of the first training data is calculated based on the upper bound of the variational contrastive logarithmic scale for all data pairs; the step of determining the upper bound of the variational contrastive logarithmic scale for each data pair in the first training data based on the updated joint distribution includes: Determine the uniform sampling value of the first training data; Based on the updated joint distribution and the uniform sampled values, the upper limit of the variational ratio logarithmic scale for each data pair in the first training data is determined.

2. The method according to claim 1, characterized in that, The step of determining the maximum likelihood function of the first training data based on the first training data includes: Determine the likelihood function for each data pair in the first training data; The largest likelihood function is determined from all the likelihood functions and used as the maximum likelihood function of the first training data.

3. The method according to claim 2, characterized in that, The likelihood function satisfies the following formula: ; Where N is the number of data pairs in the first training data. Let be the likelihood function. Let x be the joint variational distribution of x and y, where x is the phoneme encoding vector, timbre representation vector, and style representation vector, and y is the mel corresponding to the phoneme encoding vector, timbre representation vector, and style representation vector.

4. The method according to claim 1, characterized in that, The upper limit of the logarithmic scale of the variational pair satisfies the following formula: = - Where Ui is the upper bound of the variational contrast logarithm scale, and xi is the phoneme encoding vector, timbre representation vector, and style representation vector. Let be the mellow value corresponding to the phoneme encoding vector, timbre representation vector, and style representation vector, and si` be the uniformly sampled value. For the joint variational release of x and y.

5. The method according to claim 1, characterized in that, The upper limit of the logarithmic scale of the variational pair satisfies the following formula: = - N ; Ui represents the upper bound of the variational contrastive logarithmic scale, and xi represents the phoneme encoding vector, timbre representation vector, and style representation vector, respectively. Let be the mel for the phoneme encoding vector, timbre representation vector, and style representation vector, and N be the number of data pairs in the first training data. For the joint variational release of x and y.

6. The method according to claim 1, characterized in that, The mutual information of the first training data is calculated based on the following formula: = ; in, For mutual information, Ui is the upper bound of the variational contrast logarithmic scale, and N is the number of data pairs in the first training data.

7. A speech conversion method, characterized in that, The method includes: Determine the target speaker ID and target voiceprint features corresponding to the target synthesized speech; Extract the target speaker timbre representation vector based on the target speaker ID and the target voiceprint features; Determine the source phoneme encoding vector and source style representation vector of the source audio; The target speaker timbre representation vector, the source phoneme encoding vector, and the source style representation vector are concatenated to obtain a concatenated vector; The concatenated vector is used as input to the speech conversion model according to any one of claims 1-6, and the target audio is output.

8. A speech conversion model training device, characterized in that, The device includes: The first determining module is used to determine the phoneme encoding vector, timbre representation vector, and style representation vector of the training corpus; The first calculation module is used to calculate the mutual information of the phoneme encoding vector, the timbre representation vector, and the style representation vector; The cascading module is used to cascade the phoneme encoding vector, timbre representation vector, and style representation vector to obtain the target vector. The second determining module is used to determine the predicted Mel spectrum of the target vector; The second calculation module is used to calculate the loss information between the predicted Mel spectrum and the true Mel spectrum; The adjustment module is used to adjust the parameters of the speech conversion model to be trained based on the mutual information and the loss information, so as to obtain the updated speech conversion model. The training module is used to return the mutual information of the calculation of the phoneme encoding vector, timbre representation vector and style representation vector to the parameters of the speech conversion model to be trained based on the mutual information and the loss information to obtain the updated speech conversion model, until the training number is reached; The first calculation module is specifically used to sample the training corpus to obtain first training data, wherein the first training data contains multiple data pairs, each data pair including a phoneme encoding vector, a timbre representation vector, a style representation vector, and a mel corresponding to the phoneme encoding vector, timbre representation vector, and style representation vector; Based on the first training data, determine the maximum likelihood function of the first training data; Based on the maximum likelihood function, the joint distribution of the phoneme encoding vector, timbre representation vector, style representation vector, and Mel corresponding to the phoneme encoding vector, timbre representation vector, and style representation vector in the training corpus is updated. Based on the updated joint distribution, determine the upper bound of the variational contrastive logarithm scale for each data pair in the first training data. The mutual information of the first training data is calculated based on the upper bound of all variational contrastive logarithmic scales. The first calculation module is further configured to determine the uniform sampling value of the first training data; Based on the updated joint distribution and the uniform sampled values, the upper limit of the variational ratio logarithmic scale for each data pair in the first training data is determined.

Citation Information

Patent Citations

  • Voice style migration model training method and device and voice style migration method and device

    CN114203154A