Timbre feature extraction method and device, computer device and storage medium

By processing speech data through a bidirectional recurrent neural network, calculating the difference and adjusting the parameters, the problem of the inability to effectively represent timbre features in existing technologies is solved, thus improving the speech conversion effect.

CN113870875BActive Publication Date: 2026-02-13PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111130551.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-26
Publication Date
2026-02-13
Estimated Expiration
2041-09-26

AI Technical Summary

Technical Problem

In existing speech conversion technologies, methods for obtaining timbre features cannot effectively represent the speaker's timbre, resulting in poor speech conversion performance.

Method used

A bidirectional recurrent neural network is used to process speech data. The speech data is converted into a continuous vector and quantized into a discrete vector of speech text content. The difference is calculated and the network parameters are adjusted by the loss value of the objective optimization function until the preset requirements are met. The difference is then determined as the speaker's timbre feature information.

Benefits of technology

By using bidirectional recurrent neural networks, the speech text content and the speaker's timbre characteristics can be better decoupled, thus improving the speech conversion effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113870875B_ABST
    Figure CN113870875B_ABST
Patent Text Reader

Abstract

The application relates to the field of artificial intelligence, and particularly discloses a timbre feature extraction method and device, computer equipment and a storage medium. Speech data of at least two speakers is acquired, and the speech data is input into a preset bidirectional recurrent neural network to convert the speech data into a continuous vector, and the continuous vector is quantized into a speech text content discrete vector. The difference between the continuous vector and the speech text content discrete vector is calculated, and the loss value of a preset target optimization function is calculated according to the difference. When the loss value does not meet a preset requirement, the parameters of the bidirectional recurrent neural network are adjusted according to the loss value, and the bidirectional recurrent neural network with the adjusted parameters is trained using new speech data. When the loss value meets the preset requirement, the difference is determined as speaker timbre feature information associated with speaker label information. The application can obtain speaker timbre feature information that can better represent speakers, thereby improving the effect of speech conversion.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, and in particular to a timbre feature extraction method and device, computer equipment and a storage medium. BACKGROUND

[0002] In daily life, speech conversion technology is applied in fields such as driving navigation and dubbing of film and television works. Speech conversion generally refers to converting the speech of one person into the speech of another person, for example, converting the speech of a male announcer in driving navigation into the speech of a star that a driver likes.

[0003] Speech conversion essentially replaces different speakers, i.e. different timbres, without changing the content of the speech. In the prior art, the difference between the original continuous speech variable and the quantized discrete speech variable is calculated, and the expected mean value is repeatedly calculated to obtain the final speaker timbre feature.

[0004] However, the timbre feature obtained by the above-mentioned timbre feature acquisition method cannot well represent the timbre of the speaker, resulting in poor speech conversion effect. SUMMARY

[0005] Therefore, it is necessary to provide a timbre feature extraction method, device, computer equipment and storage medium to solve the problem that the timbre feature obtained by the existing speech conversion technology cannot well represent the timbre of the speaker, thereby resulting in poor speech conversion effect.

[0006] A timbre feature extraction method comprises:

[0007] Obtaining speech data of at least two speakers, wherein the speech data of at least one speaker comprises at least two speeches, and the speech data is associated with speaker label information;

[0008] Inputting the speech data into a preset bidirectional recurrent neural network to convert the speech data into a continuous vector, quantizing the continuous vector into a speech text content discrete vector, and calculating the difference between the continuous vector and the speech text content discrete vector;

[0009] According to the difference, calculating the loss value of a preset target optimization function;

[0010] When the loss value does not meet the preset requirement, adjusting the parameters of the bidirectional recurrent neural network according to the loss value, and training the bidirectional recurrent neural network with adjusted parameters using new speech data;

[0011] determine the difference value as speaker timbre feature information associated with the speaker label information when the loss value meets preset requirements.

[0012] An apparatus for extracting timbre features, comprising:

[0013] a speech data acquisition module configured to acquire speech data of at least two speakers, wherein the speech data of at least one speaker comprises at least two speeches, and the speech data is associated with speaker label information;

[0014] a difference value calculation module configured to input the speech data into a preset bidirectional recurrent neural network to convert the speech data into a continuous vector, quantize the continuous vector into a speech text content discrete vector, and calculate a difference value between the continuous vector and the speech text content discrete vector;

[0015] a loss value calculation module configured to calculate a loss value of a preset target optimization function according to the difference value;

[0016] a training module configured to, when the loss value does not meet preset requirements, adjust parameters of the bidirectional recurrent neural network according to the loss value, and train the bidirectional recurrent neural network with adjusted parameters using new speech data;

[0017] a speaker timbre feature information determination module configured to, when the loss value meets preset requirements, determine the difference value as speaker timbre feature information associated with the speaker label information.

[0018] A computer device comprising a memory, a processor, and computer readable instructions stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for extracting timbre features when executing the computer readable instructions.

[0019] One or more readable storage media storing computer readable instructions, wherein the computer readable instructions are executable by one or more processors to cause the one or more processors to perform the above-mentioned method for extracting timbre features.

[0020] The voice feature extraction method, device, computer device and storage medium, by obtaining voice data of at least two speakers, inputting the voice data into a preset bidirectional recurrent neural network to convert the voice data into a continuous vector, quantizing the continuous vector into a voice text content discrete vector, calculating a difference value between the continuous vector and the voice text content discrete vector, and calculating a loss value of a preset target optimization function according to the difference value, when the loss value does not meet a preset requirement, adjusting parameters of the bidirectional recurrent neural network according to the loss value, and training the bidirectional recurrent neural network with adjusted parameters using new voice data, and when the loss value meets the preset requirement, determining the difference value as speaker voice feature information associated with speaker label information. Compared with the voice feature extraction method in the traditional voice conversion technology, the voice data is processed by using the bidirectional recurrent neural network, the voice text content and the speaker voice feature can be well decoupled, so that the voice feature representing the speaker can be better obtained, and the voice conversion effect can be well improved. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor under the premise of the drawings.

[0022] Figure 1 is a flowchart of a voice feature extraction method in an embodiment of the present application;

[0023] Figure 2 is a flowchart of vector quantization processing of voice training data using VQ technology in an embodiment of the present application;

[0024] Figure 3 is a training diagram of a voice conversion model in an embodiment of the present application;

[0025] Figure 4 is a structural diagram of a voice feature extraction device in an embodiment of the present application;

[0026] Figure 5 is a schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0027] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be clearly and completely described below, obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0028] Artificial intelligence is a branch of computer science, and the research in this field includes robots, speech (including speech processing, speech recognition, speech synthesis, speaker recognition, speech conversion, etc.), image recognition, natural language processing and expert systems, etc. The present application relates to speech processing, speech synthesis, speaker timbre feature extraction, speech conversion, and the proposed timbre feature extraction method can well decouple the speech text content and the speaker timbre feature by processing the speech data using a bidirectional recurrent neural network, thereby obtaining a timbre that can better represent the speaker, and thus can well improve the effect of speech conversion. The speech conversion effect of the present application is good and low in cost, and can be applied to the dubbing field, such as dubbing of self-made creative videos of self-media, dubbing of self-made animations of animation enthusiasts, dubbing of film and television works, etc. It is beneficial to promote the continuous innovation and development of speech technology in the field of artificial intelligence, and has a broad market prospect.

[0029] In an embodiment, as shown in FIG. 1, a timbre feature extraction method is provided, and for ease of illustration, only the part related to the present embodiment is shown, including the following steps: Figure 1

[0030] Step S10, obtaining speech data of at least two speakers; wherein the speech data of at least one speaker includes at least two speeches, and the speech data is associated with speaker label information.

[0031] In the embodiments of the present application, the speech usually includes speech text content and speaker timbre feature information (i.e. the waveform of the speaker's voice vibration (the law of vibration)). For example, if speaker A says "I love my hometown", the speech including the speech text content "I love my hometown" and the speaker timbre feature information of speaker A will be obtained, i.e. the speech data of speaker A. It can be understood that if speaker B says "I love my hometown", the speech including the speech text content "I love my hometown" and the speaker timbre feature information of speaker B will be obtained, i.e. the speech data of speaker B. Generally, different people have different timbres when they speak, so different people's speeches can be identified by timbre.

[0032] ​The speaker label information is information indicating the identity of a speaker, which can be a character label, a figure label, a number label, a letter label, etc. For example, the speaker label information of speaker A, speaker B, and speaker C can be represented by the characters "A", "B", and "C" respectively to indicate the identities of the three speakers, or can be represented by the numbers "1", "2", and "3" respectively to indicate the identities of the three speakers.

[0033] In an embodiment of the present application, the voice data is associated with the speaker label information, which can be achieved by storing the voice data and the speaker label information in a one-to-one correspondence. Specifically, the voice data and the speaker label information can be stored in a corresponding relationship as shown in Table 1.

[0034] Table 1: Corresponding relationship table of voice data and speaker label information

[0035] Voice data Speaker label information Voice data 1 Speaker A Voice data 2 Speaker B … …

[0036] In another embodiment of the present application, the voice data is associated with the speaker label information, and each piece of voice data carries corresponding speaker label information. For example, the voice data spoken by speaker A carries the character label "A" indicating the identity of speaker A.

[0037] In step S20, the voice data is input into a preset bidirectional recurrent neural network to convert the voice data into a continuous vector, and the continuous vector is quantized into a voice text content discrete vector, and the difference between the continuous vector and the voice text content discrete vector is calculated.

[0038] In an embodiment of the present application, the voice data is converted into a continuous vector, which can be a one-dimensional array with a length of 256 dimensions, such as (1.1, 3.3, 2.5, 1.2, …). The Mel spectrum can be extracted from the voice data using extraction techniques known to those skilled in the art, for example, pre-emphasizing, framing, and windowing the audio signal, and then performing short-time Fourier transform (STFT) on each frame of signal to obtain a short-time amplitude spectrum; and then passing the short-time amplitude spectrum through a Mel filter bank to obtain a Mel spectrum.

[0039] The voice text content discrete vector is a discrete vector obtained by replacing the above-mentioned continuous vector with the nearest codebook vector by searching the codebook. For example, the codebook corresponding to the above-mentioned continuous vector (1.1, 3.3, 2.5, 1.2, …) is (1, 3, 2, 1, …). At this time, the difference between the above-mentioned continuous vector and the voice text content discrete vector can be calculated as (0.1, 0.3, 0.5, 0.2, …).

[0040] In another embodiment of the present application, the speech data can be converted into a continuous vector before being input into the preset bidirectional recurrent neural network, and the continuous vector can be quantized into a speech text content discrete vector, and then the continuous vector and the speech text content discrete vector are input into the bidirectional recurrent neural network to calculate the difference between the continuous vector and the speech text content discrete vector.

[0041] In an exemplary embodiment, the speech data includes a first speech, a second speech and a third speech; the first speech and the second speech are associated with a first speaker label information, and the third speech is associated with a second speaker label information.

[0042] It can be understood that the speakers corresponding to the first speech and the second speech are both the first speaker, and the speaker corresponding to the third speech is the second speaker.

[0043] In the step S20, the speech data is converted into a continuous vector, the continuous vector is quantized into a speech text content discrete vector, and the difference between the continuous vector and the speech text content discrete vector is calculated, which includes:

[0044] The first speech is converted into a first continuous vector, the first continuous vector is quantized into a first speech text content discrete vector, and a first difference between the first continuous vector and the first speech text content discrete vector is calculated.

[0045] The second speech is converted into a second continuous vector, the second continuous vector is quantized into a second speech text content discrete vector, and a second difference between the second continuous vector and the second speech text content discrete vector is calculated.

[0046] The third speech is converted into a third continuous vector, the third continuous vector is quantized into a third speech text content discrete vector, and a third difference between the third continuous vector and the third speech text content discrete vector is calculated.

[0047] The calculation method of the first difference, the second difference and the third difference can refer to the calculation method of the difference between the continuous vector and the speech text content discrete vector in the above embodiment, which will not be described here.

[0048] In step S30, the loss value of the preset target optimization function is calculated according to the difference.

[0049] In an embodiment, in combination with the above exemplary embodiment, the loss value of the preset target optimization function is calculated according to the first difference, the second difference and the third difference, wherein the preset target optimization function is:

[0050] L = -(y1≠y2)‖S A (x1)-SB (x1)‖+(y1==y2)‖S A (x1)-S A (x2)‖;

[0051] wherein, L is a loss value; y1 represents a first speaker; y2 represents a second speaker; S A (x1) represents a first difference value obtained by processing a first voice through a preset bidirectional recurrent neural network; S A (x2) represents a second difference value obtained by processing a second voice through a preset bidirectional recurrent neural network; S B (x1) represents a third difference value obtained by processing a third voice through a preset bidirectional recurrent neural network.

[0052] Step S40, when the loss value does not meet the preset requirement, adjusting the parameters of the bidirectional recurrent neural network according to the loss value, and training the bidirectional recurrent neural network with adjusted parameters using new voice data.

[0053] Step S50, when the loss value meets the preset requirement, determining the difference value as the speaker timbre feature information associated with the speaker label information.

[0054] In the embodiment of the present application, the preset requirement generally refers to whether the loss value is less than or equal to a preset threshold value, for example, the preset threshold value is 0.1, 0.3, etc. When the above loss value is greater than the preset requirement, it means that the loss value does not meet the preset requirement, otherwise, the loss value meets the preset requirement.

[0055] In an embodiment, when the loss value does not meet the preset requirement, the parameters of the bidirectional recurrent neural network are adjusted according to the loss value, and new voice data is continuously input to the bidirectional recurrent neural network with adjusted parameters for training. The new voice data can be 4 or 8 pieces of voice data randomly extracted from the training data set, and the number of randomly extracted voice data can be determined according to actual needs, which is not limited here. The new voice data can be voice data including the first speaker and / or the second speaker, or voice data of other speakers.

[0056] When the loss value does not meet the preset requirement, the above training step is repeated until the loss value meets the preset requirement.

[0057] In the embodiment of the present application, the parameters of the bidirectional recurrent neural network are adjusted through the loss value, so as to make the difference between the timbre feature information of the same speaker speaking different sentences as small as possible, and the difference between the timbre feature information of different speakers speaking the same sentence as large as possible, so as to obtain the timbre feature information which can well represent the speaker, and further improve the effect of voice conversion.

[0058] In an embodiment, the step S50 comprises:

[0059] calculating an average of the first difference value and the second difference value, and determining the average as the first speaker voice timbre feature information associated with the first speaker label information.

[0060] determining the third difference value as the second speaker voice timbre feature information associated with the second speaker label information.

[0061] As an example, when a loss value of a preset target optimization function calculated according to the first difference value, the second difference value and the third difference value meets a preset requirement, an average of the first difference value and the second difference value is calculated, and the average is determined as the first speaker voice timbre feature information. The third difference value is determined as the second speaker voice timbre feature information.

[0062] As another example, when a loss value of a preset target optimization function calculated according to the first difference value, the second difference value and the third difference value does not meet a preset requirement, parameters of the bidirectional recurrent neural network are adjusted according to the current calculated loss value, and new speech data is input to the bidirectional recurrent neural network with the adjusted parameters for training to obtain a new loss value. It is determined whether the new loss value meets the preset requirement. If it meets, the speaker voice timbre feature information corresponding to each speaker label information is output.

[0063] In an embodiment, after the step S50, the method further comprises:

[0064] obtaining source speech data to be converted and target speaker label information.

[0065] obtaining a corresponding relationship between the speaker label information and the speaker voice timbre feature information, and obtaining target speaker voice timbre feature information corresponding to the target speaker label information according to the corresponding relationship.

[0066] extracting a source speech text content discrete vector of the source speech data, and performing speech synthesis on the source speech text content discrete vector and the target speaker voice timbre feature information through a speech conversion model to obtain target speech data.

[0067] The source speech data generally refers to original speech data without conversion. The target speaker label information refers to the identity information of the speaker expected to be converted, for example, the target speaker is A, and the target speaker label information can be the name or code of A.

[0068] In the embodiment of the application, the corresponding relationship between the speaker label information and the speaker voice timbre feature information specifically means that the speaker label information and the speaker voice timbre feature information correspond one-to-one. For example, the label information of speaker A corresponds to the voice timbre feature information of A, and the label information of speaker B corresponds to the voice timbre feature information of B.

[0069] As an example, a correspondence table of speaker label information and speaker voice timbre feature information can be constructed in advance, as shown in Table 2 below.

[0070] Table 2 Correspondence table of speaker label information and speaker voice timbre feature information

[0071]

[0072]

[0073] The target speaker voice timbre feature information can be obtained from the correspondence table according to the target speaker label information.

[0074] In the embodiment of the present application, the source speech text content discrete vector of the source speech data can be extracted by referring to the above-mentioned "converting the speech data into a continuous vector and quantizing the continuous vector into a speech text content discrete vector" step, which will not be repeated here.

[0075] As an example, assuming that the source speech data is a speech of speaker A, and it is expected to convert the voice timbre feature of A in the speech into the voice timbre feature of B, thereby obtaining target speech data, then the voice timbre feature information of speaker B and the source speech text content discrete vector can be obtained according to the above-mentioned method, and then the source speech text content discrete vector and the voice timbre feature information of B are synthesized by the speech conversion model to obtain the target speech data.

[0076] In an embodiment, before the speech synthesis according to the source speech text content discrete vector and the target speaker voice timbre feature information to obtain the target speech data, comprising:

[0077] Obtain a plurality of speech training data, and perform vector quantization processing on the plurality of speech training data to obtain a plurality of training speech text content discrete vectors; the plurality of speech training data carries training speaker label information.

[0078] Obtain training speaker voice timbre feature information corresponding to the training speaker label information.

[0079] Train the preset generative adversarial network using the plurality of training speech text content discrete vectors and the training speaker voice timbre feature information to obtain the speech conversion model.

[0080] In an embodiment, the vector quantization processing on the plurality of speech training data to obtain a plurality of training speech text content discrete vectors comprises:

[0081] Perform normalization processing on the plurality of speech training data to obtain a plurality of training speech continuous vectors.

[0082] According to the preset codebook, a training speech text content discrete vector corresponding to the training speech continuous vector is found out.

[0083] As an example, the VQ technology as shown in FIG. 1 can be used for vector quantization processing of the plurality of speech training data. Specifically, the speech training data (i.e., Audio X in FIG. 1) is converted into a latent code (discrete vector) via an encoder (encoding vector). Due to the characteristics of the neural network, the speech training data has a corresponding latent encoding vector V (continuous vector), which is a one-dimensional array with a length of 256 dimensions. After the encoding vector is input into an input layer (i.e., IN in FIG. 1), normalization processing is performed by the IN layer to obtain a normalized vector (i.e., IN(V) in FIG. 1). Then, by searching for a codebook closest to the normalized vector, the normalized vector is replaced by the codebook, and the training speech text content discrete vector is obtained. Figure 2 Figure 2 As an example, the VQ technology as shown in FIG. 1 can be used for vector quantization processing of the plurality of speech training data. Specifically, the speech training data (i.e., Audio X in FIG. 1) is converted into a latent code (discrete vector) via an encoder (encoding vector). Due to the characteristics of the neural network, the speech training data has a corresponding latent encoding vector V (continuous vector), which is a one-dimensional array with a length of 256 dimensions. After the encoding vector is input into an input layer (i.e., IN in FIG. 1), normalization processing is performed by the IN layer to obtain a normalized vector (i.e., IN(V) in FIG. 1). Then, by searching for a codebook closest to the normalized vector, the normalized vector is replaced by the codebook, and the training speech text content discrete vector is obtained. Figure 2 Figure 2 As an example, the VQ technology as shown in FIG. 1 can be used for vector quantization processing of the plurality of speech training data. Specifically, the speech training data (i.e., Audio X in FIG. 1) is converted into a latent code (discrete vector) via an encoder (encoding vector). Due to the characteristics of the neural network, the speech training data has a corresponding latent encoding vector V (continuous vector), which is a one-dimensional array with a length of 256 dimensions. After the encoding vector is input into an input layer (i.e., IN in FIG. 1), normalization processing is performed by the IN layer to obtain a normalized vector (i.e., IN(V) in FIG. 1). Then, by searching for a codebook closest to the normalized vector, the normalized vector is replaced by the codebook, and the training speech text content discrete vector is obtained.

[0084] In an embodiment, the plurality of speech training data includes first speech training data; and the training speaker voice timbre feature information includes first training speaker voice timbre feature information and second training speaker voice timbre feature information.

[0085] The training of the preset generative adversarial network using the plurality of training speech text content discrete vectors and the training speaker voice timbre feature information to obtain the speech conversion model includes:

[0086] The generator of the preset generative adversarial network is trained using the first speech training data and the second training speaker voice timbre feature information to obtain generated speech data.

[0087] The decoder is trained using the first speech training data and the first training speaker voice timbre feature information to obtain reconstructed speech data.

[0088] The discriminator of the preset generative adversarial network is trained using the generated speech data, the reconstructed speech data, and the plurality of speech training data, and the speech conversion model is obtained after the training is completed.

[0089] As an example, the speech conversion model can be trained by referring to the training method as shown in FIG. 2. Specifically, Figure 3 Figure 3 ​​​In the diagram, the target speaker D_Vector1 represents the timbre feature information of the second training speaker, the source speaker D_Vector1 represents the timbre feature information of the first training speaker, Audio X represents the first speech training data, latent code represents the first speech text content discrete vector after the first speech training data is processed by VQ technology, Generator represents the generator of the adversarial generative network, Decoder represents the decoder, x1 represents the reconstructed speech data, x2 represents the generated speech data, and Discriminator represents the discriminator of the adversarial generative network.

[0090] Specifically, during training, the discrete vector of the first speech text content is simultaneously fed into the decoder and generator. Meanwhile, the decoder is continuously fed with the timbre feature information of the first training speaker to obtain generated speech data; the generator is continuously fed with the timbre feature information of the second training speaker to obtain reconstructed speech data. The generated speech data, reconstructed speech data, and several speech training data are used to train the discriminator of the pre-defined adversarial generative network, aiming to obtain the following discrimination results: if the discrimination result of the generated speech data is false, the corresponding speaker is the second training speaker; if the discrimination result of the reconstructed speech data is false, the corresponding speaker is the first training speaker; if the discrimination result of the speech training data is true, the corresponding speaker is the first training speaker.

[0091] In this embodiment of the invention, the above training method can standardize the generation method of the generative adversarial network, reduce the training difficulty of the speech conversion model, and improve the conversion effect of the model.

[0092] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0093] In one embodiment, a timbre feature extraction device is provided, which corresponds one-to-one with the timbre feature extraction methods described in the above embodiments. For example... Figure 4 As shown, the timbre feature extraction device includes a speech data acquisition module 11, a difference calculation module 12, a loss value calculation module 13, a training module 14, and a speaker timbre feature information determination module 15. Detailed descriptions of each functional module are as follows:

[0094] The voice data acquisition module 11 is used to acquire voice data of at least two speakers, wherein the voice data of at least one speaker includes at least two voices, and the voice data is associated with speaker tag information.

[0095] The difference calculation module 12 is configured to input the voice data into a preset bidirectional recurrent neural network to convert the voice data into a continuous vector, quantize the continuous vector into a voice text content discrete vector, and calculate a difference between the continuous vector and the voice text content discrete vector.

[0096] The loss value calculation module 13 is configured to calculate a loss value of a preset target optimization function according to the difference.

[0097] The training module 14 is configured to adjust parameters of the bidirectional recurrent neural network according to the loss value when the loss value does not meet a preset requirement, and train the bidirectional recurrent neural network with the adjusted parameters using new voice data.

[0098] The speaker timbre feature information determination module 15 is configured to determine the difference as speaker timbre feature information associated with the speaker label information when the loss value meets the preset requirement.

[0099] In an embodiment, the voice data includes a first voice, a second voice, and a third voice; the first voice and the second voice are associated with a first speaker label information, and the third voice is associated with a second speaker label information.

[0100] The difference calculation module 12 includes a first difference calculation unit, a second difference calculation unit, and a third difference calculation unit.

[0101] The first difference calculation unit is configured to convert the first voice into a first continuous vector, quantize the first continuous vector into a first voice text content discrete vector, and calculate a first difference between the first continuous vector and the first voice text content discrete vector.

[0102] The second difference calculation unit is configured to convert the second voice into a second continuous vector, quantize the second continuous vector into a second voice text content discrete vector, and calculate a second difference between the second continuous vector and the second voice text content discrete vector.

[0103] The third difference calculation unit is configured to convert the third voice into a third continuous vector, quantize the third continuous vector into a third voice text content discrete vector, and calculate a third difference between the third continuous vector and the third voice text content discrete vector.

[0104] The loss value calculation module 13 can be configured to:

[0105] calculate a loss value of a preset target optimization function according to the first difference, the second difference, and the third difference, wherein the preset target optimization function is:

[0106] L = -(y1!= y2) || S A (x1) - S B (x1) || + (y1 == y2) || S A (x1) - S A (x2) || ;

[0107] wherein, L is a loss value; y1 represents a first speaker; y2 represents a second speaker; S A (x1) represents a first difference value obtained after a first voice is processed by a preset bidirectional recurrent neural network; S A (x2) represents a second difference value obtained after a second voice is processed by a preset bidirectional recurrent neural network; S B (x1) represents a third difference value obtained after a third voice is processed by a preset bidirectional recurrent neural network.

[0108] In an embodiment, the training module 14 described above comprises a first speaker voice timbre feature information determination unit and a second speaker voice timbre feature information determination unit.

[0109] The first speaker voice timbre feature information determination unit is configured to calculate an average value of the first difference value and the second difference value, and determine the average value as first speaker voice timbre feature information associated with first speaker label information;

[0110] The second speaker voice timbre feature information determination unit is configured to determine the third difference value as second speaker voice timbre feature information associated with second speaker label information.

[0111] In an embodiment, the voice timbre feature extraction apparatus described above further comprises an acquisition module, a target speaker voice timbre feature information acquisition module, and a speech synthesis module.

[0112] The acquisition module is configured to acquire source speech data to be converted and target speaker label information.

[0113] The target speaker voice timbre feature information acquisition module is configured to acquire a correspondence between the speaker label information and speaker voice timbre feature information, and acquire target speaker voice timbre feature information corresponding to the target speaker label information according to the correspondence.

[0114] The speech synthesis module is configured to extract a source speech text content discrete vector of the source speech data, perform speech synthesis on the source speech text content discrete vector and the target speaker voice timbre feature information through a speech conversion model, and obtain target speech data.

[0115] In an embodiment, the voice timbre feature extraction apparatus described above further comprises a speech training data acquisition module, a training speaker voice timbre feature information acquisition module, and a speech conversion model training module.

[0116] The voice training data acquisition module is configured to acquire a plurality of voice training data, perform vector quantization processing on the plurality of voice training data, and obtain a plurality of training voice text content discrete vectors.

[0117] The training speaker voice timbre feature information acquisition module is configured to acquire training speaker voice timbre feature information corresponding to the training speaker label information.

[0118] The voice conversion model training module is configured to train a preset generative adversarial network using the plurality of training voice text content discrete vectors and the training speaker voice timbre feature information, to obtain the voice conversion model.

[0119] In an embodiment, the plurality of voice training data includes first voice training data, and the training speaker voice timbre feature information includes first training speaker voice timbre feature information and second training speaker voice timbre feature information.

[0120] The voice conversion model training module includes a generated voice data training unit, a reconstructed voice data training unit, and a discriminator training unit.

[0121] The generated voice data training unit is configured to train a generator of a preset generative adversarial network using the first voice training data and the second training speaker voice timbre feature information, to obtain generated voice data.

[0122] The reconstructed voice data training unit is configured to train a decoder using the first voice training data and the first training speaker voice timbre feature information, to obtain reconstructed voice data.

[0123] The discriminator training unit is configured to train a discriminator of a preset generative adversarial network using the generated voice data, the reconstructed voice data, and the plurality of voice training data, to obtain the voice conversion model after the training.

[0124] In an embodiment, the voice training data acquisition module can be configured to:

[0125] Perform normalization processing on the plurality of voice training data, to obtain a plurality of training voice continuous vectors.

[0126] According to a preset codebook, find out training voice text content discrete vectors corresponding to the training voice continuous vectors.

[0127] The specific limitations of the timbre feature extraction device can refer to the limitations of the timbre feature extraction method described above, which will not be repeated here. Each module in the above timbre feature extraction device can be realized by software, hardware and their combination. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor calls and executes the operations corresponding to the above modules.

[0128] In one embodiment, a computer device is provided, which can be a server, and its internal structure diagram can be as shown in Figure 5 The computer device includes a processor, a memory, a network interface and a database connected by a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium and an internal memory. The readable storage medium stores an operating system, computer readable instructions and a database. The internal memory provides an environment for the operation of the operating system and computer readable instructions in the readable storage medium. The database of the computer device is used to store data related to the timbre feature extraction method. The network interface of the computer device is used to communicate with external terminals through network connection. The computer readable instructions are executed by the processor to implement a timbre feature extraction method. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.

[0129] In one embodiment, a computer device is provided, which includes a memory, a processor and computer readable instructions stored on the memory and executable on the processor, and the processor executes the computer readable instructions to implement the following steps:

[0130] Obtaining speech data of at least two speakers; wherein the speech data of at least one speaker includes at least two speeches, and the speech data is associated with speaker label information;

[0131] Inputting the speech data into a preset bidirectional recurrent neural network to convert the speech data into a continuous vector, quantizing the continuous vector into a speech text content discrete vector, and calculating the difference between the continuous vector and the speech text content discrete vector;

[0132] According to the difference, calculating the loss value of the preset target optimization function;

[0133] When the loss value does not meet the preset requirement, adjusting the parameters of the bidirectional recurrent neural network according to the loss value, and training the bidirectional recurrent neural network with adjusted parameters using new speech data;

[0134] When the loss value meets the preset requirement, the difference value is determined as speaker voice tone feature information associated with the speaker label information.

[0135] In one embodiment, one or more computer readable storage media storing computer readable instructions are provided. The computer readable storage media provided in the embodiment include non-volatile readable storage media and volatile readable storage media. The computer readable instructions stored on the readable storage media are executed by one or more processors to implement the following steps:

[0136] Obtaining speech data of at least two speakers; wherein the speech data of at least one speaker includes at least two speeches, and the speech data is associated with speaker label information;

[0137] Inputting the speech data into a preset bidirectional recurrent neural network to convert the speech data into a continuous vector, quantizing the continuous vector into a discrete vector of speech text content, and calculating a difference value between the continuous vector and the discrete vector of speech text content;

[0138] According to the difference value, a loss value of a preset target optimization function is calculated;

[0139] When the loss value does not meet the preset requirement, the parameters of the bidirectional recurrent neural network are adjusted according to the loss value, and the bidirectional recurrent neural network with the adjusted parameters is trained using new speech data;

[0140] When the loss value meets the preset requirement, the difference value is determined as speaker voice tone feature information associated with the speaker label information.

[0141] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through computer readable instructions, and the computer readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer readable instructions are executed, the processes of the above-mentioned embodiments of the method can be included. Any reference to memory, storage, database or other medium used in each embodiment provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0142] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0143] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, but not to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A timbre feature extraction method characterized by, The method comprises the following steps: acquiring speech data of at least two speakers; wherein the speech data of at least one speaker comprises at least two speeches, and the speech data is associated with speaker label information; inputting the speech data into a preset bidirectional recurrent neural network to convert the speech data into continuous vectors, quantizing the continuous vectors into discrete vectors of speech text content, and calculating the difference between the continuous vectors and the discrete vectors of speech text content; calculating the loss value of a preset target optimization function according to the difference; when the loss value does not meet the preset requirement, adjusting the parameters of the bidirectional recurrent neural network according to the loss value, and training the bidirectional recurrent neural network with new speech data and the adjusted parameters; when the loss value meets the preset requirement, determining the difference as speaker timbre feature information associated with the speaker label information; the speech data comprises a first speech, a second speech and a third speech; the first speech and the second speech are associated with first speaker label information, and the third speech is associated with second speaker label information; the conversion of the speech data into continuous vectors, the quantization of the continuous vectors into discrete vectors of speech text content, and the calculation of the difference between the continuous vectors and the discrete vectors of speech text content comprise: converting the first speech into a first continuous vector, quantizing the first continuous vector into a first discrete vector of speech text content, and calculating a first difference between the first continuous vector and the first discrete vector of speech text content; converting the second speech into a second continuous vector, quantizing the second continuous vector into a second discrete vector of speech text content, and calculating a second difference between the second continuous vector and the second discrete vector of speech text content; converting the third speech into a third continuous vector, quantizing the third continuous vector into a third discrete vector of speech text content, and calculating a third difference between the third continuous vector and the third discrete vector of speech text content; the calculation of the loss value of the preset target optimization function according to the difference comprises: calculating the loss value of the preset target optimization function according to the first difference, the second difference and the third difference, wherein the preset target optimization function is: ; wherein is a loss value; represents a first speaker; represents a second speaker; represents a first difference value obtained after the first voice is processed by the preset bidirectional recurrent neural network; represents a second difference value obtained after the second speech is processed by the preset bidirectional recurrent neural network; represents a third difference value obtained after the third speech is processed by the preset bidirectional recurrent neural network.

2. The timbre feature extraction method of claim 1, wherein, when the loss value meets the preset requirement, determining the difference as speaker timbre feature information associated with the speaker label information, which comprises: calculating the average of the first difference and the second difference, and determining the average as first speaker timbre feature information associated with the first speaker label information; determining the third difference as second speaker timbre feature information associated with the second speaker label information.

3. The timbre feature extraction method of claim 1, wherein, after determining the difference as speaker timbre feature information associated with the speaker label information when the loss value meets the preset requirement, the method further comprises the following steps: acquiring source speech data to be converted and target speaker label information; acquiring the correspondence between the speaker label information and speaker timbre feature information, and acquiring target speaker timbre feature information corresponding to the target speaker label information according to the correspondence; Extract a source speech text content discrete vector of the source speech data, and perform speech synthesis on the source speech text content discrete vector and the target speaker timbre feature information through a speech conversion model to obtain target speech data.

4. The timbre feature extraction method according to claim 3, characterized by, Before the speech synthesis on the source speech text content discrete vector and the target speaker timbre feature information to obtain the target speech data, the method comprises: Obtaining a plurality of speech training data, performing vector quantization processing on the plurality of speech training data to obtain a plurality of training speech text content discrete vectors; the plurality of speech training data carries training speaker label information; Obtaining training speaker timbre feature information corresponding to the training speaker label information; Training a preset generative adversarial network using the plurality of training speech text content discrete vectors and the training speaker timbre feature information to obtain the speech conversion model.

5. The timbre feature extraction method according to claim 4, characterized by, The plurality of speech training data comprises first speech training data; and the training speaker timbre feature information comprises first training speaker timbre feature information and second training speaker timbre feature information. The training of the preset generative adversarial network using the plurality of training speech text content discrete vectors and the training speaker timbre feature information to obtain the speech conversion model comprises: Training a generator of the preset generative adversarial network using the first speech training data and the second training speaker timbre feature information to obtain generated speech data; Training a decoder using the first speech training data and the first training speaker timbre feature information to obtain reconstructed speech data; Training a discriminator of the preset generative adversarial network using the generated speech data, the reconstructed speech data, and the plurality of speech training data, and obtaining the speech conversion model after the training is completed.

6. The timbre feature extraction method of claim 4, wherein, The vector quantization processing on the plurality of speech training data to obtain the plurality of training speech text content discrete vectors comprises: Performing normalization processing on the plurality of speech training data to obtain a plurality of training speech continuous vectors; According to a preset codebook, finding out a training speech text content discrete vector corresponding to the training speech continuous vector.

7. A timbre feature extraction apparatus characterized by comprising: The method comprises: A speech data acquisition module is configured to acquire speech data of at least two speakers, wherein the speech data of at least one speaker comprises at least two speeches, and the speech data is associated with speaker label information; A difference calculation module is configured to input the speech data into a preset bidirectional recurrent neural network to convert the speech data into a continuous vector, quantize the continuous vector into a speech text content discrete vector, and calculate a difference value between the continuous vector and the speech text content discrete vector; A loss value calculation module is configured to calculate a loss value of a preset target optimization function according to the difference value; A training module is configured to adjust parameters of the bidirectional recurrent neural network according to the loss value when the loss value does not meet a preset requirement, and train the bidirectional recurrent neural network with adjusted parameters using new speech data. The speaker timbre feature information determination module is configured to determine the difference value as speaker timbre feature information associated with the speaker label information when the loss value meets a preset requirement. The voice data includes a first voice, a second voice, and a third voice; the first voice and the second voice are associated with first speaker label information, and the third voice is associated with second speaker label information. The difference value calculation module includes: The first difference value calculation unit is configured to convert the first voice into a first continuous vector, quantize the first continuous vector into a first speech text content discrete vector, and calculate a first difference value between the first continuous vector and the first speech text content discrete vector. The second difference value calculation unit is configured to convert the second voice into a second continuous vector, quantize the second continuous vector into a second speech text content discrete vector, and calculate a second difference value between the second continuous vector and the second speech text content discrete vector. The third difference value calculation unit is configured to convert the third voice into a third continuous vector, quantize the third continuous vector into a third speech text content discrete vector, and calculate a third difference value between the third continuous vector and the third speech text content discrete vector. The loss value calculation module is configured to: calculate a loss value of a preset target optimization function according to the first difference value, the second difference value, and the third difference value, wherein the preset target optimization function is: ; wherein, is a loss value; represents a first speaker; represents a second speaker; represents a first difference value obtained after the first speech is processed by a preset bidirectional recurrent neural network; represents a second difference value obtained after the second speech is processed by a preset bidirectional recurrent neural network; represents a third difference value obtained after the third speech is processed by a preset bidirectional recurrent neural network.

8. A computer device comprising a memory, a processor, and computer readable instructions stored in the memory and executable on the processor, wherein, The processor implements the timbre feature extraction method in any one of claims 1 to 6 when executing the computer readable instructions.

9. One or more readable storage media storing computer readable instructions, which are executed by one or more processors to cause the one or more processors to implement the timbre feature extraction method in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice conversion model training method and device, voice conversion model application method and device, equipment and storage medium

    CN113345454A

  • Method and a device for recognizing speech

    US6772117B1