Speech processing training program, speech processing training device, speech processing training method, speech processing program, speech processing device, and speech processing method
The speech processing system addresses the challenge of synthesizing natural-sounding speech by employing a variational autoencoder with controlled speaker feature input and sampling distributions, enhancing voice quality and phrasing through reduced reconstruction errors.
Patent Information
- Application Number
- JP2021106955
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-28
- Publication Date
- 2025-05-21
- Estimated Expiration
- 2041-06-28
AI Technical Summary
Conventional speech processing devices struggle to synthesize speech that is both natural in voice quality and phrasing, often failing to achieve a sufficiently realistic conversion.
A speech processing system utilizing a variational autoencoder with multiple sampling hierarchies, including a speaker encoder and decoder, that trains to minimize the distance between input and output acoustic features while controlling the input of speaker features at specific layers, using conditional instance normalization and varying distributions for sampling.
The system effectively converts speech to sound more natural by training the encoder and decoder to reduce reconstruction errors, resulting in a more realistic voice quality and phrasing.
Smart Images

Figure 0007680893000001 
Figure 0007680893000002 
Figure 0007680893000003
Abstract
Description
[Technical field]
[0001] The present invention relates to a speech processing training program, a speech processing training device, a speech processing training method, a speech processing program, a speech processing device, and a speech processing method. [Background technology]
[0002] A voice processing device has been developed that converts a voice uttered by a given speaker into a voice having the voice quality of another speaker. For example, a technology that applies CycleGAN, an image conversion technology, to voice conversion has been disclosed (Non-Patent Document 1). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Takuhiro Kaneko and Hirokazu Kameoka, Parallel-Data-Free Voice Conversion Using Cycle-Consistent Adversarial Networks arXiv:1711.11293,Nov. 2017 (EUSIPCO 2018) http: / / www.kecl.ntt.co.jp / people / kaneko.takuhiro / projects / cyclegan-vc / Summary of the Invention [Problem to be solved by the invention]
[0004] In a speech processing device that synthesizes and outputs the speech of another speaker from an original speaker, it is required that the voice quality and phrasing of the synthesized speech be as natural as possible. However, in the conventional speech processing device training method, there are cases where the synthesized speech cannot be made sufficiently natural. [Means for solving the problem]
[0005] One aspect of the present invention is a speech processing training program that causes a computer to function as a speech processing training device including an acoustic feature extractor that converts speech into input acoustic features, a speaker encoder that converts a speaker label of the speech into speaker features, a speech encoder including a variational autoencoder having two or more sampling hierarchies that converts the input acoustic features and the speaker features into latent representations, and a speech decoder including a variational autoencoder having two or more sampling hierarchies that generates acoustic features using at least the latent representations and the speaker features, and that trains the speech encoder, the speech decoder, and the speaker encoder to reduce the distance between the input acoustic features input to the speech encoder and the output acoustic features generated in the speech decoder.
[0006] Here, it is preferable that the speech decoder limits a hierarchical level to which speaker features are input in the two or more sampling hierarchical levels.
[0007] It is also preferable that the speech decoder does not input speaker features to layers prior to a predetermined layer in the two or more sampling layers, and inputs speaker features to layers subsequent to the predetermined layer.
[0008] It is also preferable that the audio decoder performs sampling from a posterior distribution in a layer prior to the predetermined layer in the two or more sampling layers, and performs sampling from a prior distribution in a layer subsequent to the predetermined layer.
[0009] In addition, it is preferable that the speech decoder inputs the speaker features to a conditional instance normalization layer.
[0010] Another aspect of the present invention is a speech processing program that causes a computer to function as a speech processing device including: an acoustic feature extractor that converts speech into acoustic features; a speaker encoder that converts a speaker label of the speech into speaker features; a speech encoder including a variational autoencoder having two or more sampling hierarchies that converts source acoustic features obtained by converting the speech of a source speaker in the acoustic feature extractor and source speaker features obtained by converting the speaker label of the source speaker in the speaker encoder into latent representations; a speech decoder including a variational autoencoder having two or more sampling hierarchies that generates target acoustic features using at least the latent representations and target speaker features obtained by converting the speaker label of a target speaker in the speaker encoder; and a vocoder that converts the target acoustic features generated by the speech decoder into speech.
[0011] Here, the speech encoder, the speech decoder, and the speaker encoder are trained to reduce the distance between acoustic features input to the speech encoder and acoustic features generated in the speech decoder.
[0012] It is also preferable that the speech decoder limits a hierarchical level to which the target speaker features are input in the two or more sampling hierarchical levels.
[0013] It is also preferable that the speech decoder does not input the target speaker features to a layer prior to a predetermined layer in the two or more sampling layers, and inputs the target speaker features to a layer subsequent to the predetermined layer.
[0014] It is also preferable that the audio decoder performs sampling from a posterior distribution in a layer prior to the predetermined layer in the two or more sampling layers, and performs sampling from a prior distribution in a layer subsequent to the predetermined layer.
[0015] It is also preferable that the speech decoder inputs the target speaker features to a conditional instance normalization layer.
[0016] Another aspect of the present invention is a speech processing training device comprising: an acoustic feature extractor that converts speech into input acoustic features; a speaker encoder that converts a speaker label of the speech into speaker features; a speech encoder including a variational autoencoder having two or more sampling hierarchies that converts the acoustic features and the speaker features into latent representations; and a speech decoder including a variational autoencoder having two or more sampling hierarchies that generates acoustic features using at least the latent representations and the speaker features, wherein the speech encoder, the speech decoder, and the speaker encoder are trained to reduce the distance between the input acoustic features input to the speech encoder and the output acoustic features generated in the speech decoder.
[0017] Another aspect of the present invention is a speech processing device comprising: an acoustic feature extractor that converts speech into acoustic features; a speaker encoder that converts a speaker label of the speech into speaker features; a speech encoder including a variational autoencoder having two or more sampling hierarchies that converts source acoustic features obtained by converting the speech of a source speaker in the acoustic feature extractor and source speaker features obtained by converting the speaker label of the source speaker in the speaker encoder into latent representations; a speech decoder including a variational autoencoder having two or more sampling hierarchies that generates target acoustic features using at least the latent representations and target speaker features obtained by converting the speaker label of a target speaker in the speaker encoder; and a vocoder that converts the target acoustic features generated by the speech decoder into speech.
[0018] Another aspect of the present invention is a speech processing training method comprising: a speech processing training device including an acoustic feature extractor that converts speech into input acoustic features; a speaker encoder that converts a speaker label of the speech into speaker features; a speech encoder including a variational autoencoder having two or more sampling hierarchies that converts the acoustic features and the speaker features into latent representations; and a speech decoder including a variational autoencoder having two or more sampling hierarchies that generates acoustic features using at least the latent representations and the speaker features, wherein the speech encoder, the speech decoder, and the speaker encoder are trained to reduce the distance between input acoustic features input to the speech encoder and output acoustic features generated in the speech decoder.
[0019] Another aspect of the present invention is a speech processing method for converting the speech of the source speaker to the speech of the target speaker using a speech processing device including: an acoustic feature extractor that converts speech into acoustic features; a speaker encoder that converts a speaker label of the speech into speaker features; a speech encoder including a variational autoencoder having two or more sampling hierarchies that converts source acoustic features obtained by converting the speech of the source speaker in the acoustic feature extractor and source speaker features obtained by converting the speaker label of the source speaker in the speaker encoder into latent representations; a speech decoder including a variational autoencoder having two or more sampling hierarchies that generates target acoustic features using at least the latent representations and target speaker features obtained by converting the speaker label of the target speaker in the speaker encoder; and a vocoder that converts the target acoustic features generated by the speech decoder into speech. Effect of the Invention
[0020] According to the present invention, it is possible to provide a speech processing training program, a speech processing training device, a speech processing training method, a speech processing training program, a speech processing training device, and a speech processing training method that appropriately convert a speech uttered by an arbitrary speaker into a speech uttered by a target speaker. Other objects of the embodiments of the present invention will become apparent by referring to the entire specification. [Brief description of the drawings]
[0021] [Figure 1] 1 is a diagram illustrating a configuration of a voice processing device according to an embodiment of the present invention. [Diagram 2] 1 is a functional block diagram showing a configuration of a speech processing training device according to an embodiment of the present invention. [Diagram 3] This figure shows the configuration of a variational auto-encoder. [Figure 4] This figure shows the configuration of the Nouveau Variational Auto-Encoder. [Diagram 5] FIG. 1 is a diagram showing the configuration of a neural network in each layer of a variational auto-encoder in an embodiment of the present invention. [Figure 6] FIG. 4 is a diagram for explaining a voice learning process according to the embodiment of the present invention. [Figure 7] 1 is a functional block diagram showing a configuration of a pronunciation learning device according to an embodiment of the present invention; [Figure 8] FIG. 2 is a diagram for explaining audio processing according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0022] As shown in FIG. 1, the voice processing device 100 according to the embodiment of the present invention includes a processing unit 10, a storage unit 12, an input unit 14, an output unit 16, and a communication unit 18. The processing unit 10 includes a means for performing arithmetic processing, such as a CPU. The processing unit 10 executes a voice processing learning program stored in the storage unit 12 to learn the voice processing in the present embodiment. The processing unit 10 also executes the voice processing program stored in the storage unit 12 to realize functions related to the voice processing in the present embodiment. The storage unit 12 includes a storage means such as a semiconductor memory or a memory card. The storage unit 12 is connected to the processing unit 10 so as to be accessible, and stores the voice processing learning program, the voice processing program, and information required for the processing. The input unit 14 includes a means for inputting information. The input unit 14 includes, for example, a keyboard, a touch panel, a button, or the like that receives input of information from a user. The input unit 14 also includes a voice input means that receives input of the voice of an arbitrary speaker and a predetermined target speaker. The voice input means may include, for example, a microphone, an amplifier circuit, or the like. The output unit 16 includes a user interface screen (UI) for receiving input information from an administrator and a means for outputting processing results. The output unit 16 includes, for example, a display for presenting an image. The output unit 16 also includes audio output means for outputting synthetic audio generated by the audio processing device 100. The audio output means may include, for example, a speaker, an amplifier, and the like. The communication unit 18 includes an interface for communicating information with an external terminal (not shown) via the network 102. The communication by the communication unit 18 may be wired or wireless. Note that audio information to be used for audio processing may be acquired from an external terminal via the communication unit 18.
[0023] The voice processing device 100 performs voice processing to convert a voice uttered by an arbitrary speaker into the quality of the voice of a predetermined speaker (target speaker). The voice processing device 100 also functions as a voice processing training device that performs training for the voice processing.
[0024] [Voice learning processing] 2 is a functional block diagram showing the configuration of the speech processing device 100 during speech processing training. The speech processing device 100 functions as a speech analysis unit 20, a speaker encoder 22, a speech encoder 24, a speech decoder 26, and a learning device 28. Specifically, the speech processing device 100 functions as a speech processing training device that realizes the following speech training method by executing a speech processing training program.
[0025] The voice analysis unit 20 functions as an acoustic feature extractor that acquires voice data and extracts acoustic features from the voice data. That is, the processing unit 10 of the voice processing device 100 functions as the voice analysis unit 20. The voice data may be acquired by converting the speaker's voice into voice data using a microphone constituting the input unit 14. In addition, voice data previously recorded in an external computer or the like may be received via the communication unit 18. The acquired voice data is stored in the storage unit 12.
[0026] The voice data acquisition process is performed for voices uttered by any speaker. In the voice training process, the voice encoder 24 and the voice decoder 26 are trained using voices from multiple speakers. The voices obtained from each speaker do not need to be identical in content.
[0027] Furthermore, the voice analysis unit 20 performs further voice analysis necessary for voice processing. For example, the voice analysis unit 20 performs cepstral analysis of the voice based on the frequency characteristics of the input voice, and obtains acoustic features such as the spectrum envelope (information indicating the thickness of the voice, etc.) and Mel-frequency cepstral coefficients (MFCC) including information on the fine structure, and the fundamental frequency and resonant frequency of the voice (information indicating the pitch of the voice, hoarseness, etc.). The acoustic features can be, for example, an (80×T)-dimensional Euclidean space for the length T of the voice segment. Specifically, the voice analysis unit 20 obtains acoustic features x i The acoustic features extracted by the speech analysis unit 20 are input to the speech encoder 24 and the learning device 28.
[0028] The speaker encoder 22 converts the ID of the speaker of the voice input to the voice analysis unit 20 into a speaker feature that can be used for voice processing and outputs it. The speaker encoder 22 can be configured to include an embedded module that converts the ID of the speaker into a speaker feature and outputs it. For example, when the speaker ID is i, the speaker encoder 22 converts the speaker feature y i The speaker features generated by the speaker encoder 22 are input to the speech encoder 24 and speech decoder 26.
[0029] In the training of the speech processing device 100, the acoustic feature quantity x obtained from speech uttered by multiple speakers is i and speaker feature y i Combination of (x i ,y i ) set is used.
[0030] The speech encoder 24 receives the acoustic features and speaker features and converts them into latent representations. The speech decoder 26 receives the latent representations and speaker features obtained by the speech encoder 24 and converts them into acoustic features. The latent representations represent the linguistic features of the input speech data.
[0031] As shown in FIG. 2, the speech encoder 24 and the speech decoder 26 receive the acoustic feature x i In response to the input, the speech encoder 24 converts it into a latent representation z, and the speech decoder 26 converts the latent representation z into an acoustic feature x i ^ and reconstruct the output acoustic features x i ^ is the input acoustic feature x i It is learned to restore the
[0032] In this embodiment, the speech encoder 24 and the speech decoder 26 are configured by a variational auto-encoder (VAE). The variational auto-encoder is a type of variational auto-encoder, and generates a latent representation by sampling based on a probability distribution, as shown in FIG. 3. The probability distribution is assumed to be a normal distribution defined by a mean μ and a variance σ. The variational auto-encoder is a combination of an encoder that generates a latent representation z by sampling based on a mean μ and a variance σ for an input X, and a decoder that generates an output X^ from the latent representation z. In the variational auto-encoder, the speaker encoder 22, the speech encoder 24, and the speech decoder 26 are trained so that the restoration error (restoration distance) E between the input X and the output X^ is small.
[0033] As shown in Fig. 4, a typical variational auto-encoder is composed of a single-layer neural network, but in this embodiment, it is preferable to use a Nouveau Variational Auto-Encoder (NVAE) composed of a neural network with two or more layers. In other words, the Nouveau Variational Auto-Encoder is composed of a variational auto-encoder having two or more sampling layers. For example, in the speech processing device 100, it is preferable to configure the speech encoder 24 and the speech decoder 26 with a neural network with n = 35 layers, respectively.
[0034] As shown in FIG. 5, each layer of the nouveau variational auto-encoder of the speech encoder 24 and the speech decoder 26 is composed of a combination of a Conditional-Instance-Normalization layer (CIN layer), a Convolution layer (CONV layer), and a Squeeze-and-Excitation layer (SE layer). The CIN layer is a layer provided in place of a batch normalization hypothesis (BN layer) in a general nouveau variational auto-encoder. The CIN layer is one of the normalization layers, and is a conditional instance normalization layer that performs normalization by setting different parameters for each style. In this embodiment, the CIN layer uses a speaker feature as one of the inputs and performs normalization conditioned by the input speaker feature. The Swish activation function is f(x)=x / (1+e -βx ) is an activation function. The convolution layer applies a convolution operation to the input and outputs the operation result to the next layer. The SE layer adaptively applies attention to the input based on the relationship between channels and outputs weighted features.
[0035] The speech learning process in the speech processing device 100 will be described with reference to FIG. 6. The speech encoder 24 and speech decoder 26 are each configured as a neural network with n layers. The number of layers n can be, for example, 35 layers. Each layer is configured by combining a Conditional-Instance-Normalization layer (CIN layer), a Convolution layer (CONV layer), and a Squeeze-and-Excitation layer (SE layer) shown in FIG. 5. The latent representation output from layer k (where k indicates the number of layers from 1 to n) of the speech encoder 24 is called h k The latent representation output from the layer represented by the layer number k of the speech decoder 26 is denoted by z k As shown in Fig.
[0036] In the speech encoder 24, the acoustic feature x i and speaker feature y iis input, and the latent representation h n In the next layer n-1, the latent representation h n and speaker feature y i is input, and the latent representation h n-1 Similarly, in the kth layer, the latent representation h k+1 and speaker feature y i is input, and the latent representation h k In the final stage, layer 1, the latent representation h 2 and speaker feature y i is input, and the latent representation h 1 The latent expression h 1 The latent representation z 1 In this way, in the speech encoder 24, the speaker feature y i It is preferable to include in the input:
[0037] In the speech decoder 26, the latent representation z 1 is input, and the latent representation z 2 In addition, the latent representation z k is the latent representation z k-1 , the kth layer latent representation h k and speaker feature y i The prior distribution p(z k |z k-1 ,h k ,y i ) can be obtained by sampling from the latent representation z k is the latent representation z k-1 , latent expression z k-2 ... latent expression z 1 and the kth layer latent representation h k The posterior distribution p(z k |z k-1 ,z k-2z 1 ,h k ) It is possible to obtain the distribution p(a|b) by sampling from the output. The distribution p(a|b) is a likelihood function that indicates the likelihood that a will be the output given that b is the precondition.
[0038] In the speech learning process, sampling is performed from the speech encoder 24 across layers from close to the output of the speech decoder 26 to layers far from it. That is, as shown in FIG. 6, the latent representation h k In addition, sampling from the posterior distribution is preferably performed using the speaker feature y i It is preferable not to include in the input.
[0039] That is, the speech decoder 26 stores the speaker feature y i In the input, the speaker feature y i In this case, in the layer where sampling is not performed from the speech encoder 24 but from the prior distribution, the speaker feature y i is included in the input, sampling is performed from the speech encoder 24, and in the layer where sampling is performed from the posterior distribution, the speaker feature y i It is preferable to not include in the input.
[0040] In addition, the sampling is performed using the speaker feature y i In layers that do not include i Do not enter.
[0041] In this configuration, the learning device 28 learns the acoustic feature x i and the reconstructed acoustic features x output from the speech decoder 26. i Various parameters (weighting coefficients or biases of each neuron, etc.) of the neural networks of each layer included in the speaker encoder 22, the voice encoder 24, and the voice decoder 26 are adjusted so that the error (distance) with ^ becomes small.
[0042] Here, the acoustic feature x input to the speech decoder 26 is i and the reconstructed acoustic features x output from the speech decoder 26. i The speech decoder 26 calculates the speaker feature y i The hierarchy for sampling from the prior distribution considering the speaker feature y i It is only necessary to appropriately set a stratum that is a boundary between the stratum where sampling is performed from the posterior distribution that does not take into account the above.
[0043] As described above, the acoustic feature x i and the acoustic features x reconstructed in the speech decoder 26. i The speech encoder 24 and speech decoder 26 are trained to approximate the speech represented by ^.
[0044] [Audio processing] 7 is a functional block diagram showing the configuration of the voice processing device 100 during voice processing for converting the voice uttered by the source speaker into a voice uttered by the target speaker. The voice processing device 100 functions as a voice analysis unit 20, a speaker encoder 22, a voice encoder 24, a voice decoder 26, and a vocoder 30. Specifically, the voice processing device 100 functions as a voice processing device that realizes the following voice processing by executing a voice processing program.
[0045] The speech analysis unit 20 acquires speech data of the speech uttered by the source speaker, and performs speech analysis necessary for speech processing. The acoustic features extracted by the speech analysis unit 20 are input to a speech encoder 24.
[0046] The speaker encoder 22 converts the IDs of the source speaker and the target speaker into speaker features that can be used for speech processing and outputs them. When the source speaker ID is the speaker of s, the speaker encoder 22 converts the source speaker feature y sand outputs the target speaker feature y t and outputs it to the audio decoder 26.
[0047] The speech encoder 24 receives acoustic features and source speaker features obtained from the speech of the source speaker, and converts the acoustic features and source speaker features into latent representations. The speech decoder 26 receives latent representations and target speaker features obtained by the speech encoder 24, and reconstructs acoustic features from the latent representations and target speaker features.
[0048] The voice processing in the voice processing device 100 will be described with reference to Fig. 8. The voice processing is performed using the voice encoder 24 and the voice decoder 26 that have been trained in the above-mentioned voice training process.
[0049] In the speech encoder 24, the acoustic feature x obtained from the speech of the source speaker for the layer n is s and the source speaker feature y s is input, and the latent representation h n As in the learning process, in layer k, the latent representation h k+1 and the source speaker feature y s is input, and the latent representation h k In the final stage, layer 1, the latent representation h 2 and the source speaker feature y s is input, and the latent representation h 1 The latent expression h 1 The latent representation z 1 is sampled.
[0050] In the speech decoder 26, the latent representation z 1 is input, and the latent representation z 2 In the hierarchy far from the output of the speech decoder 26, the target speaker feature yt is not included in the input, and the speech decoder 26 uses the latent representations z k-1 , latent expression z k-2 ... latent expression z 1 and the kth layer latent representation h k The posterior distribution p(z k |z k-1 ,z k-2 z 1 ,h k ) is sampled. In the layers close to the output of the speech decoder 26, sampling is not performed from the speech encoder 24, and the latent representation z k-1 and the target speaker feature y t Prior distribution p(z k |z k-1 ,y t 8 shows an example in which sampling is performed from the prior distribution in layers n-1 and n of the speech decoder 26. In this case, the sampling from the prior distribution is performed using the source speaker feature y s Instead, the target speaker feature y t It is preferable to include in the input:
[0051] Through speech processing in the speech encoder 24 and speech decoder 26, the acoustic feature x obtained from the speech of the source speaker is output from the layer n, which is the final stage of the speech decoder 26. s The acoustic features x t is constructed and output.
[0052] The vocoder 30 converts the acoustic feature x output from the speech decoder 26 into t The vocoder 30 converts the speech data into speech data and outputs it. The vocoder 30 performs the reverse process of the process performed by the speech analyzer 20 to extract the speech features from the speech data, thereby obtaining the speech features x t can be converted into audio data.
[0053] As described above, the speech processing device 100 of this embodiment can provide a speech processing device, a speech processing program, a speech processing method, and a speech learning processing device, a speech learning processing program, and a speech learning processing method that appropriately convert a speech uttered by an arbitrary speaker into the quality of a speech uttered by a target speaker. In other words, the speech processing device 100 including the trained speech encoder 24 and speech decoder 26 can realize speech processing that converts a speech uttered by a source speaker into a speech that sounds like a target speaker.
[0054] In particular, by applying the Nouveau Variational Auto-Encoder (NVAE) to the speech encoder 24 and the speech decoder 26, the speech of the source speaker can be converted into a more natural-sounding speech of the target speaker than in the past. [Explanation of symbols]
[0055] 10 processing unit, 12 memory unit, 14 input unit, 16 output unit, 18 communication unit, 20 speech analysis unit, 22 speaker encoder, 24 speech encoder, 26 speech decoder, 28 learning device, 30 vocoder, 100 speech processing device, 102 network.
Claims
1. Computer, an acoustic feature extractor that converts speech into input acoustic features; A speaker encoder that converts the speaker labels of the speech into speaker features; A speech encoder including a variational autoencoder having a sampling hierarchy of hierarchy n (where n is an integer equal to or greater than 2) that converts input acoustic features and speaker features into latent representations; A speech decoder including a variational autoencoder having a sampling hierarchy of n layers that generates acoustic features using at least a latent representation and speaker features; and functioning as a speech processing learning device having the In a sampling hierarchical level n of the first stage in the speech encoder, the input acoustic feature and the speaker feature are input, and a latent representation h is output; In each sampling hierarchical layer (m-1) (where m is an integer from n to 2) from the second stage onwards in the speech encoder, a latent representation h(m-1) is output in response to an input of the latent representation h(m-1) output from the sampling hierarchical layer (m) in the previous stage in the speech encoder and the speaker feature; In the sampling layer of the first layer 1 of the speech decoder, a latent representation h1 output from the final layer 1 of the speech encoder is received as an input and a latent representation Z1 is output; In each sampling hierarchical layer k (where k is an integer from 2 to n) from the second stage onwards in the speech decoder, a latent representation Z(k-1) output from a sampling hierarchical layer (k-1) in the preceding stage in the speech decoder and a latent representation hk output from hierarchical layer k in the speech encoder are input, and a latent representation Zk is output; The speech encoder, the speech decoder, and the speaker encoder are trained to reduce the distance between an input acoustic feature input to the speech encoder and an output acoustic feature generated from a latent representation Zn output from the speech decoder.
2. 2. The speech processing training program according to claim 1, 2. A speech processing training program, comprising: a speech decoder, the speech decoder being configured such that a hierarchy for inputting speaker features is limited in the sampling hierarchy.
3. 3. The speech processing training program according to claim 2, the speech decoder does not input speaker features to layers preceding a predetermined layer in the sampling layers, and inputs speaker features to layers following the predetermined layer.
4. 4. The speech processing training program according to claim 3, The speech decoder is characterized in that it samples from a posterior distribution in layers prior to the specified layer in the sampling layer, and samples from a prior distribution in layers subsequent to the specified layer.
5. The speech processing training program according to any one of claims 1 to 4, The speech processing training program is characterized in that the speech decoder inputs speaker features to a conditional instance normalization layer.
6. Computer, an acoustic feature extractor that converts speech into acoustic features; The speaker encoder, the speech encoder, and the speech decoder trained by the speech processing training program according to claim 1; A vocoder for converting acoustic features into speech; and make it work. inputting source acoustic features obtained by converting the speech of the source speaker in the acoustic feature extractor and source speaker features obtained by converting the speaker label of the source speaker in the speaker encoder into the speech encoder and converting them into latent representations; The latent representation and a target speaker feature obtained by converting the speaker label of the target speaker in the speaker encoder are input to the speech decoder to generate a target acoustic feature; A speech processing program, comprising: inputting the target acoustic feature generated by the speech decoder into the vocoder to convert the target acoustic feature into speech.
7. an acoustic feature extractor that converts speech into input acoustic features; A speaker encoder that converts the speaker labels of the speech into speaker features; a speech encoder including a variational autoencoder having a sampling hierarchy of n (where n is an integer equal to or greater than 2) that converts acoustic features and speaker features into a latent representation; A speech decoder including a variational autoencoder having a sampling hierarchy of n layers that generates acoustic features using at least a latent representation and speaker features; Equipped with In a sampling hierarchical level n of the first stage in the speech encoder, the input acoustic feature and the speaker feature are input, and a latent representation h is output; In each sampling hierarchical layer (m-1) (where m is an integer from n to 2) from the second stage onwards in the speech encoder, a latent representation h(m-1) is output in response to an input of the latent representation h(m-1) output from the sampling hierarchical layer (m) in the previous stage in the speech encoder and the speaker feature; In the sampling layer of the first layer 1 of the speech decoder, a latent representation h1 output from the final layer 1 of the speech encoder is received as an input and a latent representation Z1 is output; In each sampling hierarchical layer k (where k is an integer from 2 to n) from the second stage onwards in the speech decoder, a latent representation Z(k-1) output from a sampling hierarchical layer (k-1) in the preceding stage in the speech decoder and a latent representation hk output from hierarchical layer k in the speech encoder are input, and a latent representation Zk is output; The speech processing training device is characterized in that the speech encoder, the speech decoder, and the speaker encoder are trained to reduce the distance between an input acoustic feature input to the speech encoder and an output acoustic feature generated from a latent representation Zn output from the speech decoder.
8. an acoustic feature extractor that converts speech into acoustic features; The speaker encoder, the speech encoder, and the speech decoder trained by the speech processing training device according to claim 7; A vocoder for converting acoustic features into speech; Equipped with inputting source acoustic features obtained by converting the speech of the source speaker in the acoustic feature extractor and source speaker features obtained by converting the speaker label of the source speaker in the speaker encoder into the speech encoder and converting them into latent representations; The latent representation and a target speaker feature obtained by converting the speaker label of the target speaker in the speaker encoder are input to the speech decoder to generate a target acoustic feature; The speech processing device according to claim 1, wherein the target acoustic feature generated by the speech decoder is input to the vocoder to convert the target acoustic feature into speech.
9. an acoustic feature extractor that converts speech into input acoustic features; A speaker encoder that converts the speaker labels of the speech into speaker features; a speech encoder including a variational autoencoder having a sampling hierarchy of n (where n is an integer equal to or greater than 2) that converts acoustic features and speaker features into a latent representation; A speech decoder including a variational autoencoder having a sampling hierarchy of n layers that generates acoustic features using at least a latent representation and speaker features; A speech processing training device comprising: In a sampling hierarchical level n of the first stage in the speech encoder, the input acoustic feature and the speaker feature are input, and a latent representation h is output; In each sampling hierarchical layer (m-1) (where m is an integer from n to 2) from the second stage onwards in the speech encoder, a latent representation h(m-1) is output in response to an input of the latent representation h(m-1) output from the sampling hierarchical layer (m) in the previous stage in the speech encoder and the speaker feature; In the sampling layer of the first layer 1 of the speech decoder, a latent representation h1 output from the final layer 1 of the speech encoder is received as an input and a latent representation Z1 is output; In each sampling hierarchical layer k (where k is an integer from 2 to n) from the second stage onwards in the speech decoder, a latent representation Z(k-1) output from a sampling hierarchical layer (k-1) in the preceding stage in the speech decoder and a latent representation hk output from hierarchical layer k in the speech encoder are input, and a latent representation Zk is output; The speech processing training method is characterized in that the speech encoder, the speech decoder, and the speaker encoder are trained to reduce the distance between an input acoustic feature input to the speech encoder and an output acoustic feature generated from a latent representation Zn output from the speech decoder.
10. an acoustic feature extractor that converts speech into acoustic features; The speaker encoder, the speech encoder and the speech decoder trained by the speech processing training method of claim 9; A vocoder for converting acoustic features into speech; In a voice processing device comprising: inputting source acoustic features obtained by converting the speech of the source speaker in the acoustic feature extractor and source speaker features obtained by converting the speaker label of the source speaker in the speaker encoder into the speech encoder and converting them into latent representations; The latent representation and a target speaker feature obtained by converting the speaker label of the target speaker in the speaker encoder are input to the speech decoder to generate a target acoustic feature; a voice processing method comprising: inputting the target acoustic feature generated by the voice decoder to the vocoder to convert the target acoustic feature into voice;
Citation Information
Patent Citations
Voice conversion learning device, voice conversion device, method and program
JP2019144402A