Training device, conversion device, training method, conversion method, and program
The neural network-based voice conversion model addresses the constraint of identical sentence pairs in voice quality conversion by estimating gradients and incorporating phoneme features, enhancing accuracy and efficiency in converting voice attributes.
Patent Information
- Application Number
- PCT/JP2024/000968
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-24
AI Technical Summary
Existing voice quality conversion technologies using machine learning require paired voice signals of the same sentence for learning data, constraining the data preparation process.
A neural network-based voice conversion model that converts source voice features to target voice features using a probability density function, allowing for the estimation of gradients towards a nearest stationary point, and incorporates phoneme features to improve accuracy without requiring identical sentence pairs in the learning data.
Relaxes the constraints on learning data by enabling voice conversion with improved accuracy and efficiency, allowing for voice attribute changes like speaker gender without needing identical sentence pairs.
Smart Images

Figure JP2024000968_24072025_PF_FP_ABST
Abstract
Description
Learning device, conversion device, learning method, conversion method, and program
[0001] The present invention relates to a learning device, a conversion device, a learning method, a conversion method, and a program.
[0002] Voice conversion is a technology that converts only the non-linguistic and paralinguistic components (such as speaker identity and speaking style) of input speech while preserving the linguistic information (spoken sentences) and is expected to be applied to speaker identity conversion in text-to-speech synthesis, speech assistance, voice enhancement, pronunciation conversion, etc. One voice conversion technology that has been proposed is the use of machine learning (Patent Document 1 and Non-Patent Document 1).
[0003] International Publication No. 2022 / 101967
[0004] Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo, Shogo Seki, "VoiceGrad: Non-Parallel Any-to-Many Voice Conversion with Annealed Langevin Dynamics," arXiv:2010.02977 [cs.SD], Oct. 2020.
[0005] However, when using the machine learning techniques proposed so far, it was necessary to prepare a pair of a source speech signal sample and a correct speech signal sample as training data. Furthermore, the two speech signal samples contained in the training data had to be the same sentence spoken. For example, if the source speech signal sample was the result of speaking the sentence "Good morning," the corresponding correct speech signal sample also had to be the same sentence spoken. As such, conventional techniques had a constraint on the training data that had to be prepared, in that both the source speech signal sample and the correct speech signal sample had to be the same sentence spoken. The target speech is a speech having predetermined attributes, such as attributes specified by a user.
[0006] In view of the above circumstances, an object of the present invention is to provide a technology that relaxes the constraints imposed on data used for training in voice conversion technology using machine learning.
[0007] One aspect of the present invention is a learning device comprising: a first control unit that performs training of a score approximator, the score approximator being a neural network included in a speech conversion model that is a mathematical model for converting a source feature sequence, which is a sequence of speech features of a source speech signal, into a target speech feature sequence, which is a sequence of speech features of a target speech signal, which is a speech signal belonging to a target sound attribute with respect to a sound attribute, which is an attribute related to sound; the neural network receiving at least a speech feature sequence, which is a sequence of speech features obtained from a speech signal, and a phoneme feature sequence obtained from the speech feature sequence; the neural network performing estimation based on an input speech feature sequence, which is the input speech feature sequence, and an input phoneme feature sequence, which is the input phoneme feature sequence; and the score approximator being a parameterized neural network that estimates a stop point of a target feature distribution function, which is a function in a vector space representing the speech feature sequence and is a probability density function representing a distribution of the target speech feature sequence, the target feature distribution function being a function in a vector space representing the speech feature sequence and representing the distribution of the target speech feature sequence, the target feature distribution function being a probability density function representing the distribution of the target speech feature sequence, the target feature distribution function being a stop point, which is a point in the vector space, and which is the stop point nearest to the current point, which is a point in the vector space representing the input speech feature sequence.
[0008] One aspect of the present invention is a neural network included in a speech conversion model, which is a mathematical model for converting a source feature sequence, which is a sequence of speech features of a source speech signal, into a target speech feature sequence, which is a sequence of speech features of a target speech signal, which is a speech signal belonging to a target sound attribute with respect to a sound attribute, which is an attribute related to sound. The neural network is input with at least a speech feature sequence, which is a sequence of speech features obtained from a speech signal, and a phoneme feature sequence, which is obtained from the speech feature sequence, and the neural network is configured to perform a speech conversion process based on the input speech feature sequence, which is the input speech feature sequence, and the input phoneme feature sequence, which is the input phoneme feature sequence. a first control unit that performs training of a score approximator, which is a parameterized neural network that estimates a gradient at a current point of a path leading to a nearest stop point, which is a stop point of a target feature distribution function, which is a function in a vector space that represents the speech feature sequence and is a probability density function that represents a distribution of the target speech feature sequence, and which is a point in the vector space and is the nearest stop point to a current point that is a point in the vector space that represents the input speech feature sequence; and a second control unit that performs estimation using the trained score approximator obtained by a training device that includes: a first control unit that performs training of a score approximator, which is a parameterized neural network that estimates a gradient at a current point of a path leading to a nearest stop point, which is a point in the vector space that represents the input speech feature sequence.
[0009] One aspect of the present invention is a neural network included in a speech conversion model, which is a mathematical model for converting a source feature sequence, which is a sequence of speech features of a source speech signal, into a target speech feature sequence, which is a sequence of speech features of a target speech signal, which is a speech signal belonging to a target sound attribute with respect to a sound attribute, which is an attribute related to sound. The neural network is input with at least a speech feature sequence, which is a sequence of speech features obtained from a speech signal, and a phoneme feature sequence obtained from the speech feature sequence, and the input speech feature sequence is the input speech feature sequence and the input phoneme feature sequence. a first control step of training a score approximator, which is a parametrized neural network that performs estimation based on an input phoneme feature sequence that is a feature sequence, and a target feature distribution function that is a function in a vector space that represents the speech feature sequence and is a probability density function that represents a distribution of the target speech feature sequence, and that estimates a gradient at the current point of a path leading to a nearest stop point that is a point in the vector space that is nearest to the current point that is a point in the vector space that represents the input speech feature sequence.
[0010] One aspect of the present invention is a neural network included in a speech conversion model, which is a mathematical model for converting a source feature sequence, which is a sequence of speech features of a source speech signal, into a target speech feature sequence, which is a sequence of speech features of a target speech signal, which is a speech signal belonging to a target sound attribute with respect to a sound attribute, which is an attribute related to sound, and which receives as input at least a speech feature sequence, which is a sequence of speech features obtained from a speech signal, and a phoneme feature sequence obtained from the speech feature sequence, and which performs estimation based on an input speech feature sequence, which is the input speech feature sequence, and an input phoneme feature sequence, which is the input phoneme feature sequence. a target feature distribution function, which is a function in a vector space representing the speech feature sequence and is a probability density function representing a distribution of the target speech feature sequence, and a gradient at the current point of a path leading to a nearest stationary point, which is a point in the vector space and is the nearest stationary point to the current point, which is a point in the vector space representing the input speech feature sequence; and a second control step of performing estimation using the trained score approximator obtained by a training method having the first control step of training a score approximator, which is a neural network that performs
[0011] One aspect of the present invention is a program for causing a computer to function as the learning device described above.
[0012] One aspect of the present invention is a program for causing a computer to function as the above-described conversion device.
[0013] The present invention makes it possible to provide a technology that relaxes the constraints imposed on data used for learning in voice conversion technology using machine learning.
[0014] 1 is an explanatory diagram illustrating an overview of a conversion system 100 according to an embodiment. A diagram showing a network structure of a score approximator used in an experiment according to an embodiment. A first diagram showing an example of an experimental result according to an embodiment. A second diagram showing an example of an experimental result according to an embodiment. A third diagram showing an example of an experimental result according to an embodiment. A fourth diagram showing an example of an experimental result according to an embodiment. A fifth diagram showing an example of an experimental result according to an embodiment. A sixth diagram showing an example of an experimental result according to an embodiment. A seventh diagram showing an example of an experimental result according to an embodiment. A eighth diagram showing an example of an experimental result according to an embodiment. A ninth diagram showing an example of an experimental result according to an embodiment. A tenth diagram showing an example of an experimental result according to an embodiment. A eleventh diagram showing an example of an experimental result according to an embodiment. A twelfth diagram showing an example of an experimental result according to an embodiment. A thirteenth diagram showing an example of an experimental result according to an embodiment. A diagram showing an example of the hardware configuration of a learning device according to an embodiment. A flowchart showing an example of the flow of processing performed by a learning device according to an embodiment. A diagram showing an example of the hardware configuration of a conversion device according to an embodiment. A flowchart showing an example of the flow of processing performed by a conversion device according to an embodiment. A diagram showing Algorithm 1 in a modified example. FIG. 10 is a diagram showing algorithm 2 in a modified example.
[0015] 1 is an explanatory diagram illustrating an overview of a conversion system 100 according to an embodiment. The conversion system 100 includes a learning device 1 and a conversion device 2. The learning device 1 executes a learning process. The learning process is a process for learning a score approximator.
[0016] The score approximator is a neural network included in the speech conversion model. The network structure of the score approximator can be designed arbitrarily, with no constraints other than the requirement that the input and output be isomorphic. The speech conversion model is a mathematical model that converts a source feature sequence, which is a sequence of speech features of the source speech signal, into a target speech feature sequence.
[0017] The target speech feature sequence is a sequence of speech features of a target speech signal. The target speech signal is a speech signal that belongs to a target speech attribute in terms of a sound attribute, which is an attribute related to a sound.
[0018] The target sound attribute is a sound attribute that is a target. The target sound attribute is, for example, a sound attribute that is specified by a user or determined during a learning process. The sound attribute is, for example, a speaker. The sound attribute may also be, for example, the gender of the speaker. For simplicity of the following explanation, the case where the sound attribute is a speaker will be described as an example.
[0019] Since the target sound attribute is the target sound attribute, the voice conversion model converts, for example, based on the source feature sequence, into a sequence of voice features of a voice signal that indicates a voice that would be uttered by a target speaker who is different from the speaker of the voice indicated by the source voice signal, but whose paralinguistic information is the same as that indicated.
[0020] A sequence of speech features of a speech signal indicating speech that would be uttered by a speaker different from the speaker of the speech indicated by the source speech signal but that indicates the same paralinguistic information is an example of a target speech feature sequence.
[0021] Therefore, the sound attributes to which the voice signal estimated by the voice conversion model belongs are closer to the target sound attributes than the sound attributes to which the source voice signal belongs.
[0022] The speech features are features obtained from a speech signal. The speech features may be any features sufficient to constitute a speech signal, and may be obtained, for example, by Fourier transforming the speech signal. The speech features may be, for example, vocoder parameters. The speech features may be, for example, Mel-Cepstrum vocoders. Other examples of speech features will be described in the modified examples.
[0023] <Score Approximator> As described above, the learning target in the learning process is the score approximator, which is a neural network included in the voice conversion model. In fact, the score approximator is a neural network that also satisfies the following conditions:
[0024] One of the conditions that a score approximator must satisfy is that it receives at least a sequence of speech features obtained from a speech signal (hereinafter referred to as a "speech feature sequence") and a sequence of phoneme features obtained from the speech feature sequence. Note that the phoneme feature sequence is a well-known phoneme feature sequence. Therefore, the phoneme feature sequence is a sequence of features of phonemes. Therefore, the phoneme feature sequence indicates linguistic information.
[0025] The phoneme feature sequence is, for example, an output sequence (ASR output sequence) of an Automatic Speech Recognition (ASR) device. The phoneme feature sequence may also be, for example, a Bottleneck Feature (BNF) sequence of the ASR device.
[0026] One of the conditions that the score approximator must satisfy is that it performs estimation based on the above-mentioned input speech feature sequence (hereinafter referred to as "input speech feature sequence") and the above-mentioned input phoneme feature sequence (hereinafter referred to as "input phoneme feature sequence"). Note that the input speech feature sequence is, for example, a source feature sequence, but it does not necessarily have to be identical to the source feature sequence.
[0027] For example, a case will be described in which the speech conversion model is a first example model. The first example model is a speech conversion model in which unit example processing is repeated to ultimately estimate a target speech feature sequence. The unit example processing is a process in which an input phoneme feature sequence is obtained from an input speech feature sequence, a speech feature sequence is estimated based on the input speech feature sequence and the estimation result of a score approximator based on the input phoneme feature sequence, and the estimated speech feature sequence is input to the next unit example processing as a new input speech feature sequence.
[0028] In such a case, the input speech feature sequence is a source feature sequence in the unit exemplification process executed for the first time, but is the speech feature sequence estimated in the unit exemplification process executed immediately before in the unit exemplification process executed for the second time.
[0029] One of the conditions that the score approximator satisfies is that the estimation by the score approximator based on the input speech feature sequence and the input phoneme feature sequence estimates the gradient at the current point of the path toward the nearest station point, which is a station point of the target feature distribution function that is the nearest station point to the current point. The target feature distribution function is a function in the speech feature space and is a probability density function that represents the distribution of the target speech feature sequence.
[0030] The speech feature space is a vector space that represents a sequence of speech features, and is therefore a type of so-called feature space.
[0031] The current point is a point in the speech feature space, i.e., a point in the speech feature space representing the input speech feature sequence. Note that the stationary point is, for example, a local maximum point. Note that the target feature distribution function may be continuous and differentiable.
[0032] One of the conditions that the score approximator must satisfy is that it is a parameterized neural network, which is natural since the score approximator is a learning target in the learning process.
[0033] In this way, the score approximator is a neural network that estimates a so-called score function.
[0034] <Effects of inputting an input phoneme feature sequence to a score approximator> As described above, not only an input speech feature sequence but also an input phoneme feature sequence is input to the score approximator. On the other hand, what the score approximator estimates is a gradient in the speech feature space. Therefore, it may seem that inputting the input phoneme features to the score approximator is unnecessary for estimation by the score approximator. However, this is not the case.
[0035] As described above, the input phoneme feature sequence is obtained from the input speech feature sequence, and therefore represents an extracted portion of the information contained in the input speech feature sequence. Therefore, when the input speech feature sequence is input, information about the input phoneme feature sequence is also input to the score approximator.
[0036] However, unless the input phoneme feature sequence is also input, there is a possibility that the score approximator will not be able to use the input phoneme feature sequence. The process of inputting the input phoneme feature sequence to the score approximator has the technical significance of improving this point, and this process can be said to be a process of reliably providing the score approximator with information contained in the input speech feature sequence that the score approximator may not have been able to obtain if the input phoneme feature sequence had not been input.
[0037] In this way, when an input phoneme feature sequence is input to the score approximator, the score approximator can make estimations that use more information contained in the input speech feature sequence than when the input phoneme feature sequence is not used, thereby improving the accuracy of the score approximator's estimations.
[0038] To explain this in more detail, phoneme feature sequences, including ASR output sequences and BNF sequences, are sequences that contain almost no information other than the linguistic information of the input speech. Therefore, there is almost no change in the phoneme feature sequence between speech before and after speech conversion.
[0039] Therefore, by inputting the phoneme feature sequence of the speech used for training to the network during training of the score approximator, and inputting the phoneme feature sequence of the source speech during testing, speech conversion can be performed using the linguistic information of the source speech as a clue. Therefore, when an input phoneme feature sequence is input to the score approximator, the score approximator can perform more accurate estimation. In other words, the accuracy of the score approximator's estimation improves.
[0040] <More details about the speech conversion model> As described above, the speech conversion model includes a score approximator. The score approximator receives not only an input speech feature sequence but also at least an input phoneme feature sequence. Therefore, the speech conversion model includes an extractor (hereinafter referred to as a "phoneme extractor") that acquires an input phoneme feature sequence from the input speech feature sequence.
[0041] In the speech conversion model, a phoneme extractor obtains an input phoneme feature sequence each time an input speech feature is obtained. The phoneme extractor is a BNF extractor that extracts bottleneck features. A BNF extractor is obtained, for example, by training an entire neural network, in which a bottleneck layer is inserted between the encoder and decoder of an end-to-end ASR device having an encoder-decoder structure, with a speech recognition corpus, and then deleting all layers after the inserted bottleneck layer.
[0042] As described above, the speech conversion model is a mathematical model including a score approximator that converts a source feature sequence into a target speech feature sequence. Therefore, the speech conversion model is a mathematical model that converts a source feature sequence into a target speech feature sequence using a score approximator. Therefore, the speech conversion model may be any mathematical model as long as it converts a source feature sequence into a target speech feature sequence using a score approximator.
[0043] Therefore, the speech conversion model may be a mathematical model that obtains a target speech feature sequence by sequentially and repeatedly executing, for example, the execution of a speech extractor, a score function estimation process such as DSM (Denoising Score Matching) or weighted DSM, and a spatial point update process such as Langevin dynamics or simulated annealed Langevin dynamics. The score function estimation process is a process of estimating a score function at a current point. Therefore, the score function estimation process is, for example, a process of executing a score approximator. The spatial point update process is a process of updating a current point. Note that DSM is also called noise-removing score matching.
[0044] The Langevin dynamics is a technique described in detail in, for example, Reference 1. The DSM is a technique described in detail in, for example, Reference 2. The weighted DSM is a technique described in detail in, for example, Reference 3. The simulated Langevin dynamics is a technique described in detail in, for example, Reference 3.
[0045] Reference 1: M. Welling and YW Teh, “Bayesian Learning via Stochastic Gradient Langevin Dynamics,” in Proc. ICML, pp. 681-688, 2011. Reference 2: P. Vincent, “A Connection Between Score Matching and Denoising Autoencoders,” Neural Computation, Vol 23, No. 7, pp. 1661-1674, 2011. Reference 3: Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” in Advances in Neural Information Processing Systems 32, 2019, pp. 11918-11930
[0046] The speech conversion model may also be a mathematical model that obtains a target speech feature sequence by executing a de-diffusion process in a diffusion probability model. It goes without saying that the diffusion probability model itself is a well-known technique. However, it may not be easy to understand how the speech conversion model can obtain a target speech feature sequence. This is because understanding this requires a complex theory using mathematical formulas (hereinafter referred to as the "Kameoka-Kaneko-Tanaka theory"). Although this is described as a complex theory, it goes without saying that it is a theory that can be understood by following the theoretical development of the explanation.
[0047] Because the Kameoka-Kaneko-Tanaka theory is a complex theory using mathematical formulas, it will be explained in the modified example <Kameoka-Kaneko-Tanaka theory>, and here we will continue to explain the conversion system 100 using the results of the Kameoka-Kaneko-Tanaka theory. However, to briefly explain the technical idea behind obtaining a target speech feature sequence by executing a de-diffusion process in a diffusion probability model, this technology is based on the idea of treating an input speech feature sequence as the result of a diffusion process in a diffusion probability model.
[0048] According to the Kameoka-Kaneko-Tanaka theory, to construct a mathematical model for obtaining a target speech feature sequence by executing a de-diffusion process in a diffusion probability model, it is sufficient to update the parameters of the score approximator so as to minimize a first objective function in the training of the score approximator. The first objective function is an objective function in the diffusion probability model. The first objective function may be, for example, a loss expressed by equation (A30) described in the Kameoka-Kaneko-Tanaka theory, or a loss expressed by equation (A32) described in the Kameoka-Kaneko-Tanaka theory. Therefore, in the training process performed by the training device 1, the parameters of the score approximator are updated so as to minimize the first objective function, for example.
[0049] Furthermore, as explained in the Kameoka-Kaneko-Tanaka theory, when the speech conversion model is a mathematical model that uses DSM or weighted DSM and Langevin dynamics or simulated annealed Langevin dynamics, the objective function used in training the score approximator is the second objective function. The second objective function is different for each score approximator and is the sum of the difference between the data at the current point to which noise has been added and the data at the current point before noise was added, and the difference between the score function.
[0050] Therefore, the second objective function may be, for example, the loss expressed by equation (A15) described in the Kameoka-Kaneko-Tanaka theory, or the loss expressed by equation (A36) described in the Kameoka-Kaneko-Tanaka theory. Therefore, when the speech conversion model is a mathematical model that uses DSM or weighted DSM and Langevin dynamics or simulated annealed Langevin dynamics, the parameters of the score approximator are updated in training the score approximator so as to minimize the second objective function.
[0051] <Effects of the Learning Process> Here, it will be explained that the learning process performed by the learning device 1 is a technique for relaxing the constraints imposed on data used for learning in voice conversion technology using machine learning. As described above, the learning process performed by the learning device 1 involves learning a score approximator. Also, as described above, speech can be converted using a score function estimated by the score approximator. Therefore, relaxing the constraints imposed on data used for training the score approximator means relaxing the constraints imposed on data used for learning in voice conversion technology using machine learning.
[0052] Incidentally, as can be seen from the explanation of the Kameoka-Kaneko-Tanaka theory, in the training of a score approximator when score function estimation processing such as DSM or weighted DSM and spatial point update processing such as Langevin dynamics or simulated annealed Langevin dynamics are used, it is not necessarily necessary to acquire the shape of the target feature distribution function p(x) in advance as prior information. Note that the target feature distribution function is a function in the speech feature space, and is a probability density function that represents the distribution of the target speech feature sequence.
[0053] Furthermore, as explained in the explanation of the Kameoka-Kaneko-Tanaka theory, the execution of the de-diffusion process in the diffusion probability model and the execution of the score function estimation process and the spatial point update process have similar mathematical structures. Therefore, even in training a score approximator used in a mathematical model that obtains a target speech feature sequence by executing the de-diffusion process in the diffusion probability model, the shape of the target feature distribution function p(x) does not necessarily need to be acquired in advance as prior information.
[0054] In this way, the shape of the target feature distribution function p(x) does not need to be acquired in advance as prior information in the learning process performed by the learning device 1. Therefore, the learning process performed by the learning device 1 can relax the constraints imposed on the data used for learning in voice conversion technology using machine learning.
[0055] <Examples of Items Other Than Input Speech Feature Sequence and Input Phoneme Feature Sequence Input to Score Approximator> In the learning process, in addition to the input speech feature sequence and input phoneme feature sequence, information indicating target sound attributes (hereinafter referred to as "target sound attribute information") may also be input to the score approximator. The target sound attribute information is input to the voice conversion model. This can be said to be a process of explicitly inputting the target sound attributes to the score approximator.
[0056] When a trained score approximator to which target voice attribute information is also input during training is used, the target voice attribute information is also input to the trained score approximator. When target voice attribute information is input to a voice conversion model, the voice conversion model obtains a target voice feature sequence whose attribute is closest to the attribute indicated by the target voice attribute information. Hereinafter, the target voice attribute information when the voice attribute is the speaker is referred to as a speaker index.
[0057] In the training process, in addition to the input speech feature sequence and the input phoneme feature sequence, information indicating a noise level (hereinafter referred to as "noise level information") explained in <Kameoka-Kaneko-Tanaka theory> may be input to the score approximator. When a trained score approximator in which information indicating a noise level has been input during training is used, the information indicating the noise level is also input to the trained score approximator.
[0058] Therefore, for example, an input speech feature sequence, an input phoneme feature sequence, target sound attribute information, and noise level information may be input to the score approximator.
[0059] <Regarding the conversion device 2> The conversion device 2 performs estimation using the trained score approximator obtained by the learning device 1. Specifically, the estimation is estimation of a conversion destination speech feature sequence based on a speech feature sequence to be converted. Therefore, the estimation using the trained score approximator is, for example, a trained speech conversion model. The trained speech conversion model is a speech conversion model that includes a trained score approximator.
[0060] Since a target speech feature sequence is obtained through estimation, it can be said that the conversion device 2 converts the target speech feature sequence into the target speech feature sequence. Hereinafter, the process of converting the target speech feature sequence into the target speech feature sequence will be referred to as the conversion process. For the reasons described above, the conversion process can also be said to be a process of estimating a target speech feature sequence based on the target speech feature sequence.
[0061] <Regarding the Learning Phase> In the learning phase, at least a source feature sequence is input to the speech conversion model to be trained. If target sound attribute information is input to the score approximator, the target sound attribute information is also input to the speech conversion model to be trained. Furthermore, if eye noise level information is input to the score approximator, the noise level information is also input to the speech conversion model to be trained.
[0062] It goes without saying that the user may or may not prepare correct labels as data to be used for training (i.e., training data). For example, if there is another mathematical model (hereinafter referred to as a "correct label estimation model") that is different from the speech conversion model and that estimates a target speech feature sequence based on a source feature sequence with higher accuracy than the speech conversion model being trained, the user does not need to prepare correct labels.
[0063] This is because it is sufficient if the target speech feature sequence estimated by the correct label estimation model is used as the correct label in the training process. In this way, if a correct label estimation model exists, the correct label is estimated by executing the training process, so the user does not need to prepare the correct label. In this way, the correct label may or may not be prepared when executing the training process. Furthermore, since the training simply involves executing the learning object and minimizing the objective function, the information prepared by the user may be any information as long as it provides the information necessary to obtain the value of the objective function.
[0064] <About the Experiment> An example of the results of an experiment on the conversion of a speech signal using the conversion system 100 of the embodiment will be described. In the experiment, a comparison was made with other speech conversion technologies for a case where the speech conversion model was a mathematical model that obtains a target speech feature sequence by executing a de-diffusion process in a diffusion probability model.
[0065] In the experiment, the sound attribute was speaker. In the experiment, an input speech feature sequence, an input phoneme feature sequence, a speaker index, and noise level information were input to the score approximator. Furthermore, with regard to the iterative processing of Algorithm 2 described in the Kameoka-Kaneko-Tanaka theory, it was experimentally confirmed through preliminary experiments that starting from step L', rather than from step L, is effective in terms of the quality of the converted speech, which is the original speech. Therefore, in the experiment described below, L' = 11.
[0066] The hyperparameter {β l} l We will explain the values used in the experiment. In fact, the hyperparameter {β l} l It is known that the value expressed by the following formula (1) is effective for image generation tasks. 1 To prevent β from becoming too close to 1, l It is also known that it may be better to clip it so that it is below a predetermined threshold, such as 0.999.
[0067]
[0068]
[0069]
[0070] Therefore, in the experiment, the hyperparameter {β l} l The value expressed by equation (1) was used for η. Note that when l = 0, β l This is an offset value to prevent β from becoming too small. Therefore, in the experiment, η was set to 0.008. Hereinafter, the hyperparameter {β l} lThe technique for determining the value of is called noise variance scheduling.
[0071] FIG. 2 is a diagram showing the network structure of a score approximator used in experiments in this embodiment. That is, the network configuration in FIG. 2 is an example of a network configuration of a score approximator. The input and output of each layer in FIG. 2 are vector sequences, and "c" and "l" represent the number of channels and the length of the vector sequence, respectively. "Conv1d", "GLU", and "Deconv1d" represent a one-dimensional convolution, a gated linear unit (GLU), and a one-dimensional transposed convolution layer, respectively. Furthermore, "k", "c", and "s" represent the kernel size, number of output channels, and stride size of the convolution layer, respectively.
[0072] In the experiments, all convolution weights were initialized using the Glorot method with a gain of 0.5 and parameterized using the WN (Wright Normalization) method. Instead of GLU, a normalized linear layer or the like may be used as the nonlinear activation layer. Note that instead of GLU, a normalized linear layer or the like may be used as the nonlinear activation layer.
[0073] In the example network structure of FIG. 2, the noise level and speaker index were incorporated into each convolutional layer of the score approximator according to the following embedding procedure.
[0074] The embedding procedure first extracts embedding vectors from two trainable lookup tables according to the specified indexes. The embedding procedure then repeats the extracted embedding vectors in the time direction until the length matches the input of the convolutional layer. The embedding procedure finally combines the two repeated vector sequences into the input along the channel direction. This allows the noise level and speaker index to be incorporated into each convolutional layer of the score approximator.
[0075] To incorporate an ASR output sequence or BNF sequence p into a network, the ASR output sequence or BNF sequence p is first input to a strided convolutional layer h with stride size r, and the output is then connected to the input of each GLU-equipped convolutional layer along the channel direction. This incorporates the ASR output sequence or BNF sequence p into the network. The stride size r satisfies the condition that the length of the output of h matches the input of the GLU-equipped convolutional layer.
[0076] In the experiment, speech data from seven speakers from the CMU ARCTIC database was used. Specifically, speech data from four speakers was used for training and for testing assuming known speakers, and speech data from three speakers was used exclusively for testing assuming unknown speakers. An unknown speaker is a speaker whose speech data is not included in the training data.
[0077] The four speakers used for training and testing assuming known speakers were female speaker clb, male speaker bdl, female speaker slt, and male speaker rms. The three speakers used exclusively for testing assuming unknown speakers were male speaker jmk, male speaker ksp, and female speaker lnh.
[0078] The CMU ARCTIC database is a database of speech samples from multiple speakers, each of which contains the same 1,132 sentences spoken by a speaker. In the experiment, the speech samples from each speaker of the latter 132 sentences of the 1,132 sentences were used as test data. The sampling frequency of all speech signals was 16,000 Hz.
[0079] In the experiment, the input speech feature sequence was a logarithmic mel-cepstrum. Specifically, the logarithmic mel-cepstrum used was information obtained by extracting an 80-dimensional logarithmic mel spectrum from each utterance at 10 ms intervals, and then normalizing the information so that the mean and variance of each dimension component were 0 and 1, respectively. Therefore, D = 80. Note that D represents the number of channels in the vector sequence (the dimension of each vector).
[0080] In the experiment, the network configuration of the score approximator was the network configuration shown in Figure 2. The hyperparameters are explained below. Adam was used for learning the score approximator, and the learning rate was 0.001. In addition, the noise level {β l} l is a value determined by equation (1), where L=20. Also, L'=11. Therefore, for example, when the voice conversion model executes a despreading process, the number of steps in the despreading process is 11.
[0081] In the experiment, "AutoVC," an autoencoder-based voice conversion (VC) technology, "PPG-VC," a phoneme posterior probability-based VC technology, "StarGAN-VC," a VC technology based on StarGAN, and a VC technology based on DSM and Langevin dynamics (hereinafter referred to as "first KKT technology") were used as baselines for comparison with a VC technology that obtains a target speech feature sequence by executing a de-diffusion process in a diffusion probability model (hereinafter referred to as "second KKT technology"). KKT is an abbreviation of the initials of Kameoka, Kaneko, and Tanaka.
[0082] The first KKT technique is a speech conversion technique using a speech conversion model that uses the above-mentioned DSM or weighted DSM and Langevin dynamics or simulated annealed Langevin dynamics. The second KKT technique is a speech conversion technique using the above-mentioned speech conversion model that obtains a target speech feature sequence by executing a de-diffusion process in a diffusion probability model. Hereinafter, when there is no need to distinguish between the first KKT technique and the second KKT technique, they will be referred to as VoiceGrad.
[0083] In the experiment, the test set consisted of speech samples in which each speaker spoke the same sentence. Therefore, the quality of the converted speech could be evaluated by comparing it with the speech of the target speaker who also spoke the same sentence. Mel-Cepstral Distortion (MCD), calculated from the mel-cepstrum of the two pairs of speech, was used as a measure of the degree of difference in voice quality between the two pairs.
[0084] The phonemes of the source speech and the target speech do not necessarily correspond at the same time. Therefore, for each utterance, the time axis was aligned using dynamic time warping (DTW) based on the MCD criterion, and then the average MCD was calculated. Furthermore, the F0 contour of the converted speech was evaluated by calculating the correlation coefficient between the log-F0 contours of the converted speech and the target speech after time warping based on the time correspondence curve obtained using DTW based on the MCD criterion. This is called the log-F0 contour correlation coefficient (LFC).
[0085] In addition, to evaluate the intelligibility of the converted speech, the Character Error Rate (CER) [%] of the converted speech was also evaluated using the text labels of each input speech. Furthermore, to evaluate the sound quality of the converted speech, MOS prediction was performed on the converted speech using a Mean Opinion Score (MOS) predictor. This measure is called pseudo-MOS (pMOS).
[0086] <<Objective Evaluation Experiment Results>> Fig. 3 is a first diagram showing an example of experimental results in the embodiment. More specifically, Fig. 3 shows the results of a comparison between the DSM version (i.e., the first KKT technique) and the DPM version (i.e., the second KKT technique). Fig. 4 is a second diagram showing an example of experimental results in the embodiment. More specifically, Fig. 4 shows the effect of conditioning using a BNF sequence.
[0087] Fig. 5 is a third diagram showing an example of an experimental result in the embodiment. Fig. 6 is a fourth diagram showing an example of an experimental result in the embodiment. More specifically, Fig. 5 shows an example of MCD in the case of a known speaker's voice, and Fig. 6 shows an example of MCD in the case of an unknown speaker's voice.
[0088] Fig. 7 is a fifth diagram showing an example of an experimental result in the embodiment. Fig. 8 is a sixth diagram showing an example of an experimental result in the embodiment. More specifically, Fig. 7 shows an example of LFC in the case of a known speaker's voice, and Fig. 8 shows an example of LFC in the case of an unknown speaker's voice.
[0089] Fig. 9 is a seventh diagram showing an example of an experimental result in the embodiment. Fig. 10 is an eighth diagram showing an example of an experimental result in the embodiment. More specifically, Fig. 9 shows an example of CER in the case of a known speaker's voice, and Fig. 10 shows an example of CER in the case of an unknown speaker's voice.
[0090] Fig. 11 is a ninth diagram showing an example of an experimental result in the embodiment. Fig. 12 is a tenth diagram showing an example of an experimental result in the embodiment. More specifically, Fig. 11 shows an example of pMOS in the case of a known speaker's voice, and Fig. 12 shows an example of pMOS in the case of an unknown speaker's voice.
[0091] 13 is an eleventh diagram illustrating an example of an experimental result in the embodiment. More specifically, FIG. 13 illustrates the RTF of the mel spectrogram conversion process.
[0092] Figure 3 shows the MCD, LFC, CER, pMOS, and their 95% confidence intervals for speech converted using the DSM and DPM versions of VoiceGrad. Note that the DSM version of VoiceGrad refers to the first KKT technique, and the DPM version of VoiceGrad refers to the second KKT technique.
[0093] As Figure 3 shows, the DPM version outperformed the DSM version in all evaluation measures. In fact, the number of iterations required for conversion was only 11 for the DPM version, compared with 576 for the DSM version. Taking this into account, the results suggest that the DPM version was superior. However, Figure 3 also shows that both versions tended to produce speech with a relatively low CER, compared with a CER of 1.1% for the target reference speech (unprocessed speech).
[0094] The results in Figure 4 show how much the above-mentioned BNF conditioning idea (i.e., the idea of using a phoneme feature sequence as input to the score approximator) can improve each evaluation value of converted speech. The results in Figure 4 show that by configuring the score approximator to incorporate a BNF sequence, the performance of all evaluation measures, particularly CER, is significantly improved. Note that a score approximator configured to incorporate a BNF sequence refers to a score approximator to which not only a speech feature sequence but also a phoneme feature sequence is input.
[0095] Next, Figures 5 to 12 show the performance evaluation values of VoiceGrad and comparison technologies (specifically, AutoVC, PPG-VC, and StarGAN-VC) when inputting known speaker speech and unknown speaker speech. The results show that among the comparison technologies, PPG-VC excels in almost all evaluation measures and can generate very clear, high-quality speech. The results also show that PPG-VC excels, particularly in CER and pMOS. In contrast, the results show that VoiceGrad slightly outperforms PPG-VC in CER and pMOS, and significantly outperforms PPG-VC in MCD and LFC.
[0096] Figure 13 also shows the real-time factor (RTF) of the computation time required for mel-spectrogram conversion for each technique. As shown in Figure 13, the DSM version of VoiceGrad was significantly slower than the other comparison techniques, while the DPM version was almost as fast as PPG-VC. This is due to the noise variance scheduling described above, and indicates that high-quality mel-spectrograms can be generated even with a small number of despreading steps. Figure 13 also shows that conditioning with BNF sequences does not significantly affect computation time.
[0097] <<Subjective Evaluation Experiment Results>> In addition to the objective evaluation experiment, a subjective evaluation experiment of sound quality and speaker similarity was also conducted. FIGS. 14 and 15 show examples of the results of the subjective evaluation experiment. FIG. 14 is a twelfth diagram showing an example of the experimental results in the embodiment. More specifically, FIG. 14 shows the results of MOS evaluation. FIG. 15 is a thirteenth diagram showing an example of the experimental results in the embodiment. More specifically, FIG. 15 shows the results of speaker similarity evaluation.
[0098] In addition to the objective evaluation, the experiment also included a subjective evaluation of sound quality and speaker similarity. For the sound quality evaluation, the target speaker's actual speech was included in the samples, and listeners rated the naturalness of each sample on a five-point scale: "5: Excellent," "4: Good," "3: Fair," "2: Poor," and "1: Bad." Figure Y10 shows an example of the results of such a sound quality evaluation.
[0099] In the speaker similarity evaluation, listeners were presented with a pair of speech converted by each technology and the target speaker's actual speech, and the listeners rated the likelihood that both speeches were spoken by the same speaker on a four-point scale: "4: Same (sure)," "3: Same (not sure)," "2: Different (not sure)," and "1: Different (sure)." Figure Y11 shows an example of the speaker similarity results obtained in this way.
[0100] When comparing a version of VoiceGrad with and without BNF conditioning, the former showed higher performance in both sound quality and speaker similarity. Thus, the experiment confirmed the effectiveness of BNF conditioning. Among the technologies compared, PPG-VC showed the highest performance in sound quality, while StarGAN-VC showed the highest performance in speaker similarity. In contrast, the BNF-conditioned version of VoiceGrad showed better performance than PPG-VC in sound quality and better performance than StarGAN-VC in speaker similarity. Thus, the superiority of VoiceGrad was confirmed.
[0101] 16 is a diagram showing an example of the hardware configuration of the learning device 1 according to an embodiment. The learning device 1 is equipped with a first control unit 11, which is a control unit including a processor 91 such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an NPU (Neural Network Processing Unit), and a memory 92, all connected via a bus, and executes a program. By executing the program, the learning device 1 functions as a device including the first control unit 11, an interface unit 12, and a storage unit 13.
[0102] More specifically, the processor 91 reads the program stored in the storage unit 13 and stores the read program in the memory 92. The processor 91 executes the program stored in the memory 92, causing the learning device 1 to function as a device including the first control unit 11, the interface unit 12, and the storage unit 13.
[0103] The first control unit 11 controls the operation of each functional unit included in the learning device 1. The first control unit 11, for example, executes a learning process. The first control unit 11, for example, acquires information stored in the memory unit 13. Specifically, the process of acquiring information stored in the memory unit 13 is reading.
[0104] The interface unit 12 includes a communication interface for connecting the learning device 1 to an external device. The interface unit 12 communicates with the external device via wired or wireless communication. The external device is, for example, a device that transmits learning data. The interface unit 12 acquires learning data by communicating with the device that transmits the learning data. The external device is, for example, a conversion device 2. By communicating with the learning device 1 via the interface unit 12, the conversion device 2 can use the trained score approximator obtained by the learning device 1.
[0105] The interface unit 12 may be configured to include input devices such as a mouse, keyboard, or touch panel. The interface unit 12 may be configured as an interface that connects these input devices to the study device 1. In this way, the input devices of the interface unit 12 accept input of various information to the study device 1 via wired or wireless connections. Note that information does not necessarily have to be input to the communication interface of the interface unit 12, but may also be input to the input devices of the interface unit 12.
[0106] The interface unit 12 outputs, for example, various types of information. The interface unit 12 includes a display device such as a CRT (Cathode Ray Tube) display, a liquid crystal display, or an organic EL (Electro-Luminescence) display, as well as a speaker. The interface unit 12 may be configured as an interface that connects these display devices or speakers to the learning device 1. Therefore, the interface unit 12 outputs, for example, information input to an input device of the interface unit 12 as an image or sound.
[0107] The storage unit 13 is configured using a computer-readable storage medium (non-transitory computer-readable recording medium) such as a magnetic hard disk drive or a semiconductor storage device. The storage unit 13 stores various information related to the learning device 1. The storage unit 13 stores, for example, various information generated by the operation of the first control unit 11. The storage unit 13 may exist on a cloud, for example.
[0108] FIG. 17 is a flowchart showing an example of the flow of processing executed by the learning device 1 in an embodiment. The first control unit 11 executes the learning process (step S101). In the learning process, an input speech feature sequence and an input phoneme feature sequence corresponding to the input speech feature sequence are input to a score approximator to be learned. The score approximator performs estimation based on the input information. In the learning process, the score approximator is updated based on the estimation result of the score approximator. In the learning process, it is determined whether a predetermined condition for terminating learning (hereinafter referred to as a "learning termination condition") has been satisfied. The learning termination condition may be determined at any timing in the learning process, but is determined, for example, at the stage when estimation is performed by the score approximator.
[0109] The learning termination condition may be any condition related to the termination of learning, such as a condition that the score approximator estimates the result with a predetermined accuracy. The learning termination condition may be, for example, a condition that the learning object has been updated a predetermined number of times. The learning termination condition may be, for example, a condition that the change due to the update of the learning object is smaller than a predetermined change.
[0110] 18 is a diagram showing an example of the hardware configuration of the conversion device 2 in an embodiment. The conversion device 2 is equipped with a second control unit 21, which is a control unit including a processor 93 such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an NPU (Neural Network Processing Unit), and a memory 94, all connected via a bus, and executes a program. By executing the program, the conversion device 2 functions as a device including the second control unit 21, an interface unit 22, and a storage unit 23.
[0111] More specifically, the processor 93 reads the program stored in the storage unit 23 and stores the read program in the memory 94. The processor 93 executes the program stored in the memory 94, whereby the conversion device 2 functions as a device including the second control unit 21, the interface unit 22, and the storage unit 23.
[0112] The second control unit 21 controls the operation of each functional unit included in the conversion device 2. The second control unit 21 executes, for example, a conversion process. The second control unit 21 acquires, for example, information stored in the storage unit 23. Specifically, the process of acquiring the information stored in the storage unit 23 is reading.
[0113] The interface unit 22 includes a communication interface for connecting the conversion device 2 to an external device. The interface unit 22 communicates with the external device via wired or wireless communication. The external device is, for example, a device that is the source of the conversion target. The interface unit 22 acquires the conversion target by communicating with the device that is the source of the conversion target. The external device is, for example, the learning device 1. By communicating with the learning device 1 via the interface unit 22, the conversion device 2 can use the learned score approximator obtained by the learning device 1.
[0114] The interface unit 22 may be configured to include input devices such as a mouse, a keyboard, a touch panel, etc. The interface unit 22 may be configured as an interface that connects these input devices to the conversion device 2. In this way, the input devices of the interface unit 22 accept input of various information to the conversion device 2 via wired or wireless connections. Note that information does not necessarily have to be input to the communication interface of the interface unit 22, but may also be input to the input devices of the interface unit 22.
[0115] The interface unit 22 outputs, for example, various types of information. The interface unit 22 includes, for example, a display device such as a CRT (Cathode Ray Tube) display, a liquid crystal display, or an organic EL (Electro-Luminescence) display, and a speaker. The interface unit 22 may be configured as an interface that connects these display devices or speakers to the conversion device 2. Therefore, the interface unit 22 outputs, for example, information input to an input device of the interface unit 22 as an image or sound.
[0116] The storage unit 23 is configured using a computer-readable storage medium device (non-transitory computer-readable recording medium) such as a magnetic hard disk device or a semiconductor storage device. The storage unit 23 stores various information related to the conversion device 2. The storage unit 23 stores various information generated by the operation of the second control unit 21, for example. The storage unit 23 may exist on a cloud, for example.
[0117] 19 is a flowchart showing an example of the flow of processing executed by the conversion device 2 according to the embodiment. The second control unit 21 acquires a speech feature sequence to be converted (step S201). Next, the second control unit 21 executes a conversion process on the speech feature sequence acquired in step S201 (step S202).
[0118] <About solving the problem> The learning device 1 of the embodiment configured as described above executes a learning process for training a score approximator. Therefore, as described above in <About effects achieved by the learning process>, the learning device 1 can alleviate the constraints imposed on data used for training in voice conversion technology using machine learning.
[0119] Furthermore, the conversion system 100 of the embodiment configured as described above includes the learning device 1. Therefore, the conversion system 100 can alleviate the constraints imposed on data used for learning in voice conversion technology using machine learning.
[0120] Furthermore, the conversion device 2 of the embodiment configured as described above converts a speech feature sequence using the trained score approximator obtained by the training device 1. Therefore, the conversion device 2 can alleviate the constraints imposed on data used for training in voice conversion technology using machine learning.
[0121] (Variation) When comparing DSM and weighted DSM as methods for estimating a score function, the weighted DSM provides higher estimation accuracy than the DSM. This is because DSM performs estimation using a distribution with a single variance, while weighted DSM performs estimation using multiple distributions with different variances. In other words, weighted DSM uses multiple noise distributions with different noise variances σ, and therefore provides higher accuracy in estimating the score function than DSM using a single noise distribution.
[0122] As explained in the Kameoka-Kaneko-Tanaka theory, DSM assumes that the variance increases when Gaussian noise is added to the input, whereas DPM maintains the variance of the input at 1 through scaling. The more constraints there are, the higher the accuracy of estimation by a neural network, so a voice conversion model using a dediffusion process has higher estimation accuracy than a voice conversion model using DSM.
[0123] <Kameoka-Kaneko-Tanaka Theory> <<Score Matching and Langevin Dynamics>> First, we will outline the principles of score matching and Langevin dynamics and explain how they can be applied to the voice conversion problem. For any continuous and differentiable probability density function p(x), the gradient of log p(x) with respect to x is called the "score function." The gradient of log p(x) with respect to x can be expressed, for example, by the following equation (A1):
[0124]
[0125] If a score function for some data x is given, then in the manner of gradient ascent, an appropriate x is set as the initial point, and x is repeatedly updated according to the following equation (A2) so that log p(x) becomes large. The converged value can then be regarded as one of the samples according to p(x).
[0126]
[0127] where γ is a positive step size parameter, and z is a Gaussian white noise with a mean of 0 and a variance of 1. The process of executing an update rule such as equation (A2) including a noise term expressed by the following equation (A3) is called "Langevin dynamics."
[0128]
[0129] It has been shown that when the number of iterations is sufficiently large and γ is sufficiently small, the converged value of x will be a sample that conforms to p(x) under certain regularity conditions. The above idea means that even if p(x) cannot be estimated, it is possible to obtain a sample that conforms to p(x) as long as the score function ∇x log p(x) can be estimated. Therefore, when a training sample χ expressed by the following equation (A4) is given, the focus of the problem is how to estimate the score function of the training sample χ.
[0130]
[0131] Score matching is performed by using a score approximator s θ This problem is solved by optimizing (x) with respect to the parameter θ. For example, this problem can be solved by solving s θ This can be formulated as the problem of minimizing the expected squared error between (x) and ∇x log p(x).
[0132]
[0133] However, if the number of training samples χ is sufficiently large, the processing on the right side of equation (A5) can be approximated by the sample average for the training samples χ. The above objective function implicitly assumes that the target value ∇x log p(x) is observable, but several methods have been proposed that can reduce the problem to an equivalent optimization problem without assuming the specific form of p(x).
[0134] One is a technique called "shadow score matching," which utilizes the fact that equation (A5) is equal to the following equation (A6) when the constant term is removed.
[0135]
[0136] However, ∇xs θ (x) is s θ (x), and tr() represents the trace of the matrix. In this way, we can eliminate ∇x log p(x) from the objective function, and instead use tr(∇xs θ (x)) needs to be calculated. θWhen (x) is represented by a deep neural network or when x is high-dimensional, tr(∇xs θ The problem is that the calculation of (x) becomes very expensive.
[0137] In contrast, tr(∇xs θ A technique called "Denoising Score Matching (DSM)" can avoid even the calculation of (x). The idea behind DSM is to first intentionally add noise to the data x according to an appropriate distribution expressed in equation (A7), and then estimate the score function of the distribution of the noisy data expressed in equation (A8) below.
[0138]
[0139]
[0140] Here, σ represents a parameter of the appropriate distribution expressed by equation (A7). σ (x) can be viewed as a Parzen window estimator of p(x).
[0141] When the noise distribution is a Gaussian distribution (an example of a formula representing a Gaussian distribution is the following formula (A9)), the objective function to be minimized is expressed by the following formula (A10).
[0142]
[0143]
[0144] Minimize equation (A10) θ (x) is ∇x log q σ (x) is shown to be "almost certain" to match the noise variance σ 2 is small enough that q σ When (x) and p(x) are approximately the same, it also coincides with ∇x log p(x). Intuitively, the direction of the gradient of the logarithmic distribution should coincide with the direction toward x before noise is added at the point expressed by equation (A11) below.
[0145]
[0146] By the way, based on the idea of DSM, the score approximator s θ There have been attempts to achieve high-resolution image generation by representing (x) using a neural network. A challenge in applying DSM to image generation tasks is that much real-world data, including images, tends to be localized on a low-dimensional manifold in a high-dimensional space.
[0147] This fact can be a major problem when applying DSM simply, since it assumes that the data distribution is fully supported, meaning that there is no way to estimate the score at points where there is little data.
[0148] Therefore, a method called "weighted DSM" was devised. In this method, a score approximator is trained using multiple noise levels expressed by the following equation (A12), and then in test sampling, the initial q of the iterative calculation is σl This method involves gradually changing the noise level from high to low so that (x) covers the entire space and ultimately approaches the true distribution p(x).
[0149]
[0150] Here, the score approximator s is conditioned on the noise level l (lowercase L) so that the score approximator can learn different behavior depending on the noise level. θ The key to the idea is to consider (x, l).
[0151] As a learning criterion, for example, D σ1 (θ), ..., D σL As θ approaches the optimum solution, the amount expressed by the following equation (A14) becomes 1 / σ l Since the weight λ l >0 to σ 2 l It has been proposed to use the following formula (A15) as a criterion: σl (θ) represents the result of substituting σ1 for σ in equation (A10).
[0152]
[0153]
[0154]
[0155] It has also been experimentally reported that a geometric progression such as the following equation (A16) is effective for noise levels.
[0156]
[0157] Under the above settings, θ Once (x, l) is learned, q σL The sample conforming to (x) can be sampled by the iterative algorithm (Simulated Langevin Dynamics) shown in Algorithm 1 of FIG. 20, where T represents the number of iterations. However, γ l is defined by the following equation (A17), and is a step size that changes adaptively depending on the noise level, and ε is its scale parameter.
[0158]
[0159] <<Diffusion Probabilistic Model>> Next, the principles of the diffusion probabilistic model (DPM) will be outlined. Note that the diffusion probabilistic model itself is so well-known that it goes without saying that the diffusion probabilistic model itself is a well-known technique.
[0160] A data sample normalized to mean 0 and variance 1 is x 0 Then, the diffusion process of DPM is x 0 Gradually add Gaussian noise to (x 0 Based on this diffusion process assumption, the noise is gradually diffused into the x 0 The learning goal of DPM is to find the parameters θ of the inverse diffusion process model that gradually reconstructs the following. Usually, DPM refers to a latent variable model expressed by the following equation (A18).
[0161]
[0162] However, x i:j is the set {x i ,…,x j}, and x 1 ,…,x L Ha x 0 where x is the latent variable corresponding to the diffused version of 0 The joint distribution q(x 1:L |x 0 ) is called a diffusion process and is assumed to be represented by the Markov chain of the following equation (A19).
[0163]
[0164] Also, each conditional distribution q(x l |x l-1 ) is assumed to be given by the following equation (A20):
[0165]
[0166] That is, in equation (A20), the quantity expressed on the right side of the following equation (A21) is defined as H, and x l-1 After multiplying by H, the variance β l Gaussian noise is added to x l This can be interpreted as representing the process of
[0167]
[0168] This is why it is called a diffusion process. This scaling by H is x at each time step. l The sum of Gaussian noises is also Gaussian noise, so by the definition of equation (A20), the conditional distribution q(x l |x 0 ) is expressed in analytical description in equation (A22) below:
[0169]
[0170]
[0171]
[0172] The quantity expressed on the left side of the following equation (A25) represents noise. For simplicity of explanation, the following description will be given taking as an example a case where the noise is Gaussian noise that follows a standard normal distribution. When the noise is Gaussian noise that follows a standard normal distribution, equation (A22) can be rewritten as the following equation (A26).
[0173]
[0174]
[0175] By the way, the joint distribution p θ (x 0:L ) is called a de-diffusion process, and is assumed to be expressed by the Markov chain of the following equation (A28) under the assumption of the following equation (A27), similar to equation (A19).
[0176]
[0177]
[0178] Here, each conditional distribution p θ (x l-1 |x l ) is defined by the following formula (A29).
[0179]
[0180] μ θ (x l , l) is a function with parameter θ and x l This is the output of a neural network that takes l as input. For simplicity, we will use v 2 l β l The learning goal of DPM is to determine the parameter θ under the above assumptions and model.
[0181] <<Similarities and Relationships with Sample Generation Algorithms Based on Weighted DSM and Langevin Dynamics>> Up to this point, we have explained the relationship between weighted DSM and Langevin dynamics and the voice conversion problem, and have outlined the well-known probability diffusion model. Kameoka-Kaneko-Tanaka theory shows the similarities between this probability diffusion model and sample generation algorithms based on weighted DSM and Langevin dynamics. Therefore, below we will explain these similarities shown by Kameoka-Kaneko-Tanaka theory.
[0182] Based on the formula of the probability diffusion model, the learning criterion to be minimized is actually E[-log p θ (x0)] can be used as a variational upper bound.
[0183]
[0184] This variational upper bound L DPM (θ) is p θ (x 1:L |x 0 ) and q(x 1:L |x 0 ) is zero, E[-log p θ (x0)]. Therefore, L DPM Minimizing (θ) with respect to θ is p θ Not only fitting p to the data distribution, θ (x 1:L |x 0 ) and q(x 1:L |x 0 ) as close as possible to the
[0185] By the way, if the variable transformation of the following equation (A31) is used, equation (A31) can be transformed into the following equation (A32).
[0186]
[0187]
[0188] c l Ha, ha α l and ν l and the quantity expressed by the following equation (A33). l =1.
[0189]
[0190] The variable transformation in equation (A31) is μ θ (x l, l) by using a neural network. Note that, as described in the explanation of equation (A25), the quantity expressed by equation (A34) represents Gaussian noise.
[0191]
[0192] Here, in order to clarify the relationship between the DSM and the above-mentioned DPM learning rule, we will compare equation (A32) with equation (A15). 0 Using the symbols and the variable transformation of the following equation (A35), equation (A15) is transformed into the following equation (A36).
[0193]
[0194]
[0195] Comparing formula (A32) and formula (36), it is found that they are very similar to each other. It is also found that the quantity expressed by the following formula (A37) and the quantity expressed by the following formula (A38) correspond to each other.
[0196]
[0197]
[0198] According to this correspondence, it can be seen that the transformation expressed by the above formula (A34) or formula (A38) represents a score approximator. Therefore, the transformation expressed by formula (A34) or formula (A38) is also referred to as a score approximator.
[0199] The difference is that in DSM, the quantity expressed by the above equation (A35) becomes the input to the score approximator, whereas in DPM, the first argument in parentheses in the above equation (A38) (i.e., the quantity expressed by the following equation (A39)) becomes the input to the score approximator.
[0200]
[0201] This means that in DSM, the input x l It is assumed that the variance increases when Gaussian noise is added to x0 By scaling the input x l This means that the variance of is assumed to be kept at 1. As described in the explanation of equation (A25) or equation (A34), the symbol excluding the weight represented by the root in the last term of equation (A39) represents Gaussian noise.
[0202] By the above learning criteria, θ is set to p θ A sample conforming to (x) can be generated by the Markov chain of the dediffusion process (Algorithm 2) shown in Fig. 21. It can be seen that Algorithm 2 and Algorithm 1 are similar, and by comparing the two, the relationship between the Langevin dynamics and the sample generation algorithm using the dediffusion process can be found.
[0203] First, in Algorithm 1, the loop counter l varies from 1 to L, while in Algorithm 2 it varies from L to 1. However, this is merely a matter of how to assign indexes to time steps or noise levels, and is not an essential difference.
[0204] In fact, if we define l'=L-l+1 in Algorithm 1, we can change the loop counter l' from L to 1, just like Algorithm 2. Next, we will compare the update formulas of Algorithm 2 and Algorithm 1.
[0205] The relationship of the following formula (A41) can be obtained from the relationship of the following formula (A40), and therefore the following formula (A42) can be obtained.
[0206]
[0207]
[0208]
[0209] The quantity in (A34) is trained to predict the numerator of the right-hand side of (A42). Therefore, if θ is successfully learned, the quantity in (A43) below will be ∇x l log q(x l ) should be a good approximation. Therefore, the update formula in Algorithm 2 is x l ∇x llog q(x l ) direction with a step size of 1-α l After scaling with the following equation (A44), l This can be interpreted as the process of adding z. Therefore, it can be seen that this is the same process as the update formula in Algorithm 1.
[0210]
[0211]
[0212] In this way, the speech conversion model that converts a speech feature sequence through the de-diffusion process of a diffusion probability model uses a score approximator trained based on a learning criterion such as equation (A30) or equation (A32) (i.e., trained using the first objective function), and converts the speech feature sequence of the input speech into a speech feature sequence that is likely to be that of the target speaker through the de-diffusion process of Algorithm 2, using the speech feature sequence of the input speech as initial values. This is equivalent to regarding the speech feature sequence of the input speech as the speech feature sequence of the target speaker "diffused by noise."
[0213] The above is the Kameoka-Kaneko-Tanaka theory.
[0214] <About VoiceGrad> The term VoiceGrad was defined in the explanation of the experimental results above, but we will now provide additional information on this technology. VoiceGrad is, in other words, a process for executing a voice conversion model. Therefore, in VoiceGrad, the feature sequence of a source voice is converted into the feature sequence of a target voice by iteratively changing the logarithmic gradient direction of the feature sequence distribution of the target voice, starting from the feature sequence of the source voice as the initial point. One specific implementation method for this is a method using weighted DSM and simulated annealing Langevin dynamics.
[0215] In this method, it is possible to prepare a score approximator for each target speaker. However, in order to predict the scores of feature sequence distributions for K speakers using only a single score approximator, we use a network s conditioned not only by the noise level but also by the speaker index k. θIt is possible to express the score approximator by (x, l, k), where k is an element of the set {1, ..., K}, i.e., k is an integer between 1 and K, inclusive.
[0216] score approximator s θk (x, l) or s θ If (x, l, k) can be learned, the feature sequence of the input speech is set as the initial point x(0), and s θ (x, l) as s θk (x, l) or s θ By replacing (x, l, k) with (x, l, k) and then executing the above algorithm 1, the input feature sequence can be converted into the voice quality of speaker k. The above method does not depend on the feature sequence distribution of the input speech, so theoretically it can be applied to input speech by any speaker.
[0217] As the feature quantities of the speech feature sequence to be converted by VoiceGrad (i.e., speech features), any feature quantities sufficient for ultimately forming a speech waveform may be used, and several options are conceivable. One is vocoder parameters. There are various options for the type of vocoder, and for example, a mel-cepstral vocoder, which is often used in speech synthesis, may be used.
[0218] A mel-cepstral vocoder can synthesize a speech signal from mel-cepstral coefficients, fundamental frequency (F0) values, and aperiodicity indices for each short interval. Therefore, a vector combining these may be used as a feature.
[0219] For the F0 pattern (series of F0 values), a simple method can be used in which the mean and variance of the logarithmic F0 values are shifted and scaled to match those of the target speaker. Furthermore, since the aperiodicity index of the input speech can be used as is without conversion, it is also possible to use only the mel-cepstral coefficients as features. In the above experiments, a vector with mel-cepstral coefficients as elements was used as a feature.
[0220] Furthermore, many high-quality neural vocoders, including WaveNet, have been proposed, and features designed for use with these may also be used. Many of the neural vocoders proposed to date use the Mel spectrum for each short section as a feature. Therefore, the Mel spectrum may also be used as a feature in line with this.
[0221] The d-th dimension feature in the short-time frame m is expressed as x d,m Then, in learning and testing, the feature values normalized by the following formula (B1) may be used. d and ξ d and represent the mean and standard deviation of the d-th dimension feature in the voiced section, respectively.
[0222]
[0223] <Others> The learning device 1 may be implemented using multiple information processing devices connected to each other via a network. In this case, the functional units of the learning device 1 may be distributed and implemented across the multiple information processing devices.
[0224] The conversion device 2 may be implemented using a plurality of information processing devices communicably connected via a network, in which case the respective functional units of the conversion device 2 may be distributed and implemented among the plurality of information processing devices.
[0225] Note that all or part of the functions of the conversion system 100 may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). The program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, magneto-optical disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into computer systems. The program may be transmitted via a telecommunications line.
[0226] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention.
[0227] REFERENCE SIGNS LIST 100... Conversion system, 1... Learning device, 2... Conversion device, 11... First control unit, 12... Interface unit, 13... Storage unit, 21... Second control unit, 22... Interface unit, 23... Storage unit, 91... Processor, 92... Memory, 93... Processor, 94... Memory
Claims
1. A neural network included in a voice conversion model, which is a mathematical model for converting a source feature quantity series, which is a series of voice feature quantities of a source voice signal, into a target voice feature quantity series, which is a series of voice feature quantities of a target voice signal belonging to a target sound attribute with respect to a sound attribute that is an attribute related to sound; a neural network into which at least a voice feature quantity series, which is a series of voice feature quantities obtained from a voice signal, and a phoneme feature quantity series obtained from the voice feature quantity series are input; a neural network that performs estimation based on the input voice feature quantity series, which is the input voice feature quantity series, and the input phoneme feature quantity series, which is the input phoneme feature quantity series; a target feature quantity distribution function, which is a function on a vector space representing the voice feature quantity series and is a probability density function representing the distribution of the target voice feature quantity series, is a stationary point that is a point in the vector space, and is a nearest stationary point that is the nearest stationary point to the current point, which is a point in the vector space representing the input voice feature quantity series, and a neural network, which is a parameterized neural network that estimates the gradient at the current point of the path leading to the nearest stationary point. A learning device comprising a first control unit that performs learning of the score approximator.
2. In the learning, the parameters of the score approximator are updated so as to minimize the objective function in the diffusion probability model. The learning device according to claim 1.
3. The learning is weighted DSM (Denosing Score Matching). The learning device according to claim 1.
4. A neural network included in a voice conversion model, which is a mathematical model for converting a source feature quantity series, which is a series of voice feature quantities of a source voice signal, into a target voice feature quantity series, which is a series of voice feature quantities of a target voice signal that belongs to a target voice attribute with respect to a voice attribute that is an attribute related to sound; a neural network into which at least a voice feature quantity series, which is a series of voice feature quantities obtained from a voice signal, and a phoneme feature quantity series obtained from the voice feature quantity series are input; a neural network that performs estimation based on the input voice feature quantity series, which is the input voice feature quantity series, and the input phoneme feature quantity series, which is the input phoneme feature quantity series; a target feature quantity distribution function, which is a function on a vector space representing the voice feature quantity series and is a probability density function representing the distribution of the target voice feature quantity series, and is a stationary point of the vector space and is the nearest stationary point to the current point, which is a point in the vector space representing the input voice feature quantity series; a parameterized neural network that estimates the gradient at the current point of the path leading to the nearest stationary point; a first control unit that performs learning of a score approximator, which is such a neural network; and a second control unit that performs estimation using the learned score approximator obtained by a learning device including the first control unit. A conversion device comprising the above.
5. A neural network included in a voice conversion model, which is a mathematical model for converting a source feature quantity series, which is a series of voice feature quantities of a source voice signal, into a target voice feature quantity series, which is a series of voice feature quantities of a target voice signal that belongs to a target sound attribute with respect to a sound attribute that is an attribute related to sound; a neural network into which at least a voice feature quantity series, which is a series of voice feature quantities obtained from a voice signal, and a phoneme feature quantity series obtained from the voice feature quantity series are input; a neural network that performs estimation based on the input voice feature quantity series, which is the input voice feature quantity series, and the input phoneme feature quantity series, which is the input phoneme feature quantity series; a target feature quantity distribution function, which is a function on a vector space representing the voice feature quantity series and is a probability density function representing the distribution of the target voice feature quantity series, and is a stationary point of the vector space of the point, and is the nearest stationary point that is the nearest stationary point to the current point, which is the point in the vector space representing the input voice feature quantity series; a parameterized neural network that estimates the gradient at the current point of the path leading to the nearest stationary point; a first control step of learning a score approximator, which is the neural network.
6. A neural network included in a voice conversion model, which is a mathematical model for converting a source feature quantity series, which is a series of voice feature quantities of a source voice signal, into a target voice feature quantity series, which is a series of voice feature quantities of a target voice signal belonging to a target sound attribute with respect to a sound attribute that is an attribute related to sound. The neural network is a neural network into which at least a voice feature quantity series, which is a series of voice feature quantities obtained from a voice signal, and a phoneme feature quantity series obtained from the voice feature quantity series are input. The neural network is a neural network that makes an estimation based on the input voice feature quantity series, which is the input voice feature quantity series, and the input phoneme feature quantity series, which is the input phoneme feature quantity series. The neural network is a parametric neural network that estimates the gradient at the current point of the path leading to the nearest stationary point, which is a stationary point of the target feature quantity distribution function, which is a function on the vector space representing the voice feature quantity series and is a probability density function representing the distribution of the target voice feature quantity series, and is a point in the vector space and is the nearest stationary point to the current point, which is a point in the vector space representing the input voice feature quantity series. A conversion method having a second control step of performing estimation using the learned score approximator obtained by a learning method having a first control step of learning the score approximator.
7. A program for causing a computer to function as the learning device according to any one of claims 1 to 3.
8. A program for causing a computer to function as the conversion device according to claim 4.
Citation Information
Patent Citations
Voice conversion learning device, voice conversion device, method and program
JP2019144402A
Multilingual text-to-speech synthesis method
JP2021511536A