signal generation and processing device

The signal generation processing device uses multiple sub-models trained on different noise levels to enhance speech and image synthesis quality and speed by parallel learning and selective processing.

JP7720573B2Active Publication Date: 2025-08-08NAT INST OF INFORMATION & COMM TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022572997
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-28
Filing Date
2021-12-17
Publication Date
2025-08-08
Estimated Expiration
2041-12-17

AI Technical Summary

Technical Problem

Existing speech synthesis technologies using WaveGrad and DiffWave models achieve lower quality speech synthesis compared to WaveGlow, despite requiring fewer model parameters and a shorter training time, when optimized using data from a single speaker.

Method used

A signal generation processing device utilizing multiple sub-model units, each trained on different noise levels, allows parallel learning and selection based on a noise schedule to enhance processing speed and quality.

Benefits of technology

The device achieves high-quality speech and image synthesis while maintaining processing speed by distributing learning across sub-models with varying noise levels and processing accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007720573000006
    Figure 0007720573000006
  • Figure 0007720573000007
    Figure 0007720573000007
  • Figure 0007720573000008
    Figure 0007720573000008
Patent Text Reader

Abstract

Provided is a signal generation processing device that realizes a voice synthesis process or an image signal generation process which can obtain a high-quality voice signal or image signal while maintaining the speed of the voice synthesis process or image signal generation. In the signal generation processing device, first to Nth sub-model units can use respective noise levels that are included in different noise level ranges to obtain trained models by performing learning processes of learning models which are included among the first to Nth sub-model units. In other words, in the signal generation processing device, the sub-model units can carry out the respective processes in parallel, so that a high-speed learning process can be performed. Further, in the signal generation processing device, since a sub-model unit to be used can be appropriately selected at the time of a prediction process, the voice synthesis process or image generation process can be carried out with high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a processing technology for generating audio signals and image signals (for example, a vocoder technology for synthesizing audio waveforms from acoustic features). [Background technology]

[0002] In recent years, the introduction of neural networks has enabled high-quality speech synthesis in text-to-speech (TTS) technology, which synthesizes natural-sounding speech from text. Various technologies have also been developed for vocoders used in such text-to-speech synthesis. For example, various models have been proposed for neural vocoders that synthesize speech waveforms from acoustic features. Among these, the technology disclosed in Non-Patent Document 1 (hereinafter referred to as "WaveGlow") has attracted attention for its real-time, high-quality synthesis. However, WaveGlow has a problem in that it requires a huge number of model parameters and requires a long training process (e.g., approximately 20 days even when using multiple GPUs). In response to this problem, spreading stochastic neural vocoders have been developed as disclosed in Non-Patent Documents 2 and 3 (the technology disclosed in Non-Patent Document 2 is called "WaveGrad," and the technology disclosed in Non-Patent Document 3 is called "DiffWave"). These spreading stochastic neural vocoders (WaveGrad and DiffWave) use small models and few parameters to achieve high-quality speech synthesis.

[0003] The WaveGrad and DiffWave models are neural network models that input a signal in which weighted noise is added to a speech waveform signal and estimate only the added noise component. Each WaveGrad and DiffWave model is implemented using a single model. During training, WaveGrad and DiffWave also input weights (data indicating the weight values) to a single model to correspond to various real-valued weights between 0 and 1, and train the model. Furthermore, during prediction (when executing speech synthesis processing), WaveGrad and DiffWave first input only noise to the model and subtract the noise components estimated by the model from the input to obtain an estimated waveform. Next, a slightly lower level of noise is added to the estimated waveform and input to the same model (WaveGrad and DiffWave model (neural network model)). The noise components estimated by the model are then subtracted from the input to obtain a new estimated waveform. This process is repeated, gradually lowering the noise level, to ultimately obtain a clean speech waveform signal.

[0004] In waveform generation models such as WaveGrad and DiffWave, one of the key points is how to synthesize non-periodic components of speech waveforms that cannot be obtained from the input acoustic features, and in waveform generation models such as WaveGrad and DiffWave, the noise components that cannot be removed until the end correspond to the non-periodic components of speech waveforms. As a result, waveform generation models such as WaveGrad and DiffWave can achieve high-quality speech synthesis processing with fewer model parameters than WaveGlow. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] R. Prenger, R. Valle, and B. Catanzaro, "WaveGlow: A flow-based generative network for speech synthesis," in Proc. ICASSP, May 2019, pp. 3617-3621. [Non-patent document 2] N. Chen, Y. Zhang, H. Zen, RJ Weiss, M. Norouzi, and W. Chan, "WaveGrad: Estimating gradients for waveform generation," arXiv:2009.00713, 2020. [Non-patent document 3] Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, "DiffWave: A versatile diffusion model for audio synthesis," arXiv:2009.09761, 2020. Summary of the Invention [Problem to be solved by the invention]

[0006] However, when performing a training process using data from a single speaker to obtain an optimized model (trained model), there is a problem in that the speech generated by the speech synthesis waveform signal obtained by the WaveGrad or DiffWave optimized model (trained model) is of lower quality than the speech generated by the speech synthesis waveform signal obtained by the WaveGlow optimized model (trained model).

[0007] In view of the above, an object of the present invention is to provide a speech synthesis processing device that realizes speech synthesis processing capable of obtaining high-quality speech (speech signals) while maintaining the speed of speech synthesis processing.Another object of the present invention is to provide a signal processing device that realizes signal generation processing capable of obtaining high-quality signals (e.g., image signals) while maintaining the processing speed for signals other than speech signals (e.g., image signals). [Means for solving the problem]

[0008] A first invention for solving the above problem is a signal generation processing device that outputs an audio signal or an image signal from Gaussian white noise, and includes N sub-model units (N: natural number, N≧2), that is, a first sub-model unit to an Nth sub-model unit.

[0009] Each of the first to Nth sub-model units includes a learning model that receives data related to noise levels and a teacher signal of an audio signal or an image signal as input, and performs a learning process to output Gaussian white noise from a noise synthesis signal, which is a signal obtained by synthesizing the teacher signal and Gaussian white noise based on the data related to noise levels.

[0010] The first to Nth sub-model units acquire trained models by performing a training process for the training models included in the first to Nth sub-model units using noise levels included in different noise level ranges.

[0011] In this signal generation processing device, the first to Nth sub-model units can acquire trained models by performing training processes on the training models included in the first to Nth sub-model units using noise levels that fall within different noise level ranges.

[0012] That is, in this signal generation processing device, learning processing can be performed independently for each of the N sub-model units. That is, in each of the N sub-model units, learning processing can be performed as long as the teacher signal, which is the correct data, the Gaussian white noise, and the noise level that determines the ratio at which they are combined are known, so that learning processing can be performed for the N sub-model units in parallel. As a result, this signal generation processing device can achieve high-speed learning processing.

[0013] A second invention is a signal generation processing device that outputs an audio signal or an image signal corresponding to an input condition feature based on Gaussian white noise and an input condition feature, and includes N sub-model units (N: natural number, N≧2) consisting of a first sub-model unit to an Nth sub-model unit.

[0014] The first to Nth sub-model units each include a learning model that receives as input data related to noise levels, input condition features, and teacher signals of audio signals or image signals corresponding to the input condition features, and performs a learning process to output Gaussian white noise from a noise synthesis signal, which is a signal obtained by synthesizing the teacher signal and Gaussian white noise based on the data related to noise levels.

[0015] The first to Nth sub-model units acquire trained models by performing a training process for the training models included in the first to Nth sub-model units using noise levels included in different noise level ranges.

[0016] In this signal generation processing device, the first to Nth sub-model units can acquire trained models by performing training processes on the training models included in the first to Nth sub-model units using noise levels that fall within different noise level ranges.

[0017] That is, in this signal generation processing device, learning processing can be performed independently for each of the N sub-model units. That is, in each of the N sub-model units, learning processing can be performed as long as the teacher signal, which is the correct data, the corresponding input condition feature, Gaussian white noise, and the noise level that determines the ratio at which they are combined are known, so that learning processing can be performed for the N sub-model units in parallel. As a result, this signal generation processing device can achieve high-speed learning processing.

[0018] A third invention is the second invention, further comprising a control unit that sets a noise schedule.

[0019] The control unit selects a sub-model unit to be used when performing signal generation processing from among the first to Nth sub-model units based on the noise level determined based on the noise schedule, and determines the order of processing of the selected sub-model units.

[0020] The selected sub-model unit executes a prediction process using the trained model in the order determined by the control unit to obtain an audio signal or an image signal corresponding to the input condition feature.

[0021] As a result, in this signal generation processing device, during signal generation processing (prediction processing), it is possible to select the sub-model unit to be used based on the noise level determined based on the noise schedule.

[0022] A fourth invention is the third invention, wherein the first to Nth sub-model units have an order in terms of the proportion of noise components in the input noise synthesis signal, the order being an order in which the proportion of noise components in the noise synthesis signal decreases, and the sub-model units located earlier in the order have a faster processing speed than the sub-model units located later.

[0023] As a result, in this signal generation processing device, for example, when the "order of the proportion of noise components in the input noise synthesis signal" is in ascending order of the proportion of noise components in the input noise synthesis signal for the indices (1 to N) of the first to Nth submodel units, (1) the submodel units located at the front can be submodel units with a large proportion of noise components in the input noise synthesis signal and with a configuration that has a fast processing speed, and (2) the submodel units located at the rear can be submodel units with a small proportion of noise components in the input noise synthesis signal and with a slow processing speed.

[0024] Therefore, in this signal generation processing device, for example, when the first to Nth submodel units are arranged in descending order (descending order of index) from the Nth submodel unit to the first submodel unit, and the proportion of noise components in the input noise synthesis signal decreases from the Nth submodel unit to the first submodel unit, it is possible to arrange submodels having a configuration with a fast processing speed at the upstream side. At the upstream side, it is only necessary to output Gaussian white noise from a signal with a high noise component, so prediction processing is relatively easy, and by arranging submodel units having a configuration with a fast processing speed, it is possible to increase the processing speed while maintaining the accuracy of the overall signal generation processing.

[0025] A fifth invention is the third invention, wherein the first to N-th sub-model units have an order based on the proportion of noise components in the input noise synthesis signal, the order being an order in which the proportion of noise components in the noise synthesis signal decreases, and the sub-model units located later in the order have a configuration in which they have higher processing accuracy than the sub-model units located earlier. As a result, in this signal generation processing device, for example, when the "order of the proportion of noise components in the input noise synthesis signal" is in ascending order of the proportion of noise components in the input noise synthesis signal for the indices (1 to N) of the first to Nth sub-model units, (1) the sub-model units located at the rear can be sub-model units with a small proportion of noise components in the input noise synthesis signal and with a configuration having high processing accuracy, and (2) the sub-model units located at the front can be sub-model units with a large proportion of noise components in the input noise synthesis signal and with a configuration having low processing accuracy.

[0026] Therefore, in this signal generation processing device, for example, when the first to Nth submodel units are arranged in descending order (descending order of index) from the Nth submodel unit to the first submodel unit, and the proportion of noise components in the input noise synthesis signal decreases from the Nth submodel unit to the first submodel unit, it is possible to arrange submodels having a configuration with high processing accuracy at the later stage. At the later stage, Gaussian white noise must be output from a signal with a low noise component, making prediction processing difficult, and by arranging submodel units having a configuration with high processing accuracy, it is possible to increase the processing speed while maintaining the accuracy of the overall signal generation processing.

[0027] Note that a configuration with high processing accuracy may have a large number of residual layers (or a large neural network model, a large number of parameters, etc.), resulting in a large circuit size but high processing accuracy.

[0028] A sixth invention is any one of the third to fifth inventions, wherein the control unit sets the noise schedule so that the sub-model units to be used when performing signal generation processing are distributed when selecting the sub-model units to be used from the first to Nth sub-model units based on the noise level determined based on the noise schedule.

[0029] This makes it possible to prevent processing from being biased toward specific sub-model units during signal generation processing (prediction processing). In a signal generation processing device, if the processing accuracy of a sub-model unit that processes many times is poor, the prediction accuracy of that sub-model unit will affect the overall processing accuracy. Therefore, by distributing the sub-model units that perform processing, it is possible to prevent the overall processing accuracy from being significantly affected by the processing accuracy of a specific sub-model unit, and as a result, the processing accuracy of the signal generation processing as a whole is improved.

[0030] A seventh invention is the first or second invention, wherein the noise level ranges to be associated with the first to Nth sub-model units are determined based on the logarithm of the noise level, and the first to Nth sub-model units each perform a learning process using a noise level included in the noise level range associated with that unit.

[0031] This makes it easier to distribute the sub-model units that perform the processing during the prediction process in this signal generation processing device. [Effects of the Invention]

[0032] According to the present invention, it is possible to realize a speech synthesis processing device that realizes speech synthesis processing capable of obtaining high-quality speech (speech signals) while maintaining the speed of speech synthesis processing. Also, according to the present invention, it is possible to realize a signal processing device that realizes signal generation processing capable of obtaining high-quality signals (e.g., image signals) while maintaining the processing speed for signals other than speech signals (e.g., image signals). [Brief explanation of the drawings]

[0033] [Figure 1] 1 is a schematic configuration diagram of a speech synthesis processing device 100 according to a first embodiment. [Figure 2] FIG. 2 is a schematic configuration diagram of a k-th sub-model unit of the speech synthesis processing device 100 according to the first embodiment. [Figure 3] FIG. 2 is a schematic configuration diagram of a k-th sub-model unit (DiffWave model) according to the first embodiment. [Figure 4] FIG. 2 is a schematic configuration diagram of a first residual layer k_RL1 of the k-th sub-model unit (DiffWave model) according to the first embodiment. [Figure 5] FIG. 2 is a schematic configuration diagram of a k-th sub-model unit (WaveGrad model) according to the first embodiment. [Figure 6] FIG. 2 is a schematic configuration diagram of a downsampling unit of a k-th sub-model unit (WaveGrad model) according to the first embodiment. [Figure 7] FIG. 2 is a schematic configuration diagram of a linear modulation section of the k-th sub-model section (WaveGrad model) according to the first embodiment. [Figure 8] FIG. 2 is a schematic configuration diagram of an upsampling unit of the k-th sub-model unit (WaveGrad model) according to the first embodiment. [Figure 9] 1 is a schematic diagram illustrating the configuration of a k-th sub-model unit of the speech synthesis processing device 100 according to the first embodiment (during learning processing). [Figure 10] 10 is a graph showing the relationship between an index n indicating the order of processing and a converted noise level sqrt(1−α′). [Figure 11] 2 is a diagram illustrating the selector and the k-th sub-model unit of the speech synthesis processing device 100. FIG. [Figure 12] 2 is a diagram illustrating the selector and the k-th sub-model unit of the speech synthesis processing device 100. FIG. [Figure 13] 1 is a schematic diagram illustrating the configuration of a k-th sub-model unit of the speech synthesis processing device 100 according to the first embodiment (during prediction processing). [Figure 14] 1 is a schematic diagram illustrating the configuration of a k-th sub-model unit of the speech synthesis processing device 100 according to the first embodiment (during prediction processing). [Figure 15] 1 is a schematic diagram illustrating the configuration of a k-th sub-model unit of the speech synthesis processing device 100 according to the first embodiment (during prediction processing). [Figure 16] Graph (vertical axis: log scale) showing the relationship between the index n indicating the order of processing and the converted noise level sqrt(1-α'). [Figure 17] FIG. 10 is a schematic configuration diagram of a signal generation processing device 200 according to a second embodiment. [Figure 18] FIG. 10 is a schematic configuration diagram of a kth sub-model unit of a signal generation processing device 200 according to a second embodiment. [Figure 19] FIG. 10 is a schematic configuration diagram of a k-th sub-model (image model) of a k-th sub-model unit of a signal generation processing device 200 according to a second embodiment. [Figure 20] FIG. 10 is a schematic configuration diagram of a residual block layer of the k-th sub-model (image model) of the signal generation processing device 200 according to the second embodiment. [Figure 21] FIG. 10 is a schematic configuration diagram of a k-th sub-model unit of a signal generation processing device 200 according to a second embodiment (during learning processing). [Figure 22]10 is a diagram illustrating the selector and the k-th sub-model unit of the signal generation processing device 200, and clearly shows the sub-model unit used by the noise schedule. FIG. [Figure 23] FIG. 10 is a schematic configuration diagram of a k-th sub-model unit of a signal generation processing device 200 according to a second embodiment (during prediction processing). [Figure 24] A diagram showing the CPU bus configuration. DETAILED DESCRIPTION OF THE INVENTION

[0034] [First embodiment] The first embodiment will be described below with reference to the drawings.

[0035] <1.1: Configuration of speech synthesis processing device> FIG. 1 is a schematic diagram of a speech synthesis processing device 100 according to the first embodiment.

[0036] As shown in Fig. 1, the speech synthesis processing device 100 includes a control unit 1, N1 selectors (in Fig. 1, N1=10, i.e., ten selectors SEL1 to SEL10), and N1 sub-model units (in Fig. 1, N1=10, i.e., ten sub-model units, i.e., a first sub-model unit 2_1 to a tenth sub-model unit 2_10). Note that, for ease of explanation, the following description will be given assuming N1=10, but N1 may be a natural number other than "10."

[0037] The control unit 1 generates data regarding the noise schedule Noise_schedule(={β1, β2, ..., β N}, β i : real number (i: integer, 1≦i≦N), 0≦β i ≦1), and generates control signals for controlling each sub-model unit and data required by each sub-model unit based on the data Noise_schedule. The control signals and data are compiled and output to each sub-model unit as sub-model control data. Note that the sub-model control data output to the k-th sub-model unit 2_k (k: integer, 1≦k≦N1) is denoted as Ctl(sub_Mk).

[0038] Furthermore, the control unit 1 generates selection signals for controlling the N1 selectors, and outputs the generated selection signals to the corresponding selectors.

[0039] The control unit 1 performs the following processing on the data Noise_schedule related to the noise schedule to obtain noise level data α n , weighting noise level data α n (w) Get. α n =1-β n (1≦n≦N)

number

[0040] Each of the N1 selectors selects an input and an output based on a selection signal output from the control unit 1, establishing a predetermined path. The first selector of the N1 selectors (selector SEL10 in FIG. 1) is a selector with one input and two outputs (one input terminal and two output terminals), and the last selector (selector SEL0 in FIG. 1) is a selector with two inputs and one output (two input terminals and one output terminal). The other selectors are selectors with two inputs and two outputs (two input terminals and two output terminals).

[0041] Furthermore, as shown in FIG. 1, the N1 selectors are arranged so that a path having one sub-model portion and a through path are secured between two adjacent selectors.

[0042] The N1 sub-model sections (in FIG. 1, the first sub-model section 2_1 to the tenth sub-model section 2_10) each have the same configuration. Here, the configuration of the kth sub-model section (k: natural number, 1≦k≦N1) will be described.

[0043] As shown in FIG. 2, the kth sub-model unit includes an input selector SELk_in, an input data generator 11, a selector SEL_k1, a kth sub-model SubM_k (k: natural number, 1≦k≦N1), a loss evaluator 12, a noise-reduced waveform acquirer 13, an output selector SELk_out, and a buffer 14.

[0044] The input selector SELk_in receives the signal (which is the signal y n _ext) and the signal output from the buffer 14 (called signal y n _inner) and selects one of the two inputs based on the selection signal sw_in, and outputs the selected signal as the signal y n The selection signal sw_in is output to the selector SEL_k1 as "_sel." It is assumed that the selection signal sw_in is included in the sub-model control data Ctl(sub_Mk) output from the control unit 1 to the k-th sub-model unit. The selection signal sw_in included in the sub-model control data Ctl(sub_Mk) is expressed as "Ctl(sub_Mk).sw_in."

[0045] The input data generation unit 11 is a functional unit that operates in a learning mode (a mode in which a learning process is executed), and generates voice waveform data y0 (correct answer data), Gaussian white noise w_noise, and weighting noise level data α' (during learning: α'=α (w) , when predicting: α'=α n (w) ) and the weighting noise level data α n (w)Based on this, the voice waveform data y0 and the Gaussian white noise w_noise are synthesized, and the synthesized data is referred to as voice-noise synthesized data y n _gen to the selector SEL_k1. Note that the weighting noise level data α' (during learning: α'=α (w) , when predicting: α'=α n (w) ) is included in the sub-model control data Ctl(sub_Mk) output from the control unit 1 to the k-th sub-model unit. In addition, the noise level data α for weighting at the time of prediction included in the sub-model control data Ctl(sub_Mk) is n (w) , "Ctl(sub_Mk).α n (w) ". Also, the noise level data α for weighting during learning included in the sub-model control data Ctl(sub_Mk) is (w) "Ctl(sub_Mk).α (w) " is written as ".

[0046] The selector SEL_k1 receives the output from the input selector SELk_in, the output from the input data generator 11, and the mode signal mode output from the controller 1. When the mode signal mode is the "learning mode," the selector SEL_k1 selects the terminal "1," selects the output from the input data generator 11, and outputs the signal y n On the other hand, when the mode signal mode is the "prediction mode", the selector SEL_k1 selects the terminal "0" to select the output from the input selector SELk_in, and outputs the signal y n and output it to the k-th sub-model SubM_k.

[0047] The k-th sub-model SubM_k receives the signal y n and noise level data α' (during learning: α' = α (w) , when predicting: α'=α n (w)) and the acoustic feature value h. When the k-th sub-model SubM_k executes the learning process (when in the learning mode), it receives the loss evaluation data Eva_θ output from the loss evaluation unit 12. The k-th sub-model SubM_k is, for example, a model using a neural vocoder, and receives the signal y n , the noise level data α′, and the acoustic feature value h, the k-th sub-model SubM_k performs a learning process to output Gaussian white noise. n , noise level data α', and acoustic feature value h are input, and the output signal ε θ to the loss evaluation unit 12. Then, the loss evaluation unit 12 outputs the output signal ε θ According to Eva_θ, which is the data evaluating the loss between Eva_θ and Gaussian white noise w_noise, the k-th submodel SubM_k updates the parameters and outputs the output signal ε θ and Gaussian white noise w_noise are subjected to a learning process so that the difference between them falls within a predetermined range.

[0048] The k-th sub-model SubM_k constructs a model in which the parameters (optimized parameters) acquired by the learning process are set as a trained model, and performs prediction processing using the trained model during prediction (speech synthesis processing). During prediction, the k-th sub-model SubM_k calculates the signal y n , noise level data α', and acoustic feature value h are input, and the output signal ε θ is output to the noise-reduced waveform acquisition unit 13.

[0049] <A: When using the DiffWave model> The k-th sub-model SubM_k can be realized, for example, by using the architecture disclosed in Non-Patent Document 3 (referred to as the "DiffWave model").

[0050] When the k-th sub-model SubM_k is realized using a DiffWave model, as shown in FIG. 3, the k-th sub-model SubM_k includes a 1x1 convolutional layer k1, an activation unit k2, a noise level acquisition unit k3, a position encoder k4, M (M: natural number) residual layers, namely, a first residual layer k_RL1 to an M-th residual layer k_RLM, an addition unit k5, a 1x1 convolutional layer k6, an activation unit k7, and a 1x1 convolutional layer k8.

[0051] The 1x1 convolution layer k1 receives the signal y output from the selector SEL_k1. n is input, and the signal y n The signal after the convolution process is output to the activation unit k2.

[0052] The activation unit k2 receives the output from the 1x1 convolution layer k1 as input, performs activation processing (for example, processing using an activation function (ReLU function, etc.)) on the input, and outputs the signal after the activation processing as a signal y n _in(1) and output to the first residual layer k_RL1.

[0053] The noise level acquisition unit k3 receives the weighting noise level data α' as input, performs noise level conversion processing on the weighting noise level data α', and acquires a converted noise level sqrt(1-α') (sqrt(x): square root of x). Then, the noise level acquisition unit k3 outputs the acquired converted noise level sqrt(1-α') to the position encoder k4.

[0054] The position encoder k4 receives the converted noise level sqrt(1-α') output from the noise level acquisition unit k3, performs position encoding on the converted noise level sqrt(1-α'), and acquires embedded representation data α'_emb including position information.The position encoder k4 then outputs the acquired embedded representation data α'_emb to the first residual layer k_RL1 to the M-th residual layer k_RLM.

[0055] The first residual layer k_RL1 to the M-th residual layer k_RLM, which are M (M: natural number) residual layers, each have the same configuration. Here, the configuration of the first residual layer k_RL1 will be described.

[0056] As shown in FIG. 4, the first residual layer k_RL1 includes a fully connected layer k101, an extension unit k102, an addition unit k103, a bidirectionally extended convolutional layer k104, a 1x1 convolutional layer k105, an addition unit k106, an activation unit k107, a 1x1 convolutional layer k108, a 1x1 convolutional layer k109, and an addition unit k110.

[0057] The fully connected layer k101 receives the embedded representation data α'_emb output from the position encoder k4, executes fully connected layer processing on the embedded representation data α'_emb, and outputs the signal after the processing of the fully connected layer to the extension unit k102.

[0058] The extension unit k102 adds the signal y n _in(1) and add it to the signal y n If _in(1) is a vector, the signal output from the fully connected layer k101 is expanded (for example, by copying) to match the dimension of the vector, and the signal y n The number of dimensions of the extension unit k102 is set to match the number of dimensions of _in(1). Then, the extension unit k102 outputs the data after the extension process to the addition unit k103.

[0059] The adder k103 receives the signal y n The activation unit k2 performs a process of adding _in(1) (output from the activation unit k2) and the output from the extension unit k102. The addition unit k103 then outputs the signal after the addition process to the bidirectional convolutional layer k104.

[0060] The bidirectional dilated convolutional layer k104 receives the signal output from the adder k103, performs bidirectional dilated convolution on the signal, and outputs the processed signal to the adder k106.

[0061] The 1x1 convolutional layer k105 receives an acoustic feature h as input, performs convolution processing on the acoustic feature h using a 1x1 kernel, and outputs the processed signal to an adder k106.

[0062] The adder k106 receives the output from the bidirectionally extended convolutional layer k104 and the output from the 1x1 convolutional layer k105 as input, and performs processing to add the output from the bidirectionally extended convolutional layer k104 and the output from the 1x1 convolutional layer k105. The adder k106 then outputs the signal after the addition processing to the activation unit k107.

[0063] The activation unit k107 receives the output from the addition unit k106 as input, performs activation processing (for example, processing using an activation function (such as a ReLU function)) on the input, and outputs the signal after the activation processing to the 1x1 convolution layers k108 and k109.

[0064] The 1x1 convolutional layer k108 receives the output from the activation unit k107 as input, performs convolution processing on the input using a 1x1 kernel, and outputs the processed signal to the addition unit k110.

[0065] The 1x1 convolutional layer k109 receives the output from the activation unit k107 as input, performs convolution processing on the input using a 1x1 kernel, and outputs the processed signal as a signal Do(1) to the addition unit k5.

[0066] The adder k110 receives the signal y n _in(1) (output from activation unit k2) and the output from the 1x1 convolution layer k108 are added together, and the signal after the addition is called signal y n_in(2) to the second residual layer. In other words, the output of the first residual layer k_KL1 is input to the second residual layer k_KL2 (signal y n _in(2)).

[0067] The second residual layer k_RL2 to the M-th residual layer k_RLM also have the same configuration as the first residual layer k_RL1.

[0068] The adder k5 inputs the signals Do(1) to Do(M) output from the first residual layer k_RL1 to the Mth residual layer k_RLM, respectively, performs addition processing on the signals Do(1) to Do(M), and outputs the signal after addition processing as a signal Do_sum to the 1x1 convolutional layer k6.

[0069] The 1x1 convolutional layer k6 receives the signal Do_sum output from the addition unit k5, performs convolution processing on the input using a 1x1 kernel, and outputs the processed signal to the activation unit k7.

[0070] The activation unit k7 receives the output from the 1x1 convolutional layer k6 as input, performs activation processing (for example, processing using an activation function (such as a ReLU function)) on the input, and outputs the signal after the activation processing to the 1x1 convolutional layer k8.

[0071] The 1x1 convolution layer k8 receives the output from the activation unit k7 as input, performs convolution processing on the input using a 1x1 kernel, and outputs the processed signal as a signal ε θ and outputs the result to the loss evaluation unit 12 and the noise-reduced waveform acquisition unit 13.

[0072] The loss evaluation unit 12 is a functional unit that operates in the learning processing mode, and evaluates the Gaussian white noise w_noise and the signal ε output from the k-th sub-model SubM_k. θ The loss evaluation unit 12 receives the Gaussian white noise w_noise and the signal ε θ Evaluate the loss (e.g., error) of the Gaussian white noise w_noise and the signal ε θand obtains parameters (update parameters) of the k-th sub-model for making the loss estimation data Eva_θ approach the loss estimation data Eva_θ, and outputs data including the parameters (update parameters) to the k-th sub-model. During the learning process, the k-th sub-model performs parameter update processing based on the loss estimation data Eva_θ output from the loss estimation unit 12. Then, the loss estimation unit 12 calculates the Gaussian white noise w_noise and the signal ε θ When the loss between the Gaussian white noise w_noise and the signal ε θ When the change in loss with the kth submodel falls within a predetermined range, it is determined that convergence has occurred and the learning process is terminated. Then, the parameters at the end of the learning process are set in the kth submodel, thereby obtaining a trained model for the kth submodel.

[0073] The noise-reduced waveform acquisition unit 13 is a functional unit that operates in the prediction processing mode, and receives the signal y output from the selector SEL_k1. n and the signal ε output from the k-th sub-model SubM_k θ and the noise level data α output from the control unit 1. n , and weighting noise level data α n (w) The noise-reduced waveform acquisition unit 13 inputs the noise level data α n , and weighting noise level data α n (w) Based on the signal y n and signal ε θ and the noise reduction process is performed using the signal y n-1 and outputs it to the output selector SELk_out.

[0074] The output selector SELk_out receives the output from the noise-reduced waveform acquisition unit 13 and selects an output in accordance with the selection signal sw_out included in the sub-model control data Ctl(sub_Mk) output from the control unit 1. When the value of the selection signal sw_out is "0", the input signal y n-1is output to the selector arranged after the k-th sub-model unit. When the value of the selection signal sw_out is "1", the input signal y n-1 is output to the buffer 14.

[0075] The buffer 14 receives the output from the output selector SELk_out and stores the input. The buffer 14 also outputs the stored signal as a signal y n _inner and output to the input selector SELk_in.

[0076] <B: When using the WaveGrad model> The k-th sub-model SubM_k can also be realized using, for example, the architecture disclosed in Non-Patent Document 2 (referred to as the "WaveGrad model").

[0077] When the kth sub-model SubM_k is realized by the WaveGrad model, as shown in FIG. 5, the kth sub-model SubM_k includes a 5x1 convolutional layer kk1, four downsampling units kk21 to kk24, a noise level acquisition unit kk3, five linear modulation units kk31 to kk35, a 3x1 convolutional layer kk4, five upsampling units kk51 to kk55, and a 3x1 convolutional layer kk6.

[0078] The 5x1 convolution layer kk1 receives the signal y output from the selector SEL_k1. n is input, and the signal y n The signal after the convolution process is output to the downsampling unit kk21.

[0079] The four downsampling units kk21 to kk24 have the same configuration and each include a downsampling layer kk201, an activation unit kk202, a 3×1 convolutional layer kk203, an activation unit kk204, a 3×1 convolutional layer kk205, an activation unit kk206, a 3×1 convolutional layer kk207, a 1×1 convolutional layer kk208, a downsampling layer kk209, and an adder k210, as shown in FIG.

[0080] The downsampling layer kk201 performs downsampling processing on the input Din and outputs the processed signal to the activation unit kk202.

[0081] The activation units kk202, k204, and k206 each perform activation processing (for example, processing using an activation function (ReLU function or the like)) on the input, and output the signal after the activation processing to a subsequent functional unit.

[0082] The 3x1 convolutional layers kk203, kk205, and kk207 each perform convolution processing on the input using a 3x1 kernel and output the convolutionally processed signal to a subsequent functional unit. The output of the 3x1 convolutional layer kk207 is output to the adder k210.

[0083] The 1x1 convolution layer kk208 performs convolution processing on the input Din using a 1x1 kernel, and outputs the signal after the convolution processing to the downsampling layer kk209.

[0084] The downsampling layer kk209 performs downsampling processing on the output from the 1x1 convolution layer kk208, and outputs the processed signal to the adder kk210.

[0085] The adder k210 adds the output of the 3×1 convolution layer kk207 and the output of the downsampling layer kk209, and outputs the processed signal as a signal Dout. In other words, the signal Dout is output to a downstream downsampling unit.

[0086] The noise level acquisition unit k3 receives the weighting noise level data α' as input, performs noise level conversion processing on the weighting noise level data α', and acquires a converted noise level sqrt(1-α') (sqrt(x): square root of x). Then, the noise level acquisition unit k3 outputs the acquired converted noise level sqrt(1-α') to each of the five linear modulation units kk31 to kk35.

[0087] The five linear modulation units kk31 to kk35 have the same configuration. As shown in FIG. 7, each of the five linear modulation units kk31 to kk35 includes a 3x1 convolution layer kk301, an activation unit kk302, a position encoder kk303, and an adder kk304. k 304, a 3x1 convolutional layer kk305, and a 3x1 convolutional layer kk306.

[0088] The 3x1 convolution layer kk301 performs convolution processing using a 3x1 kernel on the input Din (input from the downsampling unit to the linear modulation unit), and outputs the signal after the convolution processing to the activation unit kk302.

[0089] The activation unit kk302 performs activation processing (for example, processing using an activation function (ReLU function, etc.)) on the input, and outputs the signal after the activation processing to the adder k k Output to 304.

[0090] Position encoder kk303 teeth The converted noise level sqrt(1-α') output from the noise level acquisition unit kk3 is input, and the converted noise level sqrt(1-α') is subjected to position encoding processing to obtain the embedded representation data α'_emb including the position information. k303 outputs the acquired embedded expression data α'_emb to the adder kk304.

[0091] Addition section k k 304 is a position encoder kk30 3The output from the activation unit kk301 is added to the output from the activation unit kk302, and the processed signal is output to the 3x1 convolution layers kk305 and kk306.

[0092] The 3x1 convolution layer kk305 performs convolution processing using a 3x1 kernel on the output from the addition unit kk304, and acquires the data after the convolution processing as data γ.

[0093] The 3x1 convolution layer kk306 performs convolution processing using a 3x1 kernel on the output from the addition unit kk304, and acquires the data after the convolution processing as data ξ.

[0094] Then, the linear modulation unit outputs the data including the data γ and ξ acquired as described above to the upsampling unit as output data Dout_FiLM(={γ,ξ}).

[0095] The 3x1 convolutional layer kk4 receives an acoustic feature h as input, performs convolution processing on the input using a 3x1 kernel, and outputs the signal after the convolution processing to the upsampling unit kk51.

[0096] As shown in FIG. 8 , the five upsampling units kk51 to kk55 each include an activation unit kk501, an upsampling layer kk502, a 3x1 convolution layer kk503, an affine transformation layer kk504, an activation unit kk505, a 3x1 convolution layer kk506, an upsampling layer kk507, a 1x1 convolution layer kk508, an addition unit kk509, an affine transformation layer kk510, an activation unit kk511, a 3x1 convolution layer kk512, an affine transformation layer kk513, an activation unit kk514, a 3x1 convolution layer kk515, and an addition unit kk516.

[0097] The activation unit kk501 performs activation processing (for example, processing using an activation function (such as a ReLU function)) on the input to the upsampling unit (input Din in FIG. 8), and outputs the signal after the activation processing to the upsampling layer kk502.

[0098] The upsampling layer kk502 is the activation part kk50 1 The output from is subjected to upsampling processing, and the processed signal is output to the 3x1 convolutional layer kk503.

[0099] The 3x1 convolution layer kk503 performs convolution processing using a 3x1 kernel on the output from the upsampling layer kk502, and outputs the signal after the convolution processing to the affine transformation layer kk504.

[0100] The affine transformation layer kk504 receives the output from the 3x1 convolution layer kk503 and Dout_FiLM(={γ,ξ}) output from the linear modulation unit. If the output from the 3x1 convolution layer kk503 is Di, then the affine transformation layer kk504 receives Do = HadamardDot(γ,Di) + ξ HadamardDot(x,y): Function to get the Hadamard product of x and y and obtains the data Do. The affine transformation layer kk504 then outputs the obtained data Do to the activation unit kk505.

[0101] The activation unit kk505 performs activation processing (for example, processing using an activation function (ReLU function or the like)) on the output from the affine transformation layer kk504, and outputs the signal after the activation processing to the 3x1 convolution layer kk506.

[0102] The 3x1 convolution layer kk506 performs convolution processing using a 3x1 kernel on the output from the activation unit kk505, and outputs the signal after the convolution processing to the addition unit kk509.

[0103] The upsampling layer kk507 performs upsampling processing on the input to the upsampling unit (input Din in FIG. 8), and outputs the processed signal to the 1×1 convolutional layer kk508.

[0104] The 1x1 convolution layer kk508 performs convolution processing using a 1x1 kernel on the output from the upsampling layer kk507, and outputs the signal after the convolution processing to the adder kk509.

[0105] The adder kk509 executes a process of adding the output from the 3x1 convolution layer kk506 and the output from the 1x1 convolution layer kk508, and outputs the signal after the addition process to the adder kk516 and the affine transformation layer kk510.

[0106] The affine transformation layer kk510 receives the output from the adder kk509 and Dout_FiLM(={γ,ξ}) output from the linear modulation unit. If the output from the adder kk509 is Di, then the affine transformation layer kk510 receives the output from the adder kk509 as follows: Do = HadamardDot(γ,Di) + ξ HadamardDot(x,y): Function to get the Hadamard product of x and y and obtains the data Do. The affine transformation layer kk510 then outputs the obtained data Do to the activation unit kk511.

[0107] activation part kk511 teeth , an activation process (for example, a process using an activation function (ReLU function, etc.)) is performed on the output from the affine transformation layer kk510, and the signal after the activation process is output to the 3x1 convolution layer kk512.

[0108] The 3x1 convolution layer kk512 performs convolution processing using a 3x1 kernel on the output from the activation unit kk511, and outputs the signal after the convolution processing to the affine transformation layer kk513.

[0109] The affine transformation layer kk513 receives the output from the 3x1 convolution layer kk512 and Dout_FiLM(={γ,ξ}) output from the linear modulation unit. If the output from the 3x1 convolution layer kk512 is Di, then the affine transformation layer kk513 receives Do = HadamardDot(γ,Di) + ξ HadamardDot(x,y): Function to get the Hadamard product of x and y and obtains the data Do. The affine transformation layer kk513 then outputs the obtained data Do to the activation unit kk514.

[0110] activation part kk514 teeth , an activation process (for example, a process using an activation function (ReLU function, etc.)) is performed on the output from the affine transformation layer kk513, and the signal after the activation process is output to the 3x1 convolution layer kk515.

[0111] The 3x1 convolution layer kk515 performs convolution processing using a 3x1 kernel on the output from the activation unit kk514, and outputs the signal after the convolution processing to the addition unit kk516.

[0112] Addition section kk 516 executes a process of adding the output from the addition unit kk509 and the output from the 3x1 convolution layer kk515, and outputs the signal after the addition process as a signal Dout to the functional unit in the next stage.

[0113] The 3x1 convolution layer kk6 performs convolution processing using a 3x1 kernel on the output of the upsampling unit kk55, which is the final upsampling unit, and outputs the convolution processed signal as a signal ε θ and outputs the result to the loss evaluation unit 12 and the noise-reduced waveform acquisition unit 13.

[0114] In this way, the WaveGrad model can be adopted as the k-th sub-model SubM_k.

[0115] <1.2: Operation of speech synthesis processing device> The operation of the speech synthesis processing device 100 configured as above will be described below.

[0116] The following describes the operation of the speech synthesis processing device 100, dividing it into (1) learning processing (processing during learning) and (2) prediction processing (processing during prediction). For ease of explanation, the case where the number of sub-model units (sub-models) is "10" (N1=10) will be described.

[0117] (1.2.1: Learning process) First, the learning process performed by the speech synthesis processing device 100 will be described.

[0118] The speech synthesis processing device 100 can execute a learning process independently for each sub-model unit (sub-model). That is, a sub-model to be associated with each noise level is determined, and the sub-model is trained for each noise level, thereby obtaining a trained model for each sub-model.

[0119] For example, the speech synthesis processing device 100 determines a sub-model to be associated with each noise level as follows: Note that the following describes a case where a sub-model to be subjected to learning processing is determined according to the converted noise level sqrt(1-α'). (1) When 0≦sqrt(1-α')<0.1 The learning process is performed using the first sub-model unit 2_1 (first sub-model SubM_1). (2) When 0.1≦sqrt(1-α')<0.2 The learning process is performed using the second sub-model unit 2_2 (second sub-model SubM_2). (3) When 0.2≦sqrt(1-α')<0.3 The learning process is performed using the third sub-model unit 2_3 (third sub-model SubM_3). (4) When 0.3≦sqrt(1-α')<0.4 The learning process is performed using the fourth sub-model unit 2_4 (fourth sub-model SubM_4). (5) When 0.4≦sqrt(1−α')<0.5 The learning process is performed using the fifth sub-model unit 2_5 (fifth sub-model SubM_5). (6) When 0.5≦sqrt(1−α')<0.6 The sixth sub-model unit 2_6 (sixth sub-model SubM_6) is used to perform the learning process. (7) When 0.6≦sqrt(1-α')<0.7 The seventh sub-model unit 2_7 (seventh sub-model SubM_7) is used to perform the learning process. (8) When 0.7≦sqrt(1−α')<0.8 The eighth sub-model unit 2_8 (eighth sub-model SubM_8) is used to perform the learning process. (9) When 0.8≦sqrt(1−α')<0.9 The learning process is performed using the ninth sub-model part 2_9 (ninth sub-model SubM_9). (10) When 0.9≦sqrt(1−α')<1.0 The learning process is performed using the tenth sub-model part 2_10 (tenth sub-model SubM_10).

[0120] The specific contents of the learning process of the first sub-model unit 2_1 will be explained below.

[0121] The control unit 1 determines α' (at the time of learning: α' = α') that satisfies 0≦sqrt(1-α')<0.1. (w) ) to the first sub-model unit 2_1. Furthermore, the control unit 1 sets the mode signal mode to a value indicating the learning processing mode and outputs it to the first sub-model unit 2_1. Furthermore, the control unit 1 outputs the speech waveform signal y0 (correct answer data) to the first sub-model unit 2_1.

[0122] As shown in FIG. 9, the input data generating unit 11 of the first sub-model unit 2_1 receives the speech waveform signal y0 (correct answer data), the Gaussian white noise w_noise, and the weighting noise level α′, y n _gen=α'×y0+sqrt(1-α')×w_noise and executes the process equivalent to n Then, the input data generating unit 11 obtains the obtained signal y nThe selector SEL_k1 selects the terminal "1" according to the mode signal mode, and outputs the signal y n _gen to signal y n and output it to the k-th sub-model (k=1). Also, the acoustic feature h is input to the k-th sub-model (k=1).

[0123] In the k-th submodel (k=1), the processing is performed by the functional unit shown in FIG. 3, and the signal ε θ Then, the signal ε obtained by the k-th sub-model (k=1) is θ is output to the loss evaluation unit 12.

[0124] In the learning processing mode, the loss evaluation unit 12 calculates a Gaussian white noise w_noise and a signal ε θ The loss (for example, error) between the two is evaluated using a loss function defined by the following formula, for example.

number

[0125] Then, the loss evaluation unit 12 calculates the Gaussian white noise w_noise and the signal ε θ The parameters (updated parameters) of the k-th sub-model for making the equation (Eva_θ) approach the equation (Eva_θ) are obtained, and data including the parameters (updated parameters) is output to the k-th sub-model as loss evaluation data Eva_θ.

[0126] The k-th sub-model performs parameter update processing based on the loss evaluation data Eva_θ output from the loss evaluation unit 12. The loss evaluation unit 12 then performs parameter update processing based on the Gaussian white noise w_noise and the signal ε θ When the loss between the Gaussian white noise w_noise and the signal εθ When the change in loss with the kth submodel falls within a predetermined range, it is determined that convergence has occurred and the learning process is terminated. Then, the parameters at the end of the learning process are set in the kth submodel, thereby obtaining a trained model for the kth submodel.

[0127] In this manner, the learning process of the first sub-model part 2_1 is executed.

[0128] For the second sub-model section 2_2 to the tenth sub-model section 2_10, the noise level can be changed, for example, continuously within the corresponding noise level range, and a learning process can be performed to obtain a trained model within the corresponding noise level range.

[0129] In this way, the speech synthesis processing device 100 can perform training processing independently for each of the N1 (10) sub-model units. That is, for each of the N1 (10) sub-model units, training processing can be performed as long as the speech waveform signal y0, which is the correct data, the corresponding acoustic feature h, the Gaussian white noise w_noise, and the noise level that determines the ratio at which they are combined are known. Therefore, training processing for the N1 (10) sub-model units can be performed in parallel. This allows the speech synthesis processing device 100 to achieve high-speed training processing.

[0130] Furthermore, although the above description has been given of a case where the DiffWave model is used as the sub-model, the WaveGrad model may also be used as the sub-model.

[0131] When the WaveGrad model is used as a submodel, the k-th submodel performs processing using the functional unit shown in FIG. 5, and the signal ε θ Then, the signal ε obtained by the k-th sub-model (k=1) is θ is output to the loss evaluation unit 12.

[0132] In addition, in the speech synthesis processing device 100, the N1 (10) sub-model units may all be realized using the same model (for example, all using DiffWave models, or all using WaveGrad models), or different models may be mixed to realize the N1 (10) sub-model units.

[0133] For example, the early sub-model sections of the speech synthesis processing device 100 may be implemented by adopting a WaveGrad model, which has a fast processing speed but slightly inferior speech quality, and the later sub-model sections of the speech synthesis processing device 100 may be implemented by adopting a DiffWave model, which has a slow processing speed but high speech quality. In the speech synthesis processing device 100, during prediction (speech synthesis), signals with gradually reduced noise are output from the early sub-model sections (the tenth sub-model section in FIG. 1 ) to the later sub-model sections, and finally, a speech signal (a signal with the most reduced noise components) is output from the last sub-model section (the first sub-model section in FIG. 1 ). In other words, the early sub-model sections only need to output a signal with the noise components slightly reduced from Gaussian white noise w_noise, making the prediction process relatively easy. However, the later sub-model sections must output a signal with the noise components significantly reduced from Gaussian white noise w_noise, making the prediction process more difficult. Therefore, in the speech synthesis processing device 100, by placing high-speed but low-quality submodels in the early stages and placing lower-speed but high-quality submodels towards the end, the quality of the speech signal acquired (predicted) by the speech synthesis processing device 100 can be improved.

[0134] (1.2.2: Prediction processing (speech synthesis processing)) Next, the prediction process (voice synthesis process) performed by the voice synthesis processing device 100 will be described.

[0135] For ease of explanation, the noise schedule (={β1, β2, ..., β') is calculated so that the converted noise level sqrt(1-α') is a point (a black diamond point) on the broken line Pt1 shown in FIG. 10. N}, N=6) is determined.

[0136] Fig. 10 is a graph showing the relationship between the index n indicating the processing order and the converted noise level sqrt(1-α'). In Fig. 10, the sub-model shown on the right side of the graph is the sub-model applied within the range of the converted noise level sqrt(1-α').

[0137] The noise schedule (={β1, β2, ..., β') is calculated so that the converted noise level sqrt(1-α') is the point (black diamond point) (point P6 to point P1) on the broken line Pt1 shown in Figure 10. N}, N=6) is determined, prediction processing (speech synthesis processing) is executed in the order of the seventh sub-model unit 2_7 (repetition count: 1), the second sub-model unit 2_2 (repetition count: 1), and the first sub-model unit 2_1 (repetition count: 4).

[0138] The control unit 1 determines the noise schedule (={β1, β2, ..., β N}, N=6), the sub-model part to be used is determined. Specifically, the sub-model part to be used is determined as follows. (1) The converted noise level sqrt(1-α') (=sqrt(1-α')) corresponding to point P6 (corresponding to β6) in Figure 10 n (w) ), n=6) is 0.6≦sqrt(1-α n (w) )<0.7, the sub-model part that is used first when n=N=6 is the seventh sub-model part 2_7. (2) The converted noise level sqrt(1-α') (=sqrt(1-α')) corresponding to point P5 (corresponding to β5) in Figure 10 n (w) ), n=5) is 0.1≦sqrt(1-α n (w) )<0.2, the sub-model part used when n=5 is the second sub-model part 2_2. (3) The converted noise level sqrt(1-α') (=sqrt(1-α')) corresponding to point P4 (corresponding to β4) in Figure 10 n (w)), n=4) is 0≦sqrt(1-α n (w) )<0.1, the sub-model part used when n=4 is the first sub-model part 2_1. (4) The converted noise level sqrt(1-α') (=sqrt(1-α')) corresponding to point P3 (corresponding to β3) in Figure 10 n (w) ), n=3) is 0≦sqrt(1-α n (w) )<0.1, the sub-model part used when n=3 is the first sub-model part 2_1. (5) The converted noise level sqrt(1-α') (=sqrt(1-α')) corresponding to point P2 (corresponding to β2) in Figure 10 n (w) ), n=2) is 0≦sqrt(1-α n (w) )<0.1, the sub-model part used when n=2 is the first sub-model part 2_1. (6) The converted noise level sqrt(1-α') (=sqrt(1-α')) corresponding to point P1 (corresponding to β1) in Figure 10 n (w) ), n=1) is 0≦sqrt(1-α n (w) )<0.1, the sub-model part used when n=1 is the first sub-model part 2_1.

[0139] FIG. 11 is a diagram illustrating the selector and the k-th sub-model unit of the speech synthesis processing device 100. In FIG.

[0140] FIG. 12 is a diagram illustrating the selector and the k-th sub-model unit of the speech synthesis processing device 100, and clearly shows the sub-model unit used by the noise schedule.

[0141] When the sub-model unit to be used is determined as described above, the control unit 1 controls each selector so that the selectors SEL10, SEL9, and SEL8 are switched to select the through path, and the selector SEL7 is switched to select the path to the seventh sub-model unit 2_7, as shown in FIG. 12. As a result, the Gaussian white noise w_noise (=y N , N=6) is input to the seventh sub-model section 2_7.

[0142] Moreover, the control unit 1 outputs sub-model control data Ctl(sub_M7) to the seventh sub-model unit 2_7.

[0143] FIG. 13 is a schematic diagram of the kth sub-model unit during prediction processing.

[0144] In the seventh sub-model section 2_7, the signal y shown in FIG. n _ext is signal y N (=w_noise). Then, the control unit 1 selects the terminal "0" of the input selector SELk_in, and further selects the terminal "0" of the selector SELk_1, and the signal y N (=w_noise) is input to the k-th sub-model (k=7). Also, the acoustic feature value h is input to the k-th sub-model (k=7). Also, the control unit 1 controls α n (w) (n=6) (value α calculated from noise schedule β6 n (w) (n=6)) is input to the kth submodel (k=7).

[0145] In the k-th sub-model SubM_k (k=7), the processing (processing by the trained model) is performed by the functional unit shown in Fig. 3, and the signal ε θ Then, the signal ε obtained by the k-th sub-model SubM_k (k=7) is θ is output to the noise-reduced waveform acquisition unit 13.

[0146] The noise-reduced waveform acquisition unit 13 receives the signal y output from the selector SEL_k1.n (y N (=w_noise)) and the signal ε output from the k-th submodel SubM_k (k=7) θ and the noise level data α output from the control unit 1. n , and weighting noise level data α n (w) The noise-reduced waveform acquisition unit 13 inputs the noise level data α n , and weighting noise level data α n (w) Based on the signal y n and signal ε θ Specifically, the noise-reduced waveform acquisition unit 13 executes a process corresponding to the following formula to obtain the noise-reduced signal y n-1 (n=6) are obtained.

number

[0147] The control unit 1 selects the terminal "0" of the output selector SELk_out and outputs the signal y n-1 is output to the selector SEL6.

[0148] The control unit 1 selects the output (signal y n-1 , n=6) to the through path side. Also, the control unit 1 controls the switches so that the through path is selected in the selectors SEL5, SEL4, and SEL3, as shown in FIG.

[0149] Then, the control unit 1 controls the switch in the selector SEL2 so that the path to the second sub-model unit 2_2 is selected.

[0150] Then, in the second sub-model section 2_2, the same processing as that executed in the seventh sub-model section 2_7 is executed with the input signal being the signal y5, n=5. n-1 (=y4) is acquired, and the signal y n-1 (=y4) is output to the selector SEL1.

[0151] The control unit 1 controls the switch in the selector SEL1 so that the path to the first sub-model unit 2_1 is selected.

[0152] Then, in the first sub-model unit 2_1, the same processing as that executed in the seventh sub-model unit 2_7 is executed with the input signal being the signal y4, n=4. As can be seen from the graph in FIG. 10, processing using the first sub-model unit 2_1 is also executed when n=3, 2, or 1. Therefore, as shown in FIG. 14, the control unit 1 controls the switch to select the terminal "1" of the output selector SELk_out, and outputs the signal y n-1 =y3 is output to the buffer 14.

[0153] Then, in the first sub-model unit 2_1, the same processing as that executed in the seventh sub-model unit 2_7 is executed with the input signal being signal y3, n=3. Note that the input signal y3 is the signal y3 stored in the buffer 14 when n=4, and the signal y3 is output from the buffer 14 to the input selector SELk_in (see FIG. 15). Then, the control unit 1 controls the switch to select terminal "1" of the input selector SELk_in.

[0154] Then, in the first sub-model section 2_1, the signal y n-1 (=y2) is acquired, and the signal y n-1 (=y2) is output to the buffer 14 via the selector SELk_out (which selects the terminal "1").

[0155] Next, in the first sub-model unit 2_1, the same processing as that executed in the seventh sub-model unit 2_7 is executed with the input signal being signal y2, n=2. Note that the input signal y2 is the signal y2 stored in the buffer 14 when n=3, and the signal y2 is output from the buffer 14 to the input selector SELk_in (see FIG. 15). Then, the control unit 1 controls the switch to select terminal "1" of the input selector SELk_in.

[0156] Then, in the first sub-model section 2_1, the signal y n-1 (=y1) is acquired, and the signal y n-1 (=y1) is output to the buffer 14 via the selector SELk_out (which selects the terminal "1").

[0157] Next, in the first sub-model unit 2_1, the same processing as that executed in the seventh sub-model unit 2_7 is executed with the input signal being signal y1, n=1. Note that the input signal y1 is the signal y1 stored in the buffer 14 when n=2, and the signal y1 is output from the buffer 14 to the input selector SELk_in (see FIG. 15). Then, the control unit 1 controls the switch to select terminal "1" of the input selector SELk_in.

[0158] Then, in the first sub-model section 2_1, the signal y n-1 (=y0) is acquired, and the signal y n-1 (=y0) is output to the selector SEL0 via the selector SELk_out (which selects the terminal "0").

[0159] Then, the control unit 1 controls the switch in the selector SEL0 so as to select the output from the first sub-model unit 2_1, and acquires (outputs) the signal y0.

[0160] By performing the processing in this manner, the speech synthesis processing device 100 can acquire a speech signal y0 corresponding to the acoustic feature value h.

[0161] As described above, the speech synthesis processing device 100 can execute speech synthesis processing (prediction processing) by selecting and processing a sub-model portion determined in accordance with the noise schedule.

[0162] Although the above description has been given of a case where a noise schedule according to polygonal line Ptn1 in Fig. 10 is used, the speech synthesis processing device 100 can also perform speech synthesis processing using noise schedules according to other lines (patterns Ptn2, Ptn3, Ptn4, Ptn5) in Fig. 10. Even in this case, the sub-model unit to be used can be determined according to the noise schedule, and therefore the speech synthesis processing device 100 can perform speech synthesis processing (prediction processing) by performing prediction processing using the sub-model unit determined to be used.

[0163] In addition, in the speech synthesis processing device 100, the N1 (10) sub-model units may all be realized using the same model (for example, all using DiffWave models, or all using WaveGrad models), or different models may be mixed to realize the N1 (10) sub-model units.

[0164] For example, the early sub-model sections of the speech synthesis processing device 100 may be implemented by adopting a WaveGrad model, which has a fast processing speed but slightly inferior speech quality, and the later sub-model sections of the speech synthesis processing device 100 may be implemented by adopting a DiffWave model, which has a slow processing speed but high speech quality. In the speech synthesis processing device 100, during prediction (speech synthesis), signals with gradually reduced noise are output from the early sub-model sections (the tenth sub-model section in FIG. 1 ) to the later sub-model sections, and finally, a speech signal (a signal with the most reduced noise components) is output from the last sub-model section (the first sub-model section in FIG. 1 ). In other words, the early sub-model sections only need to output a signal with the noise components slightly reduced from Gaussian white noise w_noise, making the prediction process relatively easy. However, the later sub-model sections must output a signal with the noise components significantly reduced from Gaussian white noise w_noise, making the prediction process more difficult. Therefore, in the speech synthesis processing device 100, by placing high-speed but low-quality submodels in the early stages and placing lower-speed but high-quality submodels towards the end, the quality of the speech signal acquired (predicted) by the speech synthesis processing device 100 can be improved.

[0165] For example, in the case described above (when a noise schedule based on pattern Ptn1 in Figure 10 is adopted), the WaveGrad model, which has a fast processing speed but slightly poorer voice quality, may be adopted in the early sub-model sections, the seventh sub-model section 2_7 and the second sub-model section 2_2, and the DiffWave model, which has a slower processing speed but higher voice quality, may be adopted in the late sub-model section, the first sub-model section 2_1.

[0166] This enables the speech synthesis processing device 100 to perform high-quality speech synthesis processing (prediction processing) while improving the total processing speed.

[0167] Furthermore, in the speech synthesis processing device 100, the noise schedule may be determined so that the sub-model units to be used are distributed.

[0168] For example, as shown in Fig. 16, the noise schedule (={β1, β2, ..., β N}) may be determined.

[0169] Fig. 16 shows the graph in Fig. 10 with the vertical axis in logarithmic scale. As can be seen from the graph in Fig. 10, the sub-model parts determined from the noise schedule (sub-model parts used in the prediction process) are distributed.

[0170] For example, if a noise schedule based on pattern Ptn1 is adopted, in the case of Figure 10, the sub-model used is: (A1) 7th sub-model part 2_7 (processing count: 1) (corresponding to point P6 in Figure 10) (A2) Second sub-model part 2_2 (processing count: 1 time) (corresponding to point P5 in Figure 10) (A3) First sub-model part 2_1 (processing number: 4 times) (corresponding to points P4 to P1 in Figure 10) In the case of Figure 16, the sub-model to be used is (B1) 10th submodel part 2_10 (processing count: 1) (corresponding to point P6 in Figure 16) (B2) 8th sub-model part 2_8 (processing count: 1 time) (corresponding to point P5 in Figure 16) (B3) 6th sub-model part 2_6 (processing count: 1 time) (corresponding to point P4 in Figure 16) (B4) Fourth sub-model part 2_4 (processing count: 1 time) (corresponding to point P3 in Figure 16) (B5) Second sub-model part 2_2 (processing count: 1 time) (corresponding to point P2 in Figure 16) (B6) First sub-model part 2_1 (processing count: 1) (corresponding to point P1 in Figure 16) There is no sub-model unit that executes the process multiple times, and the sub-model units used are distributed. As for the processes (B1) to (B6) above, as in the above embodiment, the control unit 1 selects the sub-model unit to be used (selected by a selector) and executes the process by each sub-model unit, thereby enabling the speech synthesis processing (prediction processing) to be executed in the speech synthesis processing device 100.

[0171] In the speech synthesis processing device 100, the accuracy of the speech synthesis processing is improved by distributing the sub-model units used. This is because if the processing accuracy of a sub-model unit that is processed many times is poor, the prediction accuracy of that sub-model unit will affect the overall processing accuracy. In the speech synthesis processing device 100, by distributing the sub-model units used, it is possible to prevent the processing accuracy of a specific sub-model unit from having a significant effect, and as a result, the processing accuracy of the speech synthesis processing as a whole is improved.

[0172] As described above, the speech synthesis processing device 100 provides multiple sub-model units according to the noise level, and can perform learning processing independently (in parallel) for the multiple sub-model units, thereby significantly reducing the time required for the learning processing.

[0173] Furthermore, the speech synthesis processing device 100 executes speech synthesis processing (prediction processing) by using, in accordance with a noise schedule, a sub-model unit that has constructed a trained model trained according to the noise level. Furthermore, the speech synthesis processing device 100 can adopt (combine) an appropriate sub-model unit according to the noise level, thereby realizing speech synthesis processing that can acquire a high-quality speech signal while maintaining the speed of the speech synthesis processing.

[0174] [Second embodiment] Next, a second embodiment will be described. Note that the same parts as those in the above embodiment are given the same reference numerals and detailed description will be omitted.

[0175] In the first embodiment, a case where voice synthesis processing is performed (a signal processing device (voice synthesis processing device) that generates a voice signal) was described, whereas in the second embodiment, a case where image generation processing is performed (a signal generation device that generates an image signal) will be described.

[0176] FIG. 17 is a schematic configuration diagram of a signal generation processing device 200 according to the second embodiment.

[0177] FIG. 18 is a schematic configuration diagram of the kth sub-model unit of the signal generation processing device 200 according to the second embodiment.

[0178] FIG. 19 is a schematic configuration diagram of the k-th sub-model (image model) of the k-th sub-model unit of the signal generation processing device 200 according to the second embodiment.

[0179] FIG. 20 is a schematic diagram illustrating the configuration of a residual block layer of the k-th sub-model (image model) of the signal generation processing device 200 according to the second embodiment.

[0180] <2.1: Configuration of signal generation processing device> FIG. 17 corresponds to FIG. 1 of the first embodiment, and the configuration will be described focusing on the differences from FIG.

[0181] The signal generation processing device 200 of the second embodiment has a configuration in which, in the speech synthesis processing device 100 of the first embodiment, the control unit 1 is replaced with a control unit 1A, and the first sub-model unit 2_1 to the tenth sub-model unit 2_10 are replaced with the first sub-model unit 2A_1 to the tenth sub-model unit 2A_10, respectively.

[0182] In the speech synthesis processing device 100 of the first embodiment, the signal y N is Gaussian white noise w_noise, that is, a signal whose signal value at time t follows a Gaussian distribution (normal distribution). However, in the signal generation processing device 200 of the second embodiment, the signal y Nis Gaussian noise w_noise that can form a two-dimensional image (for example, an image of P pixels x Q pixels (P, Q: natural numbers)), that is, if the pixel value of coordinates (x, y) on the two-dimensional image is D(x, y), the pixel value D(x, y) is a signal (a signal that can form an image) that follows a Gaussian distribution (normal distribution).

[0183] Furthermore, in the speech synthesis processing device 100 of the first embodiment, the condition input to the speech synthesis processing device 100 is an acoustic feature h, whereas in the signal generation processing device 200 of the second embodiment, the condition input to the signal generation processing device 200 is data h that specifies a label (for example, a one-hot vector or one-hot data).

[0184] 18 corresponds to FIG. 2 of the first embodiment, and since the second embodiment processes an image, there are some differences from the first embodiment (configuration of FIG. 2). The input data generation unit 11 is a functional unit that operates in a learning mode (a mode in which learning processing is executed), and generates image data y0 (correct answer data), Gaussian white noise w_noise (Gaussian white noise w_noise capable of forming a two-dimensional image), and weighting noise level data α' (during learning: α'=α (w) , when predicting: α'=α n (w) ) and the data for the time step T n (Time step T n (n: natural number, 1≦n≦N) is the noise level data α n (n: natural number, 1≦n≦N) is the time step at which processing using the weighting noise level data α′ is executed. The input data generation unit 11 combines the image data y0 and Gaussian white noise w_noise based on the weighting noise level data α′, and generates the combined data as image noise combined data y n_gen to the selector SEL_k1. It is assumed that the size of the image formed by the image data y0 and the size of the image formed by the Gaussian white noise w_noise are the same, and pixel values at the same coordinates on the two-dimensional image in the image data y0 and the Gaussian white noise w_noise are added together to execute the synthesis process of the image data y0 and the Gaussian white noise w_noise. In addition, weighting noise level data α' (during learning: α'=α (w) , when predicting: α'=α n (w) ) and the data for the time step T n is included in the sub-model control data Ctl(sub_Mk) output from the control unit 1A to the k-th sub-model unit. Also, the noise level data α for weighting at the time of prediction included in the sub-model control data Ctl(sub_Mk) is n (w) "Ctl(sub_Mk).α n (w) ". Also, the noise level data α for weighting during learning included in the sub-model control data Ctl(sub_Mk) is (w) "Ctl(sub_Mk).α (w) ". Also, the data T about the time step included in the submodel control data Ctl(sub_Mk) is n "Ctl(sub_Mk).T n " is written as ".

[0185] The k-th sub-model SubMA_k receives the signal y n and the data for the time step T nand a condition h (for example, as shown in FIGS. 17 and 18, the condition h is data indicating "ball") that specifies a label (for example, one-hot vector or one-hot data). When the k-th sub-model SubMA_k executes a learning process (when in learning mode), it receives loss evaluation data Eva_θ output from the loss evaluation unit 12. The k-th sub-model SubMA_k is, for example, a model using a neural network, and receives a signal y n (The signal y that can form an image n ) and the data for the time step T n and the condition h (data specifying the label), a learning process is performed to output Gaussian white noise (Gaussian white noise w_noise that can form a two-dimensional image). In other words, the k-th sub-model SubMA_k performs a learning process to output Gaussian white noise w_noise that can form a two-dimensional image. n (a signal that can form an image yn) and data for a time step T n and condition h are input, and the output signal ε θ (An output signal ε that can form an image) θ ) to the loss evaluation unit 12. The loss evaluation unit 12 then outputs the output signal ε θ and Gaussian white noise w_noise, the data Eva_θ is obtained. According to the data Eva_θ, the k-th sub-model SubMA_k updates the parameters and outputs the output signal ε θ and Gaussian white noise w_noise are subjected to a learning process so that the difference between them falls within a predetermined range.

[0186] The k-th sub-model SubMA_k constructs a model in which the parameters (optimized parameters) acquired by the learning process are set as a trained model, and performs prediction processing using the trained model during prediction (image signal generation processing). n and the data for the time step T n and the condition h are used as inputs to perform prediction processing, and the output signal ε θ and obtains the output signal ε θis output to the noise-reduced waveform acquisition unit 13.

[0187] The k-th sub-model SubMA_k can be realized, for example, by the configuration shown in Fig. 19. Furthermore, the residual block layer in Fig. 19 can be realized, for example, by the configuration shown in Fig. 20. Regarding the implementation of this configuration, a program related to Non-Patent Document A described later is disclosed at the following URL, so a detailed description will be omitted. (URL for publishing a program related to Non-Patent Document A): https: / / github.com / hojonathanho / diffusion The differences from the configuration shown in FIG. 5 can be briefly explained as follows.

[0188] As shown in FIG. 19, the condition h and the time step T output from the control unit 1A n are subjected to embedding and activation processes, and then combined to obtain the combined data Dset(={Dh(h), Dt(T n )}) are output to downsampling layers ka2 to ka4 and upsampling layers ka5 to ka7. Each downsampling layer downsamples its input based on the data Dset. Each upsampling layer upsamples its input based on the data Dset.

[0189] As described above, the residual block layers ka_rn1 and ka_rn2 are realized by the configuration shown in FIG. 20. In the residual block layer, the output via the multiple network layers in the residual block layer and the input to the residual block layer are added together and output. For example, as shown in FIG. 20, the multiple network layers are composed of an activation unit ka_rn_1, a normalization unit ka_rn_2, a two-dimensional convolution layer ka_rn_3, an addition unit ka_rn_4, a normalization unit ka_rn_5, an activation unit ka_rn_6, a dropout unit ka_rn_7, and a two-dimensional convolution layer ka_rn_8. Data Dset is also input to the addition unit ka_rn_4. An attention unit ka_att is arranged between the two residual block layers, and attention processing (processing by an attention mechanism) is performed on the input to obtain context data (e.g., a context vector). Then, based on the obtained context data, data y n Weighting process for _r1 (or data y n The process of adding the acquired context data to _r1 is executed.

[0190] The loss evaluation unit 12 has the same configuration and function as in the first embodiment. The input data to the loss evaluation unit 12 includes data w_noise that can form a two-dimensional image, a signal ε θ (noise signal).

[0191] The noise-reduced waveform acquisition unit 13 has the same configuration and function as in the first embodiment. The input data to the loss evaluation unit 12 is a signal y n , signal ε θ (noise signal), and the output data is also a signal y that can form a 2D image. n-1 is.

[0192] The output selector SELk_out and the buffer 14 have the same configuration and function as those in the first embodiment.

[0193] <2.2: Operation of the signal generation processing device> The operation of the signal generation processing device 200 configured as above is almost the same as in the first embodiment, and the following description will focus on the differences.

[0194] In the signal generation processing device 200 shown in FIG. 17, the number of sub-models is assumed to be "10" for convenience' sake, as in the case of the speech synthesis processing device.

[0195] (2.2.1: Learning process) As with speech, a sub-model to be associated with each noise level is determined. The method for determining the sub-model to be trained according to the converted noise level sqrt(1-α') is the same as for speech.

[0196] As shown in FIG. 21, in the signal generation processing device 200 that processes images, as explained above, it is necessary to interpret the audio waveform signal y0 as the image signal y0, the Gaussian white noise w_noise as the Gaussian noise w_noise that can form a two-dimensional image, and the acoustic feature h as the label data h that identifies the image (e.g., a ball).

[0197] Also, the time step data T n is input to the k-th sub-model (k=1), which is also different from the case of speech.

[0198] From this, the loss function is defined by the following formula (t: time step is used instead of c in the case of speech processing):

number

[0199] In this way, the signal generation processing device 200 can perform learning processing independently for each of the N1 (10) sub-model units. That is, for each of the N1 (10) sub-model units, learning processing can be performed as long as the image signal y0, which is the correct data, the corresponding condition data h, the Gaussian white noise w_noise, and the noise level that determines the ratio at which they are combined are known. Therefore, the learning processing for the N1 (10) sub-model units can be performed in parallel. This makes it possible to speed up the learning processing in the signal generation processing device 200.

[0200] In addition, in the above, a case has been described in which the neural network model of FIG. 19 is employed in the signal generation processing device 200, but this is not limited to this, and a model other than the neural network model shown in FIG. 19 may also be employed.

[0201] Furthermore, in the signal generation processing device 200, the N1 (10) sub-model units may all be realized using the same model, or different models may be mixed to realize the N1 (10) sub-model units.

[0202] For example, the early sub-model section of the signal generation processing device 200 may be realized by adopting a neural network model that has a fast processing speed but produces images of slightly lower quality, and the later sub-model section of the signal generation processing device 200 may be realized by adopting a neural network model that has a slow processing speed but produces images of higher quality.

[0203] (2.2.2: Prediction processing (image generation processing)) Next, the prediction process (image generation process) performed by the signal generation processing device 200 will be described.

[0204] For ease of explanation, the noise schedule (={β1, β2, ..., β') is set so that the converted noise levels sqrt(1-α') are equally spaced. N}, N=1000) is determined. Also, a case will be described in which, during learning, a 1000-step converted noise level (a noise level that defines 1000 steps of noise with equal level intervals) is divided into 10 parts at equal intervals, and each sub-model part (first sub-model part 2A_1 to tenth sub-model part 2A_10) is trained using the converted noise level for each of the divided 100 steps to acquire a trained model.

[0205] The control unit 1 determines a noise schedule (={β1, β2, ..., β N}, N=1000). Specifically, when executing a prediction process in 1000 steps, the sub-model parts to be used are determined as follows so that each sub-model part (first sub-model part 2A_1 to tenth sub-model part 2A_10) executes the process for 100 steps. (1) From step 1000 to step 901, the processing is executed by the tenth sub-model part 2A_10. (2) Steps 900 to 801 are processed by the ninth sub-model unit 2A_9. (3) From step 800 to step 701, the eighth sub-model unit 2A_8 executes the processing. (4) From step 700 to step 601, the seventh sub-model part 2A_7 executes the processing. (5) From step 600 to step 501, the sixth sub-model section 2A_6 executes the processing. (6) From step 500 to step 401, the processing is executed by the fifth sub-model part 2A_5. (7) From step 400 to step 301, the processing is executed by the fourth sub-model part 2A_4. (8) From step 300 to step 201, the processing is executed by the third sub-model part 2A_3. (9) From step 200 to step 101, the second sub-model part 2A_2 executes the processing. (10) From step 100 to step 1, the first sub-model part 2A_1 executes the processing.

[0206] FIG. 22 is a diagram illustrating the selector and the k-th sub-model unit of the signal generation processing device 200, and clearly shows the sub-model unit used by the noise schedule.

[0207] When the sub-model unit to be used is determined as described above, the control unit 1A controls the selectors SEL10 to SEL1 as shown in FIG. 22 so that the selectors SEL10 to SEL1 are switched to select the path to the sub-model unit, and the selector SEL0 is switched to select the path to the first sub-model unit 2A_1. As a result, the Gaussian white noise w_noise (=y N , N=1000) is input to the tenth sub-model section 2A_10.

[0208] Moreover, the control unit 1A outputs sub-model control data Ctl(sub_M10) to the tenth sub-model unit 2A_10.

[0209] FIG. 23 is a schematic diagram of the kth sub-model unit during prediction processing.

[0210] In the tenth sub-model section 2A_10, the signal y shown in FIG. n _ext is signal y N (=w_noise). Then, the control unit 1A selects the terminal "0" of the input selector SELk_in, and further selects the terminal "0" of the selector SELk_1, and the signal y N (=w_noise) is input to the kth submodel (k=10). Also, the time step data T n and condition data h are input to the k-th sub-model (k=10).

[0211] In the k-th sub-model SubMA_k (k=10), the processing (processing by the trained model) is performed by the functional unit shown in FIG. 19, and the signal ε θ Then, the signal ε obtained by the k-th sub-model SubMA_k (k=10) is θ is output to the noise-reduced waveform acquisition unit 13.

[0212] The noise-reduced waveform acquisition unit 13 receives the signal y output from the selector SEL_k1. n (y N (=w_noise)) and the signal ε output from the k-th submodel SubMA_k (k=10) θ and noise level data α output from the control unit 1A. n (n=1000), and weighting noise level data α n (w) (n=1000) is input. The noise-reduced waveform acquisition unit 13 receives the noise level data α n , and weighting noise level data α n (w) Based on the signal y n and signal ε θ Specifically, the noise-reduced waveform acquisition unit 13 executes a process corresponding to the following formula to obtain the noise-reduced signal y n-1 (n=1000)

number

[0213] The control unit 1A selects the terminal "1" of the output selector SELk_out and outputs the signal y n-1 is output to the buffer 14.

[0214] Next, in the tenth sub-model section 2A_10, the input signal is converted into the signal y 999 , n=999, the same processing as that executed by the tenth sub-model unit 2A_10 when n=1000 is executed. 999 is the signal y stored in the buffer 14 when n=1000. 1000 and the signal y 1000is output from the buffer 14 to the input selector SELk_in. Then, the control unit 1A controls the switch so that the terminal "1" of the input selector SELk_in is selected.

[0215] Then, in the tenth sub-model section 2A_10, the signal y n-1 (=y 998 ) is acquired, and the signal y n-1 (=y 999 ) is output to the buffer 14 via the selector SEL1 (which selects the terminal "1").

[0216] From n=998 to n=902, the same processing as above is executed.

[0217] Then, in the tenth sub-model section 2A_10, the input signal is converted into the signal y 901 , n=901, and the same processing as above is executed. 901 is the signal y stored in the buffer 14 when n=902. 902 and the signal y 902 is output from the buffer 14 to the input selector SELk_in. Then, the control unit 1A controls the switch so that the terminal "1" of the input selector SELk_in is selected.

[0218] Then, in the tenth sub-model section 2A_10, the signal y n-1 (=y 900 ) is acquired, and the signal y n-1 (=y 900 ) is output to the selector SEL9 via the selector SELk_out (which selects the terminal "0").

[0219] Then, the control unit 1A controls the selector SEL9 to select the path to the ninth sub-model unit 2A_9, and outputs the signal y 900 is input to the ninth sub-model section 2A_9.

[0220] In the ninth sub-model section 2A_9, the same processing as that in the tenth sub-model section 2A_10 is executed for n=900 to n=801.

[0221] Furthermore, in the eighth sub-model section 2A_8 to the first sub-model section 2A_1, the same processing as that in the tenth sub-model section 2A_10 is executed.

[0222] Then, the control unit 1A controls the switch in the selector SEL0 so as to select the output from the first sub-model unit 2A_1, and acquires (outputs) the signal y0.

[0223] By performing this processing, the signal generation processing device 200 can obtain an image signal y0 corresponding to the condition data h.

[0224] As described above, the signal generation processing device 200 can execute processing (prediction processing) to generate an image signal by selecting and processing a sub-model portion determined according to the noise schedule.

[0225] In the above, the conversion noise level Although the case where the noise levels are equally spaced has been described, the present invention is not limited to this. A noise schedule may be determined based on a logarithm of the noise level (converted noise level), and the learning process and the prediction process in the signal generation processing device 200 may be performed based on the noise schedule.

[0226] As described above, the signal generation processing device 200 provides multiple sub-model units according to the noise level, and can perform learning processing independently (in parallel) for the multiple sub-model units, thereby significantly reducing the time required for the learning processing.

[0227] Furthermore, the signal generation processing device 200 executes a process of generating an image signal (prediction process) by using, in accordance with a noise schedule, a sub-model unit that has constructed a trained model trained according to the noise level. Furthermore, the signal generation processing device 200 can adopt (combine) an appropriate sub-model unit according to the noise level, thereby generating a high-quality image signal while maintaining the speed of the signal generation process (image signal generation process).

[0228] Although the above description has been given of the case where the condition data h is input, the present invention is not limited to this, and the signal generation processing device 200 may perform processing without inputting the condition data h. In this case, a comparative experiment was conducted with the case where the technology of the following Non-Patent Document A was used. (Non-patent document A): J. Ho, A. Jain, and P. Abbeel, "Denoising diffusion probabilistic models," in Proc. NeurIPS, Dec. 2020. Specifically, an image generation neural network model was trained using 50,000 training images from CIFAR10. In the model of Non-Patent Document A, training was performed using a noise level of 1000 steps, while in the present invention (corresponding to the signal generation processing device 200 without input of condition data h), 10 submodel sections were trained, each using a noise level of 100 steps, which was obtained by equally dividing the 1000 steps into 10. During image generation (prediction processing), random noise (Gaussian white noise) was used as input, and a random image was generated each time. To verify the accuracy of these generated images, the FID (Fenchel Inception Distance) between the 50,000 generated images and the 50,000 training images was calculated. The results are as follows: FID by the method of Non-Patent Document A: 5.71 FID in the present invention: 5.50 It was thus confirmed that the present invention (corresponding to the signal generation processing device 200 without input of the condition data h) has higher image generation accuracy.

[0229] [Other embodiments] In the above embodiment of the speech synthesis processing device, the case where a DiffWave model and a WaveGrad model are used in the sub-model unit has been described, but this is not limited to this, and the speech synthesis processing device may use other models that can acquire a speech waveform corresponding to an acoustic feature from Gaussian white noise or an acoustic feature.

[0230] Furthermore, in the speech synthesis processing device and signal generation processing device described in the above embodiments, each block may be individually integrated into a single chip using a semiconductor device such as an LSI, or may be integrated into a single chip to include some or all of the blocks.

[0231] Although we refer to it as an LSI here, it may also be called an IC, system LSI, super LSI, or ultra LSI depending on the level of integration.

[0232] Furthermore, the method of integration is not limited to LSI, but may be realized by dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays), which can be programmed after LSI manufacturing, or reconfigurable processors, which allow the connections and settings of circuit cells within LSI to be reconfigured, may also be used.

[0233] Furthermore, part or all of the processing of each functional block in each of the above embodiments may be realized by a program. And part or all of the processing of each functional block in each of the above embodiments is performed by a central processing unit (CPU) in a computer. Furthermore, the programs for performing each processing are stored in a storage device such as a hard disk or ROM, and are executed in the ROM or read out to the RAM.

[0234] Each process in the above-described embodiments may be realized by hardware, software (including cases where it is realized together with an OS (operating system), middleware, or a predetermined library), or may be realized by a combination of software and hardware.

[0235] For example, when each functional unit in the above embodiment is realized by software, each functional unit may be realized by software processing using the hardware configuration shown in FIG. 24 (for example, a hardware configuration in which a CPU, GPU, ROM, RAM, input unit, output unit, communication unit, memory unit (for example, a memory unit realized by an HDD, SSD, etc.), an external media drive, etc. are connected via a bus).

[0236] Furthermore, when each functional unit of the above embodiment is realized by software, the software may be realized using a single computer having the hardware configuration shown in Figure 24, or may be realized by distributed processing using multiple computers.

[0237] Furthermore, the execution order of the processing methods in the above-described embodiments is not necessarily limited to that described in the above-described embodiments, and the execution order can be changed within the scope of the gist of the invention.

[0238] The scope of the present invention includes a computer program for causing a computer to execute the above-described method, and a computer-readable recording medium having the program recorded thereon. Examples of computer-readable recording media include flexible disks, hard disks, CD-ROMs, MOs, DVDs, DVD-ROMs, DVD-RAMs, large-capacity DVDs, next-generation DVDs, and semiconductor memories.

[0239] The computer program is not limited to one recorded on the recording medium, but may be one transmitted via a telecommunications line, a wireless or wired communication line, a network such as the Internet, or the like.

[0240] The specific configuration of the present invention is not limited to the above-described embodiment, and various changes and modifications are possible without departing from the gist of the invention.

[0241] [Note] The present invention can also be realized as follows. <Appendix 1> A speech synthesis processing device that outputs a speech signal corresponding to an acoustic feature based on Gaussian white noise and the acoustic feature, The system includes N sub-model units (N: natural number, N≧2) including a first sub-model unit to an Nth sub-model unit, each of the first to N-th sub-model units includes a learning model that receives data on noise levels, acoustic features, and speech signals corresponding to the acoustic features, and performs a learning process to output Gaussian white noise from a noise synthesis signal that is a signal obtained by synthesizing the speech signal and Gaussian white noise based on the data on noise levels; The first to Nth sub-model units acquire trained models by performing a training process on the training models included in the first to Nth sub-model units using noise levels included in different noise level ranges, respectively. Speech synthesis processor. <Appendix 2> a control unit that sets a noise schedule; the control unit selects a sub-model unit to be used when performing speech synthesis processing from among the first sub-model unit to the N-th sub-model unit, based on the noise level determined based on the noise schedule, and determines the order of processing of the selected sub-model units; the selected sub-model unit executes a prediction process using the trained model in the order determined by the control unit, thereby acquiring a speech signal corresponding to the acoustic feature. 2. A speech synthesis processing device according to claim 1. <Appendix 3> the first to N-th sub-model sections are arranged in descending order from the N-th sub-model section to the first sub-model section, and the proportion of noise components in the input noise synthesis signal decreases from the N-th sub-model section to the first sub-model section; The sub-model unit arranged on the front side has a configuration in which the processing speed is faster than that of the sub-model unit arranged on the rear side. 3. A speech synthesis processing device according to claim 2. <Appendix 4> the first to N-th sub-model sections are arranged in descending order from the N-th sub-model section to the first sub-model section, and the proportion of noise components in the input noise synthesis signal decreases from the N-th sub-model section to the first sub-model section; The sub-model unit arranged at the rear stage has a configuration with higher processing accuracy than the sub-model unit arranged at the front stage. 3. A speech synthesis processing device according to claim 2. <Appendix 5> the control unit, when selecting sub-model units to be used in performing speech synthesis processing from among the first to N-th sub-model units according to the noise level determined based on the noise schedule, sets the noise schedule so that the sub-model units to be used are distributed. 5. A speech synthesis processing device according to any one of appendices 2 to 4. <Appendix 6> a noise level range to be associated with the first to Nth sub-model units is determined based on a logarithmic value of the noise level, and the first to Nth sub-model units perform the learning process using the noise level included in the noise level range associated with each of the first to Nth sub-model units. 2. A speech synthesis processing device according to claim 1. [Explanation of symbols]

[0242] 100 Speech synthesis processing device 200 Signal generation and processing device 1, 1A control section 2_1~2_10 1st submodel section~10th submodel section 2_1A~2_10A 1st submodel section~10th submodel section SubM_k kth submodel SubMA_k kth submodel

Claims

1. A signal generation processing device that outputs an audio signal or an image signal from Gaussian white noise, The system includes N sub-model units (N: natural number, N≧2) that are a first sub-model unit to an N-th sub-model unit, each of the first to N-th sub-model units includes a learning model that receives data on noise levels and a teacher signal of an audio signal or an image signal, and performs a learning process to output Gaussian white noise from a noise synthesis signal that is a signal obtained by synthesizing the teacher signal and Gaussian white noise based on the data on noise levels; The first to N-th sub-model units acquire a trained model by performing a training process on the training model included in the first to N-th sub-model units using noise levels included in different noise level ranges, respectively. Signal generation and processing device.

2. A signal generation processing device that outputs an audio signal or an image signal corresponding to an input condition feature based on Gaussian white noise and the input condition feature, The system includes N sub-model units (N: natural number, N≧2) that are a first sub-model unit to an N-th sub-model unit, each of the first to Nth sub-model units includes a learning model that receives as input data on noise levels, input condition features, and teacher signals of audio signals or image signals corresponding to the input condition features, and performs a learning process to output Gaussian white noise from a noise synthesis signal that is a signal obtained by synthesizing the teacher signals and Gaussian white noise based on the data on noise levels; The first to N-th sub-model units acquire a trained model by performing a training process on the training model included in the first to N-th sub-model units using noise levels included in different noise level ranges, respectively. Signal generation and processing device.

3. a control unit that sets a noise schedule; the control unit selects a sub-model unit to be used when performing signal generation processing from among the first sub-model unit to the N-th sub-model unit, based on the noise level determined based on the noise schedule, and determines an order of processing the selected sub-model units; The selected sub-model unit executes a prediction process using the trained model in the order determined by the control unit, thereby acquiring a voice signal or an image signal corresponding to the input condition feature. The signal generating and processing device according to claim 2 .

4. the first to N-th sub-model units have an order based on the ratio of the noise component of the input noise synthesis signal, the order is an order in which the ratio of the noise component of the noise combined signal decreases, and a sub-model unit positioned earlier in the order has a faster processing speed than a sub-model unit positioned later. The signal generating and processing device according to claim 3 .

5. the first to N-th sub-model units have an order based on the ratio of the noise component of the input noise synthesis signal, the order is an order in which the ratio of the noise component of the noise combined signal decreases, and a sub-model unit positioned later in the order has a configuration with higher processing accuracy than a sub-model unit positioned earlier. The signal generating and processing device according to claim 3 .

6. the control unit, when selecting a sub-model unit to be used when performing signal generation processing from among the first to N-th sub-model units, sets the noise schedule so that the sub-model units to be used are distributed according to the noise level determined based on the noise schedule.

6. The signal generating and processing device according to claim 3.

7. a noise level range to be associated with the first to N-th sub-model units is determined based on a logarithmic value of the noise level, and the first to N-th sub-model units perform the learning process using the noise level included in the noise level range associated with each of the first to N-th sub-model units. The signal generating and processing device according to claim 1 or 2.

Citation Information

Patent Citations

  • Voice recognition device, voice recognition program and voice recognition method

    JP2020034683A