Signal generation processing device
By using parallel learning and prediction processing of multiple sub-models and controlling the order with a noise schedule, the problems of speech synthesis quality and speed in WaveGrad and DiffWave models are solved, achieving efficient and high-quality speech and image signal generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NAT INST OF INFORMATION & COMM TECH
- Filing Date
- 2021-12-17
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, the quality of speech synthesis waveform signals obtained by single-speaker data learning models based on WaveGrad and DiffWave is worse than that of WaveGlow, and the learning and processing time is long, making it difficult to achieve high-quality and high-speed speech synthesis processing.
Multiple sub-model units are used for learning and processing, and learning and prediction are performed through different noise level ranges. The order and selection of sub-model units are controlled by a noise timetable to achieve parallel learning and prediction processing.
It achieves high-quality speech signal synthesis processing while maintaining the speed of speech synthesis processing, and extends to high-quality generation processing of image signals, improving the efficiency and accuracy of learning processing.
Smart Images

Figure CN116686043B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to processing techniques for generating speech signals or image signals (e.g., vocoder techniques for synthesizing speech waveforms from acoustic features). Background Technology
[0002] In text-to-speech (TTS) technology, which synthesizes natural speech from text, high-quality speech has been synthesized in recent years by introducing neural networks. Various techniques have been developed for vocoders used in such TTS technologies. For example, various models have been proposed for neural network vocoders that synthesize speech waveforms from acoustic features. Among them, the technique disclosed in Non-Patent Document 1 (hereinafter referred to as "WaveGlow") has attracted much attention due to its ability to perform real-time and high-quality synthesis. However, WaveGlow suffers from the following problems: the number of model parameters is enormous, resulting in a long learning and processing time (e.g., approximately 20 days even when using multiple GPUs). In contrast, diffusion probabilistic neural network vocoders disclosed in Non-Patent Documents 2 and 3 (referred to as "WaveGrad" in Non-Patent Document 2 and "DiffWave" in Non-Patent Document 3) have been developed. These diffusion probabilistic neural network vocoders (WaveGrad and DiffWave) use small models to achieve high-quality speech synthesis with fewer parameters.
[0003] WaveGrad and DiffWave models are neural network models that take a speech waveform signal as input, weighted and noise-added, and infer only the added noise component. Each model is implemented using a single unit. In WaveGrad and DiffWave, during learning, the weights themselves (data representing the weight values) are input into a model, corresponding to weights of various values (real numbers) between 0 and 1, and learned. Additionally, in WaveGrad and DiffWave, during prediction (during speech synthesis), initially only noise is input into the model. The noise component inferred by the model is subtracted from the input to obtain the inferred waveform. Next, slightly reduced-level noise is added to the inferred waveform and input into the same model (the WaveGrad and DiffWave models (neural network models)). The noise component inferred again by this model is subtracted from the input to obtain another inferred waveform. By gradually reducing the noise level and repeatedly performing this process, a clean speech waveform signal is finally obtained.
[0004] In waveform generation models such as WaveGrad and DiffWave, one key challenge is synthesizing the non-periodic components of the speech waveform that cannot be derived from the input acoustic features. In these models, the noise components that cannot be removed in the final stages are equivalent to the non-periodic components of the speech waveform. Therefore, waveform generation models like WaveGrad and DiffWave can achieve high-quality speech synthesis with fewer model parameters than WaveGlow.
[0005] Existing technical documents
[0006] Non-patent literature
[0007] Non-patent literature 1: R. Prenge, R. Valle, and B. Catanzaro, “WaveGlow: A flow-based generative network for speech synthesis,” in Proc. ICASSP, May 2019, pp. 3617-3621.
[0008] Non-patent literature 2: N. Chen, Y. Zhang, H. Zen, RJ Weiss, M. Norouzi, and W. Chan, “WaveGrad: Estimating gradients for waveform generation,” arXiv:2009.00713,2020.
[0009] Non-patent literature 3: Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “DiffWave: Aversatile diffusion model for audio synthesis,” arXiv:2009.09761,2020. Summary of the Invention
[0010] The technical problem that the invention aims to solve
[0011] However, when using data from a single speaker for learning processing to obtain the optimal model (the learned model), the following problem exists: the quality of the speech synthesized waveform signal obtained based on the optimal model (the learned model) of WaveGrad and DiffWave is worse than the quality of the speech synthesized waveform signal obtained based on the optimal model (the learned model) of WaveGlow.
[0012] Therefore, in view of the above problems, the object of the present invention is to provide a speech synthesis processing apparatus that can achieve speech synthesis processing that maintains the speed of speech synthesis processing and obtains high-quality speech (speech signal). Furthermore, the object of the present invention is to provide a signal processing apparatus that can achieve signal generation processing that maintains the processing speed and obtains high-quality signals (e.g., image signals) for signals other than speech signals (e.g., image signals).
[0013] Technical solutions for solving the problem
[0014] The first invention for solving the above-mentioned problem is a signal generation and processing device that outputs speech signals or image signals from Gaussian white noise, which includes a first sub-model section to an Nth sub-model section as N (N: natural number, N≥2) sub-model sections.
[0015] The first to Nth sub-models each have a learning model. The learning model takes noise level-related data and a supervision signal (speech or image signal) as input and outputs Gaussian white noise from the noise synthesis signal. The noise synthesis signal is a signal synthesized from the supervision signal and Gaussian white noise based on the noise level-related data.
[0016] The first to Nth sub-models use different noise levels within a range to perform learning processing on the learning models contained in the first to Nth sub-models, thereby obtaining the learned models.
[0017] In this signal generation and processing apparatus, the first sub-model unit to the Nth sub-model unit can obtain a learned model by performing learning processing on the learned models contained in the first sub-model unit to the Nth sub-model unit using noise levels included in different noise level ranges.
[0018] In other words, in this signal generation and processing apparatus, learning processing can be performed independently for each of the N sub-model units. That is, in each of the N sub-model units, if the supervisory signal serving as the forward solution data, the Gaussian white noise, and the noise level determining the ratio at which they are synthesized are known, learning processing can be performed, thus allowing the learning processing of the N sub-model units to be performed in parallel. Therefore, high-speed learning processing can be achieved in this signal generation and processing apparatus.
[0019] The second invention is a signal generation and processing device that outputs a speech signal or image signal corresponding to the input conditional features based on Gaussian white noise and input conditional features, and it comprises a first sub-model unit to an Nth sub-model unit as N (N: natural number, N≥2) sub-model units.
[0020] The first to Nth sub-models each have a learning model. The learning model takes noise level-related data, input conditional features, and a supervision signal of speech or image signals corresponding to the input conditional features as input, and performs learning processing by outputting Gaussian white noise from a noise synthesis signal. The noise synthesis signal is a signal synthesized by combining the supervision signal and Gaussian white noise based on noise level-related data.
[0021] The first to Nth sub-models use different noise levels within a range to perform learning processing on the learning models contained in the first to Nth sub-models, thereby obtaining the learned models.
[0022] In this signal generation and processing apparatus, the first sub-model unit to the Nth sub-model unit can obtain a learned model by performing learning processing on the learned models contained in the first sub-model unit to the Nth sub-model unit using noise levels included in different noise level ranges.
[0023] In other words, in this signal generation and processing apparatus, learning processing can be performed independently for each of the N sub-model units. Specifically, in each of the N sub-model units, if the supervision signal serving as the forward solution data, its corresponding input conditional features, Gaussian white noise, and the noise level determining the ratio at which they are synthesized are known, learning processing can be performed. Therefore, learning processing for the N sub-model units can be performed in parallel. Thus, high-speed learning processing can be achieved in this signal generation and processing apparatus.
[0024] The third invention is that, in addition to the second invention, it also includes a control unit for setting a noise timetable.
[0025] The control unit selects the sub-model unit to be used for signal generation processing from the first sub-model unit to the Nth sub-model unit according to the noise level, and determines the processing order of the selected sub-model unit. The noise level is determined according to the noise time schedule.
[0026] The selected sub-model unit performs prediction processing using the learned model in the order determined by the control unit, thereby obtaining speech or image signals corresponding to the input conditional features.
[0027] Therefore, in this signal generation and processing apparatus, during signal generation and processing (prediction processing), the sub-model unit to be used can be selected according to the noise level, which is determined according to the noise time schedule.
[0028] The fourth invention is an improvement upon the third invention, in which the first to Nth sub-model units have an order in which the proportion of the noise component of the input noise synthesis signal decreases, and the sub-model units located at the beginning of this order have a structure that allows them to process faster than the sub-model units located at the end.
[0029] Therefore, in this signal generation and processing apparatus, for example, "the proportion of noise components in the input noise synthesis signal is ordered", and the proportion of noise components in the input noise synthesis signal is in ascending order with respect to the index (1 to N) of the first sub-model section to the Nth sub-model section, (1) the sub-model section located at the front can be set as a sub-model section with a large proportion of noise components in the input noise synthesis signal and a fast processing speed, and (2) the sub-model section located at the rear can be set as a sub-model section with a small proportion of noise components in the input noise synthesis signal and a slow processing speed.
[0030] Therefore, in this signal generation and processing apparatus, for example, when the first to Nth sub-model units are arranged in descending order (in descending index) from the Nth sub-model unit toward the first sub-model unit, and when the proportion of noise component in the noise synthesis signal input as it moves from the Nth sub-model unit to the first sub-model unit is low, a sub-model with a high processing speed structure can be configured on the front-end side. On the front-end side, as long as Gaussian white noise can be output from the signal with high noise component, prediction processing is easier, and by configuring a sub-model unit with a high processing speed structure, the overall accuracy of signal generation and processing can be maintained while accelerating the processing speed.
[0031] The fifth invention is in the third invention, in which the first sub-model section to the Nth sub-model section have an order in which the proportion of the noise component of the input noise synthesis signal decreases, and the sub-model section located later in this order has a structure with higher processing accuracy than the sub-model section located in front.
[0032] Therefore, in this signal generation and processing apparatus, for example, "the proportion of noise components in the input noise synthesis signal is ordered", and the proportion of noise components in the input noise synthesis signal is in ascending order with respect to the index (1 to N) of the first sub-model section to the Nth sub-model section, (1) the sub-model section located at the rear can be set as a sub-model section with a small proportion of noise components in the input noise synthesis signal and a high processing accuracy, and (2) the sub-model section located at the front can be set as a sub-model section with a large proportion of noise components in the input noise synthesis signal and a low processing accuracy.
[0033] Therefore, in this signal generation and processing apparatus, for example, when the first to Nth sub-model units are arranged in descending order (in descending index) from the Nth sub-model unit toward the first sub-model unit, and when the proportion of noise component in the noise synthesis signal input as it moves from the Nth sub-model unit to the first sub-model unit is low, sub-models with high processing accuracy can be arranged on the back-end side. On the back-end side, Gaussian white noise must be output from the signal with low noise component, making prediction processing difficult. By arranging sub-model units with high processing accuracy, the overall signal generation and processing accuracy can be maintained while accelerating the processing speed.
[0034] Furthermore, as a structure with high processing accuracy, one can cite examples such as a large number of residual layers (or a large model size and number of parameters in the neural network model), which increases the circuit size but still achieves high processing accuracy.
[0035] The sixth invention is in any of the third to fifth inventions, in which the control unit selects the sub-model unit to be used for signal generation processing from the first sub-model unit to the Nth sub-model unit according to the noise level, and sets the noise time schedule in a decentralized manner for the sub-model units used, and the noise level is determined according to the noise time schedule.
[0036] Therefore, during signal generation processing (prediction processing), it is possible to prevent bias in processing towards specific sub-model units. In signal generation processing apparatuses, if the processing accuracy of a sub-model unit that processes a large number of times is poor, the prediction accuracy of that sub-model unit will affect the overall processing accuracy. Therefore, by dispersing the sub-model units being processed, it is possible to prevent significant influence from the processing accuracy of specific sub-model units, resulting in an improvement in the overall processing accuracy of signal generation processing.
[0037] The seventh invention is that, in the first or second invention, the noise level range corresponding to the first sub-model unit to the Nth sub-model unit is determined based on the value obtained by taking the logarithm of the noise level, and the first sub-model unit to the Nth sub-model unit respectively uses the noise level contained in the noise level range corresponding to itself to perform learning processing.
[0038] Therefore, in this signal generation and processing device, during prediction processing, the sub-models being processed are easily scattered.
[0039] (Invention Effects)
[0040] According to the present invention, a speech synthesis processing apparatus can be implemented that performs speech synthesis processing while maintaining the speed of speech synthesis processing and obtaining high-quality speech (speech signal). Furthermore, according to the present invention, a signal processing apparatus can be implemented that performs signal generation processing that maintains the processing speed and obtains high-quality signals (e.g., image signals) for signals other than speech signals (e.g., image signals). Attached Figure Description
[0041] Figure 1 This is a schematic configuration diagram of the speech synthesis processing apparatus 100 according to the first embodiment.
[0042] Figure 2 This is a schematic configuration diagram of the kth sub-model unit of the speech synthesis processing apparatus 100 according to the first embodiment.
[0043] Figure 3 This is a schematic diagram of the k-th sub-model unit (DiffWave model) involved in the first embodiment.
[0044] Figure 4 This is a schematic diagram of the first residual layer k_RL1 of the k-th sub-model unit (DiffWave model) involved in the first embodiment.
[0045] Figure 5 This is a schematic diagram of the k-th sub-model (WaveGrad model) involved in the first embodiment.
[0046] Figure 6 This is a schematic diagram of the downsampling section of the k-th sub-model (WaveGrad model) according to the first embodiment.
[0047] Figure 7 This is a schematic diagram of the linear modulation section of the k-th sub-model (WaveGrad model) according to the first embodiment.
[0048] Figure 8 This is a schematic diagram of the upsampling section of the k-th sub-model (WaveGrad model) according to the first embodiment.
[0049] Figure 9 This is a schematic configuration diagram of the kth sub-model unit of the speech synthesis processing apparatus 100 according to the first embodiment (during learning processing).
[0050] Figure 10 This is a graph showing the relationship between the index n, which represents the processing order, and the transformation noise level sqrt(1-α').
[0051] Figure 11This diagram illustrates the extraction of the selector and the k-th sub-model unit from the speech synthesis processing device 100.
[0052] Figure 12 This diagram illustrates the extraction of the selector and the k-th sub-model unit from the speech synthesis processing device 100.
[0053] Figure 13 This is a schematic configuration diagram of the kth sub-model unit of the speech synthesis processing apparatus 100 according to the first embodiment (during prediction processing).
[0054] Figure 14 This is a schematic configuration diagram of the kth sub-model unit of the speech synthesis processing apparatus 100 according to the first embodiment (during prediction processing).
[0055] Figure 15 This is a schematic configuration diagram of the kth sub-model unit of the speech synthesis processing apparatus 100 according to the first embodiment (during prediction processing).
[0056] Figure 16 This is a graph showing the relationship between the index n, which represents the processing order, and the transformation noise level sqrt(1-α') (vertical axis: log scale).
[0057] Figure 17 This is a schematic configuration diagram of the signal generation and processing apparatus 200 according to the second embodiment.
[0058] Figure 18 This is a schematic configuration diagram of the kth sub-model section of the signal generation and processing apparatus 200 according to the second embodiment.
[0059] Figure 19 This is a schematic configuration diagram of the kth sub-model (image model) of the kth sub-model unit of the signal generation and processing apparatus 200 according to the second embodiment.
[0060] Figure 20 This is a schematic diagram of the residual block layer of the k-th sub-model (image model) of the signal generation and processing apparatus 200 according to the second embodiment.
[0061] Figure 21 This is a schematic configuration diagram of the kth sub-model unit of the signal generation and processing apparatus 200 according to the second embodiment (during learning processing).
[0062] Figure 22 This diagram illustrates the selector and the k-th sub-model unit extracted from the signal generation and processing device 200, and explicitly shows the sub-model unit used according to the noise schedule.
[0063] Figure 23This is a schematic configuration diagram of the kth sub-model unit of the signal generation and processing apparatus 200 according to the second embodiment (during prediction processing).
[0064] Figure 24 This is a diagram showing the structure of the CPU bus. Detailed Implementation
[0065] [First Implementation Method]
[0066] The first embodiment will now be described with reference to the accompanying drawings.
[0067] <1.1: Structure of the speech synthesis processing device>
[0068] Figure 1 This is a schematic configuration diagram of the speech synthesis processing apparatus 100 according to the first embodiment.
[0069] like Figure 1 As shown, the speech synthesis processing device 100 includes a control unit 1 and N1 selectors (in... Figure 1 In this context, N1 = 10, consisting of 10 selectors SEL1 to SEL10, and N1 sub-models (in...). Figure 1 In this context, N1 = 10, representing 10 sub-model parts (i.e., the first sub-model part 2_1 to the tenth sub-model part 2_10). Furthermore, for ease of explanation, we will assume N1 = 10 in the following explanation, but N1 can also be a natural number other than "10".
[0070] Control unit 1 inputs data related to the noise schedule: Noise_schedule(={β1, β2, ..., β...) N}、β i : Real number (i: integer, 1≤i≤N), 0≤β i ≤1), based on the data Noise_schedule, generate control signals for controlling each sub-model unit and the data required by each sub-model unit, and output the data summed from the control signals and data as sub-model control data to each sub-model unit. In addition, the sub-model control data output to the k-th sub-model unit 2_k (k: integer, 1≤k≤N1) is represented as Ctl(sub_Mk).
[0071] In addition, the control unit 1 generates selection signals for controlling N1 selectors and outputs the generated selection signals to the corresponding selectors.
[0072] Control unit 1 performs the following processing based on the noise schedule data Noise_schedule to obtain noise level data α. n Weighted noise level data α n (w) .
[0073] α n =1-β n (1≤n≤N)
[0074]
Mathematical Formula 1
[0075]
[0076] Furthermore, control unit 1 is performing the nth noise timetable (β) n During the corresponding processing, in the learning processing mode, the data located at the weighted noise level α will be processed. n (w) With noise level data α n-1 (w) The real values (continuous values) between these values are used as weighted noise level data α. (w) Output to each sub-model section.
[0077] The N1 selectors select the input and output respectively based on the selection signal output from the control unit 1, and establish a predetermined path. The first-stage selector of the N1 selectors ( Figure 1 The selector SEL10 is a 1-input, 2-output selector (one input terminal and two output terminals). The last-stage selector ( Figure 1 The selector SEL0 is a 2-input, 1-output selector (two input terminals and one output terminal). All other selectors are 2-input, 2-output selectors (two input terminals and two output terminals).
[0078] In addition, such as Figure 1 As shown, the N1 selectors are configured in such a way that a path with a sub-model section and a direct path are ensured between two adjacent selectors.
[0079] N1 sub-models (in) Figure 1 The first sub-model part 2_1 to the tenth sub-model part 2_10 have the same structure. Here, the structure of the kth sub-model part (k: a natural number, 1≤k≤N1) which is the kth sub-model part will be explained.
[0080] like Figure 2 As shown, the k-th sub-model unit includes an input selector SELk_in, an input data generation unit 11, a selector SEL_k1, a k-th sub-model SubM_k (k: a natural number, 1≤k≤N1), a loss evaluation unit 12, a noise reduction waveform acquisition unit 13, an output selector SELk_out, and a buffer 14.
[0081] The input selector SELk_in takes a signal (referred to as signal y) from the output of the selector of the preceding stage located in the k-th sub-model section (output from the terminal on the k-th sub-model section side). n_ext) and the signal output from buffer 14 (referred to as signal y) n _inner), select one of the two inputs mentioned above based on the selection signal sw_in, and use the selected signal as the signal y. n The _sel output is sent to the selector SEL_k1. Furthermore, it is assumed that the selection signal sw_in is contained in the sub-model control data Ctl(sub_Mk) output from control unit 1 to the k-th sub-model unit. Additionally, the selection signal sw_in contained in the sub-model control data Ctl(sub_Mk) is expressed as "Ctl(sub_Mk).sw_in".
[0082] The input data generation unit 11 is a functional unit that operates in learning mode (the mode in which learning processing is performed). It inputs speech waveform data y0 (forward decoding data), Gaussian white noise w_noise, and weighted noise level data α' (during learning: α' = α). (w) When making a prediction: α' = α n (w) The input data generation unit 11 calculates the weighted noise level data α. n (w) The speech waveform data y0 is synthesized with Gaussian white noise w_noise, and the synthesized data is used as the speech noise synthesis data y. n The output of _gen is sent to the selector SEL_k1. Furthermore, it is assumed that the weighted data uses noise level α' (during learning: α' = α). (w) When making a prediction: α' = α n (w) The sub-model control data Ctl(sub_Mk) output from control unit 1 to the k-th sub-model unit is included. Additionally, the weighted noise level data α from the prediction process included in the sub-model control data Ctl(sub_Mk) is also included. n (w) Expressed as "Ctl(sub_Mk).α n (w) Additionally, the weighted learning time data α contained in the sub-model control data Ctl(sub_Mk) is used with noise level data. (w) Expressed as "Ctl(sub_Mk).α (w) ".
[0083] Selector SEL_k1 receives the output from input selector SELk_in, the output from input data generation unit 11, and the mode signal mode output from control unit 1. Furthermore, when the mode signal mode is "learning mode", selector SEL_k1 selects terminal "1", selects the output from input data generation unit 11, and uses it as signal y. nThe output is sent to the k-th sub-model SubM_k. On the other hand, selector SEL_k1 selects terminal "0" when the mode signal mode is "prediction mode", selecting the output from input selector SELk_in, and using it as signal y. n Output to the k-th sub-model SubM_k.
[0084] The input of the k-th sub-model SubM_k is the signal y output from the selector SEL_k1. n Noise level data α' (during learning: α' = α) (w) When making a prediction: α' = α n (w) The k-th sub-model SubM_k, during the learning process (in learning mode), receives the loss evaluation data Eva_θ output from the loss evaluation unit 12. The k-th sub-model SubM_k is, for example, a model using a neural network vocoder, which, during the learning process, calculates the loss evaluation data based on the input signal y. n The noise level data α' and acoustic feature quantity h are used for learning by outputting Gaussian white noise. In other words, the k-th sub-model SubM_k will learn the signal y... n The noise level data α' and acoustic characteristic h are used as inputs, and the output signal ε is used as output. θ The output is sent to the loss evaluation unit 12. Furthermore, the k-th sub-model SubM_k is evaluated by the loss evaluation unit 12 based on the output signal ε. θ The parameters are updated using the data Eva_θ obtained from the loss of Gaussian white noise w_noise, and the output signal ε is adjusted accordingly. θ The learning process is performed in such a way that the difference between the noise and the Gaussian white noise w_noise converges within a specified range.
[0085] The k-th sub-model SubM_k is constructed using the parameters (optimal parameters) obtained through the above learning process as a learned model. During prediction (speech synthesis), this learned model is used for prediction. During prediction, the k-th sub-model SubM_k will use the signal y... n The noise level data α' and acoustic feature h are used as inputs, and the output signal ε is... θ Output to the noise reduction waveform acquisition unit 13.
[0086] A: The case of using the DiffWave model
[0087] The k-th sub-model SubM_k can be implemented, for example, using the architecture disclosed in Non-Patent Document 3 (referred to as the "DiffWave model").
[0088] In the case of implementing the k-th sub-model SubM_k using the DiffWave model, such as Figure 3As shown, the k-th sub-model SubM_k has a 1×1 convolutional layer k1, an activation unit k2, a noise level acquisition unit k3, a position encoder k4, a first residual layer k_RL1 to the M-th residual layer k_RLM, which are M (M: natural number) residual layers, an addition unit k5, a 1×1 convolutional layer k6, an activation unit k7, and a 1×1 convolutional layer k8.
[0089] The 1×1 convolutional layer k1 will output the signal y from the selector SEL_k1 n As input, for signal y n A convolution process using a 1×1 kernel is performed, and the signal after the convolution process is output to the activation unit k2.
[0090] The activation unit k2 takes the output from the 1×1 convolutional layer k1 as input, performs activation processing on the input (e.g., by using an activation function (ReLU function, etc.), and uses the signal after activation processing as the signal y. n _in(1) is output to the first residual layer k_RL1.
[0091] The noise level acquisition unit k3 takes the weighted noise level data α' as input, performs noise level transformation processing on the weighted noise level data α', and obtains the transformed noise level sqrt(1-α') (sqrt(x): the square root of x). Furthermore, the noise level acquisition unit k3 outputs the obtained transformed noise level sqrt(1-α') to the position encoder k4.
[0092] The position encoder k4 takes the transformed noise level sqrt(1-α') output from the noise level acquisition unit k3 as input, performs position encoding processing on the transformed noise level sqrt(1-α'), and obtains embedded representation data α'_emb including position information. Furthermore, the position encoder k4 outputs the obtained embedded representation data α'_emb to the first residual layer k_RL1 to the Mth residual layer k_RLM.
[0093] The first residual layer k_RL1 to the Mth residual layer k_RLM, which are M (M: natural number) residual layers, have the same structure. Here, the structure of the first residual layer k_RL1 will be described.
[0094] like Figure 4 As shown, the first residual layer k_RL1 includes a fully connected layer k101, an expansion layer k102, an addition layer k103, a bidirectional expanded convolutional layer k104, a 1×1 convolutional layer k105, an addition layer k106, an activation layer k107, a 1×1 convolutional layer k108, a 1×1 convolutional layer k109, and an addition layer k110.
[0095] The fully connected layer k101 takes the embedded representation data α'_emb output from the position encoder k4 as input and performs fully connected layer processing on the embedded representation data α'_emb. Furthermore, the processed signal from the fully connected layer is output to the expansion unit k102.
[0096] The expansion unit k102 expands the output from the fully connected layer k101 so that the addition unit k103 can be used to perform addition on the signal y, which is the input of the first residual layer. n The addition processing of _in(1). For example, in the signal y n When _in(1) is a vector, the signal output from the fully connected layer k101 is expanded in a way that makes it consistent with the dimension of that vector (e.g., through copying), so that it is consistent with the signal y. n The dimension of _in(1) is consistent. Moreover, the expansion unit k102 outputs the expanded data to the addition unit k103.
[0097] Adder k103 processes the signal y, which is the input to the first residual layer. n The process of adding _in(1) (the output from the activation unit k2) to the output from the expansion unit k102. Furthermore, the addition unit k103 outputs the signal after addition processing to the bidirectional convolutional layer k104.
[0098] The bidirectional dilated convolutional layer k104 takes the signal output from the adder k103 as input, performs bidirectional dilated convolution on the signal, and outputs the processed signal to the adder k106.
[0099] The 1×1 convolutional layer k105 takes the acoustic feature h as input, performs convolution processing on the acoustic feature h based on the 1×1 kernel, and outputs the processed signal to the addition unit k106.
[0100] The addition unit k106 takes the outputs from the bidirectional dilated convolutional layer k104 and the 1×1 convolutional layer k105 as inputs and adds them together. Furthermore, the addition unit k106 outputs the processed signal to the activation unit k107.
[0101] The activation unit k107 takes the output from the addition unit k106 as input, performs activation processing on the input (e.g., by using an activation function (ReLU function, etc.), and outputs the activated signal to the 1×1 convolutional layers k108 and k109.
[0102] The 1×1 convolutional layer k108 takes the output from the activation unit k107 as input, performs convolution processing based on the 1×1 kernel on the input, and outputs the processed signal to the addition unit k110.
[0103] The 1×1 convolutional layer k109 takes the output from the activation unit k107 as input, performs convolution processing based on the 1×1 kernel on the input, and outputs the processed signal as signal Do(1) to the addition unit k5.
[0104] The adder k110 processes the signal y, which is the input to the first residual layer. n The process of adding _in(1) (the output from activation unit k2) to the output from 1×1 convolutional layer k108, and using the summed signal as the signal y. n _in(2) is output to the second residual layer. That is, the output of the first residual layer k_KL1 becomes the input of the second residual layer k_KL2 (signal y). n _in(2)).
[0105] The second residual layer k_RL2 to the Mth residual layer k_RLM also have the same structure as the first residual layer k_RL1.
[0106] The addition unit k5 takes in the signals Do(1) to Do(M) output from the first residual layer k_RL1 to the Mth residual layer k_RLM, respectively, and performs addition processing on the signals Do(1) to Do(M). The signal after addition processing is output as the signal Do_sum to the 1×1 convolutional layer k6.
[0107] The 1×1 convolutional layer k6 takes the signal Do_sum output from the addition unit k5 as input, performs convolution processing based on the 1×1 kernel on the input, and outputs the processed signal to the activation unit k7.
[0108] The activation unit k7 takes the output from the 1×1 convolutional layer k6 as input, performs activation processing on the input (e.g., by using an activation function (ReLU function, etc.), and outputs the activated signal to the 1×1 convolutional layer k8.
[0109] The 1×1 convolutional layer k8 takes the output from the activation unit k7 as input, performs convolution processing based on a 1×1 kernel on this input, and uses the processed signal as the signal ε. θ The output is sent to the loss evaluation unit 12 and the noise reduction waveform acquisition unit 13.
[0110] Loss evaluation unit 12 is a functional unit that operates in the learning processing mode, and is input with Gaussian white noise w_noise and signal ε output from the k-th sub-model SubM_k. θThe loss evaluation unit 12 compares Gaussian white noise w_noise with the signal ε. θ The loss (e.g., error) is evaluated, and the Gaussian white noise w_noise is used to compare the signal ε. θ The parameters of the k-th sub-model are updated, and data containing these parameters is output as loss evaluation data Eva_θ to the k-th sub-model. During the learning process, the k-th sub-model performs parameter update processing based on the loss evaluation data Eva_θ output from the loss evaluation unit 12. Furthermore, the loss evaluation unit 12 adjusts the Gaussian white noise w_noise and the signal ε... θ When the loss converges within the specified range, or even after parameter update processing, the Gaussian white noise w_noise and the signal ε θ When the change in loss is within the specified range, convergence is determined, and the learning process ends. Furthermore, by setting the parameters at the end of the learning process to the k-th sub-model, the learned model is obtained in the k-th sub-model.
[0111] The noise reduction waveform acquisition unit 13 is a functional unit that operates in predictive processing mode, and receives the signal y output from the selector SEL_k1 as input. n The signal ε output from the k-th sub-model SubM_k θ Noise level data α output from control unit 1 n and weighted noise level data α n (w) The noise reduction waveform acquisition unit 13 acquires the noise level data α. n and weighted noise level data α n (w) Using signal y n and signal ε θ Perform noise reduction processing, and use the processed signal as signal y. n-1 Output to the output selector SELk_out.
[0112] The output selector SELk_out takes the output from the noise reduction waveform acquisition unit 13 as input and selects the output based on the selection signal sw_out contained in the sub-model control data Ctl(sub_Mk) output from the control unit 1. When the value of the selection signal sw_out is "0", the input signal y is selected. n-1 The output is sent to the selector configured in the subsequent stage of the k-th sub-model. When the selection signal sw_out is set to "1", the input signal y... n-1 Output to buffer 14.
[0113] Buffer 14 takes the output from the output selector SELk_out as input and stores and holds that input. Additionally, buffer 14 stores and holds the signal as signal y.n The _inner output is sent to the input selector SELk_in.
[0114] B: The case of using the WaveGrad model
[0115] The k-th sub-model SubM_k can also be implemented using the architecture disclosed in Non-Patent Document 2 (referred to as the "WaveGrad model").
[0116] In the case of implementing the k-th sub-model SubM_k using the WaveGrad model, such as Figure 5 As shown, the k-th sub-model SubM_k has a 5×1 convolutional layer kk1, four downsampling units kk21~kk24, a noise level acquisition unit kk3, five linear modulation units kk31~kk35, a 3×1 convolutional layer kk4, five upsampling units kk51~kk55, and a 3×1 convolutional layer kk6.
[0117] The 5×1 convolutional layer kk1 will output the signal y from the selector SEL_k1 n As input, for signal y n A convolutional process using a 5×1 kernel is performed, and the signal after convolution is output to the downsampling unit kk21.
[0118] The four downsampling units kk21 to kk24 have the same structure. For example... Figure 6 As shown, the four downsampling units kk21 to kk24 respectively include a downsampling layer kk201, an activation unit kk202, a 3×1 convolutional layer kk203, an activation unit kk204, a 3×1 convolutional layer kk205, an activation unit kk206, a 3×1 convolutional layer kk207, a 1×1 convolutional layer kk208, a downsampling layer kk209, and an addition unit kk210.
[0119] The downsampling layer kk201 performs downsampling processing on the input Din and outputs the processed signal to the activation unit kk202.
[0120] The activation units kk202, k204, and k206 perform activation processing on the input (for example, by processing it through an activation function (ReLU function, etc.)) and output the signal after activation processing to the subsequent functional unit.
[0121] The 3×1 convolutional layers kk203, kk205, and kk207 perform convolution processing on the input using a 3×1 kernel, and output the processed signal to the subsequent functional unit. Furthermore, the output of the 3×1 convolutional layer kk207 is output to the adder kk210.
[0122] The 1×1 convolutional layer kk208 performs convolution processing on the input Din using a 1×1 kernel, and outputs the signal after convolution processing to the downsampling layer kk209.
[0123] The downsampling layer kk209 performs downsampling processing on the output from the 1×1 convolutional layer kk208 and outputs the processed signal to the adder kk210.
[0124] The adder kk210 adds the output of the 3×1 convolutional layer kk207 to the downsampling layer kk209, and outputs the processed signal as the signal Dout. In other words, the signal Dout is output to the downsampling unit in the next stage.
[0125] The noise level acquisition unit k3 takes the weighted noise level data α' as input, performs noise level transformation processing on the weighted noise level data α', and obtains the transformed noise level sqrt(1-α') (sqrt(x): the square root of x). Furthermore, the noise level acquisition unit k3 outputs the obtained transformed noise level sqrt(1-α') to each of the five linear modulation units kk31 to kk35.
[0126] The five linear modulation sections kk31 to kk35 have the same structure. For example... Figure 7 As shown, the five linear modulation units kk31 to kk35 respectively include a 3×1 convolutional layer kk301, an activation unit kk302, a position encoder kk303, an adder kk304, a 3×1 convolutional layer kk305, and a 3×1 convolutional layer kk306.
[0127] The 3×1 convolutional layer kk301 performs convolution processing on the input Din (the input from the downsampling unit toward the linear modulation unit) using a 3×1 kernel, and outputs the signal after convolution processing to the activation unit kk302.
[0128] The activation unit kk302 performs activation processing on the input (for example, by processing it through an activation function (ReLU function, etc.)) and outputs the activated signal to the addition unit k304.
[0129] The position encoder kk303 takes the noise level sqrt(1-α') as input from the noise level acquisition unit kk3 and performs position encoding processing on the noise level sqrt(1-α') to obtain embedded representation data α'_emb containing position information. Furthermore, the position encoder k303 outputs the obtained embedded representation data α'_emb to the addition unit kk304.
[0130] The addition unit kk304 performs the process of adding the output from the position encoder kk303 to the output from the activation unit kk302, and outputs the processed signal to the 3×1 convolutional layers kk305 and kk306.
[0131] The 3×1 convolutional layer kk305 performs convolution processing on the output from the addition unit kk304 using a 3×1 kernel, and obtains the data after this convolution processing as data γ.
[0132] The 3×1 convolutional layer kk306 performs convolution processing on the output from the addition unit kk304 using a 3×1 kernel, and obtains the data after this convolution processing as data ξ.
[0133] Furthermore, the linear modulation unit outputs data containing the data γ and ξ obtained above as output data Dout_FiLM(={γ,ξ}) to the upsampling unit.
[0134] The 3×1 convolutional layer kk4 takes the acoustic feature h as input, performs convolution processing on the input using a 3×1 kernel, and outputs the signal after convolution processing to the upsampling unit kk51.
[0135] like Figure 8 As shown, the five upsampling units kk51 to kk55 respectively include an activation unit kk501, an upsampling layer kk502, a 3×1 convolutional layer kk503, an affine transformation layer kk504, an activation unit kk505, a 3×1 convolutional layer kk506, an upsampling layer kk507, a 1×1 convolutional layer kk508, an addition unit kk509, an affine transformation layer kk510, an activation unit kk511, a 3×1 convolutional layer kk512, an affine transformation layer kk513, an activation unit kk514, a 3×1 convolutional layer kk515, and an addition unit 516.
[0136] The activation unit kk501 responds to the input toward the upsampling unit ( Figure 8 The input Din) is activated (e.g., processed by an activation function (ReLU function, etc.) and the activated signal is output to the upsampling layer kk502.
[0137] The upsampling layer kk502 performs upsampling processing on the output from the activation unit kk501 and outputs the processed signal to the 3×1 convolutional layer kk503.
[0138] The 3×1 convolutional layer kk503 performs convolution processing on the output from the upsampling layer kk502 using a 3×1 kernel, and outputs the signal after convolution processing to the affine transformation layer kk504.
[0139] The affine transformation layer kk504 takes as input the output from the 3×1 convolutional layer kk503 and Dout_FiLM(={γ,ξ}) from the linear modulation unit. Furthermore, if the output from the 3×1 convolutional layer kk503 is set to Di, the affine transformation layer kk504 performs processing equivalent to obtaining the HadamardDot(x,y): the Hadamard product of x and y, and obtains the data Do.
[0140] Do=HadamardDot(γ,Di)+ξ
[0141] Furthermore, the affine transformation layer kk504 outputs the obtained data Do to the activation unit kk505.
[0142] The activation unit kk505 performs activation processing on the output from the affine transformation layer kk504 (e.g., by processing with an activation function (ReLU function, etc.)) and outputs the activated signal to the 3×1 convolutional layer kk506.
[0143] The 3×1 convolutional layer kk506 performs convolution processing on the output from the activation unit kk505 using a 3×1 kernel, and outputs the signal after convolution processing to the addition unit kk509.
[0144] Upsampling layer kk507 for inputs directed toward the upsampling section ( Figure 8 The input Din) is upsampled and the processed signal is output to a 1×1 convolutional layer kk508.
[0145] The 1×1 convolutional layer kk508 performs convolution processing on the output from the upsampling layer kk507 using a 1×1 kernel, and outputs the signal after convolution processing to the addition unit kk509.
[0146] The addition unit kk509 performs the process of adding the output from the 3×1 convolutional layer kk506 with the output from the 1×1 convolutional layer kk508, and outputs the signal after addition to the addition unit kk516 and the affine transformation layer kk510.
[0147] The affine transformation layer kk510 takes as input the output from the adder kk509 and the output Dout_FiLM(={γ,ξ}) from the linear modulation unit. Furthermore, if the output from the adder kk509 is set to Di, the affine transformation layer kk510 performs processing equivalent to obtaining the HadamardDot(x,y): the Hadamard product of x and y, and acquires the data Do.
[0148] Do=HadamardDot(γ,Di)+ξ
[0149] Furthermore, the affine transformation layer kk510 outputs the obtained data Do to the activation unit kk511.
[0150] The activation unit kk511 performs activation processing on the output from the affine transformation layer kk510 (e.g., by processing with an activation function (ReLU function, etc.)) and outputs the activated signal to the 3×1 convolutional layer kk512.
[0151] The 3×1 convolutional layer kk512 performs convolution processing on the output from the activation unit kk511 using a 3×1 kernel, and outputs the signal after convolution processing to the affine transformation layer kk513.
[0152] The affine transformation layer kk513 takes as input the output of the 3×1 convolutional layer kk512 and the output Dout_FiLM(={γ,ξ}) from the linear modulation unit. Furthermore, if the output from the 3×1 convolutional layer kk512 is set to Di, the affine transformation layer kk513 performs processing equivalent to obtaining the HadamardDot(x,y): the Hadamard product of x and y, and obtains the data Do.
[0153] Do=HadamardDot(γ,Di)+ξ
[0154] Furthermore, the affine transformation layer kk513 outputs the obtained data Do to the activation unit kk514.
[0155] The activation unit kk514 performs activation processing on the output from the affine transformation layer kk513 (e.g., by processing with an activation function (ReLU function, etc.)) and outputs the activated signal to the 3×1 convolutional layer kk515.
[0156] The 3×1 convolutional layer kk515 performs convolution processing on the output from the activation unit kk514 using a 3×1 kernel, and outputs the signal after convolution processing to the addition unit kk516.
[0157] The addition unit kk516 performs the process of adding the output from the addition unit kk509 with the output from the 3×1 convolutional layer kk515, and outputs the signal after addition as the signal Dout to the next functional unit.
[0158] The 3×1 convolutional layer kk6 performs convolution processing on the output of the final upsampling part, i.e., the upsampling part kk55, using a 3×1 kernel, and uses the signal after this convolution processing as the signal ε. θ The output is sent to the loss evaluation unit 12 and the noise reduction waveform acquisition unit 13.
[0159] Thus, the WaveGrad model can be used as the k-th sub-model SubM_k.
[0160] <1.2: Operation of the speech synthesis processing device>
[0161] The operation of the speech synthesis processing device 100 configured as described above will now be explained.
[0162] Hereinafter, the operation of the speech synthesis processing device 100 will be described in two parts: (1) learning processing (processing during learning) and (2) prediction processing (processing during prediction). In addition, for ease of explanation, the case in which the number of sub-model units (sub-models) is "10" (N1 = 10) will be described.
[0163] (1.2.1: Learning Processing)
[0164] First, the learning process of the speech synthesis processing device 100 will be explained.
[0165] In the speech synthesis processing apparatus 100, learning processing can be performed independently for each sub-model unit (sub-model). That is, a sub-model corresponding to each noise level is determined, and the sub-model is learned for each noise level, thereby obtaining the learned model of each sub-model.
[0166] For example, in the speech synthesis processing apparatus 100, a sub-model corresponding to each noise level is determined as follows. Furthermore, the following explains the case where the sub-model of the object to be learned is determined based on the transformed noise level sqrt(1-α').
[0167] (1) The case where 0 ≤ sqrt(1-α') < 0.1
[0168] Learning is performed using the first sub-model part 2_1 (first sub-model SubM_1).
[0169] (2) The case where 0.1 ≤ sqrt(1-α') < 0.2
[0170] The learning process is performed using the second sub-model part 2_2 (second sub-model SubM_2).
[0171] (3) The case where 0.2 ≤ sqrt(1-α') < 0.3
[0172] The learning process is performed using the third sub-model part 2_3 (the third sub-model SubM_3).
[0173] (4) The case where 0.3 ≤ sqrt(1-α') < 0.4
[0174] Learning is performed using the fourth sub-model part 2_4 (the fourth sub-model SubM_4).
[0175] (5) The case where 0.4 ≤ sqrt(1-α') < 0.5
[0176] Learning is performed using the fifth sub-model 2_5 (the fifth sub-model SubM_5).
[0177] (6) The case where 0.5 ≤ sqrt(1-α') < 0.6
[0178] Learning is performed using the sixth sub-model part 2_6 (sixth sub-model SubM_6).
[0179] (7) The case where 0.6 ≤ sqrt(1-α') < 0.7
[0180] Learning is performed using the seventh sub-model part 2_7 (the seventh sub-model SubM_7).
[0181] (8) The case where 0.7 ≤ sqrt(1-α') < 0.8
[0182] Learning is performed using the eighth sub-model 2_8 (the eighth sub-model SubM_8).
[0183] (9) The case where 0.8 ≤ sqrt(1-α') < 0.9
[0184] Learning is performed using the ninth sub-model 2_9 (the ninth sub-model SubM_9).
[0185] (10) The case where 0.9 ≤ sqrt(1-α') < 1.0
[0186] The learning process is performed using the tenth sub-model part 2_10 (the tenth sub-model SubM_10).
[0187] The following section explains the specific details of the learning process in the first sub-model section 2_1.
[0188] Control unit 1 will set α' to satisfy 0 ≤ sqrt(1-α') < 0.1 (during learning: α' = α). (w) The control unit 1 outputs the mode signal 'mode' to the first sub-model unit 2_1. Additionally, the control unit 1 sets the mode signal 'mode' to a value representing the learning processing mode and outputs it to the first sub-model unit 2_1. Furthermore, the control unit 1 outputs the speech waveform signal 'y0' (forward solution data) to the first sub-model unit 2_1.
[0189] like Figure 9 As shown, the input data generation unit 11 of the first sub-model unit 2_1 takes the speech waveform signal y0 (forward solution data), Gaussian white noise w_noise, and weighted noise level α' as input, and performs an operation equivalent to...
[0190] y n _gen=α'×y0+sqrt(1-α')×w_noise
[0191] Processing and acquiring signal y n Furthermore, the input data generation unit 11 outputs the obtained signal yn_gen to the selector SEL_k1. The selector SEL_k1 selects terminal "1" according to the mode signal mode, and outputs the signal yn_gen... n _gen as signal y n The output is fed into the k-th sub-model (k=1). Additionally, the acoustic feature h is input into the k-th sub-model (k=1).
[0192] In the k-th sub-model (k=1), through Figure 3 The functional unit shown performs processing to acquire signal ε. θ Furthermore, the signal ε obtained through the k-th sub-model (k=1) θ It is output to the loss evaluation department 12.
[0193] In the learning processing mode, the loss evaluation unit 12, for example, uses a loss function defined by the following mathematical formula to evaluate the Gaussian white noise w_noise and the signal ε. θ The loss (e.g., error) is evaluated.
[0194]
Mathematical Formula 2
[0195]
[0196] Furthermore, when using the DiffWave model as a sub-model, c = n; when using the WaveGrad model as a sub-model, c = sqrt(α). (w) ).
[0197] Furthermore, the loss evaluation unit 12 obtains the value of the loss function to compare the Gaussian white noise w_noise with the signal ε. θ The parameters of the k-th sub-model are updated, and the data containing these parameters is output as the loss evaluation data Eva_θ to the k-th sub-model.
[0198] The k-th sub-model performs parameter update processing based on the loss evaluation data Eva_θ output from the loss evaluation unit 12. Furthermore, the loss evaluation unit 12 adjusts the Gaussian white noise w_noise and the signal ε... θ When the loss converges within the specified range, or even after parameter update processing, the Gaussian white noise w_noise and the signal ε θ When the change in loss is within the specified range, convergence is determined, and the learning process ends. Furthermore, by setting the parameters at the end of the learning process to the k-th sub-model, the learned model is obtained in the k-th sub-model.
[0199] The learning process of the first sub-model section 2_1 is performed through the above steps.
[0200] For the second sub-model section 2_2 to the tenth sub-model section 2_10, the noise level can be continuously varied within the corresponding noise level range and then learned, thereby obtaining the learned model within the corresponding noise level range.
[0201] In this way, in the speech synthesis processing apparatus 100, learning processing can be performed independently for each of the N1 (10) sub-model units. That is, in each of the N1 (10) sub-model units, if the speech waveform signal y0, which serves as the forward solution data, its corresponding acoustic feature h, Gaussian white noise w_noise, and the noise level that determines the ratio at which they are synthesized are known, learning processing can be performed. Therefore, the learning processing of the N1 (10) sub-model units can be performed in parallel. Thus, high-speed learning processing can be achieved in the speech synthesis processing apparatus 100.
[0202] In addition, the above explanation focused on the case where the DiffWave model is used as a sub-model, but the WaveGrad model can also be used as a sub-model.
[0203] When the WaveGrad model is used as a sub-model, in the k-th sub-model, through Figure 5 The functional unit shown performs processing to acquire signal ε. θ Furthermore, the signal ε obtained through the k-th sub-model (k=1) θ It is output to the loss evaluation department 12.
[0204] In addition, in the speech synthesis processing device 100, all N1 (10) sub-model units can be implemented using the same model (for example, all using the DiffWave model or all using the WaveGrad model). Alternatively, N1 (10) sub-model units can be implemented by mixing different models.
[0205] For example, the initial stage sub-model unit of the speech synthesis processing device 100 can be implemented using the WaveGrad model, which has a fast processing speed but slightly lower speech quality, while the final stage sub-model unit can be implemented using the DiffWave model, which has a slow processing speed but high speech quality. In the speech synthesis processing device 100, during prediction (speech synthesis), the initial stage sub-model unit (in...) Figure 1 The tenth sub-model section outputs a signal with gradually reduced noise to the subsequent sub-model sections, ultimately exiting from the final stage sub-model section (in...). Figure 1The first sub-model unit outputs a speech signal (a signal with noise components reduced to the maximum extent). In other words, in the initial stage of the sub-model unit, it is sufficient to output a signal that slightly reduces the noise components from Gaussian white noise w_noise, making prediction processing relatively simple. However, in the final stage of the sub-model unit, a signal with significantly reduced noise components from Gaussian white noise w_noise must be output, making prediction processing difficult. Therefore, in the speech synthesis processing apparatus 100, by configuring a high-speed but low-quality sub-model in the initial stage and configuring a lower-speed but higher-quality sub-model as the final stage approaches, the quality of the speech signal obtained (predicted) by the speech synthesis processing apparatus 100 can be improved.
[0206] (1.2.2: Predictive Processing (Speech Synthesis Processing))
[0207] Next, the prediction processing (speech synthesis processing) performed by the speech synthesis processing device 100 will be described.
[0208] Furthermore, for ease of explanation, for the transformation noise level sqrt(1-α') to become Figure 10 The way the points (black diamond points) of the broken line Pt1 shown determine the noise time table (={β1, β2, ..., β...) N The case of N=6) will be explained.
[0209] Figure 10 This is a graph showing the relationship between the index n, representing the processing order, and the transformation noise level sqrt(1-α'). Furthermore, in Figure 10 In the diagram, the sub-model shown on the right side is the sub-model applied within the range of the transformed noise level sqrt(1-α').
[0210] When the transformed noise level is sqrt(1-α') Figure 10 The way the points (black diamond points) (points P6 to P1) of the broken line Pt1 shown determine the noise time table (={β1, β2, ..., β...) N In the case of N=6), prediction processing (speech synthesis processing) is performed in the order of the seventh sub-model unit 2_7 (repetition count: 1), the second sub-model unit 2_2 (repetition count: 1), and the first sub-model unit 2_1 (repetition count: 4).
[0211] Control unit 1 operates according to the determined noise schedule (={β1, β2, ..., β... N The sub-model part to be used is determined by N=6. Specifically, the sub-model part to be used is determined as follows.
[0212] (1) Due to the Figure 10The conversion noise level corresponding to point P6 (corresponding to β6) is sqrt(1-α')(=sqrt(1-α) n (w) (n=6) is 0.6≤sqrt(1-α) n (w) Since ) < 0.7, the first sub-model unit used when n = N = 6 is the seventh sub-model unit 2_7.
[0213] (2) Due to the relationship with Figure 10 The conversion noise level corresponding to point P5 (corresponding to β5) is sqrt(1-α')(=sqrt(1-α) n (w) (n=5) is 0.1≤sqrt(1-α) n (w) Since ) < 0.2, the sub-model part used when n = 5 is the second sub-model part 2_2.
[0214] (3) Due to the relationship with Figure 10 The conversion noise level corresponding to point P4 (corresponding to β4) is sqrt(1-α')(=sqrt(1-α) n (w) (n=4) is 0≤sqrt(1-α) n (w) Since ) < 0.1, the sub-model part used when n = 4 is the first sub-model part 2_1.
[0215] (4) Due to the Figure 10 The conversion noise level corresponding to point P3 (corresponding to β3) is sqrt(1-α')(=sqrt(1-α) n (w) (n=3) is 0≤sqrt(1-α) n (w) Since ) < 0.1, the sub-model part used when n = 3 is the first sub-model part 2_1.
[0216] (5) Due to Figure 10 The conversion noise level corresponding to point P2 (corresponding to β2) is sqrt(1-α')(=sqrt(1-α) n (w) (n=2) is 0≤sqrt(1-α) n (w) Since ) < 0.1, the sub-model part used when n = 2 is the first sub-model part 2_1.
[0217] (6) Due to the Figure 10 The conversion noise level corresponding to point P1 (corresponding to β1) is sqrt(1-α')(=sqrt(1-α) n (w)(n=1) is 0≤sqrt(1-α) n (w) Since ) < 0.1, the sub-model part used when n = 1 is the first sub-model part 2_1.
[0218] Figure 11 This diagram illustrates the extraction of the selector and the k-th sub-model unit from the speech synthesis processing device 100.
[0219] Figure 12 This diagram illustrates the selector and the k-th sub-model unit extracted from the speech synthesis processing device 100, and explicitly shows the sub-model unit used according to the noise timetable.
[0220] When the sub-model section to be used is determined as described above, such as Figure 12 As shown, control unit 1 controls each selector so that selectors SEL10, SEL9, and SEL8 are switched to select the direct path, and selector SEL7 is switched to select the path leading to the seventh sub-model unit 2_7. Therefore, Gaussian white noise w_noise(=y) is input to selector SEL10. N (N=6) is input into the seventh sub-model part 2_7.
[0221] In addition, the control unit 1 outputs sub-model control data Ctl(sub_M7) to the seventh sub-model unit 2_7.
[0222] Figure 13 This is a schematic diagram of the k-th sub-model part during prediction processing.
[0223] In the seventh sub-model section 2_7 Figure 13 The signal y shown n _ext becomes signal y N (=w_noise). Furthermore, by selecting terminal "0" of input selector SELk_in via control unit 1, terminal "0" of selector SELk_1 is selected, thereby causing signal y... N (=w_noise) is input into the k-th sub-model (k=7). Additionally, the acoustic feature h is input into the k-th sub-model (k=7). Furthermore, α is controlled by control unit 1. n (w) (n=6)(The value α calculated based on the noise timetable β6) n (w) (n=6)) Input the k-th sub-model (k=7).
[0224] In the k-th sub-model SubM_k (k=7), through Figure 3 The functional unit shown performs processing (processing via the learned model) to acquire the signal ε. θFurthermore, the signal ε obtained through the k-th sub-model SubM_k (k=7) θ It is output to the noise reduction waveform acquisition unit 13.
[0225] The noise reduction waveform acquisition unit 13 inputs the signal y output from the selector SEL_k1. n (y N (=w_noise)), the signal ε output from the k-th sub-model SubM_k (k=7) θ Noise level data α output from control unit 1 n and weighted noise level data α n (w) The noise reduction waveform acquisition unit 13 acquires the noise level data α. n and weighted noise level data α n (w) and using signal y n and signal ε θ Noise reduction processing is performed. Specifically, the noise reduction waveform acquisition unit 13 acquires the noise-reduced signal y by performing processing equivalent to the following mathematical formula. n-1 (n=6).
[0226]
Mathematical Expression 3
[0227]
[0228]
[0229] z ~ N(0, I) (when n>0) (I is the identity matrix)
[0230] z = 0 (when n = 0)
[0231] Furthermore, the noise reduction waveform acquisition unit 13 will use the signal y obtained above... n-1 Output to the output selector SELk_out.
[0232] Control unit 1 selects terminal "0" of output selector SELk_out and sends signal y n-1 Output to selector SEL6.
[0233] Control unit 1 selects the output (signal y) from the seventh sub-model unit in selector SEL6. n-1 Switching control is performed by outputting to the direct path side (n=6). Additionally, as... Figure 12 As shown, the control unit 1 performs switching control by selecting a direct path from selectors SEL5, SEL4, and SEL3.
[0234] Furthermore, the control unit 1 performs switching control by selecting the path to the second sub-model unit 2_2 in the selector SEL2.
[0235] Furthermore, in the second sub-model section 2_2, the input signal is set to signal y5, n=5, and the same processing as that performed in the seventh sub-model section 2_7 is executed. Thus, signal y is obtained in the second sub-model section 2_2. n-1 (=y4), and the signal y n-1 (=y4) is output to selector SEL1.
[0236] Control unit 1 switches the path to the first sub-model unit 2_1 by selecting the path in selector SEL1.
[0237] Furthermore, in the first sub-model section 2_1, the input signal is set to signal y4, n=4, and the same processing as that performed in the seventh sub-model section 2_7 is executed. Additionally, by Figure 10 As can be seen from the chart, the processing of the first sub-model part 2_1 was also performed when n=3, 2, and 1. Therefore, as Figure 14 As shown, control unit 1 switches the signal by selecting terminal "1" of output selector SELk_out. n-1 =y3 is output to buffer 14.
[0238] Furthermore, in the first sub-model section 2_1, the input signal is set to signal y3, n=3, and the same processing as that performed in the seventh sub-model section 2_7 is executed. Additionally, the input signal y3 is the signal y3 stored in buffer 14 when n=4, and signal y3 is output from buffer 14 to the input selector SELk_in (see...). Figure 15 Furthermore, the control unit 1 switches the input selector by selecting terminal "1" of the input selector SELk_in.
[0239] Furthermore, in the first sub-model section 2_1, the signal y is obtained. n-1 (=y2), and the signal y is sent via selector SELk_out (selection terminal "1"). n-1 (=y2) is output to buffer 14.
[0240] Next, in the first sub-model section 2_1, the input signal is set to signal y2, n=2, and the same processing as that performed in the seventh sub-model section 2_7 is executed. Furthermore, the input signal y2 is the signal y2 stored in buffer 14 when n=3, and signal y2 is output from buffer 14 to the input selector SELk_in (see...). Figure 15 Furthermore, the control unit 1 switches the input selector by selecting terminal "1" of the input selector SELk_in.
[0241] Furthermore, in the first sub-model section 2_1, the signal y is obtained.n-1 (=y1), the signal y n-1 (=y1) is output to buffer 14 via selector SELk_out (select terminal “1”).
[0242] Next, in the first sub-model section 2_1, the input signal is set to signal y1, n=1, and the same processing as that performed in the seventh sub-model section 2_7 is executed. Furthermore, the input signal y1 is the signal y1 stored in buffer 14 when n=2, and signal y1 is output from buffer 14 to the input selector SELk_in (see...). Figure 15 Furthermore, the control unit 1 switches the input selector by selecting terminal "1" of the input selector SELk_in.
[0243] Furthermore, in the first sub-model section 2_1, the signal y is obtained. n-1 (=y0), the signal y n-1 (=y0) is output to selector SEL0 via selector SELk_out (selector terminal “0”).
[0244] Furthermore, the control unit 1 switches the control by selecting the output from the first sub-model unit 2_1 in the selector SEL0, and acquires the (output) signal y0.
[0245] By performing this processing, a speech signal y0 equivalent to the acoustic feature h can be obtained in the speech synthesis processing device 100.
[0246] As described above, in the speech synthesis processing apparatus 100, speech synthesis processing (predictive processing) can be performed by selecting and processing a sub-model unit determined according to a noise timetable.
[0247] Furthermore, in the above, regarding the use of based on Figure 10 The case of the noise timetable of the broken line Ptn1 has been explained, but in the speech synthesis processing device 100, a method based on... Figure 10 Speech synthesis processing is performed on the noise timetable of the other lines (patterns Ptn2, Ptn3, Ptn4, Ptn5). In this case, the sub-model unit to be used can also be determined according to the noise timetable. Therefore, by using the sub-model unit to be used for prediction processing, speech synthesis processing (predictive processing) can be performed by the speech synthesis processing device 100.
[0248] In addition, in the speech synthesis processing device 100, all N1 (10) sub-model units can be implemented using the same model (for example, all using the DiffWave model or all using the WaveGrad model). Alternatively, N1 (10) sub-model units can be implemented by mixing different models.
[0249] For example, the initial stage sub-model unit of the speech synthesis processing device 100 can be implemented using the WaveGrad model, which has a fast processing speed but slightly lower speech quality, while the final stage sub-model unit can be implemented using the DiffWave model, which has a slow processing speed but high speech quality. In the speech synthesis processing device 100, during prediction (speech synthesis), the initial stage sub-model unit (in...) Figure 1 The tenth sub-model section outputs a signal with gradually reduced noise to the subsequent sub-model sections, ultimately exiting from the final stage sub-model section (in...). Figure 1 The first sub-model unit outputs a speech signal (a signal with noise components reduced to the maximum extent). In other words, in the initial stage of the sub-model unit, it is sufficient to output a signal that slightly reduces the noise components from Gaussian white noise w_noise, making prediction processing relatively simple. However, in the final stage of the sub-model unit, a signal with significantly reduced noise components from Gaussian white noise w_noise must be output, making prediction processing difficult. Therefore, in the speech synthesis processing apparatus 100, by configuring a high-speed but low-quality sub-model in the initial stage and configuring a lower-speed but higher-quality sub-model as the final stage approaches, the quality of the speech signal obtained (predicted) by the speech synthesis processing apparatus 100 can be improved.
[0250] For example, in the case described above (using a method based on...) Figure 10 In the case of the noise timetable of pattern Ptn1, the WaveGrad model, which has a fast processing speed but slightly poor speech quality, can be used in the seventh sub-model section 2_7 and the second sub-model section 2_2, which are the initial stage sub-model sections, and the DiffWave model, which has a slow processing speed but high speech quality, can be used in the first sub-model section 2_1, which is the final stage sub-model section.
[0251] Therefore, in the speech synthesis processing device 100, the overall processing speed can be improved, and high-quality speech synthesis processing (predictive processing) can be performed.
[0252] Furthermore, in the speech synthesis processing device 100, the noise timetable can also be determined in a decentralized manner based on the sub-models used.
[0253] For example, such as Figure 16 As shown, it can also be based on Figure 10 The vertical axis is set as the log scale. The transformation noise level determines the noise time table (={β1, β2, ……, β N}).
[0254] Figure 16 Is Figure 10 A chart with the vertical axis set to the log scale. From Figure 10 As can be seen from the chart, the sub-model part (the sub-model part used in the prediction process) determined according to the noise schedule is dispersed.
[0255] For example, in the case of using a noise timetable based on pattern Ptn1, in Figure 10 In this case, the sub-model part used is:
[0256] (A1) Seventh Sub-model 2_7 (Number of processing times: 1) (corresponding to) Figure 10 Point P6)
[0257] (A2) Second Sub-model 2_2 (Number of processing times: 1) (corresponding to) Figure 10 Point P5)
[0258] (A3) First Sub-model Section 2_1 (Number of Processing Times: 4) (corresponding to) Figure 10 Points P4-P1),
[0259] But Figure 16 In this case, the sub-model part used is:
[0260] (B1) Tenth Sub-model 2_10 (Number of processing times: 1) (corresponding to) Figure 16 Point P6)
[0261] (B2) Eighth Sub-model 2_8 (Number of processing times: 1) (corresponding to) Figure 16 Point P5)
[0262] (B3) Sixth Sub-model 2_6 (Number of processing times: 1) (corresponding to) Figure 16 Point P4)
[0263] (B4) Fourth Sub-model 2_4 (Number of processing times: 1) (corresponding to) Figure 16 Point P3)
[0264] (B5) Second Sub-model 2_2 (Number of processing times: 1) (corresponding to) Figure 16 Point P2)
[0265] (B6) First Sub-model Section 2_1 (Number of Processing Times: 1) (corresponding to) Figure 16 Point P1)
[0266] There is no sub-model unit that performs multiple processes, and the sub-model units used are scattered. Furthermore, regarding the processes (B1) to (B6) described above, similarly to the above embodiment, the sub-model units used are selected by the control unit 1 (by a selector), and the processing is performed by each sub-model unit, thereby enabling speech synthesis processing (predictive processing) to be performed in the speech synthesis processing apparatus 100.
[0267] In the speech synthesis processing apparatus 100, the accuracy of speech synthesis processing is improved by dispersing the sub-models used. This is because if the processing accuracy of a sub-model that is processed many times is poor, the prediction accuracy of that sub-model will affect the overall processing accuracy. In the speech synthesis processing apparatus 100, by dispersing the sub-models used, the significant impact of the processing accuracy of specific sub-models can be prevented, resulting in an improvement in the overall processing accuracy of speech synthesis.
[0268] As described above, in the speech synthesis processing apparatus 100, multiple sub-model units are set according to the noise level, and learning processing can be performed independently (in parallel) on these multiple sub-model units. Therefore, the learning processing time can be greatly reduced.
[0269] Furthermore, in the speech synthesis processing apparatus 100, speech synthesis processing (predictive processing) is performed using a sub-model unit based on a noise timetable. This sub-model unit constructs a learned model that has been trained according to the noise level. Moreover, in the speech synthesis processing apparatus 100, since appropriate sub-model units can be employed (combined) according to the noise level, speech synthesis processing that maintains the speed of speech synthesis processing while obtaining high-quality speech signals can be achieved.
[0270] [Second Implementation]
[0271] Next, the second embodiment will be described. Furthermore, the same reference numerals will be used to mark the parts that are the same as in the above embodiment, and detailed descriptions will be omitted.
[0272] In the first embodiment, the case of performing speech synthesis processing (a signal processing device for generating speech signals (speech synthesis processing device)) was described, but in the second embodiment, the case of performing image generation processing (a signal generation device for generating image signals) was described.
[0273] Figure 17 This is a schematic configuration diagram of the signal generation and processing apparatus 200 according to the second embodiment.
[0274] Figure 18 This is a schematic configuration diagram of the kth sub-model section of the signal generation and processing apparatus 200 according to the second embodiment.
[0275] Figure 19 This is a schematic configuration diagram of the kth sub-model (image model) of the kth sub-model unit of the signal generation and processing apparatus 200 according to the second embodiment.
[0276] Figure 20 This is a schematic diagram of the residual block layer of the k-th sub-model (image model) of the signal generation and processing apparatus 200 according to the second embodiment.
[0277] <2.1: Structure of the Signal Generation and Processing Device>
[0278] Figure 17 It is the same as the first embodiment. Figure 1 The corresponding diagram focuses on the relationship with Figure 1 The differences are explained in terms of composition.
[0279] The signal generation and processing apparatus 200 of the second embodiment has a structure in which the control unit 1 in the speech synthesis processing apparatus 100 of the first embodiment is replaced by the control unit 1A, and the first sub-model unit 2_1 to the tenth sub-model unit 2_10 are replaced by the first sub-model unit 2A_1 to the tenth sub-model unit 2A_10 respectively.
[0280] Furthermore, in the speech synthesis processing apparatus 100 of the first embodiment, the signal y input to the speech synthesis processing apparatus 100 is... N It is Gaussian white noise w_noise, that is, a signal whose signal value at time t follows a Gaussian distribution (normal distribution). However, in the signal generation and processing apparatus 200 of the second embodiment, the signal y input to the signal generation and processing apparatus 200 is... N It is Gaussian noise w_noise that can form a two-dimensional image (e.g., an image of P pixels × Q pixels (P, Q: natural numbers)). In other words, if the pixel value of the coordinate (x, y) on the two-dimensional image is set to D(x, y), then the pixel value D(x, y) follows a Gaussian distribution (normal distribution) (a signal that can form an image).
[0281] Furthermore, in the speech synthesis processing apparatus 100 of the first embodiment, the condition for inputting the speech synthesis processing apparatus 100 is the acoustic feature quantity h, but in the signal generation processing apparatus 200 of the second embodiment, the condition for inputting the signal generation processing apparatus 200 is the data h for determining the label (e.g., one-hot vector, one-hot data).
[0282] Figure 18 Corresponding to the first embodiment Figure 2 In the second embodiment, the image is treated as the processing object, thus exhibiting characteristics different from the first embodiment. Figure 2 The structure is different from other parts. The input data generation unit 11 is a functional unit that operates in the learning mode (the mode in which learning processing is performed). The input image data y0 (forward resolution data), Gaussian white noise w_noise (Gaussian white noise w_noise that can form a two-dimensional image), and weighted noise level data α' (during learning: α' = α) are different parts. (w) When making a prediction: α' = α n (w) ), and data T regarding the time step. n(time step T) n (n: a natural number, 1≤n≤N) represents the noise level data α used during the execution. n (The processing time step is n: a natural number, 1≤n≤N). The input data generation unit 11 synthesizes the image data y0 with Gaussian white noise w_noise based on the weighted noise level data α', and uses the synthesized data as the image noise synthesis data y0. n The output of `_gen` is sent to the selector `SEL_k1`. Furthermore, assuming the image formed from image data `y0` has the same size as the image formed from Gaussian white noise `w_noise`, pixel values at the same coordinate in both the image data `y0` and the Gaussian white noise `w_noise` are added together, thus performing the synthesis process of image data `y0` and Gaussian white noise `w_noise`. Additionally, the weighted noise level data `α'` (during learning: `α' = α`) is used. (w) When making a prediction: α' = α n (w) And data T regarding the time step. n The sub-model control data Ctl(sub_Mk) is included in the output from control unit 1A to the k-th sub-model unit. Additionally, the weighted noise level data α from the prediction process included in the sub-model control data Ctl(sub_Mk) is also included. n (w) Expressed as "Ctl(sub_Mk).α n (w) Additionally, the weighted learning time data α contained in the sub-model control data Ctl(sub_Mk) is used with noise level data. (w) Expressed as "Ctl(sub_Mk).α (w) Additionally, the time step data T contained in the sub-model control data Ctl(sub_Mk) will be... n The expression is "Ctl(sub_Mk).T n ".
[0283] The input of the k-th sub-model SubMA_k is the signal y output from the selector SEL_k1. n Data T regarding time step n and the condition h (e.g., one-hot vector, one-hot data) that serves as the data for determining the label. Figure 17 , Figure 18 As shown, for example, condition h represents the data for "ball". Furthermore, when the k-th sub-model SubMA_k performs the learning process (in learning mode), it receives the loss evaluation data Eva_θ output from the loss evaluation unit 12 as input. The k-th sub-model SubMA_k is, for example, a model using a neural network, and during the learning process, it uses the input signal y... n(signals that can form an image) n Data T regarding time step n The learning process involves processing the signal y using Gaussian white noise (w_noise, which can form a two-dimensional image) output by condition h (data that determines the label). In other words, the k-th sub-model SubMA_k processes the signal y... n (The signal yn that can form an image), data T about the time step. n And taking the condition h as input, the output signal ε θ (The output signal ε that can form an image) θ The signal ε is output to the loss evaluation unit 12. Furthermore, the loss evaluation unit 12 acquires the output signal ε. θ The k-th sub-model SubMA_k updates its parameters according to the evaluation data Eva_θ of the loss from Gaussian white noise w_noise, and makes the output signal ε... θ The learning process is performed in such a way that the difference between the noise and the Gaussian white noise w_noise converges within a specified range.
[0284] The k-th sub-model, SubMA_k, is constructed using the parameters (optimal parameters) obtained through the aforementioned learning process as a learned model. This learned model is then used for prediction (during image signal generation). During prediction, the k-th sub-model, SubMA_k, calculates the prediction result using the signal y... n Data T regarding time step n The condition h is used as input for prediction processing to obtain the output signal ε. θ and output signal ε θ Output to the noise reduction waveform acquisition unit 13.
[0285] For example, the k-th submodel SubMA_k can be obtained through... Figure 19 The structure shown is implemented. Additionally... Figure 19 The residual block layer can be, for example, through Figure 20 The structure shown is implemented. Regarding the installation of this configuration, a procedure related to Non-Patent Document A described later is disclosed in the following URL; therefore, detailed descriptions are omitted.
[0286] (URL of the procedure associated with non-patent document A):
[0287] https: / / github.com / hojonathanho / diffusion
[0288] For Figure 5 The different configurations shown are explained below.
[0289] like Figure 19As shown, condition h and time step T output from control unit 1A n After undergoing embedding and activation processing respectively, the data is combined, and the combined data is Dset(={Dh(h), Dt(T)}. n The inputs are output to downsampling layers ka2-ka4 and upsampling layers ka5-ka7. Each downsampling layer downsamples each input based on the data Dset. Similarly, each upsampling layer upsamples each input based on the data Dset.
[0290] As described above, residual block layers ka_rn1 and ka_rn2 are obtained through... Figure 20 The structure shown is implemented as follows. In the residual block layer, the outputs from multiple network layers within the residual block layer are added to the inputs directed towards the residual block layer and then output. For example, as... Figure 20 As shown, the network consists of multiple layers: an activation unit ka_rn_1, a normalization unit ka_rn_2, a 2D convolutional layer ka_rn_3, an addition unit ka_rn_4, a normalization unit ka_rn_5, an activation unit ka_rn_6, a dropout unit ka_rn_7, and a 2D convolutional layer ka_rn_8. The data Dset is also input to the addition unit ka_rn_4. An attention unit ka_att is configured between the two residual block layers to perform attention processing on the input (processing via an attention mechanism) and obtain context data (e.g., a context vector). Furthermore, the data y is processed based on the obtained context data. n _r1 performs weighted processing (or, combines the obtained context data with data y) n (Processing of adding _r1).
[0291] The loss evaluation unit 12 has the same structure and function as in the first embodiment. Furthermore, the input data for the loss evaluation unit 12 includes data w_noise, which can form a two-dimensional image, and the signal ε. θ (Noise signal).
[0292] The noise reduction waveform acquisition unit 13 has the same structure and function as in the first embodiment. Furthermore, the input data for the loss evaluation unit 12 is a signal y that can form a two-dimensional image. n , signal ε θ (Noise signal), the output data can also form a two-dimensional image signal y. n-1 .
[0293] The output selector SELk_out and buffer 14 have the same structure and function as in the first embodiment.
[0294] <2.2: Operation of the signal generation and processing device>
[0295] The operation of the signal generation and processing device 200 configured as described above is largely the same as that of the first embodiment. The following description focuses on the differences.
[0296] exist Figure 17 In the signal generation and processing apparatus 200 shown, similar to the speech synthesis processing apparatus, for convenience, the case where the number of sub-models is "10" will be explained.
[0297] (2.2.1: Learning Processing)
[0298] Similar to the case of speech, a sub-model corresponding to each noise level is determined. The method for determining the sub-model of the object to be learned based on the transformed noise level sqrt(1-α') is the same as in the case of speech.
[0299] like Figure 21 As shown, in the signal generation and processing apparatus 200 that processes images, as previously explained, it is necessary to replace the speech waveform signal y0 with the image signal y0, replace the Gaussian white noise w_noise with Gaussian noise w_noise that can form a two-dimensional image, and replace the acoustic feature h with data h that determines the label of the image (e.g., a ball).
[0300] Additionally, the time step data T n The fact that it is input into the k-th sub-model (k=1) is also different from the case of speech.
[0301] Therefore, the loss function is defined by the following mathematical formula (t: time step instead of c in speech processing).
[0302]
Mathematical Expression 4
[0303]
[0304] The k-th sub-model is learned based on the value of the loss function, and the learned model is obtained in the same way as in the first implementation method.
[0305] In this way, in the signal generation and processing apparatus 200, learning processing can be performed independently for each of the N1 (10) sub-model units. That is, in each of the N1 (10) sub-model units, if the image signal y0, which is the forward solution data, its corresponding conditional data h, Gaussian white noise w_noise, and the noise level that determines the ratio of their synthesis are known, then learning processing can be performed. Therefore, the learning processing of the N1 (10) sub-model units can be performed in parallel. Thus, high-speed learning processing can be achieved in the signal generation and processing apparatus 200.
[0306] Furthermore, the above-mentioned method is used in the signal generation and processing apparatus 200. Figure 19The neural network model has been described, but it is not limited to this; other models can also be used. Figure 19 Models other than the neural network model shown.
[0307] In addition, in the signal generation and processing apparatus 200, all N1 (10) sub-model units can be implemented using the same model, or N1 (10) sub-model units can be implemented by mixing different models.
[0308] For example, a neural network model with fast processing speed but slightly lower quality of generated images can be used to implement the sub-model section of the initial stage of the signal generation and processing device 200, while a neural network model with slow processing speed but high quality of generated images can be used to implement the sub-model section of the final stage of the signal generation and processing device 200.
[0309] (2.2.2: Predictive Processing (Image Generation Processing))
[0310] Next, the prediction processing (image generation processing) performed by the signal generation and processing apparatus 200 will be described.
[0311] Furthermore, for ease of explanation, the noise timetable (={β1, β2, ..., β') is determined at equal intervals with the transformed noise level sqrt(1-α'). N The case of N=1000 will be explained. In addition, the case in which the 1000-step conversion noise level (the noise level of 1000 steps with equal levels) is divided into 10 equal intervals during learning, and each sub-model unit (first sub-model unit 2A_1 to tenth sub-model unit 2A_10) is learned by using the conversion noise level of each 100-step division to obtain the learned model will be explained.
[0312] Control unit 1 determines a noise schedule (={β1, β2, ..., β) based on an equal interval between noise level changes. N The sub-model unit to be used is determined by (N=1000). Specifically, when performing prediction processing in 1000 steps, the sub-model unit to be used is determined by performing 100 steps of processing for each sub-model unit (first sub-model unit 2A_1 to tenth sub-model unit 2A_10).
[0313] (1) In steps 1000 to 901, the processing is performed by the tenth sub-model section 2A_10.
[0314] (2) In steps 900 to 801, processing is performed through the ninth sub-model section 2A_9.
[0315] (3) In steps 800 to 701, the processing is performed through the eighth sub-model section 2A_8.
[0316] (4) In steps 700 to 601, the processing is performed by the seventh sub-model section 2A_7.
[0317] (5) In steps 600 to 501, processing is performed through the sixth sub-model section 2A_6.
[0318] (6) In steps 500 to 401, processing is performed through the fifth sub-model section 2A_5.
[0319] (7) In steps 400 to 301, processing is performed through the fourth sub-model section 2A_4.
[0320] (8) In steps 300 to 201, processing is performed through the third sub-model section 2A_3.
[0321] (9) In steps 200 to 101, processing is performed through the second sub-model section 2A_2.
[0322] (10) In steps 100 to 1, processing is performed through the first sub-model section 2A_1.
[0323] Figure 22 This diagram illustrates the selector and the k-th sub-model unit extracted from the signal generation and processing device 200, and explicitly shows the sub-model unit used according to the noise time schedule.
[0324] As described above, if the sub-model section to be used is determined, such as Figure 22 As shown, the control unit 1A controls each selector to switch between selectors SEL10 to SEL1 to select the path leading to the sub-model section, and to switch between selectors SEL0 and the path on the side of the first sub-model section 2A_1. Therefore, Gaussian white noise w_noise(=y) is input to selector SEL10. N (N=1000) is input into the tenth sub-model part 2A_10.
[0325] In addition, the control unit 1A outputs sub-model control data Ctl(sub_M10) to the tenth sub-model unit 2A_10.
[0326] Figure 23 This is a schematic diagram of the k-th sub-model part during prediction processing.
[0327] In the tenth sub-model section 2A_10, Figure 23 The signal y shown n _ext becomes signal y N (=w_noise). Furthermore, by selecting terminal "0" of input selector SELk_in via control unit 1A, terminal "0" of selector SELk_1 is selected, thereby causing signal y... N(=w_noise) is input into the k-th sub-model (k=10). Additionally, the data T with the time step is... n The conditional data h is input into the k-th sub-model (k=10).
[0328] In the k-th sub-model SubMA_k (k=10), through Figure 19 The functional unit shown performs processing (processing via the learned model) to acquire the signal ε. θ Furthermore, the signal ε obtained through the k-th sub-model SubMA_k (k=10) θ It is output to the noise reduction waveform acquisition unit 13.
[0329] The noise reduction waveform acquisition unit 13 inputs the signal y output from the selector SEL_k1. n (y N (=w_noise)), the signal ε output from the k-th sub-model SubMA_k (k=10) θ Noise level data α output from control unit 1A n (n=1000), and weighted noise level data α n (w) (n = 1000). The noise reduction waveform acquisition unit 13 acquires the noise level data α. n and weighted noise level data α n (w) and using signal y n and signal ε θ Noise reduction processing is performed. Specifically, the noise reduction waveform acquisition unit 13 acquires the noise-reduced signal y by performing processing equivalent to the following mathematical formula. n-1 (n = 1000).
[0330]
Mathematical Expression 5
[0331]
[0332]
[0333] z ~ N(0, I) (when n>0) (I is the identity matrix)
[0334] z = 0 (when n = 0)
[0335] Furthermore, the noise reduction waveform acquisition unit 13 will use the signal y obtained above... n-1 Output to the output selector SELk_out.
[0336] Control unit 1A selects terminal "1" of output selector SELk_out and sends signal y n-1 Output to buffer 14.
[0337] Next, in the tenth sub-model section 2A_10, the input signal is set as signal y. 999 When n = 999, the same processing as that performed in the tenth sub-model section 2A_10 when n = 1000 is executed. Furthermore, the input signal y... 999 The signal y stored in buffer 14 when n=1000 1000 , signal y 1000 The output is from buffer 14 to input selector SELk_in. Furthermore, control unit 1A performs switching control by selecting terminal "1" of input selector SELk_in.
[0338] Furthermore, in the tenth sub-model section 2A_10, the signal y is obtained. n-1 (=y 998 The signal y n-1 (=y 998 The output is sent to buffer 14 via selector SEL1 (selection terminal "1").
[0339] From n=998 to n=902, perform the same processing as described above.
[0340] Furthermore, in the tenth sub-model section 2A_10, the input signal is set as signal y. 901 n = 901 and perform the same processing as described above. Additionally, the input signal y... 901 The signal y stored in buffer 14 when n=902. 902 , signal y 902 The output is from buffer 14 to input selector SELk_in. Furthermore, control unit 1A performs switching control by selecting terminal "1" of input selector SELk_in.
[0341] Furthermore, in the tenth sub-model section 2A_10, the signal y is obtained. n-1 (=y 900 The signal y n-1 (=y 900 The output is sent to selector SEL9 via selector SELk_out (selector terminal "0").
[0342] Furthermore, the control unit 1A switches the control signal y by selecting the path toward the ninth sub-model unit 2A_9 in the selector SEL9. 900 It was input into the ninth sub-model section 2A_9.
[0343] In the ninth sub-model section 2A_9, the same processing as in the tenth sub-model section 2A_10 is performed for n=900 to n=801.
[0344] Furthermore, in the eighth sub-model section 2A_8 to the first sub-model section 2A_1, the same processing as that in the tenth sub-model section 2A_10 is also performed.
[0345] Furthermore, the control unit 1A performs switching control by selecting the output from the first sub-model unit 2A_1 in the selector SEL0, and acquires the (output) signal y0.
[0346] By performing this processing, the image signal y0, which corresponds to the conditional data h, can be acquired in the signal generation and processing device 200.
[0347] As described above, the signal generation and processing apparatus 200 can perform a process (predictive processing) to generate an image signal by selecting and processing a sub-model unit determined according to a noise time schedule.
[0348] Furthermore, the above description describes the use of equally spaced levels for the transition noise level, but it is not limited to this. Alternatively, the noise timetable can be determined based on the value obtained by taking the logarithm of the noise level (transition noise level), and the learning processing and prediction processing in the signal generation and processing device 200 can be performed based on the noise timetable.
[0349] As described above, in the signal generation and processing apparatus 200, multiple sub-model units are set according to the noise level, and learning processing can be performed independently (in parallel) on these multiple sub-model units. Therefore, the learning processing time can be greatly reduced.
[0350] Furthermore, in the signal generation processing apparatus 200, image signal generation processing (predictive processing) is performed using a sub-model unit based on a noise timetable. This sub-model unit constructs a learned model that has been trained according to the noise level. Moreover, in the signal generation processing apparatus 200, since appropriate sub-model units can be used (combined) according to the noise level, the speed of signal generation processing (image signal generation processing) can be maintained while generating high-quality image signals.
[0351] Furthermore, while the above description addresses the case where condition data h is input, it is not a limitation. In the signal generation and processing apparatus 200, processing can also be performed without inputting condition data h. In this case, a comparative experiment was conducted with the case using the technology described in Non-Patent Document A.
[0352] (Non-patent document A):
[0353] J.Ho, A.Jain, and P.Abbeel, “Denoising diffusion probabilistic models,” in Proc.NeurIPS, Dec.2020.
[0354] Specifically, an image generation neural network model was learned using 50,000 CIFAR10 training images. In the model of Non-Patent Document A, learning was performed using 1,000 steps of noise level. In the present invention (equivalent to a signal generation and processing apparatus 200 without input conditional data h), 10 sub-model units were trained, each using 800,000 steps of noise level, with each step divided into 10 equal parts. During image generation (prediction processing), random images were generated each time through unconditional generation using random noise (Gaussian white noise) as input. To verify the accuracy of these generated images, the FID (Fenchel Inception Distance) between the 50,000 generated images and the 50,000 training images was calculated. The results are as follows:
[0355] FID: 5.71 in the method of non-patent literature A
[0356] FID in this invention: 5.50
[0357] It can be confirmed that the image generation accuracy of the present invention (equivalent to a signal generation and processing device 200 without input condition data h) is higher.
[0358] [Other implementation methods]
[0359] In the speech synthesis processing apparatus of the above embodiments, the use of the DiffWave model and the WaveGrad model in the sub-model unit has been described, but it is not limited to this. In the speech synthesis processing apparatus, other models that can obtain the speech waveform corresponding to the acoustic feature based on Gaussian white noise and acoustic feature quantity can also be used.
[0360] Furthermore, in the speech synthesis processing apparatus and signal generation processing apparatus described in the above embodiments, each block can be individually chipped using semiconductor devices such as LSIs, or can be chipped in a manner that includes some or all of them.
[0361] In addition, it is referred to as LSI here, but depending on the level of integration, it is sometimes also called IC, system LSI, super LSI, or premium LSI.
[0362] Furthermore, the method of integration is not limited to LSI; it can also be achieved using dedicated circuits or general-purpose processors. It can also utilize FPGAs (Field Programmable Gate Arrays) that can be programmed after LSI fabrication, or reconfigurable processors that can reconfigure the connections or settings of the circuit cells within the LSI.
[0363] Furthermore, some or all of the processing of each functional block in the above embodiments can also be implemented by a program. Moreover, some or all of the processing of each functional block in the above embodiments is performed by the central processing unit (CPU) in a computer. Additionally, the programs used to perform each process are stored in storage devices such as hard disks or ROMs, and are executed in ROMs or read into RAMs for execution.
[0364] Furthermore, the processes described in the above embodiments can be implemented either through hardware or through software (including implementations in conjunction with an OS (operating system), middleware, or specified libraries). Moreover, a hybrid approach combining software and hardware can also be used.
[0365] For example, when implementing the functional units of the above-described embodiments through software, it is also possible to use... Figure 24 The hardware configuration shown (e.g., a hardware configuration that connects the CPU, GPU, ROM, RAM, input unit, output unit, communication unit, storage unit (e.g., storage unit implemented by HDD, SSD, etc.), and external media via a driver) is implemented through software processing.
[0366] Furthermore, when each functional unit of the above-described embodiment is implemented through software, the software can use software with… Figure 24 The hardware shown can be implemented using a single computer, or it can be implemented using multiple computers through distributed processing.
[0367] Furthermore, the execution order of the processing methods in the above embodiments is not necessarily limited to the description in the above embodiments, and the execution order can be changed without departing from the spirit of the invention.
[0368] The computer program that enables a computer to perform the aforementioned method, and the computer-readable recording medium on which the program is recorded, are included within the scope of this invention. Examples of computer-readable recording media include floppy disks, hard disks, CD-ROMs, MOs, DVDs, DVD-ROMs, DVD-RAMs, high-capacity DVDs, next-generation DVDs, and semiconductor memories.
[0369] The aforementioned computer programs are not limited to being recorded on the aforementioned recording media; they can also be transmitted via telecommunication lines, wireless or wired communication lines, networks such as the Internet.
[0370] Furthermore, the specific configuration of the present invention is not limited to the aforementioned embodiments, and various changes and modifications can be made without departing from the spirit of the invention.
[0371] [appendix]
[0372] Alternatively, the present invention can also be implemented as follows.
[0373] <Appendix 1>
[0374] A speech synthesis processing device outputs a speech signal corresponding to the acoustic features based on Gaussian white noise and acoustic features, wherein...
[0375] The speech synthesis processing device includes a first sub-model unit to an Nth sub-model unit, which are N (N: a natural number, N≥2) sub-model units.
[0376] The first to Nth sub-model parts each have a learning model. The learning model takes noise level-related data, acoustic features, and a corresponding speech signal as input, and performs learning processing by outputting Gaussian white noise from a synthesized noise signal. The synthesized noise signal is a signal synthesized from the speech signal and Gaussian white noise based on the noise level-related data.
[0377] The first sub-model section to the Nth sub-model section respectively use noise levels included in different noise level ranges to perform learning processing on the learning models included in the first sub-model section to the Nth sub-model section, thereby obtaining the learned models.
[0378] <Appendix 2>
[0379] According to the speech synthesis processing apparatus described in Appendix 1, wherein...
[0380] It also has a control unit for setting noise schedules.
[0381] The control unit selects the sub-model units from the first to the Nth sub-model units for speech synthesis processing based on the noise level, and determines the processing order of the selected sub-model units. The noise level is determined according to the noise timetable.
[0382] The selected sub-model unit performs prediction processing using the learned model in the order determined by the control unit, thereby obtaining a speech signal corresponding to the acoustic feature quantity.
[0383] <Appendix 3>
[0384] According to the speech synthesis processing apparatus described in Appendix 2, wherein...
[0385] The first to Nth sub-model sections are arranged in descending order from the Nth sub-model section to the first sub-model section. As the signal moves from the Nth sub-model section to the first sub-model section, the proportion of noise components in the input noise synthesis signal decreases.
[0386] The sub-model section located on the front-end side has a structure that allows for faster processing than the sub-model section located on the rear-end side.
[0387] <Appendix 4>
[0388] According to the speech synthesis processing apparatus described in Appendix 2, wherein...
[0389] The first to Nth sub-model sections are arranged in descending order from the Nth sub-model section to the first sub-model section. As the signal moves from the Nth sub-model section to the first sub-model section, the proportion of noise components in the input noise synthesis signal decreases.
[0390] The sub-model section located on the rear stage has a structure with higher processing accuracy than the sub-model section located on the front stage.
[0391] <Appendix 5>
[0392] The speech synthesis processing apparatus according to any one of Appendices 2 to 4, wherein...
[0393] When the control unit selects a sub-model unit from the first to the Nth sub-model unit for speech synthesis processing based on the noise level, it sets the noise timetable in a distributed manner for the sub-model units used, and the noise level is determined based on the noise timetable.
[0394] <Appendix 6>
[0395] According to the speech synthesis processing apparatus described in Appendix 1, wherein...
[0396] The noise level range corresponding to the first sub-model unit to the Nth sub-model unit is determined based on the value obtained by taking the logarithm of the noise level. The first sub-model unit to the Nth sub-model unit respectively use the noise level contained in the noise level range corresponding to itself to perform the learning process.
[0397] Explanation of reference numerals in the attached figures
[0398] 100 speech synthesis processing devices
[0399] 200 Signal Generation and Processing Device
[0400] 1.1A Control Unit
[0401] 2_1~2_10 First Sub-model Section~Tenth Sub-model Section
[0402] 2_1A~2_10A First Sub-model Section~Tenth Sub-model Section
[0403] SubM_k, the k-th sub-model
[0404] SubMA_k is the k-th sub-model.
Claims
1. A signal generation and processing apparatus for outputting a speech signal or an image signal from Gaussian white noise, wherein, The signal generation and processing apparatus includes a first sub-model unit to an Nth sub-model unit, which are N sub-model units, where N is a natural number and N≥2. Each of the N sub-model units includes a diffusion model. The diffusion model takes as input a signal after adding noise to a speech signal or image signal in a weighted manner, and only infers the added noise component. The first to Nth sub-model units each have a learning model. The learning model takes as input data related to the noise time schedule and a supervision signal (either a speech signal or an image signal), and outputs Gaussian white noise from a synthesized noise signal. The synthesized noise signal is a signal synthesized from the supervision signal and Gaussian white noise based on the data related to the noise time schedule. The first sub-model section to the Nth sub-model section respectively use noise levels included in different noise level ranges to perform learning processing on the learning models included in the first sub-model section to the Nth sub-model section, thereby obtaining the learned models.
2. A signal generation and processing apparatus, comprising outputting a speech signal or image signal corresponding to the input conditional features based on Gaussian white noise and input conditional features, wherein, The signal generation and processing apparatus includes a first sub-model unit to an Nth sub-model unit, which are N sub-model units, where N is a natural number and N≥2. Each of the N sub-model units includes a diffusion model. The diffusion model takes as input a signal after adding noise to a speech signal or image signal in a weighted manner, and only infers the added noise component. The first to Nth sub-model units each have a learning model. The learning model takes as input data related to the noise time schedule, input conditional features, and a supervision signal of speech or image signals corresponding to the input conditional features. It performs learning processing by outputting Gaussian white noise from a synthesized noise signal. The synthesized noise signal is a signal synthesized from the supervision signal and Gaussian white noise based on the data related to the noise time schedule. The first sub-model section to the Nth sub-model section respectively use noise levels included in different noise level ranges to perform learning processing on the learning models included in the first sub-model section to the Nth sub-model section, thereby obtaining the learned models.
3. The signal generation and processing apparatus according to claim 2, wherein, The signal generation and processing device also includes a control unit for setting the noise time schedule. The control unit selects the sub-model units from the first to the Nth sub-model units for signal generation processing based on the noise level, and determines the processing order of the selected sub-model units. The noise level is determined according to the noise time schedule. The selected sub-model unit executes the prediction processing using the learned model in the order determined by the control unit, thereby obtaining the speech signal or image signal corresponding to the input conditional feature quantity.
4. The signal generation and processing apparatus according to claim 3, wherein, The first to the Nth sub-model units have an order in which the noise components of the input noise synthesis signal are proportioned. The order is the order in which the proportion of noise components in the synthesized noise signal decreases, and the sub-models located at the front of this order have a structure with a faster processing speed than the sub-models located at the back.
5. The signal generation and processing apparatus according to claim 3, wherein, The first to the Nth sub-model units have an order in which the noise components of the input noise synthesis signal are proportioned. The order is the order in which the proportion of noise components in the synthesized noise signal decreases, and the sub-models located later in this order have a structure with higher processing accuracy than the sub-models located in front.
6. The signal generation and processing apparatus according to any one of claims 3 to 5, wherein, When the control unit selects a sub-model unit from the first to the Nth sub-model unit for signal generation processing based on the noise level, it sets the noise timetable in a distributed manner for the sub-model units used, and the noise level is determined based on the noise timetable.
7. The signal generation and processing apparatus according to claim 1 or 2, wherein, The noise level range corresponding to the first sub-model unit to the Nth sub-model unit is determined based on the value obtained by taking the logarithm of the noise level. The first sub-model unit to the Nth sub-model unit respectively use the noise level contained in the noise level range corresponding to itself to perform the learning process.