Signal estimation method, signal estimation device, and program
The signal estimation method addresses the challenge of maintaining accuracy and reducing complexity by dividing the estimation process into stages, effectively using resampled signals to enhance target signal estimation without increasing calculation costs.
Patent Information
- Application Number
- PCT/JP2024/026799
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2026-01-29
AI Technical Summary
Existing target signal estimation methods struggle with maintaining high accuracy while minimizing the increase in calculation complexity and cost when multiple estimators are used, particularly in noisy and reverberant environments.
A signal estimation method that divides the estimation process into two stages: the first stage estimates the target signal using multiple methods, and the second stage supplements information using resampled acoustic and estimated signals, without altering the target signal extraction unit, thus minimizing additional calculation costs.
Achieves highly accurate target signal estimation by appropriately supplementing information from acoustic and estimated signals, while keeping the overall calculation cost minimal.
Smart Images

Figure JP2024026799_29012026_PF_FP_ABST
Abstract
Description
Signal estimation method, signal estimation device, and program
[0001] The present invention relates to a target signal estimation method for estimating a target signal that represents the characteristics of a target sound from an acoustic signal obtained by recording the target sound with a microphone in an environment where background noise and reverberation are present.
[0002] Non-Patent Documents 1 and 2 are known as prior art target signal estimation methods.
[0003] FIG. 1 shows the configuration of Non-Patent Document 1.
[0004] In Non-Patent Document 1, an acoustic signal x recorded by a microphone is received, and a target signal is estimated by sequentially applying a plurality of target signal estimation units 81 and 82 to the signal, and an estimated value y is output. This configuration provides a method for performing estimation with higher accuracy.
[0005] FIG. 2 shows the configuration of Non-Patent Document 2.
[0006] In Non-Patent Document 2, a recorded acoustic signal x is received, and first an estimation unit 91 extracts an estimated signal y1, which is an estimated value of the target sound. Next, a target signal estimation unit 92 receives the acoustic signal x and the estimated signal y1, estimates the target signal using these signals, and outputs an estimated value y. In Non-Patent Document 2, the target signal estimation unit 92 receives the estimated signal y1 in addition to the acoustic signal x, thereby providing a method for performing estimation with higher accuracy than when only the acoustic signal x is received.
[0007] Jean-Marie Lemercier, Julius Richter, Simon Welker, and Timo Gerkmann, "StoRM: A diffusion-based stochastic regeneration model for speech enhancement and dereverberation", IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 2724-2737, 2023.Zhong-Qiu Wang and DeLiang Wang, "Multi-microphone complex spectral mapping for speech dereverberation", Proc. IEEE ICASSP, 2020.
[0008] In Non-Patent Document 1, if a loss of the target signal occurs in the signal estimated by each of the target signal estimation units 81 and 82, it is difficult to recover the loss in the subsequent processing.
[0009] In Non-Patent Document 2, by using an acoustic signal x in addition to an estimated signal y1, even if a loss occurs in the preceding estimated signal y1, the acoustic signal x can be used to compensate for the loss. However, when two or more estimators are used, the complexity of the problem to be solved by the target signal estimator 92 increases as the number of estimators increases. Therefore, to estimate the target signal with high accuracy, more complex calculations are required, such as increasing the number of parameters in the target signal estimator 92. As a result, increasing the number of estimators requires more complex processing in the target signal estimator 92, which significantly increases the overall calculation cost. For this reason, Non-Patent Document 2 does not propose a method using two or more estimators.
[0010] An object of the present invention is to provide a signal estimation method and a signal estimation device that can estimate a target signal with high accuracy while minimizing an increase in the amount of calculation and supplementing information on the acoustic signal and the total estimated signal as appropriate.
[0011] In order to solve the above problem, according to one aspect of the present invention, a signal estimation method estimates a target signal representing characteristics of a target sound from an acoustic signal obtained by recording with a microphone. The signal estimation method includes a first estimation step of estimating the target signal from the acoustic signal by at least one method, and a second estimation step of estimating the target signal using a feature of the acoustic signal, wherein the second estimation step performs a plurality of processes, and the plurality of processes use, as auxiliary inputs, a resampled acoustic signal and the estimated signal obtained by performing a resample process on the acoustic signal and at least one estimated signal obtained by estimation in the first estimation step.
[0012] According to the present invention, highly accurate estimation can be achieved by appropriately supplementing information on the acoustic signal and the total estimated signal while minimizing an increase in the amount of calculation.
[0013] FIG. 1 is a diagram showing the configuration of Non-Patent Document 1. FIG. 2 is a diagram showing the configuration of Non-Patent Document 2. FIG. 3 is a functional block diagram of a signal estimation device according to a first embodiment. FIG. 4 is a diagram showing an example of a processing flow of a signal estimation device according to a first embodiment, Example 1, and Example 2. FIG. 5 is a functional block diagram of a target signal extraction unit 140. FIG. 6 is a functional block diagram of an auxiliary input extraction unit 120. a∈R W×H×G FIG. 1 is a functional block diagram of a signal estimation device 200 according to a first embodiment. FIG. 2 is a functional block diagram of a target signal extraction unit 240. FIG. 3 is a functional block diagram of an auxiliary input extraction unit 220. FIG. 4 is a functional block diagram of a signal estimation device 300 according to a second embodiment. FIG. 5 is a functional block diagram of a target signal extraction unit 340. FIG. 6 is a functional block diagram of an auxiliary input extraction unit 320. FIG. 7 is a diagram showing experimental results regarding a word error rate in speech recognition. FIG. 8 is a diagram showing experimental results regarding a real time factor. FIG. 9 is a diagram showing an example of the configuration of a computer to which the present method is applied.
[0014] An embodiment of the present invention will be described below. In the drawings used in the following description, components having the same function and steps performing the same processing are denoted by the same reference numerals, and duplicate explanations will be omitted. In the following description, symbols such as "^" used in the text should normally be written directly above the character immediately following them, but due to limitations in text notation, they are written immediately before the character in question. In formulas, these symbols are written in their original positions. Furthermore, unless otherwise specified, processing performed on each element of a vector or matrix is assumed to apply to all elements of that vector or matrix.
[0015] <Key Points of the First Embodiment> In this embodiment, when multiple methods for estimating a target signal are given, a method is provided that can perform estimation with higher accuracy by combining information on the estimated signal obtained by each estimation method.
[0016] Even if the number of estimation signals is increased, the increase in calculation cost can be minimized for the following reasons.
[0017] (1) There is no need to make any changes to the target signal extraction unit, and only other processing units need to be changed.
[0018] (2) The calculation cost for parts other than the target signal extraction part is overwhelmingly smaller than that for the target signal extraction part.
[0019] For the following reasons, the target signal can be estimated with high accuracy using information on the acoustic signal and all estimated signals.
[0020] Even if the information of each signal is partially lost during the multi-stage processing of the target signal extraction section, it can be compensated for as needed using the auxiliary input.
[0021] First Embodiment FIG. 3 is a functional block diagram of a signal estimation device 100 according to a first embodiment, and FIG. 4 shows the processing flow thereof.
[0022] The signal estimation device 100 includes M estimation units 101-m, a feature extraction unit 110, an auxiliary input extraction unit 120, N' (N'≦N) adaptation units 130-n, a target signal extraction unit 140, and a target signal generation unit 150, where M is an integer equal to or greater than 1, m=1, 2, ..., M.
[0023] The signal estimation device 100 receives an acoustic signal x recorded by a microphone, estimates a target signal representing the characteristics of a target sound, and outputs an estimated value y.
[0024] The M estimating units 101-m are also referred to as first estimating units 101. The first estimating unit 101 estimates a target signal from an acoustic signal using at least one method (S101) and outputs at least one estimated signal.
[0025] The configuration including the feature extraction unit 110, the auxiliary input extraction unit 120, the N' adaptation units 130-n, the target signal extraction unit 140, and the target signal generation unit 150 is also referred to as the second estimation unit 190. The second estimation unit 190 estimates the target signal using the feature of the acoustic signal (S190) and outputs the estimated value as the output value of the signal estimation device 100. The second estimation unit 190 executes a plurality of processes, and in the plurality of processes, the resampled acoustic signal and the estimated signal obtained by performing resampling on the acoustic signal and at least one estimated signal are used as auxiliary inputs.
[0026] The signal estimation device 100 is a special device configured by loading a special program into a publicly known or dedicated computer having, for example, a central processing unit (CPU), a main memory (RAM), etc. The signal estimation device 100 executes each process under the control of, for example, the central processing unit. Data input to the signal estimation device 100 and data obtained in each process are stored, for example, in the main memory, and the data stored in the main memory is read by the central processing unit as needed and used for other processes. At least a portion of each processing unit of the signal estimation device 100 may be configured with hardware such as an integrated circuit. Each storage unit included in the signal estimation device 100 may be configured with, for example, a main storage unit such as a RAM (Random Access Memory), or middleware such as a relational database or a key-value store. However, each storage unit does not necessarily have to be included within the signal estimation device 100 itself. It may also be configured with an auxiliary storage unit implemented by a semiconductor memory element such as a hard disk, optical disk, or flash memory, and provided externally to the signal estimation device 100.
[0027] Each part will be explained below.
[0028] <Estimation Unit 101-m> Each of the M estimation units 101-m receives an acoustic signal x, estimates a target signal from the acoustic signal x using a different algorithm (S101), and outputs an estimated signal ^x, which is an estimated value of the target signal. m Output.
[0029] <Feature Extraction Unit 110> The feature extraction unit 110 extracts the acoustic signal x and M estimated signals ^x m The target signal extraction unit 140 extracts and outputs a feature value h0 that matches the input size and number of channels (S110).
[0030] <Target Signal Extraction Unit 140> FIG. 5 shows a functional block diagram of the target signal extraction unit 140. As shown in FIG.
[0031] The target signal extraction unit 140 has N stages of sub-processors 140-n.
[0032] Let n=1, 2, ..., N. The sub-processing units 140-n of the N stages sequentially receive the auxiliary input a from the corresponding adaptation unit 130-n from the preceding stage to the succeeding stage. n and receives the feature quantity h from the feature quantity extraction unit 110 or sub-processing unit 140-(n-1) located in the previous stage of each sub-processing unit 140-n. n-1 Receive the feature h n and outputs the converted audio signal x n , estimated signal ^x n m and auxiliary input a n In this case, n is assumed to be equal to the number n of the sub-processing unit to which each adaptation unit corresponds. With this configuration, the target signal extraction unit 140 extracts the feature quantities h1, h2, ..., h N (S140) and output. Note that the number of parameters used in the extraction process is constant regardless of the number M of types of estimation methods or the number of target signals.
[0033] <Auxiliary Input Extraction Unit 120> FIG. 6 shows a functional block diagram of the auxiliary input extraction unit 120. As shown in FIG.
[0034] The auxiliary input extraction unit 120 includes M+1 resample processing units 120-m (m=0, 1, 2, ..., M). The auxiliary input extraction unit 120 extracts an acoustic signal x and M estimated signals ^x m The resampling processing unit 120-m performs the following resampling processing for each signal separately (auxiliary input extraction processing S120), and the resampled acoustic signal x n(1) ...x n(N') , or the estimated signal ^x n(1) m ...^x n(N') m Here, n(n') is assumed to match the number n of the sub-processing unit corresponding to the n'th matching unit, and the resampling unit 120-0 receives the audio signal x, performs resampling processing, and outputs the resampled audio signal x n(1) ...x n(N') Output.
[0035] The resampling unit 120-m calculates the feature quantity h n-1 The acoustic signal x and the estimated signal ^x are the same size as m Resample the .
[0036] The auxiliary input extraction unit 120 extracts an acoustic signal x n and the estimated signal ^x n 1...^x n M Output.
[0037] <N' number of adaptation units 130-n> Each adaptation unit 130-n adapts the feature quantity h received by the corresponding sub-processing unit 140-n. n-1 The audio signal x is resampled to have the same size as n and the estimated signal ^x n m is received from all the resampling units 120-m, and the acoustic signal x n and all estimated signals ^x n 1...^x n M are collected and converted to fit the number of input channels of the sub-processing unit 140-n (adaptation process S130), and the converted value is used as the auxiliary input a n to the corresponding sub-processing unit 140-n.
[0038] <Target signal generator 150> The target signal generator 150 generates the feature quantities h1, . . . , h2 output from the Nth stage sub-processors 140-n. N The target signal estimate y is generated and output (S150). The number of target signals is determined in accordance with the training data prepared during learning.
[0039] <Effects> With this configuration, the process of estimating a target signal from an acoustic signal recorded by a microphone and M estimated signals is divided into (i) a process of extracting the features of the target signal and (ii) a process of appropriately supplementing the auxiliary input extracted individually from the acoustic signal and each estimated signal. When the number of estimated signals increases, the former process, which accounts for the majority of the calculation cost, remains unchanged, and only the latter process is expanded. This enables highly accurate estimation while minimizing the increase in the amount of calculation and appropriately supplementing the information of the acoustic signal and all estimated signals.
[0040] Next, Example 1, which specifically illustrates the first embodiment, will be described. First, the definitions of symbols used in Example 1 will be described.
[0041] <Symbol definition> R W×H×G represents the set of all 3D real tensors whose elements are real numbers with width W, height H, and number of channels G, and C W×H×G represents the set of all three-dimensional complex tensors whose elements are complex numbers with width W, height H, and number of channels G. For example, a∈C W×H×G indicates that a is an element of it. W×H×G W×H is also called the size of the tensor.
[0042] x∈C T×F×1 is the acoustic signal recorded by a microphone, and the set of complex signal tensors C obtained by frequency division (e.g., short-time Fourier transform) of the acoustic signal is T×F×1 where T denotes the number of time frames and F denotes the number of frequency bins.
[0043] ^x m ∈C T×F×1 is an estimated signal estimated by the m-th estimation unit 101-m, and is a universal set C of tensors of complex signals obtained by frequency division (such as short-time Fourier transform) of the estimated signal. T×F×1 Indicates that the element is
[0044] y∈C T×F×1 is the target signal, and the set of complex signal tensors C obtained by frequency division of the target signal T×F×1 Indicates that the element is
[0045] h∈C W×H×G is the feature extracted by the signal estimation method. Here, W is the width of the feature, H is the height of the feature, and G is the number of channels of the feature. W × H is also called the size of the feature.
[0046] h out =Φ(h in ;θ) is the feature value h in as input and the feature value h out where θ is a trainable parameter that determines the behavior of the processing module. It is assumed that various types of processing modules can be used, such as convolutional neural networks (CNNs) and multi-layered CNNs (e.g., residual blocks (see Reference 1)).
[0047] (Reference 1) Julius Richter; Simon Welker; Jean-Marie Lemercier; Bunlong Lay; Timo Gerkmann, "Speech Enhancement and Dereverberation With Diffusion-Based Generative Models", IEEE Trans. Audio Speech, and Language Processing, vol. 31, pp. 2351-2364, 2023. <Example 1> The following description will focus on differences from the first embodiment.
[0048] FIG. 8 is a functional block diagram of a signal estimation device 200 according to the first embodiment, and FIG. 4 shows the processing flow thereof.
[0049] The signal estimation device 200 includes M estimation units 201 - m , a feature extraction unit 210 , an auxiliary input extraction unit 220 , N1 adaptation units 230 - n , a target signal extraction unit 240 , and a target signal generation unit 250 .
[0050] The signal estimation device 100 estimates an acoustic signal x∈C T×F×1 and estimates the target signal that represents the characteristics of the target sound, and calculates the estimated value y∈R T×F×1 Output.
[0051] Each part will be explained below.
[0052] <Estimation Unit 201-m> Each of the M estimation units 201-m receives an acoustic signal x, estimates a target signal from the acoustic signal x using a different algorithm (S201), and outputs an estimated signal ^x, which is an estimated value of the target signal. m Output.
[0053] <Feature Extraction Unit 210> The feature extraction unit 210 extracts an acoustic signal x∈C T×F×1 and M estimated signals ^x m ∈C T×F×1 The real and imaginary parts of each signal are stored in different channels. Here, the tensor created by connecting all signals in the channel direction is χ∈R. T×F×2(M+1) In this embodiment, the feature extraction unit 210 executes a processing module h0=Φ(χ;θ in ) and use χ as the feature h0∈R T×F×G_0 The feature value h0 is extracted from the acoustic signal x by converting it into
[0054] <Target Signal Extraction Unit 240> FIG. 9 shows a functional block diagram of the target signal extraction unit 240. As shown in FIG.
[0055] The target signal extraction unit 240 is composed of multi-layer sub-processing units 240-n (1≦n≦N). In this embodiment, the target signal extraction unit 240 is divided into three units, an encoder 240-E, a bottleneck 240-B, and a decoder 240-D, each consisting of N1, N2-N1, and N-N2 sub-processing units, respectively, in accordance with the structure of U-Net (see Reference 2).
[0056] (Reference 2) Olaf Ronneberger, Philipp Fischer, and Thomas Brox, "U-net: Convolutional networks for biomedical image segmentation", in Proc. International conference on Conference Medical Image Computing and Computer-Assisted Intervention- (MIC-CAI), 2015, pp. 234-241 (Encoder 240-E) Each sub-processing unit 240-n included in the encoder 240-E performs the following processing.
[0057] The sub-processing unit 240-n receives as input the feature quantity h from the preceding feature quantity extraction unit 210 or the sub-processing unit 240-(n-1). n-1 ∈RW_(n-1)×H_(n-1)×G_(n-1) and receives auxiliary input a n ∈RW_(n-1)×H_(n-1)×G_(n-1), where A_B is A B means that A^B is A B means that A_(BC) is A B-C means that A^(BC) is A B-C means.
[0058] The sub-processing unit 240-n is a processing module h n =Φ n (h n-1 ,a n ;θ n ) to obtain the feature value h n-1 The feature value h n ∈R W_n×H_n×G_n Here, according to the structure of U-Net, each processing module has the same or smaller feature size. n-1 ≧W n , H n-1 ≧H n Let's say.
[0059] (Bottleneck 240-B) Each sub-processing unit 240-n included in the bottleneck 240-B receives as input the feature quantity h n-1 ∈RW_(n-1)×H_(n-1)×G_(n-1) and the processing module h n =Φ n (h n-1 ;θ n ) to obtain the feature value h n-1 The feature value h n ∈R W_n×H_n×G_n Here, according to the structure of U-Net, each processing module does not change the size of the feature. n-1 =W n , H n-1 =H n Let's say.
[0060] (Decoder 240-D) Each sub-processing unit 240-n included in the decoder 240-D performs the following processing.
[0061] The sub-processing unit 240-n receives as input the feature quantity h from the previous sub-processing unit 240-(n-1). n-1 ∈RW_(n-1)×H_(n-1)×G_(n-1).
[0062] The sub-processing unit 240-n is a processing module h n =Φ n (h n-1 ;θ n ) to obtain the feature value h n-1 The feature value h n ∈R W_n×H_n×G_n Here, according to the structure of U-Net, each processing module has the same or larger feature size. n-1 ≦W n , H n-1 ≦H n Sub-processing unit 240-(N 2 +1) to 240-(N-1) are the feature values h (N_2+1) ~h (N-1) to the subsequent sub-processing unit 240-(N 2 +2) to 240-(N), and also outputs the result to the target signal generator 250. The last sub-processor 240-N generates the feature quantity h N ∈RW_N×H_N×G_N to the target signal generator 250.
[0063] <Auxiliary Input Extraction Unit 220> FIG. 10 shows a functional block diagram of the auxiliary input extraction unit 220. As shown in FIG.
[0064] The auxiliary input extraction unit 220 extracts the acoustic signal x and M estimated signals ^x m Here, m=0, 1, ..., M, where the resampling processor 220-0 processes the acoustic signal x, and the resampling processors 220-m (m=1, 2, ..., M) process the estimated signal ^x m Each resampling unit 220-m performs processing on the received acoustic signal x or the estimated signal ̂x m The real and imaginary parts of the signal x are resampled in the time and frequency directions (auxiliary input extraction process S220) to obtain N1 signals x n ∈R W_n×H_n×2 or ^x n m ∈R W_n×H_n×2 get.
[0065] At this time, the size of each resampled signal is determined by the feature quantity h received by each sub-processing unit 240-n of the encoder 240-E of the target signal extraction unit 240. n-1 ∈RW_(n-1)×H_(n-1)×G_(n-1). Various methods can be used for resampling, such as thinning, smoothing, and down sampling.
[0066] <Adaptation Unit 230 - n > The signal estimation device 200 has N1 adaptation units 230 - n , and each of the N1 adaptation units 230 - n corresponds to a sub-processing unit 240 - n constituting the encoder 240 -E of the target signal extraction unit 240 .
[0067] Each matching unit 230-n receives the feature quantity h n-1 The transformed acoustic signal x is of the same size as ∈RW_(n-1)×H_(n-1)×G_(n-1). n ∈R W_n×H_n×2 and M estimated signals ^x n m ∈R W_n×H_n×2are received from all M+1 resampling processing units 220-m and organized into one tensor in the channel direction. n ∈R W_n×H_n×2(M+1) Furthermore, each adaptation unit 230-n generates a processing module a n =Φ a_n (χ n ;θ a_n ) to obtain the acoustic signal x n ∈R W_n×H_n×2 and M estimated signals ^x n m ∈R W_n×H_n×2 From auxiliary input a n ∈RW_(n-1)×H_(n-1)×G_(n-1) is extracted (matching process S230) and handed over to the sub-processing unit 240-n.
[0068] <Target signal generator 250> The target signal generator 250 generates all the feature quantities {h n} n∈decoder and the processing module y=Φ out ({h n} n∈decoder ;θ out ) to extract an estimated value y of the target signal (S250).
[0069] <Effects> With the above-described configuration, the effects described in the first embodiment can be obtained.
[0070] <Other Configurations> The configuration of each processing unit may be modified in various ways from the above settings.
[0071] For example, based on the configuration of U-Net, the encoder 240-E and decoder 240-D of the target signal extraction unit 240 are configured with the same number of sub-processing units 240-n, and the feature quantity z extracted by the sub-processing unit 240-n of the encoder 240-E is n may be additionally input to the processing module of the sub-processing unit N-n+1 of the decoder 240-D.
[0072] In this embodiment, processing is performed on one acoustic signal x recorded by one microphone, but processing may be performed on P acoustic signals x recorded by P (P is any integer equal to or greater than 2) microphones. Furthermore, in this embodiment, processing is performed on M estimated signals estimated from one acoustic signal x, but processing may be performed on P (P is any integer equal to or greater than 2) acoustic signals x. p Alternatively, the processing may be performed on M×P estimated signals estimated from
[0073] Furthermore, the signal estimation method according to this embodiment is configured with a process (such as a neural network) capable of learning an input / output relationship from learning data, and the estimated target signal is defined by a teacher signal used during learning. Therefore, although one target signal is estimated in this embodiment, multiple target signals may be estimated by appropriately changing the teacher signal. Furthermore, the target signal may be defined as the target sound itself, or as another signal representing the characteristics of the target sound.
[0074] <About parameter learning> The parameters of each processing module are learned by i and the correct target signal y i The training signal {x i ,y i} i It is learned in advance based on the
[0075] For example, let Θ represent the total parameters of all processing modules, and let x i When input to the target signal estimation method, the output is ^y i When (Θ) is obtained, learning can be performed by updating the parameters using backpropagation or the like so as to minimize the error function E(Θ) below. Here, D(y i ,^y i (Θ)) is y i and ^y i A function that represents the distance (Θ), such as Euclidean distance, is used.
[0076] Second Embodiment The following mainly describes the differences from the first embodiment.
[0077] In the second embodiment, an example of realizing a score model that is required to apply speech enhancement based on a diffusion model (see Reference 1) to an acoustic signal is shown.
[0078] In Non-Patent Document 1, only a score model that receives only a single estimated signal is devised, and furthermore, processing when a plurality of estimated signals are received cannot be realized.
[0079] By using the score model based on the second embodiment, when an acoustic signal and a plurality of estimated signals are received, it becomes possible to achieve highly accurate processing while minimizing the increase in calculation cost.
[0080] In speech enhancement based on the diffusion model, for example, the following stochastic differential equation is solved from state number t=T to t=0 to obtain an acoustic signal x∈C T×F×1 The target sound z0∈C T×F×1 Estimate.
[0081] z t =[-γ(xz t )+g(t) 2 y t ]dt+g(t)dw (1) where z t ∈C T×F×1 is the state variable, w∈C T×F×1 is a time-reversed standard Wiener process. γ is a stiffness coefficient, and g(t) is a noise intensity function, each of which is determined based on a separately designed forward process of a diffusion model. Details are described, for example, in Reference 1.
[0082] In the above equation, y t =∇ z_t log p(z t |x)∈C T×F×1 is a tensor called a score, which is the forward process of the diffusion model, and expresses z under the given acoustic signal x. t The probability density function of p(z t |x), log p(z t |z of x) t is defined as the gradient with respect to
[0083] y tSince it is difficult to analytically calculate y t For example, the score model ^y t =φ(x,z t ,t) is x,z t ,t as input and estimate the score ^y t It is trained to output
[0084] Even when multiple estimated signals are received, speech enhancement based on the diffusion model can be achieved by solving the stochastic differential equation (1).
[0085] However, in this case, the score model receives the acoustic signal and multiple estimated signals, and calculates the score estimate ^y t You need to change it to output:
[0086] In the second embodiment, a method for constructing a score model corresponding to the above-mentioned plurality of estimated signals is provided.
[0087] Acoustic signal x∈C T×F×1 and state number t and M estimated signals ^x m ∈C T×F×1 and the state variable z at state number t t ∈C T×F×1 , estimate the score of the target sound as the target signal, and calculate the estimated value y t ∈C T×F×1 (See FIG. 11).
[0088] However, the above x, ^x m , z t , y t is a three-dimensional complex tensor C used in speech enhancement based on the diffusion model. T×F×1 The three-dimensional real tensor R can be created by dividing the complex numbers of each element into real and imaginary parts and rearranging them into two different channels. T×F×2 Let's say.
[0089] <Signal Estimation Apparatus 300> FIG. 11 is a functional block diagram of a signal estimation apparatus 300 according to the second embodiment, and FIG. 4 shows the processing flow thereof.
[0090] The signal estimation device 300 includes M estimation units 201 - m , a feature extraction unit 310 , an auxiliary input extraction unit 320 , N adaptation units 330 - n , a target signal extraction unit 340 , and a target signal generation unit 350 .
[0091] The signal estimation device 300 estimates an acoustic signal x∈C T×F×1 In addition, the state variable z t ∈C T×F×1 and state number t, estimates the above score, and outputs the estimated score as the target signal.
[0092] The M estimating units 201-m have the same configuration as in the first embodiment, and therefore their description will be omitted.
[0093] <Feature Extraction Unit 310> The feature extraction unit 310 extracts an acoustic signal x∈C T×F×1 and M estimated signals ^x m ∈C T×F×1 and the state variable z t ∈C T×F×1 That is, the feature extraction unit 310 receives the state variable z as an additional input. t The feature extraction unit 310 receives the acoustic signal x and M estimated signals ^x m and the state variable z t The real and imaginary parts of the state variable z are stored in different channels. t The tensor created by connecting t ∈R T×F×2 The feature extraction unit 310 extracts χ and ζ t In this embodiment, the feature extraction unit 310 extracts a feature h0 that matches the input size and number of channels of the target signal extraction unit 340 from the processing module h0=Φ in (χ,ζ t ;θ in ) and use χ, ζ t feature h0∈R T×F×G_0 By converting the acoustic signal x and M estimated signals ^x m and the state variable z t Extract feature h0 from
[0094] <Target Signal Extraction Unit 340> FIG. 12 shows a functional block diagram of the target signal extraction unit 340. As shown in FIG.
[0095] The target signal extraction unit 340 is composed of multiple sub-processors 340-n (1≦n≦N) similar to the target signal extraction unit 240. The sub-processor 340-n receives as input the state number t and the feature value h from the previous feature extraction unit 310 or the sub-processor 340-(n-1). n-1 ∈RW_(n-1)×H_(n-1)×G_(n-1) and receives auxiliary input a n ∈RW_(n-1)×H_(n-1)×G_(n-1) (see FIG. 12). That is, each sub-processing unit 340-n receives the state number t as an additional input.
[0096] Each sub-processing unit 340-n included in the encoder 340-E is a processing module h n =Φ n (h n-1 ,a n ,t;θ n ) to obtain the feature value h n-1 The feature value h n ∈R W_n×H_n×G_n and each sub-processing unit 340-n included in the bottleneck 340-B and the decoder 340-D is a processing module h n =Φ n (h n-1 ,t;θ n ) to obtain the feature value h n-1 The feature value h n ∈R W_n×H_n×G_n Convert to.
[0097] <Auxiliary Input Extraction Unit 320> FIG. 13 shows a functional block diagram of the auxiliary input extraction unit 320. As shown in FIG.
[0098] The auxiliary input extraction unit 320 extracts the acoustic signal x and M estimated signals ^x m M+1 resampling processing units 320-xm, which perform processing for each state variable z t In other words, the auxiliary input extraction unit 320 includes a resampler 320-z in addition to the auxiliary input extraction unit 220. The resampler 320-x-m has the same configuration as the resampler 220-m in the first embodiment.
[0099] Each resampling unit 320-z has a state variable z t and receives the state variable z t The real and imaginary parts of the signal z are resampled in the time and frequency directions (auxiliary input extraction process S320), and N1 signals z n t ∈R W_(n-1)×H_(n-1)×2 At this time, the size of each resampled signal is calculated by the feature quantity h n-1 ∈RW_(n-1)×H_(n-1)×G_(n-1).
[0100] <Adaptation Unit 330 - n > The signal estimation device 300 has N1 adaptation units 330 - n , and each of the N1 adaptation units 330 - n corresponds to a sub-processing unit 340 - n constituting the encoder 340 -E of the target signal extraction unit 340 .
[0101] Each adaptation unit 330-n, like each adaptation unit 230-n, receives the acoustic signal x n ∈R W_n×H_n×2 and M estimated signals ^x n m ∈R W_n×H_n×2 are received from all M+1 resampling units 220-m, and the signal χ n ∈R W_n×H_n×2(M+1) Each fitting unit 330-n generates the feature quantity h received by the sub-processing unit 340-n. n-1 The state variable z transformed to the same size as ∈RW_(n-1)×H_(n-1)×G_(n-1) n t ∈R W_(n-1)×H_(n-1)×2 from the resampler 320-z. Furthermore, each adaptor 330-n receives a processing module a n =Φ a_n (χ n ,z t n ;θ a_n ) to obtain the signal χ n , z n t From auxiliary input a n (adaptation process S330) and delivers it to the sub-processing unit 340-n of the encoder 3400-E.
[0102] <Target signal generator 350> The target signal generator 350 generates all the feature quantities {h n} n∈decoder and the processing module y=Φ out ({h n} n∈decoder ;θ out ) to obtain the estimated score y t is generated as an estimate of the target signal (S350) and output.
[0103] <Effects> With the above configuration, in the second embodiment, when the number of estimated signals is increased, it is possible to estimate scores corresponding to a plurality of estimated signals with higher accuracy while minimizing an increase in calculation cost.
[0104] <Other Configurations> The configuration of each processing unit may be modified in various ways in addition to the above settings.
[0105] As in the first embodiment, processing may be performed on P acoustic signals (P is any integer equal to or greater than 2) recorded by P microphones and M×P estimated signals. Alternatively, a configuration may be adopted in which multiple target signals are estimated.
[0106] <About Parameter Learning> If Θ represents the total parameters of all processing modules, the learning of the parameters of the score model can be performed in the same way as the conventional diffusion model, for example, using the error backpropagation method based on the following error function E(Θ): where v∈C T×F×M is a white noise with a multivariate complex normal distribution with mean 0 and covariance of the identity matrix, E t,(z_0,x),v is the expected value function obtained when sampling the target sound z0 and the acoustic signal x from the training data, and also sampling the state number t and the white noise v, ^y(Θ)∈C T×F×1 is the score model, x,{^x m} m ,z t The estimated score when t is input, z t is the state variable of the diffusion model at state number t determined based on x, v, and the forward process of the diffusion model, and σ(t) 2is determined according to the forward process of the diffusion model, and the state variable z t represents the variance of the white noise contained in
[0107] The error function to be used is not limited to the above definition, and any other function may be used as long as it can evaluate the estimation error of the score model.
[0108] <Experimental Results> Noise suppression and reverberation suppression were performed on an acoustic signal picked up by two microphones in a noisy and reverberant environment.
[0109] 14 and 15 show the experimental results.
[0110] 14, it can be seen that the signal-to-distortion ratio (SDR) of Example 2 is improved compared to Non-Patent Document 1. Note that the two target sound estimations in Non-Patent Document 1 utilize Complex Spectral Mapping (CSM) and Diffusion Model Speech Enhancement. The estimation section of Example 2 utilizes CSM and a dereverberation method (Weighted Prediction Error: WPE).
[0111] 15 shows the parameter size of the score model when the number of estimators is set to 0 and 3 in Example 2. It can be confirmed that even if the number of estimators is increased, the number of parameters hardly increases, and therefore the required memory size does not increase. The fact that the parameter size does not increase means that the calculation time does not increase either.
[0112] <Modifications> Furthermore, a device (terminal) for using the device, system, or method of the present invention via a network (telecommunications line) may also be included. The "device (terminal) for use" may be provided with functions (e.g., control function, decoding function, restoration function, input / output function, etc.) necessary to obtain the effects achieved by implementing the device, system, or method of the present invention. Note that a configuration including a device (terminal) for using the device or method of the present invention via a network (telecommunications line) is also referred to as a signal estimation system.
[0113] <Hardware, Programs, and Recording Media> The functions realized by the components described in this specification may be implemented in circuitry or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), a CPU (a Central Processing Unit), conventional circuits, and / or combinations thereof, programmed to realize the described functions. A processor includes transistors and other circuits and is considered to be circuitry or processing circuitry. A processor may also be a programmed processor that executes a program stored in a memory.
[0114] In this specification, a circuitry, unit, or means is hardware that is programmed to realize or performs the described functions, which may be any hardware disclosed herein or any hardware known to be programmed to realize or perform the described functions.
[0115] If the hardware is a processor considered to be a type of circuitry, the circuitry, means, or unit is a combination of the hardware and software used to configure the hardware and / or processor.
[0116] The various processes described above can be implemented by loading a program that executes each step of the above method into the recording unit 2020 of the computer 2000 shown in Figure 16, and operating the control unit 2010, input unit 2030, output unit 2040, display unit 2050, etc.
[0117] The program describing the processing contents can be recorded on a computer-readable recording medium, which may be, for example, a magnetic recording device, an optical disk, a magneto-optical recording medium, a semiconductor memory, or any other suitable recording medium.
[0118] The program may be distributed by, for example, selling, transferring, lending, etc. portable recording media such as DVDs and CD-ROMs on which the program is recorded. Furthermore, the program may be stored in a storage device of a server computer, and then transferred from the server computer to other computers via a network, thereby distributing the program.
[0119] A computer that executes such a program may first temporarily store the program recorded on a portable recording medium or transferred from a server computer in its own storage device. Then, when executing a process, the computer reads the program stored on its own recording medium and executes the process in accordance with the read program. Alternatively, the computer may read the program directly from a portable recording medium and execute the process in accordance with the program. Furthermore, the computer may execute the process in accordance with the program each time a program is transferred from a server computer to the computer. Alternatively, the server computer may not transfer the program to the computer, but may instead execute the process through a so-called ASP (Application Service Provider) service, which realizes the processing function by issuing an execution instruction and obtaining the results. Furthermore, the server computer may execute the process at the terminal using a so-called SaaS (Software as a Service) service, which allows users to use part of a server computer along with the program. In this embodiment, the program includes information used for processing by an electronic computer that is equivalent to a program (such as data that is not a direct instruction to a computer but has properties that dictate computer processing).
[0120] Furthermore, in this embodiment, the device is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.
[0121] <Other Modifications> The present invention is not limited to the above-described embodiments and modifications. For example, the various processes described above may not only be executed in chronological order as described, but may also be executed in parallel or individually depending on the processing capacity of the device that executes the processes or as needed. Other modifications are possible as long as they do not deviate from the spirit of the present invention.
Claims
1. A signal estimation method for estimating a target signal representing characteristics of a target sound from an acoustic signal obtained by recording with a microphone, the signal estimation method comprising: a first estimation step of estimating the target signal from the acoustic signal using at least one method; and a second estimation step of estimating the target signal using a feature of the acoustic signal, wherein the second estimation step performs a plurality of processes, and the plurality of processes use, as auxiliary inputs, the acoustic signal and the estimated signal after resampling, which are obtained by performing resampling on the previous acoustic signal and at least one estimated signal obtained by estimation in the first estimation step.
2. A signal estimation method according to claim 1, wherein the second estimation step comprises: a feature extraction step of receiving the acoustic signal and the at least one estimated signal and extracting features; a target signal extraction step having multiple sub-processing steps of performing conversion processing of the features; an auxiliary input extraction step of receiving the acoustic signal and the at least one estimated signal and performing resampling processing for each of the acoustic signal and the at least one estimated signal to a size corresponding to the sub-processing step of the output destination; an adaptation step of receiving from the multiple resampling processing steps the acoustic signal and the at least one estimated signal that have been resampled to match the number of channels of the input of the sub-processing step of the output destination, combining all the acoustic signals and the at least one estimated signal, and outputting the converted information to the sub-processing step of the output destination as the auxiliary input; and a target signal generation step of receiving the features output by one or more sub-processing steps and generating a target signal.
3. A signal estimation device that estimates a target signal representing characteristics of a target sound from an acoustic signal obtained by recording with a microphone, the signal estimation device comprising: a first estimation unit that estimates the target signal from the acoustic signal using at least one method; and a second estimation unit that estimates the target signal using a feature of the acoustic signal, wherein the second estimation unit performs a plurality of processes, and the plurality of processes use, as auxiliary inputs, the acoustic signal and the estimated signal after resampling, which are obtained by performing resampling on the acoustic signal and at least one estimated signal obtained by estimation with the first estimation unit.
4. A program for causing a computer to execute the signal estimation method of claim 1.
Citation Information
Patent Citations
Signal processing device, signal processing method, and program
WO2024038522A1