Acoustic signal estimation device, acoustic signal estimation method, and program
The acoustic signal estimation device addresses the challenges of audio declipping by using a data-driven sparse optimization approach within a DNN framework, achieving high accuracy for both large and small distortion signals.
Patent Information
- Application Number
- JP2022142915
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-09-08
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2042-09-08
AI Technical Summary
Existing audio declipping methods face challenges in accurately restoring signals with large distortion using sparse optimization, and in maintaining accuracy for signals with small distortion using deep neural networks (DNNs).
An acoustic signal estimation device that employs a combination of sparse optimization and DNNs, where the threshold for sparse optimization is determined in a data-driven manner using a DNN, to estimate the waveform of a pre-clip signal from a post-clip signal.
The proposed solution achieves highly accurate audio declipping for signals with large distortion compared to conventional sparse optimization methods, and maintains accuracy for signals with small distortion compared to conventional DNN-based methods.
Smart Images

Figure 0007690936000014 
Figure 0007690936000015 
Figure 0007690936000016
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technique for restoring a signal before clipping from a signal after clipping.
Background Art
[0002] Due to performance limitations of acoustic devices such as recording devices, clipping may occur during recording, where portions of an acoustic signal that exceed an amplitude limit are lost. There is Audio declipping as a technique for restoring the waveform of an original signal from the waveform of such a clipped signal. Audio declipping generally has two methods. One method is a method based on a deep neural network (DNN). While this method can achieve high restoration performance even when the distortion of the signal is large, there is a problem that the restoration performance deteriorates in the case of data that was not included in the learning data. As another method, there is a method based on sparse optimization. Different from the method based on DNN, this method can handle restoration for signals different from the learning data. That is, it is possible to perform appropriate restoration processing according to the magnitude of the distortion (in other words, according to the difficulty of the problem) (see, for example, Non-Patent Document 1).
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the case of the method based on the sparse optimization described above, since the nature of the data such as how the signal is distorted by clipping cannot be considered, there is a problem that each component in the time-frequency domain cannot be appropriately selected. On the other hand, in the case of the method based on the above-described DNN, although the nature of the data can be considered by machine learning (hereinafter also referred to as "learning"), when the clipped signal (hereinafter also referred to as "post-clip signal") to be restored has a large difference in amplitude etc. compared with the training data, there is a problem that sufficient estimation for restoration cannot be made.
[0005] Therefore, the present disclosure has been made to solve the above problems, and can realize highly accurate Audio declipping for signals with large distortion compared with conventional methods based on sparse optimization, and can realize Audio declipping that does not cause a decrease in accuracy for signals with small distortion compared with conventional methods based on DNN, and an object thereof is to provide an acoustic signal estimation device.
Means for Solving the Problems
[0006] In order to solve the above problems, an acoustic signal estimation device according to an aspect of the present disclosure is an acoustic signal estimation device that estimates the waveform of a pre-clip signal ~y, which is the signal before being clipped, from the waveform of a post-clip signal y, which is a signal clipped at a predetermined threshold value, and includes a first estimation unit, a second estimation unit, a variable update unit, and an output unit. k (k = 0, 1, 2,..., K - 1) is the number of executions of the estimation of the first estimation signal by the first estimation unit, K is a predetermined number, x [k] is the first estimation signal, v [k] is the second estimation signal, u [k] is the dual variable u, x [0] is the waveform of the post-clip signal, v [0] is the time-frequency representation of x [0] and u [0] is an arbitrary number. In this case, the first estimation unit includes the second estimation signal v [k] and the dual variable u [k]Generate a waveform to be constrained using the following as inputs, and apply a projection operator Π for constraining the generated signal to a region included in set Γ with respect to the waveform to be constrained Γ to generate a first estimated signal x [k+1] which is a new waveform. The second estimator converts the first estimated signal x [k+1] into a time-frequency representation, and executes soft thresholding using a deep neural network with the first estimated signal x [k+1] converted to this time-frequency representation and a dual variable u [k] as inputs to generate a second estimated signal v [k+1] which is a signal of a new time-frequency representation to which a sparse optimization method is applied. The variable update unit takes the dual variable u [k] the first estimated signal x converted to the time-frequency representation [k+1] and the second estimated signal v [k+1] as inputs to generate a new dual variable u [k+1] . When the number of executions k is less than K - 1, the output unit increments k by 1 and performs the processes of the first estimator, the second estimator, and the variable update unit. When the number of executions k is K - 1 or more, the generated first estimated signal x [K] is output as an estimation result of the waveform of the pre-clipping signal ŷ.
Advantages of the Invention
[0007] According to the present disclosure, while adopting a sparse optimization algorithm, since the threshold of the thresholding process for inducing sparsity is determined in a data-driven manner based on a DNN, compared with conventional methods based on sparse optimization, Audio declipping with high accuracy can be realized for signals with large distortion, and compared with conventional methods based on DNN, Audio declipping that does not cause a decrease in accuracy for signals with small distortion can be realized.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
DETAILED DESCRIPTION OF THE INVENTION
[0009] <Character Notation> In the text, the symbol "~" (tilde) should be described directly above the immediately following character, but due to text notation limitations, it is described immediately before the character. In mathematical formulas, these symbols are described in their original positions, i.e., directly above the character. For example, "~y" is represented by the following formula in a mathematical formula.
Number
[0010] (Regarding Audio declipping) When the time is t, consider the pre-clipping signal (hereinafter also referred to as the "pre-clipping signal"), which is the original signal, ~y and the signal y whose amplitude is limited by the threshold τ (hereinafter also referred to as the "post-clipping signal") as follows. [Equation] The index of the post-clipping signal y above is divided into three disjoint sets H = {t ∈ [1, T] | y[t] ≥ τ}, R = {t ∈ [1, T] | |y[t]| < τ}, L = {t ∈ [1, T] | y[t] ≤ -τ}. Audio declipping is a technique for estimating the pre-clipping signal ~y, which is the original signal, from only the signal y and the information (H, R, L) of the above index.
[0011] (Method based on sparse optimization) According to Non-Patent Document 1 above, the method based on sparse optimization is a method that uses the solution of the optimization problem expressed by the following equation as the estimation result. [Equation] Here, S is an l 1 sparsity-inducing function such as a norm ("L1 norm"), and G is a window g ∈ RT It is the discrete Gabor transform shown by the following formula using T . [Equation] Here, i is the imaginary unit, a is the time shift length, and M is the number of frequency channels. Also, Γ is the set of executable solutions shown by the following formula. m and n indicate the row and column of the determinant, respectively. In particular, m ∈ {1, …, M} is the frequency index, and n ∈ {1, …, N} is the time index. τ [Equation] Clipping generates extra harmonic components, induces sparsity in the time-frequency domain, and removes the extra harmonic components.
[0012] In the method of Non-Patent Document 1 described above, as this sparsity-inducing function, a weighted l-norm using the parabola weight w[m,n] = (m + 1) 2 / M 2 is proposed. In that case, the weighted soft threshold operator (T 1 (z))[m,n] shown by the following formula is used in the algorithm. w-soft (z))[m,n] is used in the algorithm. [Equation] Here, (·) + = max(0, ·), and λ is a hyperparameter. Originally, threshold processing that only cuts the extra components generated by clipping is desirable, but in Equation (5), since threshold processing is performed according to the predetermined value of λw[m,n], the nature of the data cannot be considered, and there is a concern that components of the original (pre-clip signal) may be largely cut or extra components may be left.
[0013] (Method based on DNN) In the method based on a deep neural network (DNN), the original signal ($\tilde{y}$) is estimated by inputting the observed signal (post - clipped signal) $y$ into the pre - trained DNN. Since the nature of the data, that is, how the signal is distorted by clipping, can be learned in a data - driven manner, it is known that the original signal can be estimated with high accuracy even when the threshold $\tau$ is small. However, when the conditions of the threshold $\tau$ are significantly different between the learning time and the inference time, a problem occurs in that accurate estimation cannot be achieved.
[0014] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Also, hereinafter, components having the same function are given the same number, and redundant explanations are omitted.
[0015] <Acoustic signal estimation device> As shown in FIG. 1, the acoustic signal estimation device 1 according to an embodiment of the present disclosure includes a first estimation unit 10, a second estimation unit 20, a variable update unit 30, and an output unit 40. The acoustic signal estimation device 1 is a device that realizes Audio declipping by estimating the pre - clipped signal ($\tilde{y}$) from the input signal, which is the post - clipped signal $y$. The acoustic signal estimation device 1 performs the acoustic signal estimation method of the present embodiment by implementing the processing flow shown in FIG. 2.
[0016] (First estimation unit 10) The first estimation unit 10 takes the second estimated signal $v$ [k] and the dual variable $u$ [k] as inputs and generates the first estimated signal $x$ [k+1] . That is, assuming that $k$ ($k = 0,1,2,\cdots,K - 1$) is the number of executions of generating the first estimated signal $x$ [k+1] , and $K$ is a predetermined number used by the output unit 40 described later, the following equation is used to generate the first estimated signal $x$ [k+1] .
Equation
[0017] In the case of the initial value (k = 0), the waveform of the clipped signal y, which is the signal to be restored, is input to x. [0] The time-frequency representation of y(x [0] ), which is obtained by applying the discrete Gabor transform shown in Equation (3) to the clipped signal y(x [0] ), is input to v. [0] An arbitrary number, such as an m×n zero matrix where all elements are 0 (zero), is input to u. [0] Hereinafter, x [0] will also be referred to as the input signal x [0] .
[0018] That is, the first estimation unit 10 first generates a signal representing the waveform of (v * u [k]- ) by G [k] (v [k]- u [k] ). Next, a projection operator Π * is applied to the generated G [0]- (v [0] ) to generate a new waveform, the first estimated signal x Γ [k+1] [k+1] (step S10).
[0019] Here, when k = 0, the above-described v [0] , u [0] are input, but when k is 1 or more, the second estimated signal v [k+1] generated by the second estimation unit 20 described later and the dual variable u [k+1] generated by the variable update unit 30 will be used.
[0020] (Second Estimation Unit 20) The second estimation unit 20 uses the first estimation signal x [k+1] and the dual variable u [k] as inputs to perform soft thresholding using a deep neural network, and generates a second estimation signal v [k+1] which is a signal of a new time-frequency representation to which the sparse optimization method is applied (step S20). The following equation is used to generate the second estimation signal v [k+1] .
Equation
[0021] That is, the second estimation unit 20 first converts the first estimation signal x [k+1] into a time-frequency representation by Gx [k+1] , adds the dual variable u [k] to this, and performs soft thresholding to calculate the soft thresholding operator T [k+1] +u [k] using a deep neural network (pre-trained model F θ ) as an input, and generates a second estimation signal v θ which is a signal of a new time-frequency representation to which the sparse optimization method is applied (step S20). [k+1]
[0022] Here, when k = 0, the above-described u [0] is input to u, but when k is 1 or more, the dual variable u [k+1] generated by the variable update unit 30 described later is used.
[0023] In the above equation (7), the following equation is used to calculate the soft thresholding operator T θ .
Equation
[0024] By passing through the process of the second estimator 20, a sparsity constraint is imposed.
[0025] The learned model F θ is a learned model composed of a multi-layer neural network. The learning method of the learned model F θ will be described later.
[0026] (Variable update unit 30) The variable update unit 30 takes the dual variable u [k] the first estimated signal x [k+1] and the second estimated signal v [k+1] as inputs and generates a new dual variable u [k+1] (step S30). That is, by generating the dual variable u [k+1] the dual variable u is updated. The following equation is used to generate the updated dual variable u [k+1] . [Equation] That is, the variable update unit 30 adds Gx [k] to u [k+1] and further subtracts u [k+1] to generate the updated dual variable u [k+1] (step S30).
[0027] Here, when k = 0, the above-mentioned u [0] is input to u, but when k is 1 or more, the dual variable u [k+1] generated by the variable update unit 30 is used.
[0028] (Output unit 40) When the number of executions k of generating the first estimated signal is less than K - 1, the output unit 40 increments k by one and causes the above-described first estimation unit 10, second estimation unit 20, and variable update unit 30 to perform processing.
[0029] Also, when the number of executions k of generating the first estimated signal is K - 1 or more, the generated first estimated signal x [K] is output as an estimation result of the waveform of the pre-clip signal ~y (step S40).
[0030] <Estimation learning device> The learning of the above-described learned model F θ is performed by the estimation learning device 300 shown in FIG. 3. The estimation learning device 300 of the present disclosure includes a clip application unit 310, a learning estimation unit 320, a loss calculation unit 330, and a parameter update unit 340. When learning acoustic data D for learning is input to the estimation learning device 300, learning of signal estimation for restoring the pre-clip signal from the post-clip signal is performed. The estimation learning device 300 performs the estimation learning method of the present embodiment by implementing the processing flow shown in FIG. 4.
[0031] (Clip application unit 310) The clip application unit 310 applies hard clip, which is a pseudo amplitude limit, to the pre-clip signal for learning input from the learning acoustic data D to generate a post-clip signal for learning (step S310). As the learning acoustic data D, for example, general-purpose data such as 5300 data of the LIBRI corpus can be used. Therefore, the learning acoustic data D is not limited to 5300 data of the LIBRI corpus.
[0032] (Learning estimation unit 320) The learning estimation unit 320 estimates the pre-clip signal for learning (estimated signal) from the post-clip signal for learning (step S320).
[0033] (Loss calculation unit 330) The loss calculation unit 330 calculates the loss between the estimated signal estimated by the learning estimation unit 320 and the pre-clip signal for learning when input from the learning acoustic data D (step S330). Examples of the loss calculation include, for example, mean-squared-error (MSE) loss for signals in the time domain. However, the loss calculation is not limited to mean-squared-error (MSE) loss. Note that the cost function is calculated only for the region where the amplitude is limited by the clip application unit 310.
[0034] (Parameter update unit 340) When the above-described loss does not satisfy a predetermined criterion, the parameters used by the learning estimation unit 320 are updated based on the loss, and the estimation by the learning estimation unit 320 is performed again. For example, based on the obtained cost, the parameters of the learning estimation unit 320 are updated by the optimization method Adam to perform the estimation.
[0035] When the above loss satisfies a predetermined criterion, the learning estimation unit 320 having the parameters used immediately before is output as the learned model F θ (step S340). Note that as the predetermined criterion, instead of judging by the loss result itself, for example, a method such as stopping learning when the parameters are updated using all the learning data a predetermined number of times, for example, 200 times, may be adopted.
[0036] <Application example of the acoustic signal estimation device 1> In order to confirm the accuracy of the above-described acoustic signal estimation device 1, the acoustic signal estimation device 1 was applied under the following conditions. FIGS. 5 and 6 show examples of execution results. In this execution result, as learning by the estimation learning device 300, 5300 data of the LIBRI corpus were used as the learning acoustic data D, Adam was used as the optimization algorithm, and learning was stopped when the parameters were updated using all the learning data 200 times, and the parameters at that time were used as the learned model F θ Also, the initial value (u[0]) of the dual variable in the acoustic signal estimation device 1 was set to 0 (zero).
[0037] In FIGS. 5 and 6, the horizontal axis (input SDR) indicates the clipped strength of the input post-clip signal y, and the vertical axis (ΔSDR) indicates the magnitude of the improvement amount.
[0038] FIG. 5 shows the SDR (Signal-to-Distortion Ratio) by the clip application unit 310 fixed to respective values of 1 db, 3 db, 5 db, 10 db, and 15 db in the learning in the estimation learning device 300, and the learned model F θ is the estimation result of the acoustic signal estimation device 1 when generated. That is, the five results (ΔSDR values) plotted at the same Input SDR value are such that for each input signal (Input SDR), one condition is that the post-clip signal y under the same condition as during learning is input, and the remaining four conditions show the results of input of unlearned post-clip signal y. In the present disclosure, if the five results are shown in one figure, the line graphs overlap and the visibility deteriorates, so the figure is divided into two. Specifically, the results of 1 db, 3 db, and 5 db are shown in FIG. 5A, and the results of 1 db, 10 db, and 15 db are shown in FIG. 5B. That is, while dividing the figure into two to prevent a decrease in visibility, regarding the result of 1 db, it is shown in both FIG. 5A and FIG. 5B to ensure ease of comparison of the mutual results.
[0039] In the conventional DNN-based method, the estimation result of the pre-clip signal ~y is based only on the nature of the learned learning acoustic data, so for unknown conditions that have not been learned, the pre-clip signal ~y cannot be accurately estimated. In the acoustic signal estimation device 1 of the present disclosure, as shown in FIGS. 5A and 5B, it can be seen that it is not greatly affected by the SDR during learning and is robust to the difference from the conditions during learning.
[0040] FIG. 6A shows the learned model F learned with a random value of SDR during learning in the estimation learning device 300 within 1 to 10 dB θThis is a diagram comparing the estimation result of the pre-clip signal ~y of the acoustic signal estimation device 1 using [technique name] with the estimation result of the pre-clip signal ~y using the sparse optimization method of the conventional technique. Here, as methods based on the sparse optimization of the conventional technique, the results calculated by (i) ASPADE, (ii) SS PEW, and (iii) PWl 1 are being compared. Also, Fig. 6B shows the result of T-UNet based on the DNN method of the conventional method as (iv) instead of (i) to (iii) in Fig. 6A.
[0041] As shown in Fig. 6A, the improvement amount △SDR results of the acoustic signal estimation device 1 for all input SDRs are higher compared to the results of the other (i) to (iii) methods based on the sparse optimization of the conventional technique. From this result, it can be seen that the method by the acoustic signal estimation device 1 has the threshold processing based on the nature of the data effectively working for Audio declipping. Also, as shown in Fig. 6B, when compared with T-UNet based on the DNN of the conventional method in (iv), it can be seen that the △SDR result of the acoustic signal estimation device 1 is larger (more effective) when the input SDR is 10 dB or more. This is presumably because the acoustic signal estimation device 1 can perform processing while considering the magnitude of distortion by imposing constraints in the time domain as in Equation (1).
[0042] [Program, Recording Medium] The above various processes can be implemented by causing the recording unit 2020 of the computer 2000 shown in Fig. 7 to read a program that executes each step of the above method and operating it on the control unit 2010, input unit 2030, output unit 2040, display unit 2050, etc.
[0043] The program describing this processing content can be recorded on a computer-readable recording medium. As a computer-readable recording medium, for example, any of a magnetic recording device, optical disk, magneto-optical recording medium, semiconductor memory, etc. may be used.
[0044] In addition, the distribution of this program is carried out, for example, by selling, transferring, lending, etc. a portable recording medium such as a DVD or CD-ROM on which the program is recorded. Further, the program may be stored in the storage device of a server computer and distributed by transferring the program from the server computer to other computers via a network.
[0045] A computer that executes such a program first stores, for example, the program recorded on a portable recording medium or the program transferred from a server computer in its own storage device. Then, at the time of executing the process, the computer reads the program stored in its own recording medium and executes the process according to the read program. As another execution form of this program, the computer may directly read the program from the portable recording medium and execute the process according to the program. Further, each time a program is transferred from the server computer to this computer, the computer may sequentially execute the process according to the received program. Also, the above-described process may be executed by a so-called ASP (Application Service Provider) type service that realizes the processing function only by the execution instruction and result acquisition without transferring the program from the server computer to this computer. Note that the program in this embodiment includes information used for processing by an electronic computer and similar to the program (data having a property of defining the processing of the computer although not a direct instruction to the computer).
[0046] In this embodiment, the present apparatus is configured by causing a computer to execute a predetermined program, but at least a part of these processing contents may be realized hardware-wise.
Explanation of Reference Numerals
[0047] 1 Acoustic signal estimation device 10 First estimation unit 20 Second Estimation Unit 30 Variable Update Unit 40 Output Unit 300 Estimation Learning Device 310 Clip Application Unit 320 Estimation Unit for Learning 330 Loss Calculation Unit 340 Parameter Update Unit D Acoustic Data for Learning F θ Trained Model T θ Soft Threshold Operator u [k] Dual Variable v [k] Second Estimation Signal x [k] First Estimation Signal ~y Signal before Clipping y Signal after Clipping
Claims
1. An acoustic signal estimation device that estimates the waveform of a pre-clip signal ~y, which is the signal before being clipped, from the waveform of a post-clip signal y that is a signal clipped at a predetermined threshold, k (k = 0, 1, 2, …, K−1) is the number of executions of the estimation of the first estimation signal by the first estimation unit, K is a predetermined number, and x [k] is the first estimation signal, and v [k] is the second estimation signal, and u [k] is the dual variable u, and x [0] is the waveform of the signal after clipping, and v [0] is x [0] 's time-frequency representation, and u [0] is an arbitrary number, in the case where The second estimated signal v [k] and the dual variable u [k] are used as inputs to generate a waveform to be constrained, and for the waveform to be constrained, a projection operator Π Γ is applied to constrain the generated signal to a region included in the set Γ, thereby generating a first estimated signal x [k+1] which is a new waveform; a first estimation unit the first estimated signal x [k+1] is converted into a time-frequency representation, and the first estimated signal x [k+1] in the time-frequency representation and the dual variable u [k] are used as inputs to perform soft thresholding processing using a deep neural network, and a second estimated signal v [k+1] which is a signal in a new time-frequency representation to which a sparse optimization method is applied is generated; a second estimation unit the dual variable u [k] and the first estimated signal x converted into the time-frequency representation [k+1] and the second estimated signal v [k+1] are used as inputs to generate a new dual variable u [k+1] a variable update unit When the number of executions k is less than K - 1, increment k by 1 and perform the processes of the first estimation unit, the second estimation unit, and the variable update unit. When the number of executions k is K - 1 or more, the generated first estimation signal x [K] An output unit that outputs as an estimation result of the waveform of the pre-clip signal ~y, An acoustic signal estimation device having the following.
2. The soft threshold processing using the deep neural network uses a learned model generated by a learning estimation device, The learning estimation device A clip application unit that applies hard clipping, which is pseudo amplitude limiting, to the input pre-clip signal for learning to generate a post-clip signal for learning, A learning estimation unit that estimates the pre-clip signal for learning from the post-clip signal for learning, A loss calculation unit that calculates the loss between the pre-clip signal for learning estimated by the learning estimation unit and the input pre-clip signal for learning, When the loss does not satisfy a predetermined criterion, the parameters used by the learning estimation unit are updated based on the loss to cause the learning estimation unit to perform estimation. When the loss satisfies the predetermined criterion, the learning estimation unit having the parameters used immediately before is output as a learned model. A parameter update unit, The acoustic signal estimation device according to claim 1, having the following.
3. G * is the adjoint operator of the operator G of the discrete Gabor transform, and when Γ is the set of executable solutions based on the predetermined threshold, the first estimated signal x [k+1] is the acoustic signal estimation device according to claim 1, which is calculated using the following formula. 【Number 11】
4. T θ When is a weighted threshold operator, the second estimated signal v [k+1] The acoustic signal estimation device according to claim 3, which is generated using the following formula. 【Number 12】
5. The new dual variable u generated by the variable update unit [k+1] The acoustic signal estimation device according to claim 4, which is generated using the following equation. 【Number 13】
6. An acoustic signal estimation method for estimating the waveform of a pre-clip signal ~y, which is the signal before being clipped, from the waveform of a post-clip signal y that is a signal clipped at a predetermined threshold, k (k = 0, 1, 2, …, K−1) is the number of times of estimation of the first estimated signal by the first estimation unit, K is a predetermined number of times, and x [k] is the first estimated signal, v [k] is the second estimated signal, u [k] is the dual variable u, x [0] is the waveform of the post-clip signal, v [0] is x [0] 's time-frequency representation, u [0] when it is an arbitrary number, The first estimator of the audio signal estimator generates a waveform to be constrained using the second estimated signal v [k] and the dual variable u [k] as inputs, and applies a projection operator ΠΓ for constraining the generated signal to a region included in the set Γ to the waveform to be constrained, thereby generating a first estimated signal x [k+1] which is a new waveform, The second estimation unit of the acoustic signal estimation device converts the first estimated signal x [k+1] into a time-frequency representation, and uses the first estimated signal x [k+1] in this time-frequency representation and the dual variable u [k] as inputs to perform soft threshold processing using a deep neural network, and generates a second estimated signal v [k+1] which is a signal in a new time-frequency representation to which a sparse optimization method is applied, The variable update unit of the acoustic signal estimation device updates the dual variable u [k] and the first estimation signal x converted into the time-frequency representation [k+1] and the second estimation signal v [k+1] as inputs, and generates a new dual variable u [k+1] and When the output unit of the acoustic signal estimation device determines that the execution count k is less than K - 1, it increments k by 1 and causes each of the first estimation unit, the second estimation unit, and the variable update unit to perform their respective processes. When the execution count k is K - 1 or more, it outputs the generated first estimated signal x [K] as the estimation result of the waveform of the pre-clipping signal ~y An acoustic signal estimation method.
7. A program for causing a computer to function as the acoustic signal estimation device according to any one of claims 1 to 5.
Citation Information
Patent Citations
Restoring device, restoring method, and program
JP2020106713A
Method and apparatus for determining a depth filter
JP2022529912A