Sound signal enhancement device, method, and program
The sound signal enhancement device and method improve estimation accuracy by employing state regularization weights to update separation and switch weights, addressing the challenge of limited time frames in conventional methods.
Patent Information
- Application Number
- JP2024115686
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Conventional sound signal enhancement methods face accuracy issues when the number of allocated time frames is small, leading to deteriorated estimation performance.
A sound signal enhancement device and method that utilizes a separation matrix application unit and a switch unit, with state regularization weights to update separation and switch weights, enhancing sound signals even in conditions with a limited number of time frames.
The method achieves high estimation accuracy for sound signals by using state regularization weights, ensuring robust performance even when the number of allocated time frames is small.
Smart Images

Figure 2026014538000001_ABST
Abstract
Description
[Technical Field]
[0001] The disclosed technology relates to a technology for acquiring a signal in which a target sound is emphasized from an observed signal in which the target sound and noise are mixed. [Background technology]
[0002] A method is known in which, for each frequency, the time frame of the recorded sound is classified into a plurality of states, a different separation matrix is estimated for each state, and a switch mechanism is used to switch and apply the separation matrix for each time frame, thereby estimating the target sound with high accuracy (see, for example, Non-Patent Document 1).
[0003] In this method, under the condition that the number of target sounds is known, time frames are classified based on the criterion of maximizing the likelihood of the recorded sounds at each time, and each separation matrix is optimized to enhance the acoustic signal. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] T. Nakatani, R. Ikeshita, K. Kinoshita, H. Sawada, N. Kamo, S. Araki, “Switching independent vector extraction and its joint optimization with weighted prediction error dereverberation”, in Proc. International Congress on Acoustics (ICA), 2022. Summary of the Invention [Problem to be solved by the invention]
[0005] However, in the conventional method, when classifying time frames, if there is a state in which the number of assigned time frames is small, the estimation accuracy may deteriorate.
[0006] The disclosed technology aims to provide a sound signal enhancement device, method, and program that can enhance a sound signal with high estimation accuracy even when there is a state in which the number of allocated time frames is small. [Means for solving the problem]
[0007] One aspect of the disclosed technology is a sound signal enhancement device that acquires a signal in which a target sound is enhanced from an observed signal in which the target sound and noise are mixed, and includes a separation matrix application unit that applies a plurality of separation matrices to the observed signal to acquire a plurality of estimated signals, and a switch unit that acquires a signal in which the target sound is enhanced using one of the plurality of estimated signals. When the plurality of separation matrices are updated using the observed signal, a state regularization weight is further used, and the state regularization weight is used to update the switch weight used by the switch. [Effects of the Invention]
[0008] According to the disclosed technology, by using a separation matrix updated using state regularization weights, it is possible to enhance sound signals with high estimation accuracy even in the presence of a state in which the number of allocated time frames is small. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of a functional configuration of a sound signal enhancing device. [Figure 2] FIG. 2 is a diagram showing an example of a processing procedure of the sound signal enhancement method. [Figure 3] FIG. 3 is a diagram illustrating an example of a functional configuration of a sound signal enhancing device according to the first modification. [Figure 4] FIG. 4 is a diagram illustrating an example of a functional configuration of a sound signal enhancing device according to the second modification. [Figure 5] FIG. 5 is a diagram illustrating an example of a functional configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the disclosed technology will be described with reference to the drawings. Note that components having the same functions in the drawings are given the same reference numerals, and redundant description will be omitted.
[0011] [Sound signal enhancement device and method] As shown in FIG. 1, the sound signal enhancement device includes, for example, an initialization unit 1, a separating matrix update unit 2, a separating matrix application unit 3, and a switch unit 4.
[0012] The sound signal enhancement method is realized, for example, by each component of the sound signal enhancement device performing the processes of steps S1 to S4 shown in FIG.
[0013] The symbol "^" used in a sentence should normally be written directly above the character immediately following it, but due to limitations in text notation, it is written immediately before the character in question. In mathematical formulas, these symbols are written in their proper position, i.e., directly above the character. For example, "^X" in a sentence is written as follows in a mathematical formula:
number
[0014] M is the number of microphones, and is a predetermined integer equal to or greater than 2.
[0015] m is the microphone number, where 1≦m≦M.
[0016] N is the number of target sounds.
[0017] n is the number of the target sound, where 1≦n≦N.
[0018] t is the number of the time frame, where t is 1≦t≦T, and T is a predetermined integer equal to or greater than 2.
[0019] f is a frequency number, where 1≦f≦F, and F is a predetermined integer equal to or greater than 2.
[0020] (·) T , (·) H denote the non-conjugate transpose and the conjugate transpose of a matrix, respectively.
[0021] {x j} j is the x for all numbers j j is a set of.
[0022] C M×M is the set of all M×M complex matrices. X∈C M×M ,X is C M×M represents an element of .
[0023] J is the number of states when classifying the time frame. It is also called the number of switch states. J is a predetermined integer equal to or greater than 2.
[0024] j is the number of the state when classifying the time frame, where 1≦j≦J.
[0025] x m (f,t)∈C 1×1 is the sound recorded by the mth microphone.
[0026] x(f,t)=[x 1 (f,t),x 2 (f,t),…,x M (f,t)] T ∈C M×1 is a vector summarizing the sounds recorded by all microphones at (time, frequency) = (f, t). The recorded sounds are sometimes called "observed signals."
[0027] ||x||2 2 is the power x of the column vector x H It is x.
[0028] ||x|| A 2 is the power x of the column vector x weighted by the matrix A. H It is Ax.
[0029] ^sj,n (f,t)∈C 1×1 is an estimate of the n-th target sound in state j. The target sound to be emphasized is a sound signal such as a speech signal or an acoustic signal.
[0030] ^s j (f,t)=[^s j,1 (f,t),,…,^s j,N (f,t)] T ∈C N×1 " is a vector summarizing the estimates of all target sounds at the time-frequency point (f, t) and state j. j (f,t) is called the "estimated signal."
[0031] ^s(f,t)∈C N×1 is the estimate of all target sounds.
[0032] z j (f,t)∈C (M-N)×1 is the noise estimate at time-frequency point (f, t), state j.
[0033] y j (f,t)=[^s j (f,t), z j (f,t)]∈C M×1 is the auxiliary estimated sound obtained at the time-frequency point (f, t) and state j.
[0034] v n (f, t) is the power (scalar) of the target sound n at the time-frequency point (f, t).
[0035] δ j (f,t) is the switch weight at time-frequency point (f,t), state j. ^s(f,t)=Σ j=1 J δ j (f,t)^s j (f,t).
[0036] Pi j (f)∈C M×M is the covariance of the recorded sound at frequency f and state j.
[0037] W j (f)∈C M×M is the separation matrix at frequency f and state j. y j (t,f)=W j H (f)x(f,t).
[0038] w j,n (f)∈C M×1 is W j It is a column vector in the nth (≦N) column of (f) and corresponds to a filter that estimates the nth target sound in state j. j,n (f,t)=w j,n H (f)x(f,t).
[0039] W j,S (f)∈C M×N is W j (f) is a submatrix of the first N columns, and w j,n (f) is a matrix summarizing all target sounds n (1≦n≦N). j (f,t)=W j,n H (f)x(f,t).
[0040] W j,Z (f)∈C M×(M-N) is W j The last MN columns of (f) are submatrices that estimate the noise at state j. j (f,t)=W j,Z H (f)x(f,t).
[0041] W j (f)=[W j,S (f), W j,Z (f)] T and W j,S (f)=[w j,1 (f),…,w j,N (f)] T is.
[0042] Next, each component of the sound signal enhancing device will be described.
[0043] <Initialization section 1> The initialization unit 1 initializes variables and parameters as shown in the following example (step S1).
[0044] The initialization unit 1 sets the state regularization weight λ state is set to a predetermined constant. An example of the predetermined constant is a predetermined real number such as 0.1.
[0045] The initialization unit 1 temporarily sets the number of states to 1, applies conventional blind source extraction (see, for example, Reference 1), and calculates the power v of each target sound using the obtained results. n (f,t), separation matrix W j (f) is initialized. The separation matrix W j (f) is assumed to be common to all j.
[0046] [Reference 1] R. Ikeshita, T. Nakatani, S. Araki, “Block coordinate descent algorithms for auxiliary-function-based independent vector extraction”, IEEE Trans. Signal Processing, 69, 3252-3267, 2021. The initialization unit 1 initializes all switch weights δ j Initialize (f,t) with random numbers.
[0047] Thereafter, the processes of the separating matrix update unit 2, separating matrix application unit 3, and switch unit 4 are repeated until a predetermined condition is met. An example of the predetermined condition is whether the number of repetitions reaches a predetermined fixed number.
[0048] <Separation matrix update section 2> Separation matrix update unit 2 is T j (f)=Σ t=1 T (δ j (f,t)+λ state ) and each separation matrix W j (f) Column vector w from 1 to N j,n (f)∈CM×1 (corresponding to a filter for estimating the n-th target sound) is updated by the following equation (step S2).
number
number
number
number
[0049] In addition, the separation matrix update unit 2 calculates the recorded sound covariance Π j (f) and the separation matrix W j (f) Submatrix W from column N+1 to column M j,Z (f)∈C M×(N+1) (Noise z j (corresponding to a matrix for estimating f, t) is updated using the following formula (step S2).
number
number
[0050] Separation matrix W updated by separation matrix update unit 2 j (f) is output to the separation matrix application unit 3 and the switch unit 4. The recorded sound covariance Π for each state j j (f) is output to the switch unit 4.
[0051] <Separation matrix application part 3> The separation matrix application unit 3 generates J (>1) separation matrices W j (f)∈C M×M and recorded sound x(f,t)∈C M×1 Receive J auxiliary estimated sounds y j (f,t)∈C M×1 (1≦j≦J) is calculated using the following formula: j (f, t) is output to the switch unit 4.
number
[0052] In this way, the separation matrix application unit 3 applies a plurality of separation matrices W j (f) is applied to generate multiple auxiliary estimated sounds y j (f, t) is obtained (step S3).
[0053] <Switch section 4> The switch unit 4 outputs a plurality of estimated signals ^s j (f, t) at each time-frequency point (f, t), the switch unit 4 acquires a signal ^s(f, t) in which the target sound is emphasized using one of the J auxiliary estimated sounds y j (f,t) and J switch weights δ j Using (f, t), the target sound ^s(f, t) is calculated using the following formula.
number
[0054] In addition, the switch unit 4 calculates the noise covariance Ω for all j. j (f) is updated using the following formula:
number
number
number
[0055] The processes of the separation matrix update unit 2, the separation matrix application unit 3, and the switch unit 4 are repeated until a predetermined condition is satisfied, and then the switch unit 4 outputs the most recently determined target sound ^s(f,t), in other words, the estimated value ^s(f,t) of the target sound.
[0056] X={x(f,t)} f,tis the set of all recorded sounds at all time-frequency points. In the conventional method, the negative logarithm likelihood function criterion L NL The following optimization criterion C is defined based on (X;Θ) conv The parameter Θ was estimated by minimizing (Θ).
number
number
number
[0057] However, at each frequency f, the switch weight δ j After fixing (f, t), estimating the j-th separation matrix based on the above optimization criterion is δ j Only for the time frame (f,t)=1, minimize θ j This corresponds to finding (f,t).
[0058] Therefore, δ jIn the state j where the number of time frames (f, t)=1 is small, the number of time frames available for parameter estimation also decreases, resulting in a decrease in the estimation accuracy of the separation matrix. j The estimation accuracy of (f,t) may decrease, and as a result, the estimation accuracy of the target sound ^s(f,t) may decrease.
[0059] Therefore, in the above embodiment, the negative logarithmic likelihood function criterion L NL (X;Θ), with state regularization weight λ state We use the following optimization criterion, which adds a state regularization term weighted by
number
[0060] By using the above optimization criteria, at each frequency f, the switch weight δ j Even when estimating the j-th separation matrix with (f,t) fixed, we can minimize θ that minimizes equation (6) for all time frames. j This means finding (f,t).
[0061] As a result, δ j Even in state j where the number of time frames is small and (f,t)=1, all time frames can be used for parameter estimation, so the estimation accuracy of the separation matrix does not decrease, and high estimation accuracy of the target sound ^s(f,t) can be achieved.
[0062] In addition, the function L(θ j (f,t),v(f,t)) corresponds to the standard for estimating parameters using all time frames without using a switching mechanism. Therefore, introducing the state regularization term has the effect of bringing the parameter estimates, especially in state j with a small number of time frames, closer to the parameter estimates estimated without using a switching mechanism.
[0063] Also, the state regularization weight λstate 0<λ state By setting a value <1, even when state regularization is performed, δ j The parameter θ is set to the time frame where (f,t)=1. j Since (f, t) can be estimated, it is possible to estimate a separating matrix suitable for each state j.
[0064] [Variation 1] The switch 4 is configured to j (f) is excluded, and the switch weight δ j (f,t) may be updated.
[0065] Specifically, the switch unit 4 determines the noise covariance Ω j In order to avoid the influence of the estimation error of (f), Ω j Using the following update formula, which excludes terms including (f), the switch weight δ j (f,t) may be updated.
number
[0066] In this case, the separation matrix application unit 3 calculates the auxiliary estimated sound y j (f,t) is calculated using the following formula, and ^s j Only the (f,t) part needs to be included.
number
[0067] Other than these points, the processing of the first modification is the same as the processing described above.
[0068] [Variation 2] In the second modification, in order to be able to distinguish between the target sound and sounds other than the target sound, a steering vector a is used, which is an estimate of the acoustic transfer function from the sound source of each target sound n to all microphones. n (f)∈C M×1 is given.
[0069] In this case, the optimization criterion C state In addition to (Θ), we have a spatial regularization term R(θ j (f,t);{a n (f)} n ) as the spatial regularization weight λ space The following optimization criteria appended with (>0) may be used:
number
[0070] As the spatial regularization term, various methods such as the method shown in Reference 2 can be used.
[0071] [Reference 2] T. Ueda, T. Nakatani, R. Ikeshita, K. Kinoshita, S. Araki, S. Makino, “Blind and spatially-regularized online joint optimization of source separation, dereverberation, and noise reduction”, IEEE / ACM Trans. Audio, Speech, and Language Processing, 32, 1157-1172, 2024. In this example, the noise estimate z j A method of using a null regularization term to guide (f, t) so that it does not include the target sound will be described. The null regularization term is defined, for example, as follows:
number
[0072] Then, the separation matrix update unit 2 updates a plurality of separation matrices so that the target sound is estimated in accordance with the target sound direction indicated by the input steering vector.
[0073] For example, in this way, the plurality of separation matrices may be updated so that the target sound is estimated in accordance with the target sound direction indicated by the input steering vector.
[0074] When non-stationary noise is included in the recorded sound, it is not possible to distinguish which sound is the target sound, and there is a possibility that sounds other than the target sound will be separated as the target sound. By updating a plurality of separation matrices so that the target sound is estimated corresponding to the target sound direction indicated by the steering vector, it is possible to estimate the desired target sound corresponding to the steering vector.
[0075] In the second modification, the initialization unit 1 performs the following process, for example.
[0076] The initialization unit 1 sets the state regularization weight λ state instead of just the spatial regularization weight λ space is set to a predetermined constant. An example of the predetermined constant is a predetermined real number such as 0.1. state and the initial value of the spatial regularization weight λ space The initial values of may be the same or different.
[0077] The initialization unit 1 calculates the power v of each target sound. n (f,t) is the power of the recorded sound ||x(f,t)||2 2 First, the number of states is set to J=1, and W1(f,t)=IM Initialize with.
[0078] First, the initialization unit 1 performs the iterative process of the second modification below with the number of states J=1, and calculates the power v of each target sound. n (f, t) and the separation matrix W1(f, t) are updated.
[0079] The initialization unit 1 returns the number of states J to the original number and sets the separation matrix W for each state j. n The separation matrix W1(f,t) obtained by updating is copied to (f).
[0080] The initialization unit 1 initializes all switch weights δ j Initialize (f,t) with random numbers. As shown in FIG. 4, the separating matrix update unit 2 includes a steering vector a n (f)∈C M×1 The separation matrix update unit 2 receives the recorded sound covariance Π j (f) is updated using the following formula:
number
[0081] [Experimental Example] We compared the performance of the conventional method using only spatial regularization with that of variant 2 for audio recordings of three people speaking simultaneously in a noisy, reverberant environment, recorded with two microphones.
[0082] The number of switch states was set to J = 3, and the steering vector of only one speaker was assumed to be known. The target sound was extracted by assuming that only the speaker whose steering vector was known was the target sound, and that the other speakers were included in the noise.
[0083] The improvement in signal-to-distortion ratio (SDR) is:
[0084] Conventional method: 5.57 dB Variation 2: 6.45 dB It can be seen that the SDR is improved compared to when only spatial regularization is used in the conventional method.
[0085] The specific configurations of the embodiments of the disclosed technology are not limited to those described above, and the specific configurations of the embodiments of the disclosed technology can be appropriately modified in design, etc., within the scope of the spirit of the embodiments of the disclosed technology.
[0086] The various processes described in the embodiments of the disclosed technology may not only be performed chronologically in the order described, but may also be performed in parallel or individually depending on the processing capacity of the device performing the processes or as needed.
[0087] For example, data may be exchanged directly between the components of the sound signal enhancing device, or may be exchanged via a storage unit (not shown).
[0088] Furthermore, a device (terminal) for using the device, system, or method of the present invention via a network (telecommunications line) may also be provided. The "device (terminal) for use" may be provided with functions (e.g., control function, decoding function, restoration function, input / output function, etc.) necessary to obtain the effects of implementing the device, system, or method of the present invention.
[0089] It goes without saying that other modifications are possible without departing from the spirit of the present invention.
[0090] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[0091] [Programs, recording media] The functions performed by the components described herein may be implemented in circuitry or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), a CPU (a Central Processing Unit), conventional circuits, and / or combinations thereof, programmed to perform the described functions. A processor includes transistors and other circuits and is considered to be circuitry or processing circuitry. A processor may also be a programmed processor that executes programs stored in memory.
[0092] In this specification, a circuitry, unit, or means is hardware that is programmed to realize or performs the described functions, which may be any hardware disclosed herein or any hardware known to be programmed to realize or perform the described functions.
[0093] If the hardware is a processor considered to be a type of circuitry, the circuitry, means, or unit is a combination of the hardware and software used to configure the hardware and / or processor.
[0094] The various processes described above can be implemented by loading a program that executes each step of the above method into the recording unit 2020 of the computer 2000 shown in Figure 5 and operating the control unit 2010, input unit 2030, output unit 2040, display unit 2050, etc.
[0095] The program describing the processing contents can be recorded on a computer-readable recording medium, which may be, for example, a magnetic recording device, an optical disk, a magneto-optical recording medium, a semiconductor memory, or any other suitable recording medium.
[0096] The program may be distributed, for example, by selling, transferring, lending, etc. a portable recording medium such as a DVD or CD-ROM on which the program is recorded. Furthermore, the program may be stored in a storage device of a server computer, and then transferred from the server computer to another computer via a network, thereby distributing the program.
[0097] A computer that executes such a program may first temporarily store the program recorded on a portable recording medium or transferred from a server computer in its own storage device. Then, when executing a process, the computer reads the program stored on its own recording medium and executes the process in accordance with the read program. Alternatively, the computer may read the program directly from a portable recording medium and execute the process in accordance with the program. Furthermore, the computer may execute the process in accordance with the program each time a program is transferred from a server computer to the computer. The server computer may not transfer the program to the computer, but may instead execute the process through a so-called ASP (Application Service Provider) service, which realizes the processing function by issuing an execution instruction and obtaining the results. Furthermore, the server computer may execute the process on a terminal using a so-called SaaS (Software as a Service) service, which allows users to use part of the server computer along with the program. In this embodiment, the program includes information used for computer processing that is equivalent to a program (such as data that is not a direct instruction to the computer but has properties that define computer processing).
[0098] Furthermore, in this embodiment, the device is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.
Claims
1. A sound signal enhancement device that acquires a signal in which a target sound is enhanced from an observation signal in which the target sound and noise are mixed, a separation matrix application unit that applies a plurality of separation matrices to the observed signals to obtain a plurality of estimated signals; a switch unit that acquires a signal in which the target sound is emphasized by using any one of the plurality of estimated signals, a state regularization weight is further used when the plurality of separation matrices are updated using the observation signals; The state regularization weight is used to update the switch weight used by the switch. Sound signal enhancement device.
2. 2. The sound signal enhancement device of claim 1, the switch unit updates the switch weights while excluding the influence of noise covariance. Sound signal enhancement device.
3. The sound signal enhancement device according to claim 1 or 2, The plurality of separation matrices are updated so that the target sound is estimated corresponding to the target sound direction indicated by the input steering vector. Sound signal enhancement device.
4. A sound signal enhancement method for acquiring a signal in which a target sound is enhanced from an observed signal in which the target sound and noise are mixed, comprising: a separation matrix application step unit in which a separation matrix application unit applies a plurality of separation matrices to the observed signals to obtain a plurality of estimated signals; a switching step of a switch unit acquiring a signal in which the target sound is emphasized by using any one of the plurality of estimated signals, a state regularization weight is further used when the plurality of separation matrices are updated using the observation signals; The state regularization weight is used to update the switch weight used by the switch. Sound signal enhancement method.
5. A program for causing a computer to execute each step of the sound signal enhancement method of claim 4.