A kalman adaptive based array microphone noise reduction method and device
By employing a Kalman adaptive array microphone noise reduction method, the input signal is filtered multiple times using a super-pointing filter, a beamforming filter, and a Kalman filter model. This solves the problem of interference noise affecting the call quality of headphones in open office scenarios, achieving efficient noise cancellation and improved voice purity.
Patent Information
- Application Number
- CN202310304623.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-27
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-03-27
AI Technical Summary
In open office settings, external noises such as keyboard sounds, typing, and talking can affect call quality. In particular, when there are interfering human voices, reducing background noise and improving call quality for headset wearers is an urgent problem to be solved.
A Kalman-adaptive array microphone noise reduction method is adopted. The input signal is filtered multiple times through a super-pointing filter, a beamforming filter, and a Kalman filter model to generate the final output signal to eliminate interference noise and improve speech purity.
It effectively eliminates interference noise, improves call quality, ensures undistorted voice, adapts to changes in headphone wearing methods and noise environments, and does not introduce adverse effects such as white noise gain.
Smart Images

Figure CN116320857B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of speech enhancement, in particular to an array microphone noise reduction method and device based on Kalman self-adaptation. BACKGROUND
[0002] In a common open office scenario, when people are talking on the phone while wearing earphones, the interference sound such as keyboard sound, knocking sound and speaking sound from the outside world will affect the call quality, especially when there are other interfering human voices around the earphone wearer, the call quality will be greatly affected, therefore, how to reduce the background noise and interfering human voices from the outside world, that is, how to reduce the interference noise and improve the call quality of the earphone wearer is a problem that needs to be solved urgently. SUMMARY
[0003] The embodiment of the present application provides an array microphone noise reduction method and device based on Kalman self-adaptation, which can improve the purity of the call voice.
[0004] An embodiment of the present application provides an array microphone noise reduction method based on Kalman self-adaptation, comprising:
[0005] obtaining input signals at each time point; wherein the input signal at each time point contains target voice and interference noise;
[0006] establishing a hyperdirectivity filter model, and then filtering the input signals at each time point according to the hyperdirectivity filter model to generate first reference signals at each time point;
[0007] establishing a beamforming filter model, and then filtering the input signals at each time point according to the beamforming filter model to generate second reference signals at each time point;
[0008] establishing a Kalman filter model and process equations and measurement equations corresponding to the Kalman filter model at each time point; generating Kalman gains at each time point according to errors corresponding to the process equations and errors corresponding to the measurement equations at each time point, so that the Kalman filter model eliminates interference noise in the first reference signals and the second reference signals at each time point according to the Kalman gains at each time point to generate final output signals at each time point.
[0009] Further, the establishment of the hyperdirectivity filter model and the filtering of the input signals at each time point according to the hyperdirectivity filter model to generate the first reference signals at each time point comprises:
[0010] generating a relative transfer function of the corresponding input signal according to the input signal at each time point;
[0011] For each time input signal, according to the input signal corresponding to the relative transfer function and pseudo coherence matrix, according to the input signal relative transfer function and pseudo coherence matrix to establish the super directivity filter model, according to the super directivity filter model for each time input signal filtering, generate corresponding first reference signal.
[0012] Further, the establishment of beam forming filter model, in turn, according to the beam forming filter model for each time input signal filtering, generate each time second reference signal, including:
[0013] The beam forming filter model is projected in the null space, and then the corresponding blocking matrix is generated;
[0014] According to the blocking matrix for each time input signal filtering, generate each time second reference signal.
[0015] Further, the process equation corresponding to each time of the Kalman filter model is established, including:
[0016] The process equation corresponding to each time of the Kalman filter model is established by the following formula:
[0017] w sc (l)= H (l)w sc (l-1)+△w(l)
[0018] Wherein, sc () represents the sidelobe cancellation filter model in the Kalman filter model at l time, A represents the state equation, H represents the conjugate transpose symbol, and △w(l) represents the error of the process equation at l time.
[0019] Further, the measurement equation corresponding to each time of the Kalman filter model is established, including;
[0020] The measurement equation corresponding to each time of the Kalman filter model is established according to the first reference signal and the second reference signal at each time; wherein the measurement equation corresponding to each time of the Kalman filter model is established by the following formula:
[0021] x bf (l)= bm H (l)w sc (l)+△s(l)
[0022] Wherein, x bf (l) represents the first reference signal at l time, x bm H(l) represents a conjugate transpose matrix of the second reference signal at the l time, H represents a conjugate transpose symbol, and △s(l) represents an error of the measurement equation at the l time.
[0023] Further, the Kalman gain at each time is generated according to the error corresponding to the process equation and the error corresponding to the measurement equation at each time, and the Kalman gain at each time comprises:
[0024] The error covariance matrix of the process equation at each time is generated according to the error corresponding to the process equation at each time;
[0025] The error covariance matrix of the measurement equation at each time is generated according to the error corresponding to the measurement equation at each time;
[0026] The Kalman gain of the Kalman filter model at each time is generated according to the error covariance matrix of the process equation and the error covariance matrix of the measurement equation at each time.
[0027] Further, the Kalman filter model eliminates the interference noise in the first reference signal and the second reference signal at each time according to the Kalman gain at each time, and the elimination of the interference noise in the first reference signal and the second reference signal at each time comprises:
[0028] The noise field of the interference noise at each time is eliminated by the Kalman gain at each time; wherein the interference noise estimation at each time comprises:
[0029] When the Kalman gain is approximately zero, the noise field of the eliminated interference noise is approximately the noise field filtered by the sidelobe cancellation filter model in the process equation;
[0030] When the Kalman gain is approximately one, the noise field of the eliminated interference noise is approximately the noise field estimated by the measurement equation.
[0031] Further, the generation of the final output signal at each time comprises:
[0032] The final output signal at each time is generated by the following formula:
[0033] e(l) = y(l) - H(l) x(l) - △s(l) bf (l)- bm H (l)w sc (l)
[0034] Wherein, e(l) represents the final output signal at the l time.
[0035] Further, after obtaining the input signal at each time, the method further comprises:
[0036] The obtained input signal at each time is de-reverberated by using a time domain deconvolution method.
[0037] Based on the above method embodiment, the application provides a device embodiment;
[0038] An embodiment of the application provides an array microphone noise reduction device based on Kalman self-adaption, which comprises a signal acquisition module, a first reference signal generation module, a second reference signal generation module and an interference signal elimination module.
[0039] The signal acquisition module is used for acquiring input signals at each moment, wherein the input signals at each moment contain target speech and interference noise.
[0040] The first reference signal generation module is used for establishing a hyperdirectivity filter model, and then filtering the input signals at each moment according to the hyperdirectivity filter model to generate first reference signals at each moment.
[0041] The second reference signal generation module is used for establishing a beamforming filter model, and then filtering the input signals at each moment according to the beamforming filter model to generate second reference signals at each moment.
[0042] The signal output module is used for establishing a Kalman filter model, process equations and measurement equations corresponding to the Kalman filter model at each moment; generating Kalman gains at each moment according to errors corresponding to the process equations at each moment and errors corresponding to the measurement equations; and making the Kalman filter model eliminate interference noise in the first reference signals and the second reference signals at each moment according to the Kalman gains at each moment to generate final output signals at each moment.
[0043] The application has the following beneficial effects:
[0044] The application provides a Kalman adaptive array microphone noise reduction method and device. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 FIG. 1 is a flowchart of a Kalman adaptive array microphone noise reduction method according to an embodiment of the application.
[0046] Figure 2 FIG. 2 is a schematic diagram of the relationship between a microphone and a noise source according to an embodiment of the application.
[0047] Figure 3 FIG. 3 is a schematic diagram of a Kalman adaptive array microphone noise reduction device according to an embodiment of the application. DETAILED DESCRIPTION
[0048] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the application.
[0049] As shown in FIG. 1, an embodiment of the application provides a Kalman adaptive array microphone noise reduction method, which comprises: Figure 1
[0050] Step S1: obtaining input signals at each time point; wherein the input signal at each time point contains target speech and interference noise;
[0051] Step S2: establishing a super-directive filter model, and then filtering the input signals at each time point according to the super-directive filter model to generate first reference signals at each time point;
[0052] Step S3: establishing a beamforming filter model, and then filtering the input signals at each time point according to the beamforming filter model to generate second reference signals at each time point;
[0053] Step S4: establishing a Kalman filter model and process equations and measurement equations corresponding to the Kalman filter model at each time point; generating Kalman gains at each time point according to errors corresponding to the process equations and errors corresponding to the measurement equations at each time point, so that the Kalman filter model eliminates interference noise in the first reference signals and the second reference signals at each time point according to the Kalman gains at each time point to generate final output signals at each time point.
[0054] For step S1, specifically, the input signals at each time point are obtained by an array microphone, and the obtained input signals are mixed signals containing target speech and external interference noise; wherein the array microphone contains multiple microphones for obtaining input signals, that is, the array microphone is composed of multiple microphones; for example: the current array microphone is a dual-microphone, and the dual-channel mixed signals obtained by the dual-microphone earphone can be expressed as:
[0055] x(t) = [x1(t), x2(t)] T
[0056] Wherein, x(t) represents the input signal picked up by the dual-microphone earphone, x1(t) represents the first input signal picked up by the dual-microphone earphone, x2(t) represents the second input signal picked up by the dual-microphone earphone, and T is the transpose symbol;
[0057] The above obtained input signals of the dual-microphone earphone can be expressed as:
[0058]
[0059] Wherein, J represents the number of sound sources picked up by the microphone, j represents the jth sound source, c j () represents the reception of the jth sound source by the microphone;
[0060] Currently, the corresponding c j (t) = [c 1j (t), 2j ()] Twherein c 1j (t) represents the first microphone of the double microphone earphone receiving the jth sound source, c 2j () represents the second microphone of the double microphone earphone receiving the jth sound source.
[0061] In a preferred embodiment, after obtaining the input signal at each time, it further comprises: using the time domain deconvolution method to de-reverberate the obtained input signal at each time;
[0062] Specifically, before the subsequent step operation of the obtained input signal, the reverberation of the input signal is eliminated by using a conventional time domain deconvolution method; wherein the conventional time domain deconvolution method usually uses a multi-channel linear prediction algorithm (MCLP) or a weighted prediction error algorithm (WPE), but in actual application process, it is not limited to the above two time domain deconvolution methods; eliminating the reverberation of the input signal can improve the accuracy of subsequent transfer function calculation and noise estimation.
[0063] For step S2, that is, establishing a super directivity filter model, and then filtering the input signal at each time according to the established super directivity filter model to generate a first reference signal at each time;
[0064] In a preferred embodiment, the establishment of the super directivity filter model and the filtering of the input signal at each time according to the super directivity filter model to generate a first reference signal at each time comprises: generating a relative transfer function corresponding to the input signal according to the input signal at each time; for the input signal at each time, generating a corresponding relative transfer function and pseudo-coherence matrix according to the input signal, establishing a super directivity filter model according to the relative transfer function and the pseudo-coherence matrix of the input signal, filtering the input signal at each time according to the super directivity filter model to generate a corresponding first reference signal;
[0065] Specifically, for the input signal at each time, a relative transfer function of the input signal to the microphone array is generated according to the input signal, which is related to the spatial position of the input signal, and the relative transfer function can be generated by the following formula:
[0066]
[0067] wherein a j represents the relative transfer function of the jth sound source to the microphone array, f represents the frequency domain unit, n represents the time frame number, a j (n, f) represents the relative transfer function of the jth sound source to the microphone array at time frame number n, i is the imaginary part, d represents the microphone spacing of the array microphone, c represents the sound velocity of the input signal, and θ represents the incidence angle of the speech incident on the microphone array.
[0068] It should be noted that the input signal contains multiple sound sources, which can be interference noise or target speech, such as Figure 2 As shown in the figure, the present application provides a schematic diagram of the relationship between the microphone and the noise source, which takes a double microphone array as an example to show the relationship between the noise source position and the microphone when the microphone is placed. Noise Source in the figure represents the noise source, which includes environmental noise and surrounding interference human voice. Array Microphone represents the array microphone arranged in front of the microphone. Head represents the earphone containing the array microphone. θ represents the incidence angle of the speech incident on the microphone array. This figure is only for understanding and example, and the number of microphones and their layout in actual application are not limited to this;
[0069] The relationship between the microphone input signal and the relative transfer function is as follows:
[0070] c j (t)= j (n,f)s j (n,f)
[0071] Where s j represents the jth sound source; the sound source can be target speech or interference noise;
[0072] According to the input signal, the corresponding pseudo-coherent matrix is generated, that is, the mean operation is performed on the signals obtained by the microphone array. The corresponding pseudo-coherent matrix can be generated by the following formula:
[0073] γ=E(X,X)
[0074] Where γ represents the pseudo-coherent matrix, X represents the signal obtained by the microphone array, and E(X,X) represents the mean operation on the obtained microphone array signal;
[0075] According to the relative transfer function of the input signal and the pseudo-coherent matrix, the super-directive filter model is established, and the corresponding super-directive filter model can be generated by the following formula:
[0076] h=γ -1 a T [ T H γ -1 a T ] -1
[0077] Where a T represents the relative transfer function of the target direction sound source to the microphone array, H is the conjugate transpose symbol, γ represents the pseudo-coherent matrix, and h represents the generated super-directive filter model;
[0078] It should be noted that in the implementation process, gamma in the above formula can represent a pseudo-coherent matrix, or a noise field model assumed in advance;
[0079] The input signal is filtered by the hyperdirectivity filter model generated above to output a corresponding first reference signal.
[0080] It should be noted that due to the change of the wearing angle of the earphone in actual use, the incidence angle of the speech incident microphone changes, and the sound path from the mouth to the earphone does not meet the far-field requirement, etc., which leads to inaccurate relative transfer function calculated by the above-mentioned geometric information, affecting the subsequent noise reduction effect. The method of real-time estimation of the relative transfer function can be used to replace the calculation of the relative transfer function, such as frame estimation according to the direction of arrival (DOA), or minimum binary estimation of the inter-channel mutual power spectral density, but not limited to the above-mentioned methods.
[0081] For step S3, i.e. establishing a beamforming filter model to filter the input signal to generate a second reference signal; in a preferred embodiment, the beamforming filter model is established, and then the input signal at each time is filtered according to the beamforming filter model to generate the second reference signal at each time, which comprises: projecting the beamforming filter model into the null space to generate a corresponding blocking matrix; filtering the input signal at each time according to the blocking matrix to generate the second reference signal at each time;
[0082] Specifically, the beamforming filter model is solved according to the beamforming filter model to satisfy the distortionless constraint condition of the target speech incident direction, i.e. the relative transfer function of the sound source transmission of the target direction to the microphone array of the beamforming filter model needs to satisfy the following formula:
[0083] g bf H a T =1
[0084] Wherein, g bf represents the beamforming filter model, a T represents the relative transfer function of the sound source transmission of the target direction to the microphone array, and H is the conjugate transpose symbol.
[0085] In the above formula, when the relative transfer function of the sound source transmission of the target direction to the microphone array is multiplied by the beamforming filter model, it means that the signal of the target direction sound source is not distorted;
[0086] The beamforming filter model is projected in a null space to generate a blocking matrix, the input signal is input into the blocking matrix generated by the beamforming filter model to block the target voice in the input signal and generate a second reference signal containing interference noise;
[0087] It should be noted that, in order to make the second reference signal as much as possible not to contain the target voice, and avoid mis-eliminating the target voice, when the blocking matrix is generated, the generated blocking matrix needs to be orthogonal to the relative transfer function;
[0088] For step S4, a Kalman filter model is established, a process equation and a measurement equation corresponding to each time are established through the established Kalman filter model, the error signals contained in the first reference signal and the second reference signal are transmitted back and forth in the Kalman filter model to minimize the error signals; the corresponding Kalman gain is generated according to the error corresponding to the process equation and the error corresponding to the measurement equation; and the interference noise in the first reference signal and the second reference signal is eliminated according to the generated Kalman gain;
[0089] In a preferred embodiment, the process equation corresponding to each time of the Kalman filter model is established, including:
[0090] w sc (l)= H (l)w sc (l-1)+△w(l)
[0091] Wherein, sc () represents a sidelobe cancellation filter model in the Kalman filter model at time l, A represents a state equation, H represents a conjugate transpose symbol, and A H (l) represents a conjugate transpose matrix of the state equation at time l, and △w(l) represents an error of the process equation at time l;
[0092] Specifically, the sidelobe cancellation filter model is a sidelobe cancellation filter model used for real-time estimation of a noise field and filtering out the noise field in a Kalman self-adaptive iteration process of the Kalman filter model;
[0093] In another preferred embodiment, the measurement equation corresponding to each time of the Kalman filter model is established, including:
[0094] The measurement equation corresponding to each time of the Kalman filter model is established according to the first reference signal and the second reference signal at each time; wherein the measurement equation corresponding to each time of the Kalman filter model is established through the following formula:
[0095] x bf(l) = x(l) - H(l) x(l) + △s(l) bm H (l) = x(l) - H(l) x(l) + △s(l) sc (l) = x(l) - H(l) x(l) + △s(l)
[0096] wherein, x bf (l) represents the first reference signal at the time l, x bm H (l) represents the conjugate transpose matrix of the second reference signal at the time l, H represents the conjugate transpose symbol, and △s(l) represents the error of the measurement equation at the time l;
[0097] In a preferred embodiment, the Kalman gain at each time is generated according to the error corresponding to the process equation and the error corresponding to the measurement equation at the time, including: generating the error covariance matrix of the process equation at the corresponding time according to the error corresponding to the process equation at each time; generating the error covariance matrix of the measurement equation at the corresponding time according to the error corresponding to the measurement equation at each time; and generating the Kalman gain of the Kalman filter model at each time according to the error covariance matrix of the process equation and the error covariance matrix of the measurement equation at each time;
[0098] Specifically, the error of the process equation and the error of the measurement equation both conform to Gaussian distribution, and the error covariance matrix of the process equation corresponding to the error of the process equation can be obtained, and the error covariance matrix of the measurement equation corresponding to the error of the measurement equation can be obtained;
[0099] The Kalman gain is obtained by the following formula:
[0100]
[0101] wherein, K(l) represents the Kalman gain at the time l, represents the error covariance matrix of the process equation, represents the error covariance matrix of the measurement equation;
[0102] In a preferred embodiment, the Kalman filter model eliminates the interference noise in the first reference signal and the second reference signal at each time according to the Kalman gain at each time, including: eliminating the noise field of the interference noise in the first reference signal and the second reference signal at each time through the Kalman gain at each time; wherein, for the interference noise estimation at each time, when the Kalman gain is approximately zero, the noise field of the eliminated interference noise is approximately the noise field filtered out by the sidelobe cancellation filter model in the process equation; and when the Kalman gain is approximately one, the noise field of the eliminated interference noise is approximately the noise field estimated by the measurement equation;
[0103] Specifically, the noise field is estimated by the following formula:
[0104] w sc (l)= sc (l)+K()( bf (l)- bm H (l)w sc (l))
[0105] wherein, w sc () represents the sidelobe cancellation filter model in the Kalman filter model at the time l, K(l) represents the Kalman gain at the time l; x bf (l)- bm H (l)w sc (l) that is △s(l) represents the error of the measurement equation at the time l;
[0106] In the above formula, when the error of the measurement equation is very large, the Kalman gain is approximately zero, at this time, the noise field of the interference noise to be cancelled is approximately the noise field estimated and filtered by the sidelobe cancellation filter model in the process equation; when the error of the process equation is very large, the Kalman gain is approximately one, at this time, the noise field of the interference noise to be cancelled is approximately the noise field estimated by the measurement equation;
[0107] After the noise field to be filtered is estimated in real time by the Kalman gain, the final error signal (i.e. the final output signal) is generated; in a preferred embodiment, the final output signal at each time is generated, including:
[0108] The final output signal at each time is generated by the following formula:
[0109] e(l)= bf (l)- bm H (l)w sc (l)
[0110] wherein, e(l) represents the final output signal at the time l;
[0111] It should be noted that in the process of estimating and filtering the noise field by the Kalman gain, the product of the covariance matrix of the error signal (i.e. the final output signal) is minimized, which means that a more accurate noise field can be estimated.
[0112] The above embodiments of the present application have the following beneficial effects by implementing the present application:
[0113] 1. Compared with traditional beamforming technology, the super-directional filter model of the present invention does not need to take a large directivity from the beginning. Instead, it improves the noise reduction effect through the adaptive process of the subsequent Kalman filter model. Therefore, implementing the above embodiments of the present invention can effectively ensure that the speech is not distorted while improving the noise reduction level, and will not introduce adverse effects such as white noise gain.
[0114] 2. It can narrow down the range of input voice signals that can be acquired;
[0115] 3. In actual use of headphones, the noise field can be adaptively estimated in real time in response to changes in headphone wearing method or noise scene, avoiding the impact of changes in headphone wearing method or noise scene on noise estimation.
[0116] 4. It can solve the defect that the noise reduction performance of the super-directional filter model is heavily dependent on the number of microphones. In the above embodiments of the present invention, the array microphone with dual microphones can also achieve a very good noise reduction effect.
[0117] Based on the above method embodiments, the present invention provides corresponding apparatus embodiments.
[0118] like Figure 3 As shown, an embodiment of the present invention provides a Kalman adaptive array microphone noise reduction device, including: a signal acquisition module, a first reference signal generation module, a second reference signal generation module, and an interference signal cancellation module;
[0119] The signal acquisition module is used to acquire the input signal at each time moment; wherein, the input signal at each time moment includes the target speech and interference noise;
[0120] The first reference signal generation module is used to establish a superpointing filter model, and then filter the input signal at each time step according to the superpointing filter model to generate the first reference signal at each time step.
[0121] The second reference signal generation module is used to establish a beamforming filter model, and then filter the input signal at each time step according to the beamforming filter model to generate the second reference signal at each time step.
[0122] The signal output module is used to establish a Kalman filter model and the process equations and measurement equations corresponding to the Kalman filter model at each time step; and to generate the Kalman gain at each time step based on the errors corresponding to the process equations and measurement equations at each time step, so that the Kalman filter model can eliminate the interference noise in the first reference signal and the second reference signal at each time step based on the Kalman gain at each time step, and generate the final output signal at each time step.
[0123] It should be noted that the apparatus embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. In addition, the connection between the modules in the apparatus embodiments provided by the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement it without creative labor.
[0124] Those skilled in the art can clearly understand that, for the convenience and brevity, the specific working process of the apparatus described above can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0125] The terminal device can be a desktop computer, a notebook computer, a palm computer, a cloud server and other computing devices. The terminal device can include, but is not limited to, a processor and a memory.
[0126] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, and connects all parts of the terminal device through various interfaces and lines.
[0127] The memory can be used to store the computer program, and the processor realizes various functions of the terminal device by running or executing the computer program stored in the memory and calling data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required by a function, and the like; and the data storage area can store data created according to the use of the mobile phone and the like. In addition, the memory can include a high-speed random access memory, and can also include a nonvolatile memory, for example, a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory devices.
[0128] The storage medium is a computer readable storage medium, and the computer program is stored in the computer readable storage medium. When the computer program is executed by the processor, the steps of each method embodiment described above can be realized. The computer program includes computer program code, which can be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer readable medium can include any entity or device, a recording medium, a U disk, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. that can carry the computer program code. It should be noted that the content included in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in a jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include an electrical carrier signal and a telecommunication signal.
[0129] The above is the preferred embodiment of the present application. It should be noted that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which are also considered within the scope of protection of the present application.
Claims
1. A Kalman adaptive based array microphone noise reduction method, characterized in that, The method comprises the following steps: acquiring input signals at each time point, wherein the input signal at each time point comprises target speech and interference noise; establishing a superdirectivity filter model, and then filtering the input signals at each time point according to the superdirectivity filter model to generate first reference signals at each time point; establishing a beamforming filter model, and then filtering the input signals at each time point according to the beamforming filter model to generate second reference signals at each time point; establishing a Kalman filter model and process equations and measurement equations corresponding to the Kalman filter model at each time point; generating Kalman gains at each time point according to errors corresponding to the process equations and errors corresponding to the measurement equations, so that the Kalman filter model eliminates interference noise in the first reference signals and the second reference signals at each time point according to the Kalman gains at each time point to generate final output signals at each time point.
2. The Kalman filter based adaptive array microphone noise reduction method of claim 1, wherein, The step of establishing the superdirectivity filter model and then filtering the input signals at each time point according to the superdirectivity filter model to generate the first reference signals at each time point comprises the following steps: generating a relative transfer function of the corresponding input signal according to the input signal at each time point; for the input signal at each time point, generating a corresponding relative transfer function and a pseudo-coherence matrix according to the input signal, establishing a superdirectivity filter model according to the relative transfer function and the pseudo-coherence matrix of the input signal, and filtering the input signals at each time point according to the superdirectivity filter model to generate the corresponding first reference signals.
3. The Kalman filter based adaptive array microphone noise reduction method of claim 2, wherein, The step of establishing the beamforming filter model and then filtering the input signals at each time point according to the beamforming filter model to generate the second reference signals at each time point comprises the following steps: performing zero-space projection on the beamforming filter model to generate a corresponding blocking matrix; filtering the input signals at each time point according to the blocking matrix to generate the second reference signals at each time point.
4. The Kalman filter based adaptive array microphone noise reduction method of claim 1, wherein, The step of establishing the process equations at each time point corresponding to the Kalman filter model comprises the following steps: the process equations at each time point corresponding to the Kalman filter model are established by the following formula: w sc (l) = A H (l)w sc (l-1) + Aw(l) wherein w sc (l) represents a sidelobe cancellation filter model in the Kalman filter model at time l, A represents a state equation, H represents a conjugate transpose symbol, and Δw(l) represents an error of the process equation at time l.
5. The Kalman filter based adaptive array microphone noise reduction method of claim 4, wherein, The step of establishing the measurement equations at each time point corresponding to the Kalman filter model comprises the following steps: the measurement equations at each time point corresponding to the Kalman filter model are established according to the first reference signals and the second reference signals at each time point; wherein the measurement equations at each time point corresponding to the Kalman filter model are established by the following formula: x bf (l) = x bm H (l) w sc (l) + Δs(l) where x bf (l) denotes a first reference signal at time l, x bm H (l) denotes a conjugate transpose matrix of a second reference signal at time l, H denotes a conjugate transpose symbol, and Δs(l) denotes an error of the measurement equation at time l.
6. The Kalman adaptive based array microphone noise reduction method of claim 5, wherein, The step of generating the Kalman gains at each time point according to errors corresponding to the process equations and errors corresponding to the measurement equations comprises the following steps: generating an error covariance matrix of the process equation at the corresponding time point according to the error corresponding to the process equation at each time point; generating an error covariance matrix of the measurement equation at the corresponding time point according to the error corresponding to the measurement equation at each time point; generating the Kalman gains of the Kalman filter model at each time point according to the error covariance matrix of the process equation and the error covariance matrix of the measurement equation at each time point.
7. The Kalman adaptive based array microphone noise reduction method of claim 6, wherein, The step that the Kalman filter model eliminates interference noise in the first reference signals and the second reference signals at each time point according to the Kalman gains at each time point comprises: The noise field of the interference noise in the first reference signal and the second reference signal at each moment is eliminated by the Kalman gain at each moment; wherein the interference noise estimation at each moment includes: When the Kalman gain is approximately zero, the noise field of the eliminated interference noise is approximately the noise field filtered by the sidelobe cancellation filter model in the process equation; When the Kalman gain is approximately one, the noise field of the eliminated interference noise is approximately the noise field estimated by the measurement equation.
8. The Kalman adaptive based array microphone noise reduction method of claim 5, wherein, The final output signal at each moment is generated, including: The final output signal at each moment is generated by the following formula: e(l) = x bf (l) - x bm H (l) w sc (l) Wherein e(l) represents the final output signal at moment l.
9. The Kalman adaptive based array microphone noise reduction method of claim 1, wherein, After obtaining the input signal at each moment, it further includes: The acquired input signal at each moment is de-reverberated by the time domain deconvolution method.
10. A Kalman filter based adaptive array microphone noise reduction device, characterized in that, It includes: A signal acquisition module, a first reference signal generation module, a second reference signal generation module, and an interference signal elimination module; The signal acquisition module is used to acquire the input signal at each moment; wherein the input signal at each moment contains target speech and interference noise; The first reference signal generation module is used to establish a hyperdirectivity filter model, and then filter the input signal at each moment according to the hyperdirectivity filter model to generate the first reference signal at each moment; The second reference signal generation module is used to establish a beamforming filter model, and then filter the input signal at each moment according to the beamforming filter model to generate the second reference signal at each moment; The signal output module is used to establish a Kalman filter model and the process equation and the measurement equation corresponding to the Kalman filter model at each moment; generate the Kalman gain at each moment according to the error corresponding to the process equation and the error corresponding to the measurement equation at each moment, so that the Kalman filter model eliminates the interference noise in the first reference signal and the second reference signal at each moment according to the Kalman gain at each moment, and generates the final output signal at each moment.
Citation Information
Patent Citations
Method for dereverberation of an acoustic signal
US20090117948A1
Speech enhancement method and apparatus, and device and computer-readable storage medium
WO2022105571A1