Voice dereverberation method and device, computer device, readable storage medium and program product

By decomposing the filter coefficient matrix using a two-stage Kalman filter method, the storage resource requirements are reduced, the problem of voice quality degradation caused by reverberation components in indoor calls is solved, and the efficiency and quality of voice processing are improved.

CN119785809BActive Publication Date: 2025-11-04ZHUHAI JIELI TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411881211.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-11-04
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

In indoor call scenarios, the sound signal picked up by the sound pickup device includes not only direct sound components but also reverberant sound components, which leads to a decrease in speech quality and intelligibility. Existing multi-channel linear prediction methods suffer from high computational complexity and excessive storage resource consumption.

Method used

A two-stage Kalman filtering method is adopted, which decomposes the filter coefficient matrix into a low-dimensional matrix through the Kronecker integral solution. The full diagonal Kalman filtering is used to reduce the dimension of the error covariance matrix, variance matrix and Kalman gain coefficients, and only the diagonal elements are stored, thus reducing the storage resource requirements.

Benefits of technology

It effectively reduces the consumption of storage resources, improves the computational efficiency and storage resource utilization of speech dereverberation, and enhances speech quality and intelligibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785809B_ABST
    Figure CN119785809B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of audio processing, and provides a speech dereverberation method and device, computer equipment, a readable storage medium and a program product. The method comprises the following steps: performing two-stage filter dereverberation operation according to a plurality of speech frequency domain delay signals, a first-stage observation equation and a second-stage observation equation to obtain an estimated value of a second-stage current frame speech frequency domain dereverberation signal, so as to obtain a speech time domain dereverberation signal of a current frame; when performing the current-stage filter dereverberation operation, obtaining a current-stage state space observation value of the current frame according to the plurality of speech frequency domain delay signals, the current-stage observation equation and spatial regression coefficients required in the current stage; obtaining a priori estimation value of a current-stage error covariance diagonal matrix of the current frame according to a posteriori estimation value of a previous frame error covariance diagonal matrix in the current stage; and obtaining an estimated value of a current-stage speech frequency domain dereverberation signal in the current stage according to the current-stage state space observation value. The method can reduce the occupied storage resources.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio processing, in particular to a speech dereverberation method and device, computer equipment, computer readable storage medium and computer program product. BACKGROUND

[0002] In an indoor communication scenario, such as an indoor voice communication or a remote video conference, due to the reflection of indoor objects or walls, in addition to the direct sound component, the sound signal picked up by the sound pickup device also contains a reverberation sound component. The late reflection sound in the reverberation sound component can seriously damage the speech quality and intelligibility, causing the performance of the sound processing algorithm of the sound pickup system to be greatly reduced. Therefore, speech dereverberation has important practical significance.

[0003] Speech dereverberation can be based on a spectral enhancement method, an inverse filtering method or a multi-channel linear prediction method. Among these methods, the multi-channel linear prediction method is one of the most popular techniques, which first predicts the reverberation component in the signals received by the microphone array, and then subtracts the estimated reverberation component from the received signals to achieve the purpose of dereverberation. In practical applications, the multi-channel linear prediction filter is usually estimated in an adaptive manner. In order to reduce the computational complexity, the filter coefficient matrix can be decomposed into multiple low-dimensional matrices by the method of Kronecker product decomposition, but there is still a problem of updating multiple sub-matrices, which leads to the need to cache the above sub-matrices in the algorithm processing process, resulting in the problem of excessive storage resource occupation. SUMMARY

[0004] Therefore, it is necessary to provide a speech dereverberation method, device, computer equipment, computer readable storage medium and computer program product to solve the above technical problems.

[0005] In a first aspect, the present application provides a speech dereverberation method, comprising:

[0006] obtaining a plurality of frames of speech frequency domain delay signals according to the speech frequency domain signal of the current frame and the speech frequency domain signals of a plurality of historical frames;

[0007] performing two-stage filter dereverberation operations according to the plurality of frames of speech frequency domain delay signals, a first-stage observation equation and a second-stage observation equation to obtain an estimated value of a second-stage current frame speech frequency domain dereverberation signal, and then performing a posteriori update operation to form a current frame speech time domain dereverberation signal; the first-stage observation equation and the second-stage observation equation are obtained by Kronecker product decomposition of spatial regression coefficients in an initial observation equation;

[0008] wherein, in the filtering dereverberation operation, a state space observation value corresponding to the current frame is obtained according to the multi-frame speech frequency domain delay signal, a current stage observation equation and a spatial regression coefficient required by the current stage; a priori estimation value of a current frame error covariance diagonal matrix of the current stage is updated according to a posteriori estimation value of a previous stage frame error covariance diagonal matrix of the current stage; the current frame error covariance diagonal matrix of the current stage is obtained by approximately full diagonalization of a current frame error covariance matrix of the current stage; and an estimation value of a current frame speech frequency domain dereverberation signal of the current stage is updated according to the state space observation value of the current stage.

[0009] If the current stage is the second stage, the spatial regression coefficient required by the current stage is a current frame spatial regression coefficient estimation value of the first stage; the current frame spatial regression coefficient estimation value of the first stage is obtained by updating an estimation value of a variance diagonal matrix of a current frame speech frequency domain dereverberation signal of the first stage according to an estimation value of the current frame speech frequency domain dereverberation signal of the first stage; the variance diagonal matrix is obtained by approximately full diagonalization of a variance matrix; a Kalman gain coefficient of the current frame of the first stage is updated according to a state space observation value corresponding to the current frame of the first stage, a diagonal multiplication observation value corresponding to the current frame of the first stage, the priori estimation value of the current frame error covariance diagonal matrix of the first stage and the estimation value of the variance diagonal matrix of the current frame speech frequency domain dereverberation signal of the first stage; the diagonal multiplication observation value corresponding to the current frame of the first stage is obtained by multiplying a state space conjugate transpose observation value corresponding to the current frame of the first stage with the state space observation value corresponding to the current frame of the first stage and then approximately full diagonalizing the multiplication result; and the current frame spatial regression coefficient estimation value of the first stage is updated according to the estimation value of the current frame speech frequency domain dereverberation signal of the first stage, the Kalman gain coefficient of the current frame of the first stage and a previous frame spatial regression coefficient estimation value of the first stage.

[0010] In one of the embodiments, if the current stage is the first stage, the spatial regression coefficient required by the current stage is a previous frame spatial regression coefficient estimation value of the second stage; the previous frame spatial regression coefficient estimation value of the second stage is obtained by:

[0011] updating an estimation value of a variance diagonal matrix of a previous frame speech frequency domain dereverberation signal of the second stage according to an estimation value of the previous frame speech frequency domain dereverberation signal of the second stage; the variance diagonal matrix is obtained by approximately full diagonalization of a variance matrix;

[0012] The two-stage upper frame Kalman gain coefficient is updated according to the two-stage state space observation value corresponding to the upper frame, the two-stage diagonal multiplication observation value corresponding to the upper frame, the prior estimation value of the two-stage upper frame error covariance diagonal matrix, and the estimation value of the two-stage upper frame speech frequency domain dereverberation signal variance diagonal matrix; the two-stage diagonal multiplication observation value corresponding to the upper frame is obtained by multiplying the two-stage state space observation value corresponding to the upper frame and the two-stage state space conjugate transpose observation value corresponding to the upper frame and then performing approximate full diagonalization;

[0013] The two-stage upper frame space regression coefficient estimation value is updated according to the estimation value of the two-stage upper frame speech frequency domain dereverberation signal, the two-stage upper frame Kalman gain coefficient, and the two-stage upper upper frame space regression coefficient estimation value.

[0014] In one of the embodiments, the obtaining of the multi-frame speech frequency domain delay signal according to the speech frequency domain signal of the current frame and the speech frequency domain signals of the historical multi-frames comprises:

[0015] The speech time domain signal of the current frame is obtained by using a microphone array.

[0016] The speech time domain signal is windowed to obtain a windowed speech time domain signal.

[0017] The windowed speech time domain signal is subjected to Fourier transform to obtain the speech frequency domain signal of the current frame.

[0018] The multi-frame speech frequency domain signal is obtained according to the speech frequency domain signal of the current frame and the speech frequency domain signals of the historical multi-frames.

[0019] The multi-frame speech frequency domain delay signal is obtained according to the time delay coefficient and the multi-frame speech frequency domain signal.

[0020] In one of the embodiments, the speech time domain dereverberation signal of the current frame is formed according to the estimation value of the two-stage current frame speech frequency domain dereverberation signal.

[0021] The speech frequency domain dereverberation signal of the current frame is obtained by performing posterior filtering update according to the estimation value of the two-stage current frame speech frequency domain dereverberation signal, the two-stage state space observation value corresponding to the current frame, and the two-stage current frame Kalman gain coefficient.

[0022] The speech time domain dereverberation signal of the current frame is obtained by performing inverse Fourier transform on the speech frequency domain dereverberation signal of the current frame.

[0023] In one of the embodiments, the prior estimation value of the current frame error covariance diagonal matrix of the current stage is updated according to the posterior estimation value of the upper frame error covariance diagonal matrix of the current stage.

[0024] modeling the space regression coefficient required in the current stage as a first-order Markov process, and performing Kronecker integral decomposition to obtain a state equation of the current stage;

[0025] obtaining a variance of the current frame complex Gaussian disturbance noise in the current stage according to the current frame complex Gaussian disturbance noise in the state equation of the current stage;

[0026] obtaining a posteriori estimation value of the upper frame error covariance diagonal matrix in the current stage;

[0027] updating a priori estimation value of the current frame error covariance diagonal matrix in the current stage according to a sum of the posteriori estimation value of the upper frame error covariance diagonal matrix in the current stage and the variance of the current frame complex Gaussian disturbance noise in the current stage.

[0028] In one embodiment, the obtaining the posteriori estimation value of the upper frame error covariance diagonal matrix in the current stage comprises:

[0029] multiplying the state space observation value in the current stage corresponding to the upper frame and the Kalman gain coefficient in the current stage to obtain a multiplication matrix in the current stage;

[0030] subtracting the unit matrix in the current stage from the multiplication matrix in the current stage to obtain a subtraction matrix in the current stage;

[0031] approximately full-diagonalizing the subtraction matrix in the current stage to obtain a subtraction diagonal matrix in the current stage;

[0032] obtaining the posteriori estimation value of the upper frame error covariance diagonal matrix in the current stage according to the subtraction diagonal matrix in the current stage and the priori estimation value of the upper frame error covariance diagonal matrix in the current stage.

[0033] In one embodiment, obtaining the one-stage diagonal multiplication observation value corresponding to the current frame according to the state space observation value corresponding to the current frame and the state space conjugate transpose observation value corresponding to the current frame comprises:

[0034] obtaining the state space conjugate transpose observation value corresponding to the current frame according to the state space observation value corresponding to the current frame;

[0035] multiplying the state space observation value corresponding to the current frame and the state space conjugate transpose observation value corresponding to the current frame to obtain a multiplication observation value corresponding to the current frame;

[0036] approximately full-diagonalizing the multiplication observation value corresponding to the current frame to obtain the one-stage diagonal multiplication observation value corresponding to the current frame.

[0037] In a second aspect, the present application further provides a voice dereverberation device, comprising:

[0038] a frequency domain delay signal obtaining module, configured to obtain a plurality of frames of speech frequency domain delay signals according to the speech frequency domain signal of the current frame and speech frequency domain signals of a plurality of historical frames;

[0039] a time domain dereverberation signal obtaining module, configured to perform two-stage filter dereverberation operation according to the plurality of frames of speech frequency domain delay signals, a first-stage observation equation and a second-stage observation equation to obtain an estimated value of a second-stage current frame speech frequency domain dereverberation signal, thereby forming the speech time domain dereverberation signal of the current frame; the first-stage observation equation and the second-stage observation equation are obtained by Kronecker integral decomposition of spatial regression coefficients in an initial observation equation;

[0040] a current-stage filter dereverberation operation module, configured to, when performing the current-stage filter dereverberation operation, obtain a current-stage state space observation value corresponding to the current frame according to the plurality of frames of speech frequency domain delay signals, a current-stage observation equation and spatial regression coefficients required in the current stage; update a priori estimation value of a current-stage current frame error covariance diagonal matrix to obtain a posteriori estimation value of the current-stage current frame error covariance diagonal matrix according to the posteriori estimation value of a previous-stage current frame error covariance diagonal matrix; the current-stage current frame error covariance diagonal matrix is obtained by approximate full diagonalization of a current-stage current frame error covariance matrix; and update the estimated value of the current-stage current frame speech frequency domain dereverberation signal to obtain an updated estimated value of the current-stage current frame speech frequency domain dereverberation signal according to the current-stage state space observation value;

[0041] a spatial regression coefficient obtaining module, configured to, if the current stage is the second stage, the spatial regression coefficients required in the current stage are first-stage current frame spatial regression coefficient estimation values; the first-stage current frame spatial regression coefficient estimation values are obtained by the following steps: updating a variance diagonal matrix estimation value of a first-stage current frame speech frequency domain dereverberation signal to obtain an updated estimated value of the first-stage current frame speech frequency domain dereverberation signal according to the estimated value of the first-stage current frame speech frequency domain dereverberation signal; the variance diagonal matrix is obtained by approximate full diagonalization of a variance matrix; updating a first-stage current frame Kalman gain coefficient to obtain an updated estimated value of the first-stage current frame spatial regression coefficient according to a first-stage state space observation value corresponding to the current frame, a first-stage diagonal multiplication observation value corresponding to the current frame, a priori estimation value of the first-stage current frame error covariance diagonal matrix and the variance diagonal matrix estimation value of the first-stage current frame speech frequency domain dereverberation signal; the first-stage diagonal multiplication observation value corresponding to the current frame is obtained by multiplying a first-stage state space observation value corresponding to the current frame and a first-stage state space conjugate transpose observation value corresponding to the current frame and then performing approximate full diagonalization; and updating the first-stage current frame spatial regression coefficient estimation value to obtain an updated estimated value of the first-stage current frame spatial regression coefficient according to the estimated value of the first-stage current frame speech frequency domain dereverberation signal, the first-stage current frame Kalman gain coefficient and a previous-stage spatial regression coefficient estimation value.

[0042] In a third aspect, the present application provides a computer device. The computer device comprises a memory and a processor, the memory stores a computer program, and the processor executes the method described above.

[0043] In a fourth aspect, the present application provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program is executed by a processor to perform the method described above.

[0044] In a fifth aspect, the present application provides a computer program product. The computer program product comprises a computer program, and the computer program is executed by a processor to perform the method described above.

[0045] The voice dereverberation method, device, computer device, computer readable storage medium and computer program product obtain a plurality of frames of voice frequency domain delay signals according to the voice frequency domain signal of the current frame and the voice frequency domain signals of a plurality of historical frames; perform two-stage filtering dereverberation operations according to the plurality of frames of voice frequency domain delay signals, a first-stage observation equation and a second-stage observation equation to obtain an estimated value of the second-stage voice frequency domain dereverberation signal of the current frame, and then perform a posteriori update operations to form the voice time domain dereverberation signal of the current frame; the first-stage observation equation and the second-stage observation equation are obtained by Kronecker decomposition of a spatial regression coefficient in an initial observation equation; during the current stage filtering dereverberation operation, a current stage state space observation value corresponding to the current frame is obtained according to the plurality of frames of voice frequency domain delay signals, the current stage observation equation and a spatial regression coefficient required in the current stage; a priori estimation of a current stage current frame error covariance diagonal matrix is obtained by updating according to a posteriori estimation of a previous stage current frame error covariance diagonal matrix; the current stage current frame error covariance diagonal matrix is obtained by approximately full diagonalization of a current stage current frame error covariance matrix; the estimated value of the current stage voice frequency domain dereverberation signal of the current frame is obtained by updating according to the current stage state space observation value; if the current stage is the second stage, the spatial regression coefficient required in the current stage is a first-stage current frame spatial regression coefficient estimation value; the first-stage current frame spatial regression coefficient estimation value is obtained by the following steps: the estimated value of the first-stage current frame voice frequency domain dereverberation signal is updated to obtain an estimated value of a first-stage current frame voice frequency domain dereverberation signal variance diagonal matrix; the variance diagonal matrix is obtained by approximately full diagonalization of a variance matrix; a first-stage current frame Kalman gain coefficient is obtained by updating according to a first-stage state space observation value corresponding to the current frame, a first-stage diagonal multiplication observation value corresponding to the current frame, a priori estimation of the first-stage current frame error covariance diagonal matrix and the estimated value of the first-stage current frame voice frequency domain dereverberation signal variance diagonal matrix; the first-stage diagonal multiplication observation value corresponding to the current frame is obtained by multiplying a first-stage state space conjugate transpose observation value corresponding to the current frame with the first-stage state space observation value corresponding to the current frame and then performing approximately full diagonalization; the first-stage current frame spatial regression coefficient estimation value is obtained by updating according to the estimated value of the first-stage current frame voice frequency domain dereverberation signal, the first-stage current frame Kalman gain coefficient and a previous stage spatial regression coefficient estimation value.The application utilizes full-diagonal Kalman filtering to replace the original Kalman filtering, and performs approximate full-diagonalization on the error covariance matrix of the current frame of the current stage to obtain the diagonal matrix of the error covariance of the current frame of the current stage, performs approximate full-diagonalization on the variance matrix to obtain the diagonal matrix of the variance, multiplies the one-stage state space observation value corresponding to the current frame by the one-stage state space conjugate transpose observation value corresponding to the current frame and performs approximate full-diagonalization to obtain the one-stage diagonal multiplication observation value corresponding to the current frame, and updates the one-stage Kalman gain coefficient of the current frame according to the one-stage state space observation value corresponding to the current frame, the one-stage diagonal multiplication observation value corresponding to the current frame, the prior estimate value of the diagonal matrix of the error covariance of the current frame of the current stage, and the estimate value of the variance diagonal matrix of the current frame of the current stage. The dimensions of the error covariance matrix, the variance matrix and the Kalman gain coefficient are reduced, only the diagonal elements of the error covariance matrix, the variance matrix and the Kalman gain coefficient need to be stored, and the amount of parameters to be stored is greatly reduced, thereby reducing the occupied storage resources. BRIEF DESCRIPTION OF DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application or the related art. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other related drawings can also be obtained without creative labor.

[0047] Figure 1 An application environment diagram of the speech dereverberation method in an embodiment;

[0048] Figure 2 A reverberation application environment diagram of the speech dereverberation method in an embodiment;

[0049] Figure 3 A flow diagram of the speech dereverberation method in an embodiment;

[0050] Figure 4 A flow diagram of the speech dereverberation method in another embodiment;

[0051] Figure 5 A flow diagram of the one-stage filtering dereverberation operation in an embodiment;

[0052] Figure 6 A flow diagram of the two-stage filtering dereverberation operation in an embodiment;

[0053] Figure 7 A structural block diagram of the speech dereverberation device in an embodiment;

[0054] Figure 8Fig. 1 is a schematic diagram of an internal structure of a computer device in an embodiment. DETAILED DESCRIPTION

[0055] For the purposes of the present application, the technical solutions and advantages thereof will be more clearly apparent from the following detailed description in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely intended to explain the present application and are not intended to limit the present application.

[0056] An embodiment of the present application provides a speech dereverberation method, which can be executed by a computer device, as shown in Fig. 1. The computer device can obtain a current frame of speech frequency domain signals and a plurality of historical frames of speech frequency domain signals, and then obtain a current frame of speech time domain dereverberation signals. It should be understood that the computer device can be implemented by a server, or by a terminal, or by an interactive system of a terminal and a server. The speech dereverberation method provided by the present application is applicable to a microphone array system having a single or multiple microphones. Assuming that there is a uniform linear array composed of M omnidirectional microphones in a reverberation environment, the application scenario of the speech dereverberation method is as shown in Fig. 2. In an embodiment, the speech dereverberation method comprises the steps shown in Fig. 3. Figure 1 Figure 2 Figure 3

[0057] In step S301, a plurality of frames of speech frequency domain delay signals are obtained according to the current frame of speech frequency domain signals and the plurality of historical frames of speech frequency domain signals.

[0058] The current frame of speech time domain signals and the plurality of historical frames of speech time domain signals can be subjected to Fourier transform to obtain the current frame of speech frequency domain signals and the plurality of historical frames of speech frequency domain signals.

[0059] Assuming that the speech reverberation signal is generated by a multi-channel autoregressive (MAR) process, according to the approximate relationship of the convolution transfer function, the speech time domain signal received by the microphone array can be expressed as formula (1) after short-time Fourier transform.

[0060] (1)

[0061] wherein, represents the speech frequency domain signal of the microphone array, represents the autoregressive coefficient, represents the pure speech signal, is a frequency band index, is a frame number index, and the superscript represents a transposition operation, ​​​​representing a time delay coefficient, used to distinguish early reflected sound from late reflected sound, the time delay coefficient is related to frame length, representing a linear convolution length.

[0062] The multi-frame speech frequency domain signal can be obtained according to the speech frequency domain signal of the current frame and the speech frequency domain signals of the historical multi-frames; and the multi-frame speech frequency domain delay signal can be obtained according to the time delay coefficient D and the multi-frame speech frequency domain signal.

[0063] In step S302, two-stage filtering dereverberation operations are performed according to the multi-frame speech frequency domain delay signal, a first-stage observation equation and a second-stage observation equation, to obtain an estimated value of the second-stage current-frame speech frequency domain dereverberation signal, and then a speech time domain dereverberation signal of the current frame is formed through a posteriori update operation; the first-stage observation equation and the second-stage observation equation are obtained by Kronecker decomposition of the spatial regression coefficient in the initial observation equation.

[0064] The following variables can be defined:

[0065] (2)

[0066] (3)

[0067] wherein, represents an identity matrix of , is a Kronecker product, is a column vector stacking operation. Formula (2) represents the stacking of the historical multi-frame speech frequency domain signals, and the Kronecker product is used to ensure that the same multi-frame speech frequency domain delay signal is used for different channels, and the spatial regression coefficient in formula (3) corresponds to formula (2) one by one. Formula (1) can be rewritten in the following form by using formula (2) and (3):

[0068] (4)

[0069] According to the adaptive filtering theory, formula (4) can be regarded as an observation equation in the Kalman filtering process, and the observation equation of formula (4) is used as an initial observation equation.

[0070] The spatial regression coefficient in the initial observation equation as shown in formula (4) can be decomposed by Kronecker decomposition, and the spatial regression coefficient can be decomposed into two low-dimensional coefficients by using the properties of Kronecker decomposition, let wherein, and are vectors with lengths of and respectively, and the initial observation equation can be rewritten by substituting into formula (1), to obtain the initial observation rewriting equation as shown in formula (5).

[0071] (5)

[0072] According to the property of the Kronecker product, equivalent relations as shown in formula (6) and formula (7) can be obtained.

[0073] (6)

[0074] (7)

[0075] Substitute formula (6) and formula (7) into formula (5), and define and . The observation equations as shown in formula (8) and formula (9) can be obtained.

[0076] (8)

[0077] (9)

[0078] According to formula (8) and formula (9), two different observation equations can be obtained, which can correspond to two different sets of filtering update processes respectively. Although there is a difference between the observation equations, the two sets of filtering update processes do not have a sequence. For a certain update process, the other set of update processes can be regarded as an update process after the observation equation changes. Without loss of generality, formula (8) can be taken as a one-stage observation equation, and formula (9) can be taken as a two-stage observation equation, or formula (9) can be taken as a one-stage observation equation, and formula (8) can be taken as a two-stage observation equation, and the two stages are related to each other.

[0079] According to the multi-frame speech frequency domain delay signal and the one-stage observation equation, a one-stage filtering dereverberation operation can be performed to obtain a two-stage required spatial regression coefficient. According to the multi-frame speech frequency domain delay signal, the two-stage observation equation and the two-stage required spatial regression coefficient, a two-stage filtering dereverberation operation can be performed to obtain an estimated value of a two-stage current frame speech frequency domain dereverberation signal, and then a post-update operation is performed to form a speech time domain dereverberation signal of the current frame.

[0080] In step S303, when performing the filtering dereverberation operation of the current stage, a state space observation value corresponding to the current frame is obtained according to the multi-frame speech frequency domain delay signal, the observation equation of the current stage and the spatial regression coefficient required by the current stage; a priori estimation value of the current frame error covariance diagonal matrix of the current stage is obtained by updating the posteriori estimation value of the last frame error covariance diagonal matrix of the current stage; the current frame error covariance diagonal matrix of the current stage is obtained by approximately full diagonalizing the current frame error covariance matrix of the current stage; and the estimated value of the current frame speech frequency domain dereverberation signal of the current stage is obtained by updating the state space observation value of the current stage.

[0081] If the current stage is the first stage, in the first stage filtering dereverberation operation, according to the multi-frame speech frequency domain delay signals , the first stage observation equation and the spatial regression coefficients required by the first stage, the first stage state space observation value corresponding to the current frame is obtained , as shown in equation (10).

[0082] (10)

[0083] , the first stage observation equation as shown in equation (8) is calculated .

[0084] The prior estimate of the first stage current frame error covariance diagonal matrix can be updated according to the posterior estimate of the first stage upper frame error covariance diagonal matrix , as shown in equation (11).

[0085] (11)

[0086] , wherein represents the variance of the first stage current frame complex Gaussian disturbance noise; the first stage current frame error covariance diagonal matrix is obtained by approximately full diagonalization of the first stage current frame error covariance matrix , as shown in equation (12).

[0087] (12)

[0088] , wherein represents the identity matrix of .

[0089] The estimate of the first stage current frame speech frequency domain dereverberation signal can be updated according to the first stage state space observation value , as shown in equation (13).

[0090] (13)

[0091] , wherein represents the current frame speech frequency domain signal; represents the first stage upper frame spatial regression coefficient estimate.

[0092] If the current stage is the second stage, in the second stage filtering dereverberation operation, according to the multi-frame speech frequency domain delay signals , the second stage observation equation and the spatial regression coefficients required by the second stage, the second stage state space observation value corresponding to the current frame is obtained , as shown in equation (14).​​

[0093] (14)

[0094] wherein the two-stage observation equation as shown in equation (9) is calculated .

[0095] The prior estimate of the two-stage current frame error covariance diagonal matrix can be updated according to the posterior estimate of the two-stage upper frame error covariance diagonal matrix , as shown in equation (15).

[0096] (15)

[0097] wherein represents the variance of the two-stage current frame complex Gaussian disturbance noise; the two-stage current frame error covariance diagonal matrix is obtained by approximately full diagonalizing the two-stage current frame error covariance matrix , as shown in equation (16).

[0098] (16)

[0099] wherein represents the identity matrix of .

[0100] The estimate of the two-stage current frame speech frequency domain dereverberation signal can be updated according to the two-stage state space observation value , as shown in equation (17).

[0101] (17) wherein

[0102] represents the current frame speech frequency domain signal; represents the two-stage upper frame spatial regression coefficient estimate.

[0103] ​​Step S304, if the current stage is the second stage, the spatial regression coefficient required in the current stage is the first stage current frame spatial regression coefficient estimate value; the first stage current frame spatial regression coefficient estimate value is obtained by the following steps: updating the first stage current frame speech frequency domain dereverberation signal variance diagonal matrix estimate value according to the first stage current frame speech frequency domain dereverberation signal estimate value; the variance diagonal matrix is obtained by approximately full diagonalization of the variance matrix; updating the first stage current frame Kalman gain coefficient according to the first stage state space observation value corresponding to the current frame, the first stage diagonal multiplication observation value corresponding to the current frame, the first stage current frame error covariance diagonal matrix prior estimate value and the first stage current frame speech frequency domain dereverberation signal variance diagonal matrix estimate value; the first stage diagonal multiplication observation value corresponding to the current frame is obtained by multiplying the first stage state space observation value corresponding to the current frame and the first stage state space conjugate transpose observation value corresponding to the current frame and performing approximately full diagonalization; updating the first stage current frame spatial regression coefficient estimate value according to the first stage current frame speech frequency domain dereverberation signal estimate value, the first stage current frame Kalman gain coefficient and the first stage previous frame spatial regression coefficient estimate value.

[0104] If the current stage is the second stage, the spatial regression coefficient required in the current stage is the first stage current frame spatial regression coefficient estimate value .

[0105] The first stage current frame spatial regression coefficient estimate value is obtained by the following steps: updating the first stage current frame speech frequency domain dereverberation signal variance diagonal matrix estimate value according to the first stage current frame speech frequency domain dereverberation signal estimate value . as shown in equation (18).

[0106] (18)

[0107] wherein, denotes a forgetting factor, and the forgetting factor can be set as . denotes the first stage previous frame speech frequency domain dereverberation signal variance diagonal matrix estimate value; the variance diagonal matrix is obtained by approximately full diagonalization of the variance matrix as shown in equation (19).

[0108] (19)

[0109] updating the first stage current frame Kalman gain coefficient according to the first stage state space observation value corresponding to the current frame , the first stage diagonal multiplication observation value corresponding to the current frame , the first stage current frame error covariance diagonal matrix prior estimate value an estimate of a diagonal matrix of variances of the one-stage current frame speech de-reverberation signal , to obtain an estimate of a one-stage current frame Kalman gain coefficient , as shown in equation (20).

[0110] (20)

[0111] wherein a one-stage diagonally multiplied observation corresponding to the current frame is obtained by multiplying a one-stage state space observation corresponding to the current frame and a one-stage state space conjugate transpose observation corresponding to the current frame , as shown in equation (21).

[0112] (21)

[0113] an estimate of a diagonal matrix of variances of the one-stage current frame speech de-reverberation signal , a one-stage current frame Kalman gain coefficient and a one-stage previous frame space regression coefficient estimate , to obtain a one-stage current frame space regression coefficient estimate , as shown in equation (22).

[0114] (22)

[0115] In the speech de-reverberation method, the full diagonal Kalman filter is used to replace the original Kalman filter, the one-stage current frame error covariance diagonal matrix is obtained by approximately full diagonalizing the one-stage current frame error covariance matrix, the variance diagonal matrix is obtained by approximately full diagonalizing the variance matrix, the one-stage diagonally multiplied observation corresponding to the current frame is obtained by multiplying the one-stage state space observation corresponding to the current frame and the one-stage state space conjugate transpose observation corresponding to the current frame, the one-stage current frame Kalman gain coefficient is obtained according to the one-stage state space observation corresponding to the current frame, the one-stage diagonally multiplied observation corresponding to the current frame, the prior estimate of the one-stage current frame error covariance diagonal matrix and the estimate of the variance diagonal matrix of the one-stage current frame speech de-reverberation signal, which reduces the dimensions of the error covariance matrix, the variance matrix and the Kalman gain coefficient, only the diagonal elements of the error covariance matrix, the variance matrix and the Kalman gain coefficient need to be stored, which greatly reduces the amount of parameters to be stored, thereby reducing the occupied storage resources.

[0116] In one embodiment, if this stage is stage one, the spatial regression coefficients required for this stage are the estimated values ​​of the spatial regression coefficients of the previous frame in stage two. The specific steps to obtain the estimated values ​​of the spatial regression coefficients of the previous frame in stage two are as follows: Based on the estimated value of the dereverberated signal in the audio domain of the previous frame in stage two, update the estimated value of the variance diagonal matrix of the dereverberated signal in the audio domain of the previous frame in stage two; the variance diagonal matrix is ​​obtained by approximately fully diagonalizing the variance matrix; based on the second-stage state space observations corresponding to the previous frame, the second-stage diagonally multiplied observations corresponding to the previous frame, and the second-stage error covariance pair... The prior estimate of the diagonal matrix and the estimate of the variance diagonal matrix of the dereverberation signal in the audio domain of the second-stage upper frame are used to update the Kalman gain coefficient of the second-stage upper frame. The second-stage diagonal multiplication observation corresponding to the previous frame is obtained by multiplying the second-stage state space observation corresponding to the previous frame and the second-stage state space conjugate transpose observation corresponding to the previous frame and performing an approximate full diagonalization. The second-stage upper frame spatial regression coefficient estimate is obtained by updating the estimate of the dereverberation signal in the audio domain of the second-stage upper frame, the Kalman gain coefficient of the second-stage upper frame, and the spatial regression coefficient estimate of the frame two steps ahead of the second-stage upper frame.

[0117] If this stage is stage one, then the spatial regression coefficients required for this stage are the estimated spatial regression coefficients of the previous frame in stage two. If the current frame is the initial frame, then the estimated value of the spatial regression coefficients of the previous frame in the second stage is the estimated value of the initial spatial regression coefficients in the second stage. Let the estimated value of the initial spatial regression coefficients in the second stage be... ,in It represents a vector consisting entirely of 1s.

[0118] The specific steps to obtain the estimated values ​​of the spatial regression coefficients of the second-stage upper frame are as follows: Based on the estimated values ​​of the dereverberation signal in the audio domain of the second-stage upper frame... The estimated value of the variance diagonal matrix of the dereverberation signal in the audio domain of the second-stage upper frame is obtained by updating. As shown in equation (23).

[0119] (twenty three)

[0120] in, Representing the forgetting factor, it can be made ; This represents the estimated value of the variance diagonal matrix of the dereverberation signal in the audio domain of the previous frame in the second stage; the variance diagonal matrix is ​​the variance matrix. The result obtained by approximating full diagonalization is shown in equation (24).

[0121] (twenty four)

[0122] Based on the two-stage state space observations corresponding to the previous frame , the two-stage diagonal multiplied observation value corresponding to the upper frame , the prior estimate value of the two-stage upper frame error covariance diagonal matrix , the estimate value of the two-stage upper frame speech frequency domain dereverberation signal variance diagonal matrix , the two-stage upper frame Kalman gain coefficient is updated , as shown in equation (25).

[0123] (25)

[0124] , the two-stage diagonal multiplied observation value corresponding to the upper frame is multiplied by the two-stage state space observation value corresponding to the upper frame , and the two-stage state space conjugate transpose observation value corresponding to the upper frame , and is approximately full diagonalized, as shown in equation (26).

[0125] (26)

[0126] , the two-stage upper frame speech frequency domain dereverberation signal estimate value , the two-stage upper frame Kalman gain coefficient , and the two-stage upper upper frame space regression coefficient estimate value , the two-stage upper frame space regression coefficient estimate value is updated , as shown in equation (27).

[0127] (27)

[0128] In this embodiment, if the current stage is the first stage, the space regression coefficient required in the current stage is the two-stage upper frame space regression coefficient estimate value, which can be updated according to the two-stage upper frame speech frequency domain dereverberation signal estimate value, the two-stage upper frame Kalman gain coefficient, and the two-stage upper upper frame space regression coefficient estimate value.

[0129] In one of the embodiments, a plurality of frames of speech frequency domain delay signals are obtained according to a speech frequency domain signal of a current frame and a plurality of frames of speech frequency domain signals of historical frames, and the specific steps are as follows: a speech time domain signal of the current frame is obtained through a microphone array; the speech time domain signal is windowed to obtain a windowed speech time domain signal; the windowed speech time domain signal is subjected to Fourier transform to obtain the speech frequency domain signal of the current frame; the plurality of frames of speech frequency domain signals are obtained according to the speech frequency domain signal of the current frame and the plurality of frames of speech frequency domain signals of the historical frames; and the plurality of frames of speech frequency domain delay signals are obtained according to a time delay coefficient and the plurality of frames of speech frequency domain signals.

[0130] The two-stage upper frame speech frequency domain dereverberation signal estimate value can be obtained according to the two-stage upper frame speech frequency domain dereverberation signal estimate value Figure 2The illustrated microphone array acquires a speech time domain signal of a current frame. The speech time domain signal can be windowed to obtain a windowed speech time domain signal. The windowed speech time domain signal can be subjected to Fourier transform to obtain a speech frequency domain signal of the current frame.

[0131] The speech frequency domain signal of the current frame is sequentially buffered. When the number of buffered frames is greater than a preset delay frame number, a historical multi-frame speech frequency domain signal is obtained.

[0132] The multi-frame speech frequency domain signal is obtained according to the speech frequency domain signal of the current frame and the historical multi-frame speech frequency domain signal.

[0133] The multi-frame speech frequency domain delay signal is obtained by signal delay of the multi-frame speech frequency domain signal according to a time delay coefficient D.

[0134] In the embodiment, the windowed speech time domain signal is subjected to Fourier transform to obtain a relatively accurate speech frequency domain signal of the current frame. The multi-frame speech frequency domain signal is obtained according to the speech frequency domain signal of the current frame and the historical multi-frame speech frequency domain signal. The multi-frame speech frequency domain delay signal is obtained according to the time delay coefficient and the multi-frame speech frequency domain signal.

[0135] In one of the embodiments, the speech time domain dereverberation signal of the current frame is formed according to the estimated value of the two-stage current frame speech frequency domain dereverberation signal, and the specific steps are as follows: the speech frequency domain dereverberation signal of the current frame is obtained by performing posterior filtering update according to the estimated value of the two-stage current frame speech frequency domain dereverberation signal, the two-stage state space observation value corresponding to the current frame, and the two-stage current frame Kalman gain coefficient; and the speech time domain dereverberation signal of the current frame is obtained by performing inverse Fourier transform on the speech frequency domain dereverberation signal of the current frame.

[0136] The estimated value of the two-stage current frame speech frequency domain dereverberation signal , the two-stage state space observation value corresponding to the current frame , and the two-stage current frame Kalman gain coefficient are subjected to posterior filtering update to obtain the speech frequency domain dereverberation signal of the current frame , as shown in equation (28).

[0137] (28)

[0138] The speech frequency domain dereverberation signal of the current frame is subjected to inverse Fourier transform to obtain the speech time domain dereverberation signal of the current frame.

[0139] In the embodiment, the two-stage current frame speech frequency domain dereverberation signal estimation value, the two-stage state space observation value corresponding to the current frame and the two-stage current frame Kalman gain coefficient are used for post-filtering update, and further dereverberation operation is performed to obtain a more accurate current frame speech frequency domain dereverberation signal. The current frame speech frequency domain dereverberation signal is subjected to inverse Fourier transform, so as to obtain a current frame speech time domain dereverberation signal.

[0140] In one embodiment, the prior estimation value of the current frame error covariance diagonal matrix of the current stage is updated according to the posterior estimation value of the error covariance diagonal matrix of the previous frame of the current stage. The specific steps are as follows: the spatial regression coefficient required in the current stage is modeled as a first-order Markov process, and Kronecker integral decomposition is performed to obtain the state equation of the current stage; the variance of the complex Gaussian disturbance noise of the current frame of the current stage in the state equation of the current stage is obtained; the posterior estimation value of the error covariance diagonal matrix of the previous frame of the current stage is obtained; and the prior estimation value of the error covariance diagonal matrix of the current frame of the current stage is updated according to the sum of the posterior estimation value of the error covariance diagonal matrix of the previous frame of the current stage and the variance of the complex Gaussian disturbance noise of the current frame of the current stage.

[0141] According to the multi-channel autoregressive process, the spatial regression coefficient required in the current stage can be modeled as a first-order Markov process, and Kronecker integral decomposition is performed to obtain two different state equations as shown in formula (29) and formula (30).

[0142] (29)

[0143] (30)

[0144] wherein, and represent the state transition matrix of the first-stage filtering and the second-stage filtering respectively, and the unit matrix is usually selected, and represent the complex Gaussian disturbance noise in the state transition process of the first-stage filtering and the second-stage filtering respectively, and the variances of the complex Gaussian disturbance noise are represented by and respectively, and the update of the disturbance noise directly affects the final effect of the filtering algorithm, and and can be estimated by formula (31) and formula (32) respectively,

[0145] (31)

[0146] (32)

[0147] wherein denotes a positive minimum value, representing a change in the disturbance; denotes a mathematical expectation.

[0148] If the current stage is a first stage, the first stage state equation is obtained according to the first stage current frame complex Gaussian disturbance noise , and a variance of the first stage current frame complex Gaussian disturbance noise is obtained according to the first stage current frame complex Gaussian disturbance noise .

[0149] An a posteriori estimation value of the first stage upper frame error covariance diagonal matrix is obtained according to the first stage upper frame error covariance diagonal matrix ; and a priori estimation value of the first stage current frame error covariance diagonal matrix is obtained according to a sum of the a posteriori estimation value of the first stage upper frame error covariance diagonal matrix and the variance of the first stage current frame complex Gaussian disturbance noise , as shown in equation (11).

[0150] If the current stage is a second stage, the second stage state equation is obtained according to the second stage current frame complex Gaussian disturbance noise , and a variance of the second stage current frame complex Gaussian disturbance noise is obtained according to the second stage current frame complex Gaussian disturbance noise .

[0151] An a posteriori estimation value of the second stage upper frame error covariance diagonal matrix is obtained according to the second stage upper frame error covariance diagonal matrix ; and a priori estimation value of the second stage current frame error covariance diagonal matrix is obtained according to a sum of the a posteriori estimation value of the second stage upper frame error covariance diagonal matrix and the variance of the second stage current frame complex Gaussian disturbance noise , as shown in equation (16).

[0152] In the embodiment, the spatial regression coefficient required in the current stage is modeled as a first-order Markov process, and the state equation of the current stage is obtained; a variance of the current stage current frame complex Gaussian disturbance noise is obtained according to the current stage current frame complex Gaussian disturbance noise in the state equation of the current stage; and a priori estimation value of the current stage current frame error covariance diagonal matrix is obtained according to a sum of an a posteriori estimation value of the current stage upper frame error covariance diagonal matrix and the variance of the current stage current frame complex Gaussian disturbance noise.

[0153] ​​In one of the embodiments, the posterior estimation of the diagonal matrix of the error covariance of the current stage is obtained, and the specific steps are as follows: the state space observation value corresponding to the previous stage of the current stage is multiplied by the Kalman gain coefficient of the previous stage, to obtain a multiplication matrix of the current stage; a unit matrix of the current stage is subtracted from the multiplication matrix of the current stage, to obtain a subtraction matrix of the current stage; the subtraction matrix of the current stage is approximately full diagonalized, to obtain a subtraction diagonal matrix of the current stage; and the posterior estimation of the diagonal matrix of the error covariance of the current stage is obtained according to the subtraction diagonal matrix of the current stage and the prior estimation of the diagonal matrix of the error covariance of the previous stage.

[0154] If the current stage is a first stage, the state space observation value corresponding to the previous stage of the first stage is multiplied by the Kalman gain coefficient of the previous stage, to obtain a multiplication matrix of the first stage. A unit matrix of the first stage is subtracted from the multiplication matrix of the first stage, to obtain a subtraction matrix of the first stage. .

[0155] The subtraction matrix of the first stage is approximately full diagonalized, to obtain a subtraction diagonal matrix of the first stage, as shown in equation (33).

[0156] (33)

[0157] The posterior estimation of the diagonal matrix of the error covariance of the first stage is obtained according to the subtraction diagonal matrix of the first stage and the prior estimation of the diagonal matrix of the error covariance of the previous stage, as shown in equation (34).

[0158] (34)

[0159] If the current stage is a second stage, the state space observation value corresponding to the previous stage of the second stage is multiplied by the Kalman gain coefficient of the previous stage, to obtain a multiplication matrix of the second stage. A unit matrix of the second stage is subtracted from the multiplication matrix of the second stage, to obtain a subtraction matrix of the second stage. .

[0160] The subtraction matrix of the second stage is approximately full diagonalized, to obtain a subtraction diagonal matrix of the second stage, as shown in equation (35).

[0161] (35)

[0162] ​​​​​​​​​​According to the prior estimation value of the two-stage subtraction diagonal matrix and the two-stage upper frame error covariance diagonal matrix , the posterior estimation value of the two-stage upper frame error covariance diagonal matrix is obtained , as shown in formula (36).

[0163] (36)

[0164] In this embodiment, the two-stage subtraction diagonal matrix is obtained according to the state space observation value corresponding to the upper frame in the current stage, the upper frame Kalman gain coefficient in the current stage and the unit matrix in the current stage; and the posterior estimation value of the two-stage upper frame error covariance diagonal matrix is obtained according to the two-stage subtraction diagonal matrix and the prior estimation value of the two-stage upper frame error covariance diagonal matrix.

[0165] In one embodiment, the one-stage diagonal multiplication observation value corresponding to the current frame is obtained according to the one-stage state space observation value corresponding to the current frame and the one-stage state space conjugate transpose observation value corresponding to the current frame, and the specific steps are as follows: the one-stage state space conjugate transpose observation value corresponding to the current frame is obtained according to the one-stage state space observation value corresponding to the current frame; the one-stage multiplication observation value corresponding to the current frame is obtained by multiplying the one-stage state space observation value corresponding to the current frame and the one-stage state space conjugate transpose observation value corresponding to the current frame; and the one-stage diagonal multiplication observation value corresponding to the current frame is obtained by approximately full diagonalizing the one-stage multiplication observation value corresponding to the current frame.

[0166] The one-stage state space observation value corresponding to the current frame is obtained, and the one-stage state space conjugate transpose observation value corresponding to the current frame is obtained. The one-stage state space observation value corresponding to the current frame and the one-stage state space conjugate transpose observation value corresponding to the current frame are multiplied to obtain the one-stage multiplication observation value corresponding to the current frame; and the one-stage diagonal multiplication observation value corresponding to the current frame is obtained by approximately full diagonalizing the one-stage multiplication observation value corresponding to the current frame.

[0167] In this embodiment, the one-stage multiplication observation value corresponding to the current frame is obtained by multiplying the one-stage state space observation value corresponding to the current frame and the one-stage state space conjugate transpose observation value corresponding to the current frame; and the one-stage diagonal multiplication observation value corresponding to the current frame is obtained by approximately full diagonalizing the one-stage multiplication observation value corresponding to the current frame.

[0168] In order to better understand the above method, an application embodiment of the voice dereverberation method of the present application is described in detail as follows, and the specific steps are as follows Figure 4The voice dereverberation method provided in the application is applicable to a microphone array system with a single or multiple microphones. Assuming that there is a uniform linear array with M omnidirectional microphones in a reverberation environment, the application scenario of the voice dereverberation method is as shown in Figure 2 .

[0169] In step S401, a voice time-domain signal of a current frame is obtained through a microphone array.

[0170] In step S402, the voice time-domain signal is subjected to windowing processing to obtain a windowed voice time-domain signal; and the windowed voice time-domain signal is subjected to Fourier transform to obtain a voice frequency-domain signal of the current frame.

[0171] In step S403, the voice frequency-domain signal of the current frame is sequentially buffered, and it is determined whether the current number of buffered frames is greater than a preset delay frame. If yes, a plurality of historical voice frequency-domain signals are obtained, and otherwise, the step S401 is returned. According to the voice frequency-domain signal of the current frame and the plurality of historical voice frequency-domain signals, a plurality of voice frequency-domain signals are obtained. According to a time delay coefficient D, the plurality of voice frequency-domain signals are subjected to signal delay to obtain a plurality of voice frequency-domain delay signals.

[0172] In step S404, a one-stage filtering dereverberation operation can be performed according to the plurality of voice frequency-domain delay signals and a one-stage observation equation to obtain a spatial regression coefficient required in a two-stage, and a specific flow is as shown in Figure 5 . The one-stage observation equation and a two-stage observation equation are obtained by Kronecker integral decomposition of a spatial regression coefficient in an initial observation equation.

[0173] In step S405, a two-stage filtering dereverberation operation can be performed according to the plurality of voice frequency-domain delay signals, the two-stage observation equation and the spatial regression coefficient required in the two-stage to obtain an estimated value of a two-stage voice frequency-domain dereverberation signal of the current frame, and a specific flow is as shown in Figure 6 .

[0174] In step S406, a posterior filtering update is performed according to the estimated value of the two-stage voice frequency-domain dereverberation signal of the current frame, a two-stage state space observation value corresponding to the current frame and a two-stage Kalman gain coefficient of the current frame to obtain a voice frequency-domain dereverberation signal of the current frame; and an inverse Fourier transform is performed on the voice frequency-domain dereverberation signal of the current frame to obtain a voice time-domain dereverberation signal of the current frame . The voice time-domain dereverberation signal of the current frame represents a voice time-domain dereverberation signal of the current frame of an independent frequency band of the i th microphone in the microphone array.

[0175] ​The speech frequency domain de-reverberation signal of the current frame of each frequency band of one of the microphones in the microphone array can be independently filtered, and then the speech time domain de-reverberation signal of the current frame of one of the microphones in the microphone array can be obtained by using the overlap-add method .

[0176] In this embodiment, first, the autoregressive coefficients to be estimated in the Kalman filtering process are decomposed into two vectors with lower dimensions by using the Kronecker decomposition, and the original Kalman update process is changed into two filtering de-reverberation operations with mutual connection. The multiplication value of the current frame error covariance matrix in the filtering de-reverberation operation, the variance matrix and the one-stage state space observation value corresponding to the current frame, and the one-stage state space conjugate transpose observation value corresponding to the current frame is approximately full-diagonalized, which greatly reduces the amount of parameters to be stored, thereby reducing the occupied storage resources.

[0177] In addition, the space-time correlation is decoupled to some extent through the Kronecker decomposition. In the multi-channel de-reverberation scene, by reasonably designing the two-stage filtering weights of the Kronecker decomposition, for example, decomposing into “channel number” dimension and “prediction order” dimension, the filtering de-reverberation operation can be regarded as two-stage filtering in the spatial dimension and the time dimension respectively, and finally combining them through the Kronecker product. When the autoregressive coefficients related to the channel tend to be stable, only the autoregressive coefficients related to the prediction can be updated.

[0178] Further, in the two-stage filtering de-reverberation operation, the autoregressive coefficient estimation of each stage can be selected according to the current device's computing and storage capabilities, such as the original Kalman filter, the block-diagonal or full-diagonal Kalman filter, or a combination of the two. At the same time, the length of the autoregressive coefficient Kronecker decomposition can also be flexibly selected. In addition to the selection mode of space-time decoupling, it can also be used as a two-stage iterative process with only one dimension reduced.

[0179] It should be understood that, although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.

[0180] Based on the same inventive concept, the embodiments of the present application further provide a voice dereverberation device for implementing the voice dereverberation method described above. The device provides a solution to the implementation scheme as described in the above method, and therefore the specific limitations in one or more voice dereverberation device embodiments provided below can refer to the limitations of the voice dereverberation method described above, which will not be repeated here.

[0181] In one exemplary embodiment, as shown in Figure 7 A voice dereverberation device is provided, wherein:

[0182] The frequency domain delay signal obtaining module 701 is configured to obtain a plurality of frames of voice frequency domain delay signals according to the voice frequency domain signal of the current frame and the voice frequency domain signals of the historical multiple frames;

[0183] The time domain dereverberation signal obtaining module 702 is configured to perform two-stage filter dereverberation operation according to the plurality of frames of voice frequency domain delay signals, a first-stage observation equation and a second-stage observation equation to obtain an estimated value of the second-stage voice frequency domain dereverberation signal of the current frame, thereby forming the voice time domain dereverberation signal of the current frame; the first-stage observation equation and the second-stage observation equation are obtained by Kronecker integral decomposition of the spatial regression coefficient in the initial observation equation;

[0184] The current-stage filter dereverberation operation module 703 is configured to, when performing the current-stage filter dereverberation operation, obtain a current-stage state space observation value corresponding to the current frame according to the plurality of frames of voice frequency domain delay signals, a current-stage observation equation and a spatial regression coefficient required in the current stage; update the prior estimate value of the current-frame error covariance diagonal matrix in the current stage to obtain a posterior estimate value of the current-frame error covariance diagonal matrix in the current stage; the current-stage current-frame error covariance diagonal matrix is obtained by approximately full diagonalization of the current-stage current-frame error covariance matrix; and update the estimated value of the current-stage voice frequency domain dereverberation signal of the current frame to obtain an estimated value of the current-stage voice frequency domain dereverberation signal of the current frame according to the current-stage state space observation value;

[0185] The spatial regression coefficient obtaining module 704 is configured to: if the current stage is the first stage, the spatial regression coefficient required by the current stage is a spatial regression coefficient estimate of a previous frame in the first stage; the spatial regression coefficient estimate of the previous frame in the first stage is obtained by the following steps: updating a variance diagonal matrix estimate of a speech frequency domain dereverberation signal of the previous frame in the first stage according to an estimate of the speech frequency domain dereverberation signal of the previous frame in the first stage; the variance diagonal matrix is obtained by approximately full diagonalization of a variance matrix; updating a Kalman gain coefficient of the previous frame in the first stage according to a state space observation of the previous frame in the first stage, a diagonal multiplication observation of the previous frame in the first stage, a priori estimate of an error covariance diagonal matrix of the previous frame in the first stage, and the variance diagonal matrix estimate of the speech frequency domain dereverberation signal of the previous frame in the first stage; the diagonal multiplication observation of the previous frame in the first stage is obtained by multiplying a state space observation of the previous frame in the first stage and a conjugate transpose observation of the state space of the previous frame in the first stage, and then performing approximately full diagonalization; and updating the spatial regression coefficient estimate of the previous frame in the first stage according to the estimate of the speech frequency domain dereverberation signal of the previous frame in the first stage, the Kalman gain coefficient of the previous frame in the first stage, and a spatial regression coefficient estimate of a previous previous frame in the first stage.

[0186] In one of the embodiments, the spatial regression coefficient obtaining module 704 is further configured to: if the current stage is the first stage, the spatial regression coefficient required by the current stage is a spatial regression coefficient estimate of a previous frame in the second stage; the spatial regression coefficient estimate of the previous frame in the second stage is obtained by the following steps: updating a variance diagonal matrix estimate of a speech frequency domain dereverberation signal of the previous frame in the second stage according to an estimate of the speech frequency domain dereverberation signal of the previous frame in the second stage; the variance diagonal matrix is obtained by approximately full diagonalization of a variance matrix; updating a Kalman gain coefficient of the previous frame in the second stage according to a state space observation of the previous frame in the second stage, a diagonal multiplication observation of the previous frame in the second stage, a priori estimate of an error covariance diagonal matrix of the previous frame in the second stage, and the variance diagonal matrix estimate of the speech frequency domain dereverberation signal of the previous frame in the second stage; the diagonal multiplication observation of the previous frame in the second stage is obtained by multiplying a state space observation of the previous frame in the second stage and a conjugate transpose observation of the state space of the previous frame in the second stage, and then performing approximately full diagonalization; and updating the spatial regression coefficient estimate of the previous frame in the second stage according to the estimate of the speech frequency domain dereverberation signal of the previous frame in the second stage, the Kalman gain coefficient of the previous frame in the second stage, and a spatial regression coefficient estimate of a previous previous frame in the second stage.

[0187] In one of the embodiments, the frequency domain delay signal obtaining module 701 is further configured to: obtain a speech time domain signal of a current frame through a microphone array; perform windowing processing on the speech time domain signal to obtain a windowed speech time domain signal; perform Fourier transform on the windowed speech time domain signal to obtain a speech frequency domain signal of the current frame; obtain a plurality of frames of speech frequency domain signals according to the speech frequency domain signal of the current frame and a plurality of frames of historical speech frequency domain signals; and obtain a plurality of frames of speech frequency domain delay signals according to the time delay coefficient and the plurality of frames of speech frequency domain signals.

[0188] In one of the embodiments, the time domain dereverberation signal obtaining module 702 is further configured to: perform posterior filtering update on a two-stage current frame speech frequency domain dereverberation signal estimation value, a two-stage state space observation value corresponding to the current frame, and a two-stage current frame Kalman gain coefficient to obtain a speech frequency domain dereverberation signal of the current frame; and perform inverse Fourier transform on the speech frequency domain dereverberation signal of the current frame to obtain a speech time domain dereverberation signal of the current frame.

[0189] In one of the embodiments, the device further comprises a priori estimation value obtaining module configured to: model a spatial regression coefficient required in the current stage as a first-order Markov process, and perform Kronecker integral decomposition to obtain a current stage state equation; obtain a variance of a current stage current frame complex Gaussian disturbance noise according to the current stage current frame complex Gaussian disturbance noise in the current stage state equation; obtain a posteriori estimation value of a current stage previous frame error covariance diagonal matrix; and update to obtain a priori estimation value of a current stage current frame error covariance diagonal matrix according to a sum of the posteriori estimation value of the current stage previous frame error covariance diagonal matrix and the variance of the current stage current frame complex Gaussian disturbance noise.

[0190] In one of the embodiments, the priori estimation value obtaining module is further configured to: multiply a current stage state space observation value corresponding to a previous frame and a current stage previous frame Kalman gain coefficient to obtain a current stage multiplication matrix; subtract a current stage unit matrix from the current stage multiplication matrix to obtain a current stage subtraction matrix; perform approximate full diagonalization on the current stage subtraction matrix to obtain a current stage subtraction diagonal matrix; and obtain a posteriori estimation value of a current stage previous frame error covariance diagonal matrix according to the current stage subtraction diagonal matrix and a priori estimation value of a current stage previous previous frame error covariance diagonal matrix.

[0191] In one of the embodiments, the spatial regression coefficient obtaining module 704 is further configured to: obtain a one-stage state space conjugate transpose observation corresponding to the current frame according to the one-stage state space observation corresponding to the current frame; multiply the one-stage state space observation corresponding to the current frame and the one-stage state space conjugate transpose observation corresponding to the current frame to obtain a one-stage multiplied observation corresponding to the current frame; and perform approximate full diagonalization on the one-stage multiplied observation corresponding to the current frame to obtain a one-stage diagonal multiplied observation corresponding to the current frame.

[0192] The modules in the voice dereverberation device can be implemented by software, hardware, or a combination thereof. The modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in the computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the modules.

[0193] In one example embodiment, a computer device, which can be a server, is provided, and an internal structure diagram of the computer device can be as shown in Figure 8 The computer device includes a processor, a memory, an input / output interface, and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store data of the embodiments of the voice dereverberation method. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through a network connection. The computer program is executed by the processor to implement a voice dereverberation method.

[0194] Those skilled in the art can understand that Figure 8 The structure shown in the above

[0195] In one embodiment, a computer device is also provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0196] In an embodiment, a computer readable storage medium is provided, having stored thereon a computer program which, when executed by a processor, implements the steps of any of the method embodiments described above.

[0197] In an embodiment, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the steps of any of the method embodiments described above.

[0198] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0199] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.

[0200] The technical features of the above embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.

[0201] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.

Claims

1. A speech dereverberation method, characterized in that, The method includes: Based on the audio domain signal of the current frame and the audio domain signals of multiple historical frames, the multi-frame audio domain delay signal is obtained. Based on the multi-frame speech domain delay signal, the first-stage observation equation, and the second-stage observation equation, a two-stage filtering and dérarization operation is performed to obtain the estimated value of the second-stage speech domain dérarization signal of the current frame. Then, through a posterior update operation, the speech time domain dérarization signal of the current frame is formed. The first-stage observation equation and the second-stage observation equation are obtained by Kronecker integral solutions of the spatial regression coefficients in the initial observation equation. Specifically, during the current stage of filtering and de-reverberation operation, the current stage state space observation value corresponding to the current frame is obtained based on the multi-frame speech domain delay signal, the current stage observation equation, and the spatial regression coefficients required for the current stage. The prior estimate of the current stage error covariance diagonal matrix is ​​updated based on the posterior estimate of the previous frame error covariance diagonal matrix. The current stage error covariance diagonal matrix is ​​obtained by approximately fully diagonalizing the current stage error covariance matrix. Finally, the estimated value of the current stage speech domain de-reverberation signal is updated based on the current stage state space observation value. If this stage is a two-stage process, then the spatial regression coefficients required for this stage are the estimated spatial regression coefficients of the current frame in the first stage. These estimated spatial regression coefficients are obtained through the following steps: based on the estimated value of the dereverberated signal in the speech domain of the current frame in the first stage, the estimated value of the variance diagonal matrix of the dereverberated signal in the speech domain of the current frame in the first stage is updated; the variance diagonal matrix is ​​obtained by approximately fully diagonalizing the variance matrix; based on the state space observations corresponding to the current frame in the first stage, the diagonally multiplied observations corresponding to the current frame in the first stage, and the prior estimate of the error covariance diagonal matrix of the current frame in the first stage... The Kalman gain coefficient of the current frame in the first stage is updated by combining the estimated value of the variance diagonal matrix of the dereverberation signal in the audio domain of the current frame in the first stage with the estimated value of the variance diagonal matrix of the current frame in the first stage. The diagonal multiplication observation value of the current frame in the first stage is obtained by multiplying the state space observation value of the current frame in the first stage with the conjugate transpose observation value of the state space of the current frame in the first stage and performing an approximate full diagonalization. The estimated value of the spatial regression coefficient of the current frame in the first stage is updated by combining the estimated value of the dereverberation signal in the audio domain of the current frame in the first stage, the Kalman gain coefficient of the current frame in the first stage, and the estimated value of the spatial regression coefficient of the previous frame in the first stage.

2. The method according to claim 1, characterized in that, If this stage is stage one, then the spatial regression coefficients required for this stage are the estimated values ​​of the spatial regression coefficients of the previous frame in stage two; obtaining the estimated values ​​of the spatial regression coefficients of the previous frame in stage two specifically includes: Based on the estimated value of the dereverberation signal in the audio domain of the second-stage previous frame, the estimated value of the variance diagonal matrix of the dereverberation signal in the audio domain of the second-stage previous frame is updated; the variance diagonal matrix is ​​obtained by approximating the full diagonalization of the variance matrix. The Kalman gain coefficients of the previous frame are updated based on the two-stage state space observations corresponding to the previous frame, the two-stage diagonal multiplication observations corresponding to the previous frame, the prior estimate of the two-stage previous frame error covariance diagonal matrix, and the estimate of the variance diagonal matrix of the two-stage previous frame speech domain dereverberation signal. The two-stage diagonal multiplication observations corresponding to the previous frame are obtained by multiplying the two-stage state space observations corresponding to the previous frame and the two-stage state space conjugate transpose observations corresponding to the previous frame and then performing an approximate full diagonalization. Based on the estimated value of the dereverberation signal in the audio domain of the second-stage upper frame, the estimated value of the Kalman gain coefficient of the second-stage upper frame, and the estimated value of the spatial regression coefficient of the second-stage upper frame, the estimated value of the spatial regression coefficient of the second-stage upper frame is updated.

3. The method according to claim 1, characterized in that, The step of obtaining multi-frame audio domain delay signals based on the audio domain signals of the current frame and the audio domain signals of historical frames includes: The speech time-domain signal of the current frame is obtained through a microphone array; The speech time-domain signal is windowed to obtain the windowed speech time-domain signal; Perform a Fourier transform on the windowed speech time-domain signal to obtain the speech audio domain signal of the current frame; Based on the audio domain signal of the current frame and the audio domain signals of multiple historical frames, multi-frame audio domain signals are obtained. Based on the time delay coefficient and the multi-frame audio domain signal, the multi-frame audio domain delay signal is obtained.

4. The method according to claim 1, characterized in that, Based on the estimated values ​​of the speech dereverberation signal in the audio domain of the current frame in the second stage, the speech temporal dereverberation signal of the current frame is formed, including: Based on the estimated value of the speech domain dereverberation signal of the current frame in the second stage, the second-stage state space observation value corresponding to the current frame, and the Kalman gain coefficient of the current frame in the second stage, the posterior filtering update is performed to obtain the speech domain dereverberation signal of the current frame. Perform an inverse Fourier transform on the speech audio domain dereverberation signal of the current frame to obtain the speech time domain dereverberation signal of the current frame.

5. The method according to claim 1, characterized in that, The step of updating the prior estimate of the current frame error covariance diagonal matrix based on the posterior estimate of the previous frame error covariance diagonal matrix includes: The spatial regression coefficients required for this stage are modeled as a first-order Markov process, and the Kronecker integral is performed to obtain the state equation for this stage. Based on the complex Gaussian perturbation noise of the current frame in the state equation of this stage, the variance of the complex Gaussian perturbation noise of the current frame in this stage is obtained. Obtain the posterior estimate of the diagonal matrix of the previous frame error covariance in this stage; Based on the sum of the posterior estimate of the error covariance diagonal matrix of the previous frame in this stage and the variance of the complex Gaussian perturbation noise of the current frame in this stage, the prior estimate of the error covariance diagonal matrix of the current frame in this stage is updated.

6. The method according to claim 5, characterized in that, The step of obtaining the posterior estimate of the diagonal matrix of the previous frame error covariance in this stage includes: Multiply the current state space observation corresponding to the previous frame with the Kalman gain coefficient of the previous frame in the current stage to obtain the multiplication matrix of the current stage; Subtract the identity matrix of this stage from the multiplication matrix of this stage to obtain the subtraction matrix of this stage; The subtraction matrix of this stage is approximately fully diagonalized to obtain the subtraction diagonal matrix of this stage; Based on the prior estimates of the subtracted diagonal matrix of this stage and the error covariance diagonal matrix of the previous frame in this stage, the posterior estimate of the error covariance diagonal matrix of the previous frame in this stage is obtained.

7. The method according to claim 1, characterized in that, Based on the one-stage state space observations corresponding to the current frame, and the one-stage state space conjugate transpose observations corresponding to the current frame, the one-stage diagonal multiplication observations corresponding to the current frame are obtained, including: Based on the one-stage state space observations corresponding to the current frame, obtain the one-stage state space conjugate transpose observations corresponding to the current frame; Multiply the one-stage state space observation corresponding to the current frame and the one-stage state space conjugate transpose observation corresponding to the current frame to obtain the one-stage multiplied observation corresponding to the current frame. The one-stage multiplication observation corresponding to the current frame is approximated as a full diagonalization to obtain the one-stage diagonal multiplication observation corresponding to the current frame.

8. A speech de-reverberation device, characterized in that, The device includes: The frequency domain delay signal acquisition module is used to obtain multi-frame audio domain delay signals based on the audio domain signal of the current frame and the audio domain signals of multiple historical frames. The temporal dereverberation signal acquisition module is used to perform two-stage filtering dereverberation operations based on the multi-frame speech domain delay signal, the first-stage observation equation, and the second-stage observation equation to obtain the estimated value of the second-stage speech domain dereverberation signal of the current frame, thus forming the speech temporal dereverberation signal of the current frame; the first-stage observation equation and the second-stage observation equation are obtained by Kronecker integral solutions of the spatial regression coefficients in the initial observation equation; The current-stage filtering and de-reverberation module is used to, during the current-stage filtering and de-reverberation operation, obtain the current-stage state-space observation value corresponding to the current frame based on the multi-frame speech domain delay signal, the current-stage observation equation, and the spatial regression coefficients required for the current stage; update the prior estimate of the current-stage error covariance diagonal matrix based on the posterior estimate of the previous frame's error covariance diagonal matrix; the current-stage error covariance diagonal matrix is ​​obtained by approximately fully diagonalizing the current-stage error covariance matrix; and update the estimate of the current-stage speech domain de-reverberation signal based on the current-stage state-space observation value. The spatial regression coefficient acquisition module is used to obtain the estimated spatial regression coefficients of the current frame in the first stage if the current stage is a two-stage stage. The estimated spatial regression coefficients of the current frame in the first stage are obtained through the following steps: based on the estimated value of the dereverberated signal in the speech domain of the current frame in the first stage, the estimated value of the variance diagonal matrix of the dereverberated signal in the speech domain of the current frame in the first stage is updated; the variance diagonal matrix is ​​obtained by approximately fully diagonalizing the variance matrix; based on the state space observations corresponding to the current frame in the first stage, the diagonally multiplied observations corresponding to the current frame in the first stage, and the error covariance of the current frame in the first stage... The prior estimate of the diagonal matrix and the estimate of the variance diagonal matrix of the current frame speech domain dereverberation signal in the first stage are used to update the Kalman gain coefficient of the current frame in the first stage. The first-stage diagonal multiplication observation corresponding to the current frame is obtained by multiplying the first-stage state space observation corresponding to the current frame and the first-stage state space conjugate transpose observation corresponding to the current frame and performing an approximate full diagonalization. The estimated spatial regression coefficient of the current frame in the first stage is updated based on the estimate of the speech domain dereverberation signal in the first stage, the Kalman gain coefficient of the current frame in the first stage, and the estimated spatial regression coefficient of the previous frame in the first stage.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and system for dereverberation based on Kalman filtering

    CN108172231A

  • Adaptive delay compensation for acoustic echo cancellation

    EP2493167A1