A method and device for dereverberation of speech signal
By obtaining the multi-channel voice signal and the filter coefficients at the previous moment, determining the filter coefficients at the current moment according to the update step size, and using this coefficient for dereverberation operation, the problems of large amount of calculation and poor effect in the prior art are solved, and real-time dereverberation and effect improvement of the multi-channel voice signal are achieved.
Patent Information
- Application Number
- CN202110909090.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-09
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2041-08-09
AI Technical Summary
The existing multi-channel voice signal dereverberation technology solution has a large amount of calculation, resulting in poor results and is difficult to apply in real time.
By obtaining the multi-channel voice signal and the filter coefficients at the previous moment, the filter coefficients at the current moment are determined according to the update step size, and the coefficients are used for dereverberation operations, reducing the calculation amount and improving real-time performance.
Real-time dereverberation of multi-channel voice signals is realized, improving the dereverberation effect and reducing the calculation amount.
Smart Images

Figure CN113611322B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech signal processing, and in particular to a speech signal dereverberation method and device. Background Art
[0002] Usually, when collecting or recording sound signals, the sound receiver not only receives the part of the sound waves emitted by the required sound source that arrives directly, but also receives the sound waves emitted by the sound source and transmitted through other channels, as well as the unnecessary sound waves (i.e., background noise) generated by other sound sources in the environment. In acoustics, the reflected wave with a delay time of more than about 50ms is called an echo, and the effect produced by the remaining reflected waves is called reverberation. The reverberation phenomenon will affect the reception effect of the desired sound signal, resulting in a deterioration in the performance of the acoustic receiving system. For example, reverberation will cause a significant decrease in the performance of the speech recognition system. Therefore, how to reduce the impact of reverberation on the sound receiving system, i.e., dereverberation, is a very important topic.
[0003] With the continuous promotion of microphone array technology, it is widely used in speech processing equipment, and the dereverberation problem has become the dereverberation of multi-channel speech signals. However, most of the existing dereverberation technical solutions estimate the reverberation components for each channel separately, resulting in poor dereverberation effect. However, due to the complexity and large amount of calculation of the multi-channel joint dereverberation algorithm, the multi-channel joint dereverberation algorithm is restricted by the application scenario. Therefore, how to improve the dereverberation of multi-channel speech signals is a technical problem that needs to be solved urgently. Summary of the invention
[0004] In view of this, the embodiments of the present application provide a method and device for dereverberation of a speech signal, so as to make full use of multi-channel speech signals for dereverberation, reduce the amount of calculation, and improve the dereverberation effect.
[0005] To achieve the above purpose, the technical solutions provided by the embodiments of the present application are as follows:
[0006] In a first aspect of an embodiment of the present application, a method for dereverberation of a speech signal is provided, the method comprising:
[0007] For a multi-channel speech signal at any time, obtain a first multi-channel speech signal corresponding to the i-th time and a first filter coefficient corresponding to the i-1-th time, where i=1, ..., n, n is a positive integer not less than 1, the first multi-channel speech signal is an M*1 matrix, the first filter coefficient is an M*M matrix, and M is the number of microphones;
[0008] Determine the second filter coefficient corresponding to the i-th moment according to the update step size, the first multi-channel speech signal and the first filter coefficient;
[0009] A dereverberated speech signal is obtained according to the first multi-channel speech signal and the second filter coefficient.
[0010] In a specific implementation manner, determining the second filter coefficient corresponding to the i-th moment according to the update step size, the first multi-channel signal, and the first filter coefficient includes:
[0011] Obtaining the update step length corresponding to the i-th moment, wherein the update step length is negatively correlated with i;
[0012] Determining a second multi-channel voice signal according to the first multi-channel voice signal, the order of the filter and the time delay of the filter;
[0013] Multiplying the update step size, the second multi-channel speech signal, and the conjugate transpose of the second multi-channel speech signal to obtain a first matrix;
[0014] Multiplying the update step size, the second multi-channel speech signal, and the conjugate transpose of the first multi-channel speech signal to obtain a second matrix;
[0015] A second filter coefficient is obtained according to the first matrix, the second matrix and the first filter coefficient.
[0016] In a specific implementation manner, obtaining a second filter coefficient according to the first matrix, the second matrix, and the first filter coefficient includes:
[0017] Adding the first matrix to the identity matrix to obtain a third matrix;
[0018] The third matrix is multiplied by the first filter coefficient and then the second matrix is subtracted to obtain the second filter coefficient.
[0019] In a specific implementation manner, obtaining a dereverberated speech signal according to the first multi-channel speech signal and the second filter coefficient includes:
[0020] Obtaining a third multi-channel signal by multiplying the first multi-channel speech signal, the delay of the filter and the second filter coefficient;
[0021] A difference between the first multi-channel signal and the third multi-channel signal is determined as a dereverberation signal.
[0022] In a specific implementation, the method further includes:
[0023] Inputting the dereverberation signal into a compensation model to obtain a compensation coefficient, wherein the compensation model is obtained by training according to training samples, and the training samples include a undernerized speech signal to be trained and a dereverberated speech signal to be trained;
[0024] A compensated dereverberation speech signal is obtained according to the compensation coefficient and the dereverberation speech signal.
[0025] In a specific implementation manner, the step of obtaining a compensated dereverberation speech signal according to the compensation coefficient and the dereverberation speech signal includes:
[0026] The compensation coefficient is multiplied by the dereverberated speech signal to obtain a compensated dereverberated speech signal.
[0027] In a specific implementation, when i=1, the first filter coefficient corresponding to the 0th time is determined by the unit matrix and the random diagonal matrix.
[0028] In a second aspect of an embodiment of the present application, a speech signal dereverberation device is provided, the device comprising:
[0029] A first acquisition unit is used to acquire, for a multi-channel speech signal at any moment, a first multi-channel speech signal corresponding to the i-th moment and a first filter coefficient corresponding to the i-1-th moment, wherein i=0, ..., n, n is a positive integer not less than 1, the first multi-channel speech signal is an M*1 matrix, the first filter coefficient is an M*M matrix, and M is the number of microphones;
[0030] a determining unit, configured to determine the second filter coefficient corresponding to the i-th moment according to an update step size, the first multi-channel speech signal, and the first filter coefficient;
[0031] The second acquisition unit is used to obtain a dereverberation speech signal according to the first multi-channel speech signal and the second filter coefficient.
[0032] In a third aspect of the embodiments of the present application, a voice device is provided, including: a processor, a memory;
[0033] The memory is used to store computer-readable instructions or computer programs;
[0034] The processor is used to read the computer-readable instructions or the computer program so that the device implements the speech signal dereverberation method described in the first aspect.
[0035] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, comprising instructions or a computer program, which, when executed on a computer, enables the computer to execute the speech signal dereverberation method described in the first aspect above.
[0036] It can be seen that the embodiments of the present application have the following beneficial effects:
[0037] When the embodiment of the present application performs a dereverberation operation on a speech signal at a certain moment, the first multi-channel speech signal corresponding to the moment and the filter coefficient corresponding to the previous adjacent moment, i.e., the first filter coefficient, are first obtained. The second filter coefficient corresponding to the current moment is determined according to the update step size, the first multi-channel speech signal and the first filter coefficient. Among them, the update step size is used to update the first filter coefficient to obtain the second filter coefficient. After obtaining the second filter coefficient, the first multi-channel speech signal is dereverberated using the second filter coefficient to obtain a dereverberated speech signal. It can be seen that when the embodiment of the present application performs a dereverberation operation on a multi-channel speech signal, the filter coefficient of the previous moment is used to determine the filter coefficient of the current moment, without using a large number of other additional parameters, reducing the amount of calculation, and achieving real-time dereverberation of the speech signal, thereby improving the quality of the speech signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A flow chart of a method for dereverberation of a speech signal provided in an embodiment of the present application;
[0039] Figure 2 A structural diagram of a speech signal dereverberation device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0040] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0041] To facilitate understanding of the technical solutions of the embodiments of the present application, the technical terms involved in the embodiments of the present application will be explained below.
[0042] Microphone arrays are systems composed of a certain number of acoustic sensors (usually microphones) used to sample and process the spatial characteristics of the sound field. Multi-channel refers to the multiple input multiple output (MIMO) of the speech processing system, that is, multi-channel data collection and processing is achieved with the help of microphone arrays.
[0043] Reverberation refers to the phenomenon that when sound waves propagate in a room, they are reflected by obstacles such as walls, ceilings, and floors. When the sound source stops making sound, the sound waves will be reflected and absorbed multiple times in the room before finally disappearing. This phenomenon is called reverberation. Appropriate reverberation will make the sound mellow, pleasant, and appealing. Too long a reverberation time will make the sound unclear and hard to hear. Too much reverberation will cause overlapping and masking of phonemes, which will seriously affect the effect of speech recognition, especially long-distance speech recognition.
[0044] The inventors have found through research that most traditional dereverberation technologies estimate the reverberation component of each channel separately to achieve dereverberation. In addition, since joint dereverberation of speech signals of each channel requires a large amount of calculation, the multi-channel joint dereverberation solution cannot be applied in real time.
[0045] Based on this, a method for dereverberation of a speech signal provided in an embodiment of the present application, for a multi-channel speech signal, obtains a multi-channel speech signal corresponding to each moment, that is, a first multi-channel speech signal, and a filter coefficient corresponding to the adjacent previous moment, that is, a first filter coefficient. The first filter coefficient is used to estimate the reverberation component in the multi-channel speech signal at the previous moment, so as to dereverberate the multi-channel speech signal at the previous moment. After obtaining the first multi-channel speech signal corresponding to the current moment and the first filter coefficient at the previous moment, the second filter coefficient corresponding to the current moment is determined according to the update step size, the first multi-channel speech signal and the first filter coefficient. After obtaining the second filter coefficient, the first multi-channel speech signal is processed using the second filter coefficient to obtain a dereverberation speech signal. That is, the embodiment of the present application calculates the filter coefficient at the current moment by using the filter coefficient at the previous moment, without introducing other calculation amounts, improving the calculation speed, reducing the delay, and thus being able to perform real-time dereverberation on the speech signal at the current moment, thereby improving the dereverberation effect.
[0046] For ease of understanding, a speech signal dereverberation method provided in an embodiment of the present application will be described below with reference to the accompanying drawings.
[0047] See also Figure 1 , which is a flow chart of a speech signal dereverberation method provided in an embodiment of the present application, such as Figure 1 As shown, the method may include:
[0048] S101: Obtain a first multi-channel speech signal corresponding to the i-th moment and a first filter coefficient corresponding to the i-1-th moment.
[0049] In this embodiment, for a multi-channel speech signal at any moment, the multi-channel speech signal corresponding to the moment and the first filter coefficient corresponding to the previous moment are obtained. Wherein, i=0,…,n, n is a positive integer not less than 1. That is, n observation points are pre-set, and a dereverberation operation is performed for each observation point. Wherein, the multi-channel speech signal is an M*1 matrix, the first filter coefficient is an M*M matrix, and M is the number of microphones. For ease of understanding, y(i) represents the first multi-channel signal corresponding to the i-th moment.
[0050] Among them, the order of the dereverberation filter coefficient for the multi-channel speech signal is L, and the filter coefficient corresponding to the i-th moment can be expressed as: G L(i)=[G0(i),G1(i),.....,G L-1 (i)],G l (i) where l = 0, ..., L-1, G l (i) is an M×M matrix, and the filter coefficient corresponding to the i-1th moment can be expressed as G L (i-1)=[G0(i-1),G1(i-1),.....,G L-1 (i-1)].
[0051] S102: Determine a second filter coefficient corresponding to the i-th moment according to the update step size, the first multi-channel speech signal and the first filter coefficient.
[0052] In this embodiment, after obtaining the first multi-channel speech signal corresponding to the i-th moment and the first filter coefficient corresponding to the i-1-th moment, the second filter coefficient corresponding to the i-th moment is determined according to the update step size, the first multi-channel speech signal and the first filter coefficient. The update step size is used to update the first filter coefficient to obtain the second filter coefficient. The update step size can be set according to the actual application situation, and the value is (0,1).
[0053] Specifically, the update step size can also be set to a variable related to i, which changes with the change of i. Among them, the update step size is negatively correlated with i. For example, the update step size is α(i)=0.85 / (i+1). Based on this, the present embodiment provides a method for determining the second filter coefficient, specifically: obtaining the update step size at the i-th moment; determining the second multi-channel voice signal according to the first multi-channel voice signal, the order of the filter and the delay of the filter; multiplying the update step size, the second multi-channel voice signal and the conjugate transpose of the second multi-channel voice signal to obtain a first matrix; multiplying the update step size, the second multi-channel voice signal and the conjugate transpose of the first multi-channel voice signal to obtain a second matrix; obtaining the second filter coefficient according to the first matrix, the second matrix and the first filter coefficient.
[0054] The second filter coefficient is obtained according to the first matrix, the second matrix and the first filter coefficient, including: adding the first matrix to the unit matrix to obtain a third matrix; multiplying the third matrix by the first filter coefficient and then subtracting the second matrix to obtain the second filter coefficient.
[0055] For easier understanding, see formula (1):
[0056] G L (i) = (I + α (i) * x L (i)*x L H (i))*G L (i-1)-α(i)*xL (i)*y H (i) (1)
[0057] Among them, α(i) represents the update step size at the i-th moment, x L (i) represents the second multi-channel speech signal, G L (i-1) represents the filter coefficient corresponding to the i-1th moment, y(i) represents the first multi-channel speech signal at the i-th moment, and H represents the conjugate transpose of the matrix.
[0058] From formula (1), we can know that α(i)*x L (i)*x L H (i) represents the first matrix, α(i)*x L (i)*y H (i) represents the second matrix, and I represents the unit matrix.
[0059] Among them, the second multi-channel speech signal x L (i) It can be expressed as:
[0060] x L (i) = [y(i-τ) T ,.....,y(i-τ-L+1) T ] T
[0061] Wherein, τ represents the delay of the filter, T represents the matrix transpose, and L represents the order of the filter. It should be noted that when i=1, the second filter coefficient corresponding to the first moment is:
[0062] G L (1) = (I + α (1) * x L (1)*x L H (1))*G L (0)-α(1)*x L (1)*y H (1)
[0063] Among them, G L (0) is obtained after initialization, which can be determined by the unit matrix and the random matrix. Specifically, the sum of the unit matrix and the random matrix is determined as the filter coefficient corresponding to the 0th moment.
[0064] S103: Obtain a dereverberated speech signal according to the first multi-channel speech signal and the second filter coefficient.
[0065] After the second filter coefficient corresponding to the i-th moment is obtained through the above steps, the first multi-channel speech signal is dereverberated using the second filter coefficient to obtain a dereverberated speech signal.
[0066] Specifically, a third multi-channel voice signal is obtained by multiplying the first multi-channel voice signal, the delay of the filter and the second filter coefficient; and the difference between the first multi-channel voice signal and the third multi-channel voice signal is determined as the dereverberation voice signal. The third multi-channel voice signal is the reverberation estimation, that is, the reverberation voice signal mixed in the first multi-channel voice signal. Specifically, refer to formula (2):
[0067]
[0068] Where, e(i) represents the dereverberation signal, y(i) represents the first multi-channel speech signal, G L (i) represents the second filter coefficient, y(i-τ-l) represents the speech signal corresponding to the delay when the first multi-channel speech signal passes through the filter, Represented as a third multi-channel speech signal.
[0069] Since, the filter coefficient can be expressed as:
[0070] G L (i)=[G0(i) T ,G1(i) T ,.....,G L-1 (i) T ] T
[0071] The speech signal after the filter can be expressed as:
[0072] x L (i) = [y(i-τ) T ,.....,y(i-τ-L+1) T ] T
[0073] Therefore, the dereverberated signal can be expressed as:
[0074] e(i)=y(i)-G L (i) H *x L (i)
[0075] It can be seen that for a multi-channel speech signal at any time, the above formula (1) and formula (2) can be used to perform a dereverberation operation to obtain a dereverberated speech signal.
[0076] It should be noted that, in order to further perform denoising on the obtained dereverberation speech signal to obtain a purer speech signal, the dereverberation speech signal can also be input into a compensation model to obtain a compensation coefficient; and a compensated dereverberation speech signal is obtained according to the compensation coefficient and the dereverberation speech signal. That is, the dereverberation speech signal is further dereverberated by the compensation coefficient. Specifically, the compensation coefficient is multiplied by the dereverberation speech signal to obtain a compensated dereverberation speech signal. Specifically, refer to formula (3):
[0077] e(i)=δ(i)*(y(i)-G L (i) H *x L (i)) (3)
[0078] Wherein, δ(i) represents the compensation coefficient, and e(i) represents the speech signal after dereverberation.
[0079] The compensation model is obtained based on training samples, which include undisturbed speech signals to be trained and dereverberated speech signals to be trained. That is, during training, a large number of undisturbed speech signals and dereverberated speech signals are obtained, and the two are input into the initial network model to obtain the initial network model to learn the corresponding relationship between the two, thereby obtaining the compensation model. The compensation model can be a neural network model, such as a long short-term memory neural network (Long-Short Term Memory, LSTM).
[0080] It can be seen that, through the method provided by the embodiment of the present application, when a de-reverberation operation is performed on a speech signal at a certain moment, the first multi-channel speech signal corresponding to the moment and the filter coefficient corresponding to the adjacent previous moment, that is, the first filter coefficient, are first obtained. The second filter coefficient corresponding to the current moment is determined according to the update step size, the first multi-channel speech signal and the first filter coefficient. Among them, the update step size is used to update the first filter coefficient to obtain the second filter coefficient. After obtaining the second filter coefficient, the first multi-channel speech signal is de-reverberated using the second filter coefficient to obtain a de-reverberated speech signal. It can be seen that when the embodiment of the present application performs a de-reverberation operation on a multi-channel speech signal, the filter coefficient of the previous moment is used to determine the filter coefficient of the current moment, without the need to use a large number of other additional parameters, thereby reducing the amount of calculation, and achieving real-time de-reverberation of the speech signal, thereby improving the quality of the speech signal.
[0081] Based on the above method embodiment, the embodiment of the present application also provides a speech dereverberation device, which will be described below in conjunction with the accompanying drawings.
[0082] See also Figure 2, which is a schematic diagram of the structure of a speech dereverberation device provided in an embodiment of the present application, such as Figure 2 As shown, the device may include: a first acquiring unit 201 , a determining unit 202 , and a second acquiring unit 203 .
[0083] The first acquisition unit 201 is used to acquire, for a multi-channel speech signal at any moment, a first multi-channel speech signal corresponding to the i-th moment and a first filter coefficient corresponding to the i-1-th moment, where i=0,…,n, n is a positive integer not less than 1, the first multi-channel speech signal is an M*1 matrix, the first filter coefficient is an M*M matrix, and M is the number of microphones;
[0084] A determination unit 202, configured to determine a second filter coefficient corresponding to the i-th moment according to an update step size, the first multi-channel speech signal, and the first filter coefficient;
[0085] The second acquisition unit 203 is configured to obtain a dereverberated speech signal according to the first multi-channel speech signal and the second filter coefficient.
[0086] In a specific implementation method, the determination unit 202 is specifically used to obtain the update step size corresponding to the i-th moment, and the update step size is negatively correlated with i; determine the second multi-channel voice signal according to the first multi-channel voice signal, the order of the filter and the delay of the filter; multiply the update step size, the second multi-channel voice signal and the conjugate transpose of the second multi-channel voice signal to obtain a first matrix; multiply the update step size, the second multi-channel voice signal and the conjugate transpose of the first multi-channel voice signal to obtain a second matrix; obtain the second filter coefficient according to the first matrix, the second matrix and the first filter coefficient.
[0087] In a specific implementation, the determination unit 202 is specifically configured to add the first matrix to the unit matrix to obtain a third matrix; multiply the third matrix by the first filter coefficient and then subtract the second matrix to obtain the second filter coefficient.
[0088] In a specific implementation, the second acquisition unit 203 is specifically configured to obtain a third multi-channel signal by multiplying the first multi-channel speech signal, the delay of the filter and the second filter coefficient; and determine the difference between the first multi-channel signal and the third multi-channel signal as the dereverberation signal.
[0089] In a specific implementation, the device further includes: a third acquisition unit and a fourth acquisition unit;
[0090] A third acquisition unit, configured to input the dereverberation signal into a compensation model to obtain a compensation coefficient, wherein the compensation model is obtained by training according to training samples, wherein the training samples include an underner speech signal to be trained and a dereverberation speech signal to be trained;
[0091] A fourth acquisition unit is used to obtain a compensated dereverberation speech signal according to the compensation coefficient and the dereverberation speech signal.
[0092] In a specific implementation manner, the fourth acquisition unit is specifically configured to multiply the compensation coefficient by the dereverberation speech signal to obtain a compensated dereverberation speech signal.
[0093] In a specific implementation, when i=1, the first filter coefficient corresponding to the 0th time is determined by the unit matrix and the random diagonal matrix.
[0094] It should be noted that the implementation of each unit in this embodiment can refer to the above method embodiment, and this embodiment will not be repeated here.
[0095] In addition, an embodiment of the present application further provides a voice device, including: a processor, a memory;
[0096] The memory is used to store computer-readable instructions or computer programs;
[0097] The processor is used to read the computer-readable instructions or the computer program so that the device implements the speech signal dereverberation method as described.
[0098] An embodiment of the present application provides a computer-readable storage medium, including instructions or computer programs, which, when executed on a computer, enable the computer to execute the above-mentioned method for dereverberation of speech signals.
[0099] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system or device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0100] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0101] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0102] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0103] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for dereverberation of a speech signal, characterized in that: The method comprises: For a multi-channel speech signal at any time, obtain a first multi-channel speech signal corresponding to the i-th time and a first filter coefficient corresponding to the i-1-th time, where i=1, ..., n, n is a positive integer not less than 1, the first multi-channel speech signal is an M*1 matrix, the first filter coefficient is an M*M matrix, and M is the number of microphones; Determine the second filter coefficient corresponding to the i-th moment according to the update step size, the first multi-channel speech signal and the first filter coefficient; Obtain a dereverberated speech signal according to the first multi-channel speech signal and the second filter coefficient; The step of determining the second filter coefficient corresponding to the i-th moment according to the update step size, the first multi-channel signal, and the first filter coefficient includes: Obtaining the update step length corresponding to the i-th moment, wherein the update step length is negatively correlated with i; Determining a second multi-channel voice signal according to the first multi-channel voice signal, the order of the filter and the time delay of the filter; Multiplying the update step size, the second multi-channel speech signal, and the conjugate transpose of the second multi-channel speech signal to obtain a first matrix; Multiplying the update step size, the second multi-channel speech signal, and the conjugate transpose of the first multi-channel speech signal to obtain a second matrix; Obtaining a second filter coefficient according to the first matrix, the second matrix, and the first filter coefficient; wherein obtaining the second filter coefficient according to the first matrix, the second matrix, and the first filter coefficient comprises: Adding the first matrix to the identity matrix to obtain a third matrix; The third matrix is multiplied by the first filter coefficient and then the second matrix is subtracted to obtain the second filter coefficient.
2. The method according to claim 1, characterized in that The step of obtaining a dereverberated speech signal according to the first multi-channel speech signal and the second filter coefficient comprises: Obtaining a third multi-channel signal by multiplying the first multi-channel speech signal, the delay of the filter and the second filter coefficient; A difference between the first multi-channel signal and the third multi-channel signal is determined as a dereverberation signal.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: Inputting the dereverberation signal into a compensation model to obtain a compensation coefficient, wherein the compensation model is obtained by training according to training samples, and the training samples include a undernerized speech signal to be trained and a dereverberated speech signal to be trained; A compensated dereverberation speech signal is obtained according to the compensation coefficient and the dereverberation speech signal.
4. The method according to claim 3, characterized in that The step of obtaining a compensated dereverberation speech signal according to the compensation coefficient and the dereverberation speech signal comprises: The compensation coefficient is multiplied by the dereverberated speech signal to obtain a compensated dereverberated speech signal.
5. The method according to claim 1, characterized in that When i=1, the first filter coefficient corresponding to the 0th time is determined by the unit matrix and the random diagonal matrix.
6. A speech signal dereverberation device, characterized in that: The device comprises: A first acquisition unit is used to acquire, for a multi-channel speech signal at any moment, a first multi-channel speech signal corresponding to the i-th moment and a first filter coefficient corresponding to the i-1-th moment, where i=0, ..., n, n is a positive integer not less than 1, the first multi-channel speech signal is an M*1 matrix, the first filter coefficient is an M*M matrix, and M is the number of microphones; a determining unit, configured to determine the second filter coefficient corresponding to the i-th moment according to an update step size, the first multi-channel speech signal, and the first filter coefficient; A second acquisition unit, configured to obtain a dereverberated speech signal according to the first multi-channel speech signal and the second filter coefficient; The determination unit is specifically used to obtain the update step length corresponding to the i-th moment, and the update step length is negatively correlated with i; determine the second multi-channel voice signal according to the first multi-channel voice signal, the order of the filter and the delay of the filter; multiply the update step length, the second multi-channel voice signal and the conjugate transpose of the second multi-channel voice signal to obtain a first matrix; multiply the update step length, the second multi-channel voice signal and the conjugate transpose of the first multi-channel voice signal to obtain a second matrix; add the first matrix and the unit matrix to obtain a third matrix; multiply the third matrix and the first filter coefficient and then subtract the second matrix to obtain the second filter coefficient.
7. A voice device, characterized in that: Including: processor, memory; The memory is used to store computer-readable instructions or computer programs; The processor is used to read the computer-readable instructions or the computer program so that the device implements the speech signal dereverberation method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, comprising instructions or a computer program, which, when executed on a computer, enables the computer to execute the speech signal dereverberation method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-channel speech enhancement method and device
CN113030862A
Method and apparatus for removing noise ofmulti-channel voice signal
KR1020070050694A