Speech enhancement method and device based on double microphone array, equipment and medium

By calculating the coherence information between two microphones in the speech enhancement system, the problem of redundant computing resources in the prior art is solved, achieving efficient speech enhancement under limited resource conditions and ensuring clear extraction of the target speech.

CN122392555APending Publication Date: 2026-07-14ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG GEELY HLDG GRP CO LTD
Filing Date
2026-05-20
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

In existing speech enhancement systems, the computation process of spatial filters and post-filters is redundant, resulting in a waste of computational resources and difficulty in efficiently suppressing environmental noise and background noise.

Method used

By calculating the coherence information between the two microphones, it is uniformly applied to update the noise covariance matrix of the spatial filter and the gain calculation of the post-filter, thereby realizing the reuse of coherence information and reducing the computational resource requirements.

Benefits of technology

It saves computing resources, improves the computational efficiency of the speech enhancement system, and ensures clear extraction of the target speech under limited resource conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122392555A_ABST
    Figure CN122392555A_ABST
Patent Text Reader

Abstract

The specification provides a speech enhancement method based on a dual microphone array, the method comprising: acquiring surrounding acoustic signals collected by two microphones in a microphone array respectively. Based on the acoustic signals collected by the two microphones respectively, coherence information between the two microphones is determined. Based on the coherence information, a speech activity state is judged, and in the case that the speech activity state is a speech missing state, a noise covariance matrix of the spatial filter is updated; and a gain function value of the post-filter is determined based on the coherence information. The two collected acoustic signals are subjected to beamforming processing by the spatial filter to obtain acoustic signals of at least one beam direction. The acoustic signals of any beam direction are subjected to filtering processing by the post-filter to obtain an estimation result of a speech component contained in the acoustic signals of the any beam direction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of acoustic signal processing technology, and in particular to speech enhancement methods, apparatuses, devices and media based on dual-microphone arrays. Background Technology

[0002] In daily life, target speech signals are easily interfered with by background noise such as environmental noise, echoes, or conversations during the acquisition process, leading to the failure of target speech recognition. To solve this problem, speech enhancement technology has emerged. Its core task is to accurately extract target speech from a specific direction from a noisy background and suppress irrelevant interference to the greatest extent possible.

[0003] An efficient speech enhancement system typically relies on microphone array hardware to collect ambient acoustic signals, uses spatial filters to pick up sound from a specific direction, and then uses a post-filter to further filter the sound from that direction to obtain clean speech components, ensuring that the system outputs clear, pure, and high-quality speech from that specific direction.

[0004] The main function of the spatial filter is to determine whether the noise covariance matrix needs to be updated based on the speech activity detection results, and to identify and suppress the sound signal in the direction of interference by updating the noise covariance matrix when necessary. The post-filter, on the other hand, refines and filters out the residual noise after spatial filtering by calculating the gain function value.

[0005] In existing solutions, the logic for determining whether the noise covariance matrix of the spatial filter needs to be updated and the logic for calculating the gain function value of the post-filter are often regarded as two completely decoupled calculation processes, each relying on independent parameters for calculation, resulting in redundant computing resources. Summary of the Invention

[0006] To overcome the problems existing in related technologies, this specification provides a speech enhancement method, apparatus, device and medium based on a dual-microphone array.

[0007] According to a first aspect of the embodiments of this specification, a speech enhancement method based on a dual-microphone array is provided. The method is applied to a speech enhancement system, the speech enhancement system including a microphone array, a spatial filter, and a post-filter. The method includes: Acquire the ambient acoustic signals collected by two microphones in the microphone array; Based on the acoustic signals collected by the two microphones respectively, the coherence information between the two microphones is determined; Based on the coherence information, the speech activity state is determined. If the speech activity state is a speech absence state, the noise covariance matrix of the spatial filter is updated. The gain function value of the post-filter is determined based on the coherence information. The speech absence state indicates that the acoustic signal does not contain any speech components. The two acquired acoustic signals are processed by beamforming through the spatial filter to obtain an acoustic signal with at least one beam orientation. The acoustic signal from any beam direction is filtered by the post-filter to obtain an estimate of the speech components contained in the acoustic signal from any beam direction.

[0008] According to a second aspect of the embodiments of this specification, a speech enhancement device based on a dual-microphone array is provided. The device is applied to a speech enhancement system, the speech enhancement system including a microphone array, a spatial filter, and a post-filter. The device includes: The acoustic signal acquisition module is used to acquire the ambient acoustic signals collected by the two microphones in the microphone array. The coherence information calculation module is used to determine the coherence information between the two microphones based on the acoustic signals collected by the two microphones respectively. The parameter calculation module is used to determine the speech activity state based on the coherence information, update the noise covariance matrix of the spatial filter when the speech activity state is a speech absence state, and determine the gain function value of the post-filter based on the coherence information; the speech absence state indicates that the acoustic signal does not contain speech components. A spatial filtering module is used to perform beamforming processing on the two acquired acoustic signals through the spatial filter to obtain an acoustic signal with at least one beam orientation. The post-filtering module is used to filter the acoustic signal in any beam direction through the post-filter to obtain an estimation result of the speech components contained in the acoustic signal in any beam direction.

[0009] According to a third aspect of the embodiments of this specification, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method as described in the first aspect.

[0010] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method as described in the first aspect.

[0011] The technical solutions provided in the embodiments of this specification may include the following beneficial effects: In the embodiments described in this specification, the coherence information between two microphones is calculated and simultaneously applied to speech activity detection required for updating the noise covariance matrix of the spatial filter, as well as to the gain calculation of the post-filter. By calculating the coherence information once and applying it to both the spatial filter and the post-filter, the calculation results of the coherence information are reused, saving computational resources and reducing the computational resource requirements of the speech enhancement system.

[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description

[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this specification and, together with the description, serve to explain the principles of this specification.

[0014] Figure 1 This is a schematic diagram illustrating a general speech enhancement system according to an exemplary embodiment of this specification.

[0015] Figure 2 This is a flowchart illustrating a speech enhancement method based on a dual-microphone array according to an exemplary embodiment of this specification.

[0016] Figure 3 This specification is a schematic diagram illustrating the placement of an in-vehicle microphone array according to an exemplary embodiment.

[0017] Figure 4 This is a simplified schematic diagram of the placement of an in-vehicle microphone array facing the left side of the second row, according to an exemplary embodiment of this specification.

[0018] Figure 5 This is a schematic diagram illustrating an improved speech enhancement system according to an exemplary embodiment of this specification.

[0019] Figure 6 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of this specification.

[0020] Figure 7 This is a block diagram illustrating a voice enhancement device based on a dual-microphone array according to an exemplary embodiment of this specification. Detailed Implementation

[0021] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.

[0022] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0023] It should be understood that although the terms first, second, third, etc., may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0024] In daily life, target speech signals are easily interfered with by background noise such as environmental noise, echoes, or conversations during the acquisition process, leading to the failure of target speech recognition. To solve this problem, speech enhancement technology has emerged. Its core task is to accurately extract target speech from a specific direction from a noisy background and suppress irrelevant interference to the greatest extent possible.

[0025] An efficient speech enhancement system typically relies on microphone array hardware to collect ambient acoustic signals, uses spatial filters to pick up sound from a specific direction, and then uses a post-filter to further filter the sound from that direction to obtain clean speech components, ensuring that the system outputs clear, pure, and high-quality speech from that specific direction.

[0026] A single microphone can only pick up the intensity and frequency of sound, but cannot determine the direction of the sound source. A microphone array is a device composed of multiple microphones. Because it has multiple microphones, a microphone array can locate the sound source by the time difference and intensity difference of sound waves arriving at different microphones. A spatial filter picks up sound from a specific beam direction and suppresses sound from other directions to meet the need to focus on the direction of the sound source in the target area. The beam direction can be a preset fixed direction, for example, fixing the direction of sound pickup in a certain direction of the microphone array. Alternatively, the direction from which the target speech is emitted can be determined as the beam direction, in which case the direction from which the target speech is emitted must be determined first.

[0027] The main function of a spatial filter is to pick up the acoustic signal of the beam direction. However, the acoustic signal of the beam direction still contains a lot of non-speech noise components. Therefore, a post-filter is still needed to further filter the acoustic signal output by the spatial filter in order to obtain clean speech components.

[0028] like Figure 1 As shown, an exemplary voice enhancement system 1 includes a microphone array 01, a spatial filter 02, and a post-filter 03. The microphone array 01 can be a dual-microphone array, that is, it contains two microphones.

[0029] Two microphones in microphone array 01 collect ambient sound and process it into acoustic signals, which are then input into spatial filter 02. Spatial filter 02 performs beamforming processing on the two collected acoustic signals to obtain acoustic signals from at least one beam direction. The acoustic signals from each beam direction are then input into post-filter 03. Post-filter 03 filters the acoustic signal from any beam direction to obtain an estimate of the speech components contained in the acoustic signal from that beam direction.

[0030] A common application scenario for voice enhancement systems is voice interaction between users and electronic devices. For example, voice interaction methods may include, but are not limited to, users controlling electronic devices through voice commands, users making remote calls through the electronic devices, and electronic devices dictating users' conversations.

[0031] Electronic devices include, but are not limited to, smart home devices, in-vehicle systems, wearable devices, computers, mobile phones, and other devices.

[0032] Taking a cockpit scenario as an example, when a driver attempts to interact with the in-vehicle system via voice commands while driving at high speed, the cabin is filled with intense ambient noise and conversations from other passengers. Assuming a microphone array is positioned in front of the driver, capturing ambient acoustic signals, these signals include environmental noise, the driver's voice commands, and conversations from other passengers. By focusing the beam direction in the driver's direction and using a spatial filter to pick up the acoustic signals emanating from that direction while suppressing conversations from other passengers and ambient noise, the system can concentrate on the acoustic signals from the driver's direction. Then, a post-filter further filters out residual ambient noise from this direction, retaining only the driver's voice, thus ensuring the in-vehicle system can accurately recognize the voice commands.

[0033] Taking smart home devices as an example, a user issues a voice command to a smart speaker from one side of the living room, while a television on the other side is playing a news program. The microphone array on the smart speaker collects ambient acoustic signals, which simultaneously contain the voice command and the sound of the news program. By automatically locating the user's sound source, the target direction is determined. Then, a spatial filter picks up the acoustic signal from that target direction, and a post-filter further processes the noise reduction of the acoustic signal from that target direction, retaining only the user's voice component, thus ensuring that the smart speaker can accurately recognize the user's voice command.

[0034] like Figure 2 As shown, Figure 2 This is a flowchart illustrating a speech enhancement method based on a dual-microphone array according to an exemplary embodiment, which can be applied to... Figure 1 The speech enhancement system 1 includes steps 201-205: Step 201: Acquire the ambient acoustic signals collected by the two microphones in the microphone array.

[0035] Step 202: Determine the coherence information between the two microphones based on the acoustic signals collected by each of the two microphones.

[0036] Step 203: Determine the speech activity state based on the coherence information; if the speech activity state is a speech absence state, update the noise covariance matrix of the spatial filter; and determine the gain function value of the post-filter based on the coherence information; the speech absence state indicates that the acoustic signal does not contain speech components.

[0037] Step 204: Perform beamforming processing on the two acquired acoustic signals using the spatial filter to obtain an acoustic signal with at least one beam orientation.

[0038] Step 205: Filter the acoustic signal of any beam direction using the post-filter to obtain an estimation result of the speech components contained in the acoustic signal of any beam direction.

[0039] In this embodiment, the coherence information of the two microphones is calculated and simultaneously applied to speech activity detection required to update the noise covariance matrix of the spatial filter, as well as to the gain calculation of the post-filter. By calculating the coherence information once and applying it to both the spatial filter and the post-filter, the calculation results of the coherence information are reused, saving computational resources and reducing the computational resource requirements of the speech enhancement system.

[0040] Next, we will take the vehicle's cockpit system as an example to introduce this solution.

[0041] To address the challenges of in-vehicle space and cost, automakers commonly use dual-microphone arrays when designing microphone arrays, taking into account factors such as placement and wiring. Furthermore, considering the computational and storage resources allocated to voice applications by the automotive chip, the microphone arrays near each seat in the cabin are also predominantly dual-microphone. This approach leverages the spatial filtering capabilities of the microphone array while mitigating the constraints of the automotive chip's hardware resources. By considering in-vehicle space, cost, and chip availability, and without increasing costs or the complexity of microphone array wiring, this approach reduces voice distortion in the target direction, obtains clearer voice signals from that direction, and improves the quality of in-vehicle voice interaction scenarios.

[0042] This solution only requires two microphones during operation, which reduces the number of microphones needed and makes it directly applicable to hardware scenarios with dual-microphone arrays commonly used in vehicles.

[0043] Regarding the deployment of microphone arrays in the vehicle cockpit, such as... Figure 3 As shown, microphone arrays can be deployed in the vehicle cabin 3 for the driver, front passenger, and rear left and right passengers, respectively. Each microphone array can be specifically used to acquire the voice of a person in a specific location. For example, microphone array 1 is specifically used to acquire the driver's voice, microphone array 2 is specifically used to acquire the front passenger's voice, microphone array 3 is specifically used to acquire the voice of the passenger in the left seat, and microphone array 4 is specifically used to acquire the voice of the passenger in the right seat.

[0044] In one embodiment, a microphone array can be deployed anywhere in the vehicle cabin 3. For example, the dual microphone array in the first row can be placed near the reading lights, in the ceiling diagonally above the driver and passenger seats, or near the left and right door frames. The dual microphone array in the second row can be placed near the left and right door frames.

[0045] For any microphone array, the beam orientation used by the spatial filter for beamforming can be singular, and this beam orientation can be determined based on a fixed spatial relationship between the microphone array and the seat it serves. For example, for a driver's seat microphone array, the beam orientation can be determined based on the spatial relationship between the driver's seat and the microphone array in that driver's seat. Similarly, for a passenger seat microphone array, the beam orientation can be determined based on the spatial relationship between the passenger seat and the microphone array in that passenger seat.

[0046] In this embodiment, by fixing the beam orientation, signal pickup is performed in a fixed direction each time, omitting the computational logic of determining the pickup direction based on sound source localization each time, thus reducing algorithm complexity from the algorithm design stage. Furthermore, the fixed beam orientation and the limitation of the number of beam orientations to one make it suitable even considering the movement of passengers in their seats, especially in scenarios where the relative positions of passengers and the microphone array do not change significantly. This approach can still meet the needs of vehicle-based voice recognition scenarios while reducing computational resources.

[0047] like Figure 4 As shown, the explanation focuses only on the microphone array facing the left side of the second row within cockpit 3. It is understood that the explanation is independent of deployment location and can be applied to microphone arrays in other locations within cockpit 3, as well as to electronic devices in other application scenarios.

[0048] In one embodiment, the spatial filter can be an MVDR filter. Of course, any spatial filter based on the noise covariance matrix principle can also be used as the spatial filter in this scheme.

[0049] Combination Figure 5 The working principle of speech enhancement system 1 will be explained using an MVDR filter as an example. Figure 4 In this example, mic1 is used as the reference microphone. Of course, mic2 can also be used as the reference microphone. In subsequent formulas that involve calculations for only one microphone, mic1 will be used as the reference microphone.

[0050] Both microphones mic1 and mic2 in microphone array 01 are used to collect acoustic signals generated inside the cockpit.

[0051] In Equation 1, the acoustic signal as a whole includes the target speech signal and the noise signal. The target speech signal can be understood as the speech component contained in the acoustic signal at the beam direction to be picked up. The noise signal can be understood as the acoustic signal outside the beam direction, as well as the non-speech component contained in the acoustic signal at the beam direction. The goal of the speech enhancement system is to extract the target speech signal from the acoustic signal. For any acoustic signal acquired by a microphone, it is modeled as Equation 1: Formula 1 in, Indicates mic Acoustic signals collected, This indicates the target speech signal contained within the acoustic signal. This indicates the noise signal contained within the acoustic signal. This represents a time frame, and it is assumed that the speech signal and the noise signal are uncorrelated. The clean speech picked up by the reference microphone is set as the target speech signal.

[0052] Indicates the first The relative delay between microphone 2 (mic2) and reference microphone 1 (mic1) can be expressed as Equation 2: Formula 2 in, Indicates the distance between the two microphones. Indicates the speed of sound. The incident angle of the target speech, i.e. the beam orientation, can be determined based on the fixed spatial relationship between the microphone array and the seats served by the microphone array. After measurement, it is used as a preset value and does not need to be recalculated subsequently.

[0053] Equation 1 is transformed using a Fourier transform (for example, a Short-Time Fourier Transform (STFT)) to obtain the expression in the frequency domain, as shown in Equation 3: Formula 3 in, , , They are , , Frequency domain result after STFT (STFT involves windowing, so the relevant description is omitted here). It is the imaginary unit. It is frequency The corresponding angular frequency, This represents discrete frequency bands. Equation 3 can be expressed in vector form as shown in Equation 4: in, A vector representation of the acoustic signal captured by the microphone. It is a vector representation of the noise signal. It is the steering vector, superscript This represents the transpose operation. MVDR is a transpose operation on the microphone input vector. Perform a linear filtering operation to obtain the beamforming output. ,Right now Let be the acoustic signal at any beam orientation. As shown in formula (5): Formula 5 The weights for MVDR beamforming are obtained by minimizing the output noise power under the constraint of distortion-free speech signal in the beam orientation, and their expression is shown in Equation 6: superscript This indicates the Hermitian Transposition. It is the covariance matrix of the noise signal, i.e., the noise covariance matrix, with dimensions of . , To represent the expected value, superscript This represents matrix inversion. As shown in Equation 6, the covariance matrix of the noise signal needs to be estimated in order to obtain the weights of MVDR.

[0054] and These represent scenario assumptions for a speech absence state and a speech presence state, respectively. The speech absence state indicates that the acoustic signal does not contain speech components, while the speech presence state indicates that the acoustic signal contains speech components.

[0055] Formula 7 can be used to achieve the following: The estimation is as follows: When the speech activity state is a speech absence state, the noise covariance matrix of the spatial filter is updated; when the speech activity state is a speech presence state, the noise covariance matrix of the spatial filter is not updated.

[0056] Formula 7 This represents the time smoothing factor, which is a preset constant. (Subscript) Indicates the current time frame. This represents the previous time frame. The determination of whether speech is missing or present can be achieved through the Voice Activity Detection (VAD) unit 05. Specifically, the VAD unit 05 determines the speech activity state; if the speech activity state is missing, the noise covariance matrix is ​​updated; if the speech activity state is present, no update is made to the noise covariance matrix. Beamforming is a spatial filtering process and cannot completely suppress noise. Therefore, a post-filtering unit 03 needs to be added after MVDR to further improve interference suppression capabilities. The gain function of the post-filter 03 is denoted as... Then, the estimation result of the speech component contained in the acoustic signal of beam orientation obtained by post-filtering can be expressed by Equation 8.

[0057] Formula 8 in, Indicates the target speech signal The estimation results.

[0058] As shown in Formulas 1-8 above, the application of spatial filter 02 and post-filter 03 involves speech activity state detection and the calculation of the gain function of post-filter 03. Furthermore, the speech activity state detection result serves the noise covariance matrix, used to determine whether to update the noise covariance matrix of the current frame.

[0059] The innovation of this solution lies in simultaneously applying the coherence information between the two microphones to the speech activity detection unit 05 and the post-filter 03, thereby reducing the algorithm complexity. For example... Figure 5 As shown, the coherence calculation unit 04 is used to calculate the coherence information between the two microphones. The speech activity detection unit 05 is used to determine the speech activity state based on the coherence information, and the determination result is used to decide whether to update. Post-filter 03 is used to calculate based on coherence information. .

[0060] Figure 5 The inverse Fourier transform in the context can be exemplified by ISTFT (Inverse Short-Time Fourier Transform), which is used to convert frequency domain signals to time domain signals (ISTFT involves windowing, which is omitted here). The target speech signal The estimation is the result of estimating the speech components contained in the acoustic signal of the beam azimuth.

[0061] The coherence information between the two microphones can be represented by Equation 9: Formula 9 in, It is the cross-power spectral density (CSD) of the acoustic signals input from the two microphones. and These are the power spectral density (PSD) of the acoustic signals input from the dual microphones, calculated using Equation 10: Formula 10 Among them, the superscript " " indicates taking the conjugate. Using the ergodic assumption, we can obtain the recursive average representation of Equation 10, as shown in Equation 11: Formula 11 The time smoothing factor is a preset constant.

[0062] In one embodiment, the coherence information is related to the estimation of the signal-to-noise ratio of the two microphone inputs (assuming that the signal-to-noise ratios of the two microphones are comparable), as shown in Equation 12.

[0063] Formula 12 in, The true signal-to-noise ratio The true signal-to-noise ratio (SNR) can be estimated by the ratio between the PSD of the speech signal and the PSD of the noise signal, expressed as: . Indicates the coherence of speech signals. It indicates the coherence of the noise signal.

[0064] Formulas 9 and 12 can be used to correlate the coherence information from the dual-microphone input with the signal-to-noise ratio (SNR). The SNR value can be used as a criterion for determining VAD (Voice-to-Audio Difference): when the SNR is greater than a certain noise threshold... (Preset value) is considered to be in state, otherwise in The state can be represented by the following formula 13: Formula 13 That is, the signal-to-noise ratio (SNR) of the acoustic signal collected by any microphone can be determined based on coherence information; and the state of speech activity can be judged based on the SNR.

[0065] The noise threshold can be a preset fixed value. Alternatively, it can be set as a threshold that varies with the frequency band, matching the noise threshold to the frequency. For example, in the high-frequency band (between 7000 and 8000 Hz), human voice components are scarce, and the corresponding noise threshold can be set to a larger value, making the judgment of the speech activity state more likely to be in a speech absence state. Therefore, before each judgment on the relationship between the signal-to-noise ratio (SNR) and the noise threshold, the frequency of the acoustic signal collected by any microphone can be determined, and a noise threshold matching the frequency can be obtained. The noise threshold obtained based on the current matching is then compared with the SNR. This dynamic threshold that varies with the frequency band improves the adaptability to the distribution patterns of human voices, achieving accurate judgment of the speech activity state.

[0066] Of course, to accurately estimate the state of speech activity, the energy value of the acoustic signal collected by any microphone can be determined, and the speech activity state can be judged based on both the energy value and the signal-to-noise ratio (SNR). For example, the speech activity state can be judged based on the energy value alone, or it can be judged based on the SNR alone. Then, the two judgment results are weighted to obtain the final judgment result of the speech activity state. The weight value can be positively correlated with the confidence value of the judgment result.

[0067] By determining the speech activity state through the cross-reference of features from multiple dimensions, the detection accuracy of speech activity state can be improved.

[0068] Of course, the signal-to-noise ratio can also be used for post-filtering, such as in the implementation based on Wiener filtering, as shown in Equation 14: Formula 14 As shown in Equation 14, the post-filter can be determined based on the signal-to-noise ratio. The gain function value. In Equation 16, a minimum protection value can be set to avoid excessive gain function of the post-filter on certain frequency bands. The minimum value protection is zero. .

[0069] By setting a minimum protection value, excessive damage to the speech can be prevented without significantly sacrificing noise reduction performance, providing users with a more stable, natural, and realistic speech enhancement result.

[0070] This solution utilizes coherence information to determine only the same signal-to-noise ratio value, which can be applied simultaneously to determine the state of speech activity and the calculation of the gain function value of the post-filter. This reduces the complexity of the algorithm and is suitable for scenarios with limited computing resources, such as vehicles.

[0071] Next, this specification discloses an implementation method for signal-to-noise ratio estimation based on coherence information. There are multiple implementation methods; only one is presented here. , The real part is recorded as The imaginary part is denoted as Define the relational variables as shown in Formula 15: The relational variables in Formula 15 only use the coherence information between the two microphones and the preset incident direction information, that is, the beam orientation information. It can be obtained through the following formula 16: in Take the absolute value. Note that because the true signal-to-noise ratio (SNR) is defined as the ratio between the PSD of the speech signal and the PSD of the noise signal, the minimum true SNR is 0. It can be a non-negative value.

[0072] Corresponding to the embodiments of the foregoing methods, this specification also provides embodiments of the apparatus and the terminal to which it is applied.

[0073] Figure 6 This is a schematic diagram illustrating the structure of an electronic device according to an exemplary embodiment. Figure 6 As shown, at the hardware level, the electronic device 600 includes a processor 602, an internal bus 604, a network interface 606, memory 608, and non-volatile memory 610, and may also include other hardware required for business operations. One or more embodiments of this specification can be implemented in software, for example, the processor 602 reads the corresponding computer program from the non-volatile memory 610 into memory 608 and then runs it. Of course, in addition to software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to each logic module, but can also be hardware or logic devices.

[0074] Figure 7 This is a block diagram illustrating a voice enhancement device based on a dual-microphone array according to an exemplary embodiment of this specification. Figure 7 As shown, this device can be applied to, for example Figure 6 The illustrated electronic device 600 is used to implement the technical solution of this specification. The device is also applied to a speech enhancement system, which includes a microphone array, a spatial filter, and a post-filter. The device includes: The acoustic signal acquisition module 702 is used to acquire the surrounding acoustic signals collected by the two microphones in the microphone array.

[0075] The coherence information calculation module 704 is used to determine the coherence information between the two microphones based on the acoustic signals collected by the two microphones respectively.

[0076] The parameter calculation module 706 is used to determine the speech activity state based on the coherence information, update the noise covariance matrix of the spatial filter when the speech activity state is a speech absence state, and determine the gain function value of the post-filter based on the coherence information; the speech absence state indicates that the acoustic signal does not contain speech components.

[0077] The spatial filtering module 708 is used to perform beamforming processing on the two acquired acoustic signals through the spatial filter to obtain an acoustic signal with at least one beam orientation.

[0078] The post-filtering module 710 is used to filter the acoustic signal of any beam direction through the post-filter to obtain an estimation result of the speech component contained in the acoustic signal of any beam direction.

[0079] Optionally, the parameter calculation module 706 is specifically used to determine the signal-to-noise ratio (SNR) of the acoustic signal acquired by any microphone based on the coherence information; determine the speech activity state based on the SNR; or, determine the SNR of the acoustic signal acquired by any microphone based on the coherence information; determine the energy value of the acoustic signal acquired by any microphone; and determine the speech activity state based on the energy value and the SNR.

[0080] Optionally, the parameter calculation module 706 is specifically used to determine the gain function value of the post-filter based on the signal-to-noise ratio value.

[0081] Optionally, the parameter calculation module 706 is specifically used to determine the frequency of the acoustic signal collected by any microphone and obtain a noise threshold that matches the frequency; if the signal-to-noise ratio is not greater than the noise threshold, then the speech activity state is determined to be a speech absence state.

[0082] Optionally, the speech activity state also includes a speech presence state, which indicates that the acoustic signal contains speech components. The parameter calculation module 706 is further configured not to update the noise covariance matrix of the spatial filter when the speech activity state is a speech presence state.

[0083] Optionally, the spatial filter includes an MVDR filter.

[0084] Optionally, the microphone array is deployed at any location in the vehicle cabin, the number of beam orientations is one, and the beam orientation is determined based on a fixed spatial relationship between the microphone array and the seat served by the microphone array.

[0085] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0086] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0087] This specification also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the aforementioned dual-microphone array-based speech enhancement methods provided in this application.

[0088] Specifically, computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks.

[0089] This specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the aforementioned dual-microphone array-based speech enhancement methods.

Claims

1. A speech enhancement method based on a dual-microphone array, characterized in that, The method is applied to a speech enhancement system, the speech enhancement system including a microphone array, a spatial filter, and a post-filter, and the method includes: Acquire the ambient acoustic signals collected by two microphones in the microphone array; Based on the acoustic signals collected by the two microphones respectively, the coherence information between the two microphones is determined; Based on the coherence information, the speech activity state is determined. If the speech activity state is a speech absence state, the noise covariance matrix of the spatial filter is updated. The gain function value of the post-filter is determined based on the coherence information. The speech absence state indicates that the acoustic signal does not contain any speech components. The two acquired acoustic signals are processed by beamforming through the spatial filter to obtain an acoustic signal with at least one beam orientation. The acoustic signal from any beam direction is filtered by the post-filter to obtain an estimate of the speech components contained in the acoustic signal from any beam direction.

2. The method according to claim 1, characterized in that, The determination of the voice activity state based on the coherence information includes: The signal-to-noise ratio of the acoustic signal acquired by any microphone is determined based on the coherence information. The state of voice activity is determined based on the signal-to-noise ratio value; or, The signal-to-noise ratio of the acoustic signal acquired by any microphone is determined based on the coherence information. Determine the energy value of the acoustic signal collected by any of the microphones; The state of speech activity is determined based on the energy value and the signal-to-noise ratio value.

3. The method according to claim 2, characterized in that, Determining the gain function value of the post-filter based on the coherence information includes: The gain function value of the post-filter is determined based on the signal-to-noise ratio value.

4. The method according to claim 2, characterized in that, The step of determining the voice activity state based on the signal-to-noise ratio value includes: Determine the frequency of the acoustic signal collected by any microphone, and obtain a noise threshold that matches the frequency; If the signal-to-noise ratio is not greater than the noise threshold, then the speech activity state is determined to be a speech missing state.

5. The method according to claim 1, characterized in that, The speech activity state also includes a speech presence state, which indicates that the acoustic signal contains speech components. The method further includes: When the speech activity state is a speech presence state, the noise covariance matrix of the spatial filter is not updated.

6. The method according to claim 1, characterized in that, The spatial filter includes an MVDR filter.

7. The method according to claim 1, characterized in that, The microphone array is deployed at any location in the vehicle cabin, the number of beam orientations is one, and the beam orientation is determined based on a fixed spatial relationship between the microphone array and the seat served by the microphone array.

8. A voice enhancement device based on a dual-microphone array, characterized in that, The device is applied to a speech enhancement system, the speech enhancement system including a microphone array, a spatial filter, and a post-filter, and the device includes: The acoustic signal acquisition module is used to acquire the ambient acoustic signals collected by the two microphones in the microphone array. The coherence information calculation module is used to determine the coherence information between the two microphones based on the acoustic signals collected by the two microphones respectively. The parameter calculation module is used to determine the speech activity state based on the coherence information, update the noise covariance matrix of the spatial filter when the speech activity state is a speech absence state, and determine the gain function value of the post-filter based on the coherence information; the speech absence state indicates that the acoustic signal does not contain speech components. A spatial filtering module is used to perform beamforming processing on the two acquired acoustic signals through the spatial filter to obtain an acoustic signal with at least one beam orientation. The post-filtering module is used to filter the acoustic signal in any beam direction through the post-filter to obtain an estimation result of the speech components contained in the acoustic signal in any beam direction.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method as described in any one of claims 1-7.