Audio processing method and device, electronic equipment and storage medium

CN116634319BActive Publication Date: 2026-09-29BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210134718.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-14
Publication Date
2026-09-29
Estimated Expiration
2042-02-14

AI Technical Summary

Technical Problem

但是,相关技术中通过对麦克风采集的音频进行处理而定向拾取的音频质量和稳定性均较差

Benefits of technology

本公开通过获取音频录制设备的至少两个麦克风中每个麦克风所采集的原始音频信号,并根据分频点分别对每个麦克风采集的原始音频信号进行分频处理,得到每个麦克风对应的低频信号和高频信号,然后根据目标声源方向和每个麦克风的端射方向,对每个麦克风对应的低频信号进行正向滤波,并将所述低频信号的正向滤波结果与所述高频信号合并,得到有效信号,以及根据目标声源方向和每个麦克风的端射方向,对每个麦克风对应的低频信号进行反向滤波,并将所述低频信号的反向滤波结果与所述高频信号合并,得到噪声信号,最后使用所述噪声信号对所述有效信号进行噪声消除处理,得到目标音频信号。由于正反向滤波只针对低频信号,且有效信号和噪声信号都结合了滤波结果和高频信号,因此实现了对低频信号的定向波束形成,同时避免滤波过程对高频信号造成失真影响,提高了定向拾取的目标音频信号的保真率和稳定性,进而使音频录制设备在语音通话、人机语音交互等场景下提高用户的使用体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116634319B_ABST
    Figure CN116634319B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an audio processing method, device, electronic equipment and storage medium. The method is applied to an audio recording device having at least two microphones. The method comprises: obtaining original audio signals collected by each of the at least two microphones; performing frequency division processing on the original audio signals collected by each of the at least two microphones according to a frequency division point, to obtain low-frequency signals and high-frequency signals corresponding to each microphone; performing forward filtering and reverse filtering on the low-frequency signals corresponding to each microphone according to a target sound source direction and an end-on direction of each microphone, determining effective signals according to a result of the forward filtering and high-frequency signals corresponding to at least one of the microphones, and determining noise signals according to a result of the reverse filtering and high-frequency signals corresponding to at least one of the microphones; and performing noise cancellation processing on the effective signals using the noise signals to obtain target audio signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of audio processing technology, specifically to an audio processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] Current mobile phones, headsets, and other devices can all record audio, a function that can be used in scenarios such as voice calls and human-computer voice interaction. Mobile phones and headsets contain microphone microarrays, each microphone in which can collect audio. This audio needs to be processed through delay estimation, beamforming, and noise cancellation to achieve directional audio pickup. However, the audio quality and stability of directional pickup achieved through processing the audio collected by the microphones in related technologies are generally poor. Summary of the Invention

[0003] To overcome the problems existing in the related technologies, this disclosure provides an audio processing method, apparatus, electronic device, and storage medium to solve the defects in the related technologies.

[0004] According to a first aspect of the present disclosure, an audio processing method is provided, applied to an audio recording device having at least two microphones, the method comprising: Acquire the raw audio signal captured by each of the at least two microphones; The original audio signal collected by each microphone is divided into low-frequency and high-frequency signals according to the frequency division point to obtain the low-frequency signal and high-frequency signal corresponding to each microphone. Based on the direction of the target sound source and the end-fire direction of each microphone, the low-frequency signal corresponding to each microphone is subjected to forward filtering and reverse filtering respectively. The effective signal is determined based on the result of the forward filtering and the high-frequency signal corresponding to at least one of the microphones. The noise signal is determined based on the result of the reverse filtering and the high-frequency signal corresponding to at least one of the microphones. The effective signal is a mixed signal of the target audio signal and the noise signal. The effective signal is then subjected to noise cancellation processing using the noise signal to obtain the target audio signal.

[0005] In one embodiment, the step of performing forward filtering and reverse filtering on the low-frequency signals corresponding to each microphone based on the direction of the target sound source and the end-fire direction of each microphone includes: The low-frequency signal corresponding to the target microphone is input as the target signal to the spatial filter, and the low-frequency signals corresponding to the other microphones are input as interference signals to the spatial filter for spatial filtering to obtain the positive filtering result. The end-fire direction of the target microphone is matched with the direction of the target sound source. The low-frequency signal corresponding to the far-end microphone is used as the target signal and input to the spatial filter. The low-frequency signals corresponding to the other microphones are used as interference signals and input to the spatial filter for spatial filtering to obtain the result of reverse filtering. The end-fire direction of the far-end microphone is opposite to the end-fire direction of the target microphone.

[0006] In one embodiment, the step of inputting the low-frequency signal corresponding to the target microphone as the target signal into the spatial filter, and simultaneously inputting the low-frequency signals corresponding to the other microphones as interference signals into the spatial filter for spatial filtering, includes: The low-frequency signal corresponding to the target microphone is used as the target signal input to the spatial filter; based on the distance between each of the other microphones and the target microphone, the low-frequency signals corresponding to each of the other microphones are used as interference signals at various levels and input to the spatial filter for spatial filtering; and / or, The step of inputting the low-frequency signal corresponding to the far-end microphone as the target signal into the spatial filter, and inputting the low-frequency signals corresponding to the other microphones as interference signals into the spatial filter for spatial filtering includes: The low-frequency signal corresponding to the remote microphone is input into the spatial filter as the target signal. Based on the distance between each of the other microphones and the remote microphone, the low-frequency signals corresponding to each of the other microphones are input into the spatial filter as interference signals at various levels for spatial filtering.

[0007] In one embodiment, determining a valid signal based on the result of the forward filtering and a high-frequency signal corresponding to at least one of the microphones, and determining a noise signal based on the result of the reverse filtering and a high-frequency signal corresponding to at least one of the microphones, includes: The result of the forward filtering is combined with the high-frequency signal corresponding to the target microphone to obtain the effective signal; The result of the inverse filtering is combined with the high-frequency signal corresponding to the remote microphone to obtain the noise signal.

[0008] In one embodiment, it also includes: Select the direction of the target sound source according to the recording mode selection command.

[0009] In one embodiment, the target sound source direction includes multiple sub-directions; The step of performing forward filtering and reverse filtering on the low-frequency signals corresponding to each microphone according to the direction of the target sound source and the end-fire direction of each microphone, determining the effective signal based on the result of the forward filtering and the high-frequency signal corresponding to at least one of the microphones, and determining the noise signal based on the result of the reverse filtering and the high-frequency signal corresponding to at least one of the microphones, includes: For each sub-direction of the target sound source direction, determine the effective signal and the noise signal corresponding to the sub-direction; The step of using the noise signal to perform noise cancellation processing on the effective signal to obtain the target audio signal includes: Based on the noise signal corresponding to each sub-direction, noise cancellation processing is performed on the effective signal corresponding to each sub-direction to obtain the target audio sub-signal corresponding to each sub-direction; The target audio signal is determined based on the target audio sub-signal corresponding to each of the sub-directions.

[0010] In one embodiment, the target sound source direction includes a first sub-direction and a second sub-direction, wherein the first sub-direction and the second sub-direction are opposite.

[0011] In one embodiment, it also includes: The frequency division point is determined based on the distance between adjacent microphones; The step of performing frequency division processing on the original audio signal collected by each microphone according to the frequency division point to obtain the low-frequency signal and high-frequency signal corresponding to each microphone includes: If the frequency division point is less than the preset frequency division point threshold, the original audio signal collected by each microphone is divided according to the frequency division point to obtain the low-frequency signal and high-frequency signal corresponding to each microphone.

[0012] According to a second aspect of the present disclosure, an audio processing apparatus is provided, applied to an audio recording device, the audio recording device having at least two microphones, the apparatus comprising: An acquisition module is used to acquire the raw audio signal collected by each of the at least two microphones; The frequency division module is used to perform frequency division processing on the original audio signal collected by each microphone according to the frequency division point, so as to obtain the low-frequency signal and high-frequency signal corresponding to each microphone; The filtering module is used to perform forward filtering and reverse filtering on the low-frequency signals corresponding to each microphone according to the direction of the target sound source and the end-fire direction of each microphone, and to determine the effective signal according to the result of the forward filtering and the high-frequency signal corresponding to at least one of the microphones, and to determine the noise signal according to the result of the reverse filtering and the high-frequency signal corresponding to at least one of the microphones, wherein the effective signal is a mixed signal of the target audio signal and the noise signal; The noise cancellation module is used to perform noise cancellation processing on the effective signal using the noise signal to obtain the target audio signal.

[0013] In one embodiment, when the filtering module performs forward filtering and reverse filtering on the low-frequency signals corresponding to each microphone based on the direction of the target sound source and the end-fire direction of each microphone, it is specifically used for: The low-frequency signal corresponding to the target microphone is input as the target signal to the spatial filter, and the low-frequency signals corresponding to the other microphones are input as interference signals to the spatial filter for spatial filtering to obtain the positive filtering result. The end-fire direction of the target microphone is matched with the direction of the target sound source. The low-frequency signal corresponding to the far-end microphone is used as the target signal and input to the spatial filter. The low-frequency signals corresponding to the other microphones are used as interference signals and input to the spatial filter for spatial filtering to obtain the result of reverse filtering. The end-fire direction of the far-end microphone is opposite to the end-fire direction of the target microphone.

[0014] In one embodiment, when the filtering module is used to input the low-frequency signal corresponding to the target microphone as the target signal into the spatial filter, and input the low-frequency signals corresponding to the other microphones as interference signals into the spatial filter for spatial filtering, it is specifically used for: The low-frequency signal corresponding to the target microphone is used as the target signal input to the spatial filter; based on the distance between each of the other microphones and the target microphone, the low-frequency signals corresponding to each of the other microphones are used as interference signals at various levels and input to the spatial filter for spatial filtering; and / or, The filtering module is used to input the low-frequency signal corresponding to the far-end microphone as the target signal into the spatial filter, and to input the low-frequency signals corresponding to the other microphones as interference signals into the spatial filter for spatial filtering. Specifically, it is used for: The low-frequency signal corresponding to the remote microphone is input into the spatial filter as the target signal. Based on the distance between each of the other microphones and the remote microphone, the low-frequency signals corresponding to each of the other microphones are input into the spatial filter as interference signals at various levels for spatial filtering.

[0015] In one embodiment, when the filtering module is used to determine a valid signal based on the result of the forward filtering and a high-frequency signal corresponding to at least one of the microphones, and to determine a noise signal based on the result of the reverse filtering and a high-frequency signal corresponding to at least one of the microphones, it is specifically used for: The result of the forward filtering is combined with the high-frequency signal corresponding to the target microphone to obtain the effective signal; The result of the inverse filtering is combined with the high-frequency signal corresponding to the remote microphone to obtain the noise signal.

[0016] In one embodiment, a direction selection module is further included, for: Select the direction of the target sound source according to the recording mode selection command.

[0017] In one embodiment, the target sound source direction includes multiple sub-directions; The filtering module is specifically used for: For each sub-direction of the target sound source direction, determine the effective signal and the noise signal corresponding to the sub-direction; The noise cancellation module is specifically used for: Based on the noise signal corresponding to each sub-direction, noise cancellation processing is performed on the effective signal corresponding to each sub-direction to obtain the target audio sub-signal corresponding to each sub-direction; The target audio signal is determined based on the target audio sub-signal corresponding to each of the sub-directions.

[0018] In one embodiment, the target sound source direction includes a first sub-direction and a second sub-direction, wherein the first sub-direction and the second sub-direction are opposite.

[0019] In one embodiment, a frequency division point determination module is further included, for: The frequency division point is determined based on the distance between adjacent microphones; The frequency division module is specifically used for: If the frequency division point is less than the preset frequency division point threshold, the original audio signal collected by each microphone is divided according to the frequency division point to obtain the low-frequency signal and high-frequency signal corresponding to each microphone.

[0020] According to a third aspect of the present disclosure, an electronic device is provided, the electronic device including a memory and a processor, the memory being configured to store computer instructions executable on the processor, and the processor being configured to execute the computer instructions based on the audio processing method described in the first aspect.

[0021] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.

[0022] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: This disclosure acquires the raw audio signal from each of at least two microphones in an audio recording device, and performs frequency division processing on the raw audio signal from each microphone according to a frequency division point to obtain a low-frequency signal and a high-frequency signal corresponding to each microphone. Then, based on the direction of the target sound source and the end-fire direction of each microphone, the low-frequency signal corresponding to each microphone is forward filtered, and the forward filtering result of the low-frequency signal is combined with the high-frequency signal to obtain an effective signal. Additionally, based on the direction of the target sound source and the end-fire direction of each microphone, the low-frequency signal corresponding to each microphone is reverse filtered, and the reverse filtering result of the low-frequency signal is combined with the high-frequency signal to obtain a noise signal. Finally, the noise signal is used to perform noise cancellation processing on the effective signal to obtain the target audio signal. Since the forward and reverse filtering only targets the low-frequency signal, and both the effective signal and the noise signal combine the filtering result and the high-frequency signal, directional beamforming of the low-frequency signal is achieved, while avoiding distortion of the high-frequency signal during the filtering process. This improves the fidelity and stability of the directional target audio signal, thereby enhancing the user experience of the audio recording device in scenarios such as voice calls and human-computer voice interaction. Attached Figure Description

[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0024] Figure 1 This is a flowchart illustrating an exemplary embodiment of the audio processing method disclosed herein; Figure 2 This is a schematic diagram of a frequency divider shown in an exemplary embodiment of the present disclosure; Figure 3 This is a schematic diagram of a spatial filter for two microphones illustrating an exemplary embodiment of this disclosure; Figure 4 This is a schematic diagram of a spatial filter for three microphones shown in an exemplary embodiment of this disclosure; Figure 5 This is a schematic diagram of an adaptive noise cancellation filter illustrated in an exemplary embodiment of this disclosure; Figure 6 This is a flowchart illustrating an audio processing method according to another exemplary embodiment of this disclosure; Figure 7 This is a schematic diagram of the structure of an audio processing apparatus shown in an exemplary embodiment of the present disclosure; Figure 8 This is a structural block diagram of an electronic device illustrated in an exemplary embodiment of the present disclosure. Detailed Implementation

[0025] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0026] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0027] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0028] In related technologies, the distance between microphones in mobile phones, headphones, and other devices is relatively large, resulting in unclear directional sound pickup and causing significant distortion during audio signal processing. Furthermore, the noise reduction and blind source separation algorithms used in these devices are not directional, failing to effectively filter out interference signals from non-target directions. These algorithms typically require noise detection based on the characteristics of the speech signal, showing some effectiveness in enhancing or extracting the target speech signal, but exhibiting weak enhancement effects on non-speech target signals.

[0029] Based on this, in a first aspect, at least one embodiment of this disclosure provides an audio processing method, please refer to the appendix. Figure 1 The diagram illustrates the process of the method, including steps S101 and S104.

[0030] This audio processing method is applied to audio recording devices such as mobile phones and headphones. The audio recording device has at least two microphones, which can form a microphone array or be set individually. Each microphone has a different end-fire direction, which is the direction of the end of the audio recording device corresponding to that microphone. For example, if a mobile phone has two microphones at the top and two at the bottom, the end-fire direction of the top microphone is the top direction, and the end-fire direction of the bottom microphone is the bottom direction. Each microphone can be used to collect audio. This audio processing method is used to process the collected audio to directionally pick up audio from the direction of the target sound source. Since the audio collected by the microphone is not used as the final output or saved audio, the audio collected by the microphone is referred to as the raw audio signal below.

[0031] In step S101, the raw audio signal captured by each of the at least two microphones is acquired.

[0032] Microphones in audio recording devices such as mobile phones and headphones can capture raw audio signals in real time or in specific modes. For example, a mobile phone microphone can capture raw audio signals in call mode or human-computer interaction mode, while a headphone microphone can capture raw audio signals when the host device to which the headphone is connected is in headphone mode.

[0033] The original audio signal can be a time-domain signal, that is, an audio signal in time-domain form.

[0034] In step S102, the original audio signal collected by each microphone is divided according to the frequency division point to obtain the low-frequency signal and high-frequency signal corresponding to each microphone.

[0035] The crossover point is the highest frequency at which the spatial filter can achieve directivity. Therefore, if the frequency exceeds the crossover point, the spatial filter will not have a significant beam pointing effect and will introduce additional signal distortion. The crossover point can be determined in advance based on the distance between adjacent microphones, for example, by determining the crossover point according to the following formula. (Hz): Where c is the speed of sound and d is the distance between adjacent microphones; then, if the frequency division point is less than the preset frequency division point threshold, step S102 is executed, that is, the original audio signal collected by each microphone is divided according to the frequency division point to obtain the low-frequency signal and high-frequency signal corresponding to each microphone. For common voice, music and other sound signals, their spectral energy is mainly distributed below 10kHz, so the frequency division point threshold can be set to 10kHz.

[0036] Frequency division processing can be understood as identifying the portion of the original audio signal above the division point as a high-frequency signal and the portion below the division point as a low-frequency signal. In one possible embodiment, the original audio signal can be frequency divided as follows: the original audio signal is input into a frequency division filter for frequency division processing to obtain low-frequency and high-frequency signals. The frequency division filter is designed based on the division point, and the frequency division filter can be, for example,... Figure 2 The form shown includes at least two cascaded low-pass filters (LP) and at least two cascaded high-pass filters (HP). The frequency division process can be represented by the following function: ,in, It is the input signal of the crossover, i.e., the original audio signal. This represents the frequency division processing function. They are respectively based on The low-frequency signal and high-frequency signal are obtained by dividing the frequency at the dividing point.

[0037] When the audio recording device has two microphones, the two microphones can be divided into two frequency groups: , ,in, yes Low-frequency signals and high-frequency signals, yes Low-frequency signals and high-frequency signals.

[0038] By using frequency division processing, spatial directivity is achieved within the effective frequency range, while avoiding distortion caused by beam nonlinearity in the high-frequency range.

[0039] In step S103, based on the direction of the target sound source and the end-fire direction of each microphone, the original audio signals collected by each microphone are subjected to forward filtering and reverse filtering respectively. Based on the result of the forward filtering and the high-frequency signal corresponding to at least one microphone, an effective signal is determined, and based on the result of the reverse filtering and the high-frequency signal corresponding to at least one microphone, a noise signal is determined. The effective signal is a mixed signal of the target audio signal and the noise signal.

[0040] Forward filtering can be considered a spatial filtering process performed along the direction of the target sound source, while inverse filtering can be considered a spatial filtering process performed in the opposite direction of the target sound source. Both forward and inverse filtering can be implemented using spatial filters. The filtering function of a spatial filter is... It can be expressed using the following formula: ( ); These are the low-frequency compensation filter coefficients. It is a function related to the distance between adjacent microphones, therefore the filter function in the spatial filter is generated based on the distance between adjacent microphones; ( ) is the steering vector of the microphone array; The filtering process of a spatial filter can be represented by the following formula: ; Y is the frequency domain output signal of the spatial filter, which can be converted into the time domain signal y by inverse Fourier transform; S is the input signal of the spatial filter, which is a vector composed of the target signal in frequency domain form and the interference signals in frequency domain form at each level. The target signal and the interference signal are both the original audio signals collected by the microphone.

[0041] In one possible scenario, the audio recording device has two microphones, and the spatial filter is as follows: Figure 3 As shown, the low-frequency signals corresponding to the two microphones , The inputs are fed into the spatial filter, and the output is the directional filtering result.

[0042] In the scenario with two microphones, the low-frequency compensation filter coefficient of the spatial filter. It can be: Where j is a negative unit vector. Represents digital angular frequency. Indicates a delay. ; In the scenario with two microphones, the steering vector of the microphone array ( ) can be: Where c is the speed of sound; In the scenario with two microphones, the input signal S of the spatial filter can be a vector [S1 S2], where S1 is the low-frequency signal corresponding to one of the microphones. The signal is converted from the time domain to the frequency domain, and S2 is the low-frequency signal corresponding to another microphone. The signal is converted from the time domain to the frequency domain, where S1 is the target signal and S2 is the interference signal.

[0043] In another possible scenario, the audio recording device has three microphones, and the spatial filter is as follows: Figure 4 As shown, the low-frequency signals corresponding to the three microphones , , The inputs are fed into the spatial filter, and the output is the directional filtering result.

[0044] In this scenario with three microphones, the low-frequency compensation filter coefficient of the spatial filter... It can be: ; In a three-microphone scenario, the steering vector of the microphone array ( ) can be: ,in, .

[0045] In this scenario with three microphones, the input signal S of the spatial filter can be a vector [S1 S2 S3], where S1 is the low-frequency signal corresponding to the first microphone. The signal converted from the time domain to the frequency domain, S2 is the low-frequency signal corresponding to the second microphone. The signal converted from the time domain to the frequency domain, S3 is the low-frequency signal corresponding to the third microphone. The signal is converted from the time domain to the frequency domain, and S1 is the target signal, S2 is the first-level interference signal, and S3 is the second-level interference signal.

[0046] Based on the structure and parameters of the spatial filter described above, this step can perform forward filtering on the low-frequency signals corresponding to each microphone in the following manner: the low-frequency signal corresponding to the target microphone is input as the target signal into the spatial filter, and the low-frequency signals corresponding to the other microphones are input as interference signals into the spatial filter for spatial filtering to obtain the result of forward filtering.

[0047] In this context, the end-fire direction of the target microphone matches the direction of the target sound source. That is, the target microphone is the microphone located at the end of the audio recording device corresponding to the direction of the target sound source. Specifically, a preset angle range on both sides of the microphone's end-fire direction can be set as the matching range. When the direction of the target sound source is within the matching range of a certain microphone, the end-fire direction of that microphone matches the direction of the target sound source. For example, a mobile phone has microphones at its top and bottom. When recording with the bottom of the phone facing the sound source, the target sound source direction is the bottom, so the microphone located at the bottom, with its end-fire direction at the bottom, is the target microphone. As another example, an earphone has microphones at its head and tail. When a user wears the earphone for a voice call, the target sound source direction is the tail of the earphone, so the microphone located at the tail, with its end-fire direction at the tail, is the target microphone. Users can select a target microphone by operating the audio recording device according to the direction of the target sound source and the recording method. For example, users can select different target microphones by selecting different recording modes; or, if the user does not select a target microphone, the audio recording device can identify the direction of the target sound source and automatically determine the target microphone based on the end-fire direction of each microphone.

[0048] Optionally, the low-frequency signal corresponding to the target microphone is used as the target signal input to the spatial filter, and the low-frequency signals corresponding to each of the other microphones are used as interference signals at various levels input to the spatial filter for spatial filtering based on the distance between each of the other microphones and the target microphone.

[0049] When the audio recording device has two microphones, the low-frequency signal corresponding to the other microphone (other than the target microphone) can be used as an interference signal input to the spatial filter. The input signal S of the spatial filter can then be a vector [S1 S2], where S1 is the low-frequency signal corresponding to one of the microphones. The signal is converted from the time domain to the frequency domain, and S2 is the low-frequency signal corresponding to another microphone. The signal is converted from the time domain to the frequency domain, where S1 is the target signal and S2 is the interference signal.

[0050] When the audio recording device has at least three microphones, the low-frequency signals corresponding to each of the remaining microphones can be input into the spatial filter as interference signals of various levels based on the distance between each microphone and the target microphone. That is, the smaller the distance to the target microphone, the higher the interference level. For example, each microphone can be numbered along the direction from the target microphone to the far-end microphone. The target microphone is numbered 1, and the interference level of the original audio signal collected by microphone number 2 is 1, and so on for the other microphones. The input signal S of the spatial filter can then be a vector [S1 …… Sn], where S1 is the low-frequency signal corresponding to the target microphone. The signal converted from the time domain to the frequency domain, Sn is the low-frequency signal corresponding to the far-end microphone. The signal is converted from the time domain to the frequency domain. The far-end microphone is the microphone furthest from the target microphone. The end-fire direction of the far-end microphone is opposite to that of the target microphone, and n≥3.

[0051] Based on the structure and parameters of the spatial filter described above, this step can perform reverse filtering on the low-frequency signals corresponding to each microphone in the following way: the low-frequency signal corresponding to the far-end microphone is used as the target signal and input into the spatial filter, and the low-frequency signals corresponding to the other microphones are used as interference signals and input into the spatial filter for spatial filtering to obtain the result of reverse filtering.

[0052] In this context, the end of the audio recording device containing the far-end microphone is opposite to the end of the audio recording device corresponding to the direction of the target sound source. For example, a mobile phone has microphones at both the top and bottom. When recording with the bottom of the phone facing the sound source, the target sound source direction is the bottom, so the microphone located at the top, i.e., with its end-to-end direction pointing towards the top, is the far-end microphone. As another example, an earphone has microphones at both the head and tail. When a user is wearing the earphone for a voice call, the target sound source direction is the tail of the earphone, so the microphone located at the head, i.e., with its end-to-end direction pointing towards the head, is the far-end microphone.

[0053] The input signal of the spatial filter in the forward filtering process can be changed in direction and used as the input signal of the spatial filter in the reverse filtering process.

[0054] Optionally, the low-frequency signal corresponding to the remote microphone is input into the spatial filter as the target signal, and the low-frequency signals corresponding to each of the other microphones are input into the spatial filter as interference signals of each level for spatial filtering based on the distance between each of the other microphones and the remote microphone.

[0055] When the audio recording device has two microphones, the original audio signal captured by the other microphone (other than the far-end microphone) can be used as an interference signal input to the spatial filter. The input signal S of the spatial filter can then be a vector [S2 S1], where S1 is the low-frequency signal corresponding to the target microphone. The signal is converted from the time domain to the frequency domain, and S2 is the low-frequency signal corresponding to the far-end microphone. The signal is converted from the time domain to the frequency domain, and S1 is the interference signal and S2 is the target signal.

[0056] When the audio recording device has at least three microphones, the low-frequency signals corresponding to each of the remaining microphones can be input into the spatial filter as interference signals of various levels based on the distance between each microphone and the far-end microphone. That is, the smaller the distance to the far-end microphone, the higher the interference level. For example, the interference level of the nearest microphone is 1, and so on for the other microphones. The input signal S of the spatial filter can then be a vector [Sn…S1], where S1 is the low-frequency signal corresponding to the target microphone. The signal converted from the time domain to the frequency domain, Sn is the low-frequency signal corresponding to the far-end microphone. The signal is converted from the time domain to the frequency domain. The far-end microphone is the microphone furthest from the target microphone. The end-fire direction of the far-end microphone is opposite to that of the target microphone, and n≥3.

[0057] In this step, the result of the forward filtering and the high-frequency signal corresponding to the target microphone can be combined to obtain the effective signal. For example, the above combination can be performed according to the following formula: Where y1 is the valid signal, This is the result of forward filtering. It is the raw audio signal captured by the target microphone.

[0058] In this step, the result of the inverse filtering can be combined with the high-frequency signal corresponding to the far-end microphone to obtain a noise signal. For example, the above combination can be performed using the following formula: Where y2 is the noise signal, It is the result of inverse filtering. It is the raw audio signal captured by the remote microphone.

[0059] In step S104, the noise signal is used to perform noise cancellation processing on the effective signal to obtain the target audio signal.

[0060] Such as Figure 5 The adaptive noise cancellation filter shown performs noise cancellation processing. This filter uses the adaptive LMS algorithm to achieve noise cancellation. The target audio signal is the audio signal that the device ultimately acquires, stores, and transmits.

[0061] The effective signal y1 can be used as the input signal x(t) of the adaptive noise cancellation filter, and the noise signal y2 can be used as the noise reference signal n(t) of the adaptive noise cancellation filter. Then the adaptive noise cancellation filter can output the target audio signal ys1(t).

[0062] It is important to note that if the target sound source direction is reversed, the target microphone and the far-end microphone are swapped, thus converting the original valid signal y1 into a noise signal and the original noise signal y2 into a valid signal. The valid signal y2 can then be used as the input signal x(t) of the adaptive noise cancellation filter, and the noise signal y1 can be used as the noise reference signal n(t) of the adaptive noise cancellation filter. The adaptive noise cancellation filter can then output the target audio signal ys2(t) after the target sound source direction is reversed.

[0063] This disclosure acquires the raw audio signal from each of at least two microphones in an audio recording device, and performs frequency division processing on the raw audio signal from each microphone according to a frequency division point to obtain a low-frequency signal and a high-frequency signal corresponding to each microphone. Then, based on the direction of the target sound source and the end-fire direction of each microphone, the low-frequency signal corresponding to each microphone is forward filtered, and the forward filtering result of the low-frequency signal is combined with the high-frequency signal to obtain an effective signal. Additionally, based on the direction of the target sound source and the end-fire direction of each microphone, the low-frequency signal corresponding to each microphone is reverse filtered, and the reverse filtering result of the low-frequency signal is combined with the high-frequency signal to obtain a noise signal. Finally, the noise signal is used to perform noise cancellation processing on the effective signal to obtain the target audio signal. Since the forward and reverse filtering only targets the low-frequency signal, and both the effective signal and the noise signal combine the filtering result and the high-frequency signal, directional beamforming of the low-frequency signal is achieved, while avoiding distortion of the high-frequency signal during the filtering process. This improves the fidelity and stability of the directional target audio signal, thereby enhancing the user experience of the audio recording device in scenarios such as voice calls and human-computer voice interaction.

[0064] In addition, the audio processing method provided by this disclosure has a significant directional effect of directional sound pickup, which avoids distortion during the processing of audio signals; moreover, this disclosure obtains noise signals by directional sound pickup in non-target directions, without the need for noise estimation, so it can effectively filter out interference signals in non-target directions, and is not limited to speech signals, but is also effective for music, singing and other signals.

[0065] In some embodiments of this disclosure, the target sound source direction can be determined in advance based on the recording mode selection instruction. The recording mode selection instruction can be generated based on user operations. For example, a user can access the recording mode settings interface and click on the identifier of a specific recording mode to generate the corresponding recording mode selection instruction. Each recording mode has a corresponding target sound source direction. For instance, if an audio recording device such as a mobile phone has a microphone at the top and a microphone at the bottom, the audio recording device can have a top recording mode and a bottom recording mode. In the top recording mode, the target sound source direction is the top direction, the target microphone is the top microphone, and the far-end microphone is the bottom microphone. Conversely, in the bottom recording mode, the target sound source direction is the bottom direction, the target microphone is the bottom microphone, and the far-end microphone is the top microphone.

[0066] It should be noted that the recording mode can also be a multi-directional recording mode, in which case the target sound source direction can include multiple sub-directions. For example, the audio recording device with microphones at the top and bottom can have a bidirectional recording mode, in which the target sound source direction includes two sub-directions: the top direction and the bottom direction.

[0067] When the target sound source direction includes multiple sub-directions, when executing step S103, that is, according to the target sound source direction and the end-fire direction of each microphone, the low-frequency signal corresponding to each microphone is subjected to forward filtering and reverse filtering respectively, and the effective signal is determined according to the result of the forward filtering and the high-frequency signal corresponding to at least one of the microphones, and the noise signal is determined according to the result of the reverse filtering and the high-frequency signal corresponding to at least one of the microphones, the effective signal and the noise signal corresponding to each sub-direction of the target sound source direction can be determined separately. In other words, step S103 can be executed for each sub-direction to determine its corresponding effective signal and noise signal. The specific execution details of step S103 have been described in detail in the above embodiments, and will not be repeated here.

[0068] Based on the effective signal and noise signal determined for each sub-direction, when performing step S104, i.e., using the noise signal to perform noise cancellation processing on the effective signal to obtain the target audio signal, noise cancellation processing can first be performed on the effective signal corresponding to each sub-direction according to the noise signal corresponding to each sub-direction to obtain the target audio sub-signal corresponding to each sub-direction; then, the target audio signal can be determined according to the target audio sub-signal corresponding to each sub-direction. For example, the target audio sub-signals corresponding to each sub-direction can be added together to obtain the target audio signal.

[0069] In one example, the target sound source direction includes a first sub-direction and a second sub-direction, and the first sub-direction and the second sub-direction are opposite. For example, the target sound source direction described in step S104 is the first sub-direction in this example, and the opposite direction of the target sound source direction described in step S104 is the second sub-direction in this example. Then, the target audio signal ys1(t) of the target sound source direction in step S104 can be used as the target audio sub-signal of the first sub-direction in this example, and the target audio signal ys2(t) of the opposite direction of the target sound source direction in step S104 can be used as the target audio sub-signal of the second sub-direction in this example. Then, the target audio sub-signals of the two sub-directions are processed according to the following formula to obtain the target audio signal in this embodiment. : .

[0070] In some embodiments of this disclosure, when the frequency division point is greater than or equal to a preset frequency division point threshold, the original audio signal collected by the microphone may not be frequency divided. Instead, based on the direction of the target sound source and the end-fire direction of each microphone, the original audio signal collected by each microphone is subjected to forward filtering and reverse filtering respectively; then, the result of the reverse filtering is used to perform noise cancellation processing on the result of the forward filtering to obtain the target audio signal.

[0071] The forward and reverse filtering processes are the same as those described in step S103 for the forward and reverse filtering of low-frequency signals, and will not be repeated here.

[0072] In this embodiment, when the frequency division point is greater than or equal to the preset frequency division threshold, the original audio signal collected by the microphone is not divided. Then, the original audio signal is bidirectionally filtered, and the forward filtering result is used as the effective signal, while the reverse filtering result is used as the noise signal. Noise is then eliminated to obtain the target audio signal, which is the audio signal that the device finally acquires, stores, and transmits.

[0073] Please refer to the appendix. Figure 6 This example illustrates a complete flow of audio processing according to an embodiment of the present disclosure. Figure 6As can be seen from the diagram, the audio recording device in this embodiment has two microphones, designated as a target microphone and a far-end microphone based on the direction of the target sound source. The original audio signal collected by the target microphone is sig1, and the original audio signal collected by the far-end microphone is sig2. Since the frequency division point determined by the distance between the two microphones is less than the frequency division point threshold, sig1 and sig2 are respectively input into a frequency division filter. Frequency division is performed at the aforementioned frequency division point to obtain the high-frequency signal sig1_h corresponding to sig1 and the low-frequency signal sig1_l corresponding to sig1. The high-frequency signal sig2_h and the low-frequency signal sig2_l corresponding to sig2 are used. Next, sig1_l and sig2_l are forward-filtered, and the result of the forward filtering is combined with sig1_h to obtain the effective signal. Then, sig1_l and sig2_l are reverse-filtered to obtain the noise signal, and the result of the reverse filtering is combined with sig2_h to obtain the noise signal. Finally, the effective signal is used as the noisy input signal, and the noise signal is used as the noise reference signal for adaptive noise cancellation, thus obtaining the target signal, i.e., the target audio signal. Since the forward and reverse filtering only targets the low-frequency signal, and both the effective signal and the noise signal combine the filtering results and the high-frequency signal, directional beamforming of the low-frequency signal is achieved. At the same time, the filtering process avoids distortion of the high-frequency signal. In other words, through frequency division processing, spatial directivity is achieved within the effective frequency range, while avoiding distortion caused by beam nonlinearity in the high-frequency part. This improves the fidelity and stability of the directional pickup target audio signal, thereby enhancing the user experience of audio recording devices in scenarios such as voice calls and human-computer voice interaction.

[0074] According to a second aspect of the present disclosure, an audio processing apparatus is provided, applied to an audio recording device, the audio recording device having at least two microphones. Please refer to the appendix. Figure 7 The device includes: Acquisition module 701 is used to acquire the raw audio signal collected by each of the at least two microphones; The frequency division module 702 is used to perform frequency division processing on the original audio signal collected by each microphone according to the frequency division point to obtain the low-frequency signal and high-frequency signal corresponding to each microphone. The filtering module 703 is used to perform forward filtering and reverse filtering on the low-frequency signals corresponding to each microphone according to the direction of the target sound source and the end-fire direction of each microphone, and to determine the effective signal according to the result of the forward filtering and the high-frequency signal corresponding to at least one of the microphones, and to determine the noise signal according to the result of the reverse filtering and the high-frequency signal corresponding to at least one of the microphones, wherein the effective signal is a mixed signal of the target audio signal and the noise signal; The noise cancellation module 704 is used to perform noise cancellation processing on the effective signal using the noise signal to obtain the target audio signal.

[0075] In some embodiments of this disclosure, when the filtering module performs forward filtering and reverse filtering on the low-frequency signals corresponding to each microphone according to the direction of the target sound source and the end-fire direction of each microphone, it is specifically used for: The low-frequency signal corresponding to the target microphone is input as the target signal to the spatial filter, and the low-frequency signals corresponding to the other microphones are input as interference signals to the spatial filter for spatial filtering to obtain the positive filtering result. The end-fire direction of the target microphone is matched with the direction of the target sound source. The low-frequency signal corresponding to the far-end microphone is used as the target signal and input to the spatial filter. The low-frequency signals corresponding to the other microphones are used as interference signals and input to the spatial filter for spatial filtering to obtain the result of reverse filtering. The end-fire direction of the far-end microphone is opposite to the end-fire direction of the target microphone.

[0076] In some embodiments of this disclosure, when the filtering module is used to input the low-frequency signal corresponding to the target microphone as the target signal into the spatial filter, and to input the low-frequency signals corresponding to other microphones as interference signals into the spatial filter for spatial filtering, it is specifically used for: The low-frequency signal corresponding to the target microphone is used as the target signal input to the spatial filter; based on the distance between each of the other microphones and the target microphone, the low-frequency signals corresponding to each of the other microphones are used as interference signals at various levels and input to the spatial filter for spatial filtering; and / or, The filtering module is used to input the low-frequency signal corresponding to the far-end microphone as the target signal into the spatial filter, and to input the low-frequency signals corresponding to the other microphones as interference signals into the spatial filter for spatial filtering. Specifically, it is used for: The low-frequency signal corresponding to the remote microphone is input into the spatial filter as the target signal. Based on the distance between each of the other microphones and the remote microphone, the low-frequency signals corresponding to each of the other microphones are input into the spatial filter as interference signals at various levels for spatial filtering.

[0077] In some embodiments of this disclosure, when the filtering module is used to determine a valid signal based on the result of the forward filtering and a high-frequency signal corresponding to at least one of the microphones, and to determine a noise signal based on the result of the reverse filtering and a high-frequency signal corresponding to at least one of the microphones, it is specifically used for: The result of the forward filtering is combined with the high-frequency signal corresponding to the target microphone to obtain the effective signal; The result of the inverse filtering is combined with the high-frequency signal corresponding to the remote microphone to obtain the noise signal.

[0078] In some embodiments of this disclosure, a direction selection module is also included, for: Select the direction of the target sound source according to the recording mode selection command.

[0079] In some embodiments of this disclosure, the target sound source direction includes multiple sub-directions; The filtering module is specifically used for: For each sub-direction of the target sound source direction, determine the effective signal and the noise signal corresponding to the sub-direction; The noise cancellation module is specifically used for: Based on the noise signal corresponding to each sub-direction, noise cancellation processing is performed on the effective signal corresponding to each sub-direction to obtain the target audio sub-signal corresponding to each sub-direction; The target audio signal is determined based on the target audio sub-signal corresponding to each of the sub-directions.

[0080] In some embodiments of this disclosure, the target sound source direction includes a first sub-direction and a second sub-direction, wherein the first sub-direction and the second sub-direction are opposite.

[0081] In some embodiments of this disclosure, a frequency division point determination module is also included, for: The frequency division point is determined based on the distance between adjacent microphones; The frequency division module is specifically used for: If the frequency division point is less than the preset frequency division point threshold, the original audio signal collected by each microphone is divided according to the frequency division point to obtain the low-frequency signal and high-frequency signal corresponding to each microphone.

[0082] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments of the method in the first aspect, and will not be elaborated upon here.

[0083] According to a third aspect of the embodiments of this disclosure, please refer to the appendix. Figure 8 The diagram illustrates, for example, a block diagram of an electronic device. For instance, device 800 could be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0084] Reference Figure 8The device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0085] Processing component 802 typically controls the overall operation of device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.

[0086] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of this data include instructions for any application or method operating on device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0087] The power supply component 806 provides power to the various components of the device 800. The power supply component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 800.

[0088] Multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, swipe, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0089] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.

[0090] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0091] Sensor assembly 814 includes one or more sensors for providing status assessments of various aspects of device 800. For example, sensor assembly 814 can detect the on / off state of device 800, the relative positioning of components such as the display and keypad of device 800, image detection of changes in the position of device 800 or a component of device 800, the presence or absence of user contact with device 800, orientation or acceleration / deceleration of device 800, and temperature changes of device 800. Sensor assembly 814 may also include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0092] Communication component 816 is configured to facilitate wired or wireless communication between device 800 and other devices. Device 800 can access wireless networks based on communication standards, such as WiFi, 2G or 3G, 4G or 5G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0093] In an exemplary embodiment, device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the power supply method of the aforementioned electronic device.

[0094] Fourthly, in exemplary embodiments, this disclosure also provides a non-transitory computer-readable storage medium including instructions, such as a memory 804 including instructions, which can be executed by a processor 820 of device 800 to complete the power supply method of the electronic device. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0095] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0096] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. An audio processing method, characterized in that, Applied to an audio recording device having at least two microphones, the method includes: Acquire the raw audio signal captured by each of the at least two microphones; The original audio signal collected by each microphone is divided into low-frequency and high-frequency signals according to the frequency division point to obtain the low-frequency signal and high-frequency signal corresponding to each microphone. Based on the direction of the target sound source and the end-fire direction of each microphone, the low-frequency signal corresponding to each microphone is subjected to forward filtering and reverse filtering respectively. The effective signal is determined based on the result of the forward filtering and the high-frequency signal corresponding to at least one of the microphones. The noise signal is determined based on the result of the reverse filtering and the high-frequency signal corresponding to at least one of the microphones. The effective signal is a mixed signal of the target audio signal and the noise signal. The effective signal is subjected to noise cancellation processing using the noise signal to obtain the target audio signal; The step of performing forward filtering on the low-frequency signals corresponding to each microphone based on the direction of the target sound source and the end-fire direction of each microphone includes: The low-frequency signal corresponding to the target microphone is input as the target signal to the spatial filter, and the low-frequency signals corresponding to the other microphones are input as interference signals to the spatial filter for spatial filtering to obtain the positive filtering result. The end-fire direction of the target microphone is matched with the direction of the target sound source. The step of performing reverse filtering on the low-frequency signals corresponding to each microphone based on the direction of the target sound source and the end-fire direction of each microphone includes: The low-frequency signal corresponding to the far-end microphone is used as the target signal and input to the spatial filter. The low-frequency signals corresponding to the other microphones are used as interference signals and input to the spatial filter for spatial filtering to obtain the result of reverse filtering. The end-fire direction of the far-end microphone is opposite to the end-fire direction of the target microphone.

2. The audio processing method according to claim 1, characterized in that, The step of inputting the low-frequency signal corresponding to the target microphone as the target signal into the spatial filter, and inputting the low-frequency signals corresponding to the other microphones as interference signals into the spatial filter for spatial filtering includes: The low-frequency signal corresponding to the target microphone is used as the target signal input to the spatial filter. Based on the distance between each of the other microphones and the target microphone, the low-frequency signals corresponding to each of the remaining microphones are used as interference signals at various levels and input to the spatial filter for spatial filtering; and / or, The step of inputting the low-frequency signal corresponding to the far-end microphone as the target signal into the spatial filter, and inputting the low-frequency signals corresponding to the other microphones as interference signals into the spatial filter for spatial filtering includes: The low-frequency signal corresponding to the remote microphone is input into the spatial filter as the target signal. Based on the distance between each of the other microphones and the remote microphone, the low-frequency signals corresponding to each of the other microphones are input into the spatial filter as interference signals at various levels for spatial filtering.

3. The audio processing method according to claim 1, characterized in that, The step of determining a valid signal based on the result of the forward filtering and a high-frequency signal corresponding to at least one of the microphones, and determining a noise signal based on the result of the reverse filtering and a high-frequency signal corresponding to at least one of the microphones, includes: The result of the forward filtering is combined with the high-frequency signal corresponding to the target microphone to obtain the effective signal; The result of the inverse filtering is combined with the high-frequency signal corresponding to the remote microphone to obtain the noise signal.

4. The audio processing method according to claim 1, characterized in that, Also includes: Select the direction of the target sound source according to the recording mode selection command.

5. The audio processing method according to claim 1, characterized in that, The target sound source direction includes multiple sub-directions; The step of performing forward filtering and reverse filtering on the low-frequency signals corresponding to each microphone according to the direction of the target sound source and the end-fire direction of each microphone, determining the effective signal based on the result of the forward filtering and the high-frequency signal corresponding to at least one of the microphones, and determining the noise signal based on the result of the reverse filtering and the high-frequency signal corresponding to at least one of the microphones, includes: For each sub-direction of the target sound source direction, determine the effective signal and the noise signal corresponding to the sub-direction; The step of using the noise signal to perform noise cancellation processing on the effective signal to obtain the target audio signal includes: Based on the noise signal corresponding to each sub-direction, noise cancellation processing is performed on the effective signal corresponding to each sub-direction to obtain the target audio sub-signal corresponding to each sub-direction; The target audio signal is determined based on the target audio sub-signal corresponding to each of the sub-directions.

6. The audio processing method according to claim 5, characterized in that, The target sound source direction includes a first sub-direction and a second sub-direction, and the first sub-direction and the second sub-direction are opposite.

7. The audio processing method according to claim 1, characterized in that, Also includes: The frequency division point is determined based on the distance between adjacent microphones; The step of performing frequency division processing on the original audio signal collected by each microphone according to the frequency division point to obtain the low-frequency signal and high-frequency signal corresponding to each microphone includes: If the frequency division point is less than the preset frequency division point threshold, the original audio signal collected by each microphone is divided according to the frequency division point to obtain the low-frequency signal and high-frequency signal corresponding to each microphone.

8. An audio processing apparatus, characterized in that, Applied to an audio recording device having at least two microphones, the device includes: An acquisition module is used to acquire the raw audio signal collected by each of the at least two microphones; The frequency division module is used to perform frequency division processing on the original audio signal collected by each microphone according to the frequency division point, so as to obtain the low-frequency signal and high-frequency signal corresponding to each microphone; The filtering module is used to perform forward filtering and reverse filtering on the low-frequency signals corresponding to each microphone according to the direction of the target sound source and the end-fire direction of each microphone, and to determine the effective signal according to the result of the forward filtering and the high-frequency signal corresponding to at least one of the microphones, and to determine the noise signal according to the result of the reverse filtering and the high-frequency signal corresponding to at least one of the microphones, wherein the effective signal is a mixed signal of the target audio signal and the noise signal; A noise cancellation module is used to perform noise cancellation processing on the effective signal using the noise signal to obtain the target audio signal; The filtering module is used to perform forward filtering on the low-frequency signals corresponding to each microphone according to the direction of the target sound source and the end-fire direction of each microphone. Specifically, it is used for: The low-frequency signal corresponding to the target microphone is input as the target signal to the spatial filter, and the low-frequency signals corresponding to the other microphones are input as interference signals to the spatial filter for spatial filtering to obtain the positive filtering result. The end-fire direction of the target microphone is matched with the direction of the target sound source. The filtering module is used to perform reverse filtering on the low-frequency signals corresponding to each microphone according to the direction of the target sound source and the end-fire direction of each microphone. Specifically, it is used for: The low-frequency signal corresponding to the far-end microphone is used as the target signal and input to the spatial filter. The low-frequency signals corresponding to the other microphones are used as interference signals and input to the spatial filter for spatial filtering to obtain the result of reverse filtering. The end-fire direction of the far-end microphone is opposite to the end-fire direction of the target microphone.

9. The audio processing apparatus according to claim 8, characterized in that, The filtering module is used to input the low-frequency signal corresponding to the target microphone as the target signal into the spatial filter, and to input the low-frequency signals corresponding to the other microphones as interference signals into the spatial filter for spatial filtering. Specifically, it is used for: The low-frequency signal corresponding to the target microphone is used as the target signal input to the spatial filter. Based on the distance between each of the other microphones and the target microphone, the low-frequency signals corresponding to each of the remaining microphones are used as interference signals at various levels and input to the spatial filter for spatial filtering; and / or, The filtering module is used to input the low-frequency signal corresponding to the far-end microphone as the target signal into the spatial filter, and to input the low-frequency signals corresponding to the other microphones as interference signals into the spatial filter for spatial filtering. Specifically, it is used for: The low-frequency signal corresponding to the remote microphone is input into the spatial filter as the target signal. Based on the distance between each of the other microphones and the remote microphone, the low-frequency signals corresponding to each of the other microphones are input into the spatial filter as interference signals at various levels for spatial filtering.

10. The audio processing apparatus according to claim 8, characterized in that, The filtering module is used to determine a valid signal based on the result of the forward filtering and a high-frequency signal corresponding to at least one of the microphones, and to determine a noise signal based on the result of the reverse filtering and a high-frequency signal corresponding to at least one of the microphones, specifically for: The result of the forward filtering is combined with the high-frequency signal corresponding to the target microphone to obtain the effective signal; The result of the inverse filtering is combined with the high-frequency signal corresponding to the remote microphone to obtain the noise signal.

11. The audio processing apparatus according to claim 8, characterized in that, It also includes a direction selection module, used for: Select the direction of the target sound source according to the recording mode selection command.

12. The audio processing apparatus according to claim 8, characterized in that, The target sound source direction includes multiple sub-directions; The filtering module is specifically used for: For each sub-direction of the target sound source direction, determine the effective signal and the noise signal corresponding to the sub-direction; The noise cancellation module is specifically used for: Based on the noise signal corresponding to each sub-direction, noise cancellation processing is performed on the effective signal corresponding to each sub-direction to obtain the target audio sub-signal corresponding to each sub-direction; The target audio signal is determined based on the target audio sub-signal corresponding to each of the sub-directions.

13. The audio processing apparatus according to claim 12, characterized in that, The target sound source direction includes a first sub-direction and a second sub-direction, and the first sub-direction and the second sub-direction are opposite.

14. The audio processing apparatus according to claim 8, characterized in that, It also includes a frequency division point determination module, used for: The frequency division point is determined based on the distance between adjacent microphones; The frequency division module is specifically used for: If the frequency division point is less than the preset frequency division point threshold, the original audio signal collected by each microphone is divided according to the frequency division point to obtain the low-frequency signal and high-frequency signal corresponding to each microphone.

15. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory being used to store computer instructions that can be executed on the processor, and the processor being used to execute the computer instructions based on the audio processing method according to any one of claims 1 to 7.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice noise reduction device

    CN101853667A

  • Signal processing method, equipment and device

    CN112785998A