A method and apparatus for processing an audio signal

By using N beam processing and adaptive filtering techniques in multi-person conversation scenarios, the main beam signal and interference beam signals are determined, solving the problem of insufficient interference signal suppression capability in existing technologies and improving audio signal quality and voice interaction experience.

CN115691531BActive Publication Date: 2026-01-13HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110867113.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-29
Publication Date
2026-01-13
Estimated Expiration
2041-07-29

AI Technical Summary

Technical Problem

Existing audio signal processing methods have poor ability to suppress interference signals in multi-person conversation scenarios, which affects the quality of voice interaction.

Method used

N beams are used to perform beam processing on the audio signal to be processed, the main beam signal and the interference beam signal are determined, and they are processed by adaptive filtering technology to obtain the target sound source enhancement signal.

Benefits of technology

It effectively suppresses interference signals in audio signals, improving audio signal quality and voice interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115691531B_ABST
    Figure CN115691531B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a kind of processing method and device of audio signal, it is related to media technical field, the method comprises: using N beams respectively to the audio signal to be handled is carried out beam processing, obtains N beam signals, N is greater than or equal to 2 integer;The audio signal to be handled is in target sound pickup range, and the audio signal including at least one sound source is collected by microphone, the coverage range of the N beams is in target sound pickup range;And determine N beam signals main beam signal and N-1 interference beam signals, the main beam signal is the beam signal of target sound source;And the adaptive filtering is carried out to main beam signal and N-1 interference beam signals, to obtain target sound source enhancement signal.Using the method provided in the embodiments of the present application, the interference signal in audio signal can be effectively suppressed, and the quality of audio signal is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of media technology, and in particular to a method and apparatus for processing audio signals. Background Technology

[0002] More and more electronic devices (such as mobile phones, tablets, large-screen TVs, and in-vehicle devices) are integrating microphone arrays, enabling voice interaction with users. During this voice interaction, processing the audio signals captured by the microphone array to improve the quality of the interaction is crucial.

[0003] For example, in a multi-person conversation scenario, there may be one or more people speaking. The audio signals collected by the electronic device include multiple audio signals. Among them, the speech content of the main speaker (i.e., the main speaker's audio signal or data) needs to be focused on. Therefore, when the audio signal collected by the microphone is sent to the other end of the call, the audio signal needs to be processed to enhance the main speaker's audio signal and suppress other interfering audio signals.

[0004] Currently, beam-directed sound source enhancement technology can be used to enhance sound sources in a specific direction and reduce interference from sound sources in other directions. In one implementation, a beamformer, such as a minimum variance distortionless response (MVDR) beamformer, can be used to beamform the acquired audio signal, and then the target beam signal (i.e., the beam signal corresponding to the speaker) in the beamformed signal is filtered.

[0005] It should be understood that different beamformers and filtering methods result in different audio signal processing effects, and existing audio signal processing methods generally have poor ability to suppress interference signals. Summary of the Invention

[0006] This application provides an audio signal processing method and apparatus that can effectively suppress interference signals in audio signals and improve the quality of audio signals.

[0007] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:

[0008] In a first aspect, embodiments of this application provide an audio signal processing method applicable to electronic devices having an audio signal acquisition device (microphone). The method includes: performing beam processing on the audio signal to be processed using N beams to obtain N beam signals, where N is an integer greater than or equal to 2; the audio signal to be processed is an audio signal including at least one sound source acquired by the microphone within a target pickup range, and the coverage area of ​​the N beams is within the target pickup range; and determining a main beam signal and N-1 interfering beam signals among the N beam signals, wherein the main beam signal is the beam signal where the target sound source is located; and performing adaptive filtering on the main beam signal and the N-1 interfering beam signals to obtain a target sound source enhancement signal.

[0009] In the audio signal processing method provided in this application embodiment, after beam processing is performed on the audio signal to be processed, adaptive filtering is performed on the main beam signal and N-1 interference beam signals among the N beam signals. This can more thoroughly filter out the interference signals in the corresponding audio signal of the target sound source, that is, it can effectively suppress the interference signals in the audio signal and improve the quality of the audio signal.

[0010] Furthermore, the audio signal processing method provided in this application embodiment can be applied to interactive scenarios such as multi-person conversations (e.g., video conversations or voice conversations). Since the audio signal processing method provided in this application embodiment can effectively suppress interference signals in the audio signal and improve the quality of the audio signal, it can also improve the voice interaction experience.

[0011] In one possible implementation, the specific method for adaptively filtering the main beam signal and N-1 interfering beam signals to obtain the target sound source enhancement signal includes: using the main beam signal as a reference signal, adaptively filtering the N-1 interfering beam signals to obtain filtered N-1 interfering beam signals, wherein the audio signal corresponding to the target sound source in the filtered N-1 interfering beam signals is filtered out; then using the filtered N-1 interfering beam signals as a reference signal, adaptively filtering the main beam signal to obtain the target sound source enhancement signal, wherein the audio signal corresponding to the interfering sound source in the target sound source enhancement signal is filtered out.

[0012] The first adaptive filtering results in either no audio signal corresponding to the target sound source, or very little audio signal corresponding to the target sound source, in the filtered interference signal. The second adaptive filtering results in either no audio signal corresponding to the interfering beam, or very little audio signal corresponding to the interfering beam, in the filtered main beam signal. In summary, the above two-stage adaptive filtering effectively suppresses interference signals in the audio signal.

[0013] In one possible implementation, the filtered N-1 interfering beam signals satisfy the following:

[0014]

[0015] in, Let y represent the filtered i-th interfering beam signal. i (t,f) represents the i-th interfering beam signal before filtering. This represents the filter coefficients for adaptive filtering of N-1 interfering beam signals, where y0(t,f) represents the main beam signal; where i takes values ​​of 1, 2, ..., N-1, t is the current time, and f is the frequency; or

[0016] The enhanced signal of the above target sound source satisfies:

[0017]

[0018] in, Let yt represent the target sound source enhancement signal, and y0(t,f) represent the main beam signal before filtering. This represents the i-th filter coefficient used for adaptive filtering of the main beam signal, with the i-th filtered interference beam signal as the reference signal. Let f represent the filtered i-th interference beam signal; where i takes the values ​​1, 2, ..., N-1, t is the current time, and f is the frequency.

[0019] In one possible implementation, the filter coefficients for adaptive filtering of N-1 interfering beam signals are... satisfy:

[0020]

[0021] Where w(t-1,f) represents the filter coefficients at the previous time step, μ(t,f) represents the control coefficient update flag, and g1 represents the filter gain of the adaptive filter; or

[0022] The i-th filter coefficient for adaptive filtering of the main beam signal mentioned above satisfy:

[0023]

[0024] in, μ represents the filter coefficient at the previous time step. i (t, f) represents the control coefficient update flag, g 2i This represents the filter gain of the adaptive filter.

[0025] In one possible implementation, the control coefficient update identifier μ(t, f) described above satisfies:

[0026]

[0027]

[0028] Where α is the energy threshold of the beam signal, θ is the azimuth angle of the current sound source, and θ l The azimuth angle of the left boundary of the main beam, θ r The azimuth angle of the right boundary of the main beam.

[0029] Optionally, the value of α can be set according to the actual situation. The principle for selecting the value of α is that when there is a sound source within the main beam range and it is the main sound source, the ratio between the maximum value of the energy of the main beam signal and the maximum value of the energy of the interference beam signal is greater than α.

[0030] The above-mentioned control coefficient update identifier μ i (t, f) satisfies:

[0031]

[0032] Where θ is the azimuth angle of the current sound source, θ il Let θ be the azimuth angle of the left boundary of the i-th interfering beam. ir Let be the azimuth angle of the right boundary of the i-th interfering beam.

[0033] In one possible implementation, one beam corresponds to a set of beam coefficients, and the beam coefficients of the i-th beam out of N beams satisfy the following objective function:

[0034] min|η i (θ d ,f)Q(θ,f)||

[0035] The constraints of the objective function are:

[0036] |||η i (θ d ,f)Q(θ,f)||-1|<ε θ∈[θ i_l ,θ i_r ], and ||η i (θ d ,f)|| 2 <γ;

[0037] Where, η i (θ d Let Q(θ,f) be a set of beam coefficients corresponding to the i-th beam, and let Q(θ,f) be the array steering vector of the microphone array. d Let θ be the center angle of the i-th beam, and θ be the azimuth angle of the sound source. i_l Let θ be the azimuth angle of the left boundary of the i-th beam.i_r Let θ be the azimuth angle of the right boundary of the i-th beam, ε be the beam tolerance factor, γ be the beam robustness factor, and f be the frequency.

[0038] In this embodiment, the beamwidth and center angle can be preset or adjusted in real time according to the location of the sound source within the target pickup range (also known as the target area). The objective function described above means that audio signals originating outside the coverage area of ​​the i-th beam are attenuated to the maximum extent after processing by the i-th beam. The constraint 1 (|||η) above... i (θ d ,f)Q(θ,f)||-1|<ε θ∈[θ i_l ,θ i_r The meaning of ]) is: Audio signals from within the coverage area of ​​the i-th beam are preserved to the greatest extent possible after processing by the i-th beam. The above constraint 2 (||η) i (θ d ,f)|| 2 The meaning of <γ) is: for the i-th beam, after the white noise is processed by the i-th beam, the attenuation coefficient of the white noise is less than the preset value (i.e., the beam robustness factor γ). Usually, this value of γ is less than 1, and the smaller the value, the better.

[0039] In one possible implementation, the specific method of using N beams to perform beam processing on the audio signal to be processed to obtain N beam signals includes: using a set of beam coefficients corresponding to each of the N beams to perform beam processing on the audio signal to be processed to obtain N beam signals.

[0040] Assume the acquired audio signal to be processed is X, X = [x1 x2 ... x m ], where x m This represents the signal acquired by the m-th microphone. Taking the i-th beam as an example, the audio signal to be processed is processed using a set of filter coefficients corresponding to the i-th beam to obtain the i-th beam signal, which is Y. i Y i Satisfy: Y i =η i (θ d ,f)·X, where, η i (θ d f) represents the beam coefficient of the i-th beam, and · represents the dot product operation.

[0041] In one possible implementation, the main beam signal is the beam signal corresponding to the beam covering the location of the target sound source, and the interference beam signal is the beam signal corresponding to the beam not covering the location of the target sound source.

[0042] In one possible implementation, the audio signal processing method provided in this application embodiment further includes: determining the location of a target sound source, the location of which is used to determine the main beam signal.

[0043] In one possible implementation, the specific method for determining the location of the target sound source includes: acquiring current image information of the target pickup range, which includes at least one person; and performing lip movement detection on the at least one person based on the acquired image information to determine the lip movement state of the at least one person; the lip movement state being either present or absent; and determining the location of the person whose lip movement state is present as the location of the target sound source.

[0044] Optionally, artificial intelligence image processing techniques can be used to process one or more image frames to obtain the lip movement state of the person in the image. For example, machine learning algorithms or deep learning algorithms can be used to train a model for detecting lip movement state (e.g., a neural network model), and then the image to be detected can be input into the lip movement detection model to obtain the lip movement state of the person in the image.

[0045] In one possible implementation, the specific method for determining the location of the target sound source includes: performing sound source localization processing on the audio signal to be processed to determine the direction of arrival (DOA) of the sound source in the audio signal to be processed; and determining the location of the sound source located in the DOA as the location of the target sound source. Optionally, the DOA of the audio signal to be processed can be determined based on the DOAs of multiple audio frames of the audio signal to be processed within a preset time period, for example, determining the DOA of the audio signal to be processed as the DOA of the audio signal to be processed with the highest statistical value within the preset time period.

[0046] In one possible implementation, the specific method for determining the location of the target sound source includes: determining the location of the beam corresponding to the first beam signal among N beam signals as the location of the target sound source, wherein the first beam signal is the beam signal with the highest energy among the N beam signals.

[0047] Secondly, embodiments of this application provide an electronic device, including: a beam processing module, a determination module, and a filtering module. The beam processing module is used to perform beam processing on an audio signal to be processed using N beams, resulting in N beam signals, where N is an integer greater than or equal to 2; the audio signal to be processed is an audio signal including at least one sound source, collected by a microphone within the target pickup range, and the coverage area of ​​the N beams is within the target pickup range; the determination module is used to determine the main beam signal and N-1 interfering beam signals among the N beam signals, where the main beam signal is the beam signal containing the target sound source; the filtering module is used to adaptively filter the main beam signal and the N-1 interfering beam signals to obtain an enhanced signal for the target sound source.

[0048] In one possible implementation, the filtering module is specifically used to use the main beam signal as a reference signal to adaptively filter N-1 interfering beam signals to obtain filtered N-1 interfering beam signals, in which the audio signal corresponding to the target sound source is filtered out; and using the filtered N-1 interfering beam signals as a reference signal to adaptively filter the main beam signal to obtain a target sound source enhancement signal, in which the audio signal corresponding to the interfering sound source is filtered out.

[0049] In one possible implementation, the filtered N-1 interfering beam signals satisfy the following:

[0050]

[0051] in, Let y represent the filtered i-th interfering beam signal. i (t,f) represents the i-th interfering beam signal before filtering. This represents the filter coefficients for adaptive filtering of N-1 interfering beam signals, where y0(t,f) represents the main beam signal; where i takes values ​​of 1, 2, ..., N-1, t is the current time, and f is the frequency; or

[0052] The enhanced signal of the above target sound source satisfies:

[0053]

[0054] in, Let yt represent the target sound source enhancement signal, and y0(t,f) represent the main beam signal before filtering. This represents the i-th filter coefficient used for adaptive filtering of the main beam signal, with the i-th filtered interference beam signal as the reference signal. Let f represent the filtered i-th interference beam signal; where i takes the values ​​1, 2, ..., N-1, t is the current time, and f is the frequency.

[0055] In one possible implementation, the filter coefficients for adaptive filtering of N-1 interfering beam signals are... satisfy:

[0056]

[0057] Where w(t-1,f) represents the filter coefficients at the previous time step, μ(t,f) represents the control coefficient update flag, and g1 represents the filter gain of the adaptive filter; or

[0058] The i-th filter coefficient for adaptive filtering of the main beam signal above satisfies:

[0059]

[0060] in, μ represents the filter coefficient at the previous time step. i (t, f) represents the control coefficient update flag, g 2i This represents the filter gain of the adaptive filter.

[0061] In one possible implementation, the control coefficient update identifier μ(t, f) described above satisfies:

[0062]

[0063]

[0064] Where α is the energy threshold of the beam signal, θ is the azimuth angle of the current sound source, and θ l The azimuth angle of the left boundary of the main beam, θ r The azimuth angle of the right boundary of the main beam.

[0065] Optionally, the value of α can be set according to the actual situation. The principle for selecting the value of α is that when there is a sound source within the main beam range and it is the main sound source, the ratio between the maximum value of the energy of the main beam signal and the maximum value of the energy of the interfering beam signal is greater than α. Or

[0066] The above-mentioned control coefficient update identifier μ i (t, f) satisfies:

[0067]

[0068] Where θ is the azimuth angle of the current sound source, θ il Let θ be the azimuth angle of the left boundary of the i-th interfering beam. ir Let be the azimuth angle of the right boundary of the i-th interfering beam.

[0069] In one possible implementation, one beam corresponds to a set of beam coefficients, and the beam coefficients of the i-th beam out of N beams satisfy the following objective function:

[0070] min||η i (θ d ,f)Q(θ,f)||

[0071] The constraints of the objective function are:

[0072] |||η i (θ d ,f)Q(θ,f)||-1|<ε θ∈[θ i_l ,θ i_r ], and ||ηi (θ d ,f)|| 2 <γ;

[0073] Where, η i (θ d Let Q(θ,f) be a set of beam coefficients corresponding to the i-th beam, and let Q(θ,f) be the array steering vector of the microphone array. d Let θ be the center angle of the i-th beam, and θ be the azimuth angle of the sound source. i_l Let θ be the azimuth angle of the left boundary of the i-th beam. i_r Let θ be the azimuth angle of the right boundary of the i-th beam, ε be the beam tolerance factor, γ be the beam robustness factor, and f be the frequency.

[0074] In one possible implementation, the main beam signal is the beam signal corresponding to the beam covering the location of the target sound source, and the interference beam signal is the beam signal corresponding to the beam not covering the location of the target sound source.

[0075] In one possible implementation, the electronic device provided in this application embodiment further includes a sound source localization module, which is used to determine the location of a target sound source, and the location of the target sound source is used to determine the main beam signal.

[0076] In one possible implementation, the electronic device provided in this application embodiment further includes an image acquisition module and a detection module. The image acquisition module is used to acquire image information of a target sound pickup range, which includes at least one person. The detection module is used to perform lip movement detection on at least one person based on the acquired image information to determine the lip movement state of at least one person. The lip movement state is either present or absent. The sound source localization module is specifically used to determine the position of the person whose lip movement state is present as the position of the target sound source.

[0077] The sound source localization module is specifically used to perform sound source localization processing on the audio signal to be processed, determine the direction of arrival of the sound source in the audio signal to be processed, and determine the position of the sound source located in the direction of arrival as the position of the target sound source.

[0078] In one possible implementation, the sound source localization module is specifically used to determine the position of the beam corresponding to the first beam signal among the N beam signals as the position of the target sound source. The first beam signal is the beam signal with the largest energy among the N beam signals.

[0079] Thirdly, embodiments of this application provide an electronic device, including a memory and at least one processor connected to the memory. The memory is used to store instructions, and after the instructions are read by the at least one processor, the method described in the first aspect and any of its possible implementations is executed.

[0080] Fourthly, a computer-readable storage medium having a computer program stored thereon, characterized in that, when executed by a processor, the computer program implements the method described in the first aspect and any of its possible implementations.

[0081] Fifthly, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the method described in the first aspect and any of its possible implementations.

[0082] Sixthly, embodiments of this application provide a chip including a memory and a processor. The memory is used to store computer instructions. The processor is used to retrieve and execute the computer instructions from the memory to perform the method described in the first aspect and any of its possible implementations.

[0083] It should be understood that the beneficial effects achieved by the second to sixth aspects of the technical solutions and the corresponding possible implementations of the embodiments of this application can be referred to the above-described technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description

[0084] Figure 1 A schematic diagram of a multi-person voice conversation scenario provided for an embodiment of this application;

[0085] Figure 2 A schematic diagram of the hardware structure of a mobile phone provided in an embodiment of this application;

[0086] Figure 3 A schematic diagram illustrating an audio signal processing method provided in an embodiment of this application;

[0087] Figure 4 A schematic diagram of a wide beam provided for an embodiment of this application;

[0088] Figure 5 A schematic diagram of a wide-view multi-person conversation scenario provided for an embodiment of this application;

[0089] Figure 6 A schematic diagram of a narrow beam provided in an embodiment of this application;

[0090] Figure 7 A schematic diagram of a narrow-angle multi-person conversation scenario provided for an embodiment of this application;

[0091] Figure 8 A schematic diagram of a linear microphone array provided for an embodiment of this application;

[0092] Figure 9 A schematic diagram illustrating another audio signal processing method provided in an embodiment of this application;

[0093] Figure 10 A schematic diagram illustrating another audio signal processing method provided in an embodiment of this application;

[0094] Figure 11 A schematic diagram illustrating another audio signal processing method provided in an embodiment of this application;

[0095] Figure 12 A schematic diagram illustrating another audio signal processing method provided in an embodiment of this application;

[0096] Figure 13 A schematic diagram illustrating another audio signal processing method provided in an embodiment of this application;

[0097] Figure 14 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0098] Figure 15 This is a schematic diagram of the structure of another electronic device provided in an embodiment of this application. Detailed Implementation

[0099] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0100] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0101] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple beams means two or more beams.

[0102] It should be noted that in the embodiments of this application, audio signal can also be referred to as audio data. That is to say, the words signal and data have the same meaning in the embodiments of this application, and the following embodiments will not make a distinction.

[0103] First, some concepts involved in the audio signal processing method and apparatus provided in the embodiments of this application will be explained.

[0104] Beamforming, also known as beamforming or spatial filtering, involves beamforming data from an array of multiple array elements (e.g., antenna elements). The beamformed data becomes directional in a predetermined direction. For ease of understanding, the unprocessed data is referred to as the data to be processed. This data includes data collected by sensors on multiple array elements. Data collected by a single array element's sensors can include data transmitted from different directions or angles. The directional nature of the processed data in the predetermined direction can be understood as: enhancing data from the predetermined direction and suppressing data from other directions (non-predetermined directions).

[0105] Beamforming is essentially a weighted processing of data. The beamforming process described above can also be seen as a type of filtering process, and the weights in the weighting process are the filtering coefficients (or beam coefficients) of the beam. In other words, the data is weighted using a pre-designed beam (corresponding to this set of beam coefficients) to obtain the beam signal.

[0106] In the audio field, the aforementioned multiple array elements can be multiple microphones, which together form a microphone array. The microphone array can be a linear array or a circular array, etc. The data to be processed is audio data (also referred to as audio signals) collected by multiple microphones. Beamforming processing of the audio signals involves weighted summation of the multiple microphone signals to suppress interference signals from non-target directions and enhance the sound signal from the target direction. The target direction can be the direction of interest to the user or the direction of the sound source; this application does not limit this.

[0107] For example, refer to Figure 1 The scenario shown is a multi-person voice conversation. Figure 1 The audio device shown serves as one end of a voice conversation, and its microphone array (including multiple microphones) Figure 1 The microphone array shown is a linear array, capable of capturing all audio signals within the target pickup range. The audio device processes the audio signals and then sends them to the other end of the voice conversation. Figure 1 In the audio device, the target pickup range of the audio device includes attendee 1, attendee 2 and attendee 3. The target pickup range also includes a fan. Optionally, the target pickup range may also include other devices or things that can generate sound. This application embodiment does not limit this.

[0108] Combination Figure 1During a voice conversation, when one of the three participants speaks as the main speaker, the audio signal captured by the microphone array includes the main speaker's audio signal. If other participants are speaking quietly nearby, the captured audio signal will also include their audio signals, as well as noise from the fan. The other participant in the voice conversation is primarily focused on the main speaker's content. Therefore, from their perspective, the audio signals of other participants and the fan noise in the audio signal captured by the voice device are interference signals. The voice device needs to suppress audio signals other than the main speaker's audio signal (i.e., other participants' audio signals, fan noise, and other environmental noise) as much as possible when processing the captured audio signal.

[0109] Assumption Figure 1 In this example, participant 2 is the current speaker. Six beams, designated beam 1 through beam 6, are set within the target pickup range of the audio device. After the audio device acquires audio signals (which can be referred to as the audio signals to be processed) including the sounds of participant 1, participant 2, participant 3, and the fan, these signals are processed by beamforming through the six beams to obtain six beam signals. Specifically, beamforming through beam 4 enhances the audio signals in the direction of beam 4 while suppressing audio signals in other directions. This means that the audio signal of participant 2 can be enhanced, while the audio signals of participant 1, participant 2, or the fan can be suppressed. Therefore, users can enhance and suppress audio signals according to their actual needs.

[0110] In this embodiment, the beam may include a wide beam, a narrow beam, or other types of beam. Beamforming techniques may include fixed beamforming, adaptive beamforming, and so on.

[0111] Fixed beamforming involves pre-setting beamwidth and center angle, and then establishing a beam model (objective function and constraints) based on the set center angle and beamwidth. Adaptive beamforming, on the other hand, optimizes the weight set (i.e., calculates the optimal weights) using an adaptive algorithm under a certain optimal criterion. Adaptive beamforming can adapt to various environmental changes, adjusting the weight set to near the optimal position in real time. Currently, commonly used adaptive beamforming techniques include minimum variance distortionless response (MVDR) beamforming and linearly constrained minimum variance (LCMV) beamforming. A device using MVDR beamforming is called an MVDR beamformer, and similarly, a device using LCMV beamforming is called an LCMV beamformer. Different beamformers correspond to different beam models, which include the objective function and constraints on the objective function.

[0112] The beam signal after beamforming is then subjected to post-beam processing to obtain the audio signal of the target area. Post-beam processing methods include Kalman filtering, normalized least mean square (NLMS) filtering, and recursive least square (RLS) filtering.

[0113] It should be understood that different beamformers and filtering methods produce varying effects on audio signal processing, and existing audio signal processing methods generally suffer from poor interference suppression capabilities. For example, existing beam post-processing methods cannot effectively filter interference signals present in the beam signal.

[0114] To address the problem that existing technologies cannot effectively suppress interference signals in audio signals, this application provides an audio signal processing method and apparatus. The method is applied to an electronic device with a microphone. Within the target pickup range, the electronic device acquires an audio signal (referred to as the audio signal to be processed) containing at least one sound source through the microphone. Then, it performs beamforming processing (i.e., beamforming) on ​​the audio signal to be processed using N beams, obtaining N beam signals. The coverage area of ​​these N beams is within the target pickup range, where N is an integer greater than or equal to 2. The electronic device then determines the main beam signal and N-1 interfering beam signals from the N beam signals. The main beam signal is the beam signal containing the target sound source. Furthermore, the electronic device performs adaptive filtering on the main beam signal and the N-1 interfering beam signals to obtain an enhanced signal for the target sound source. The technical solution provided by this application can effectively suppress interference signals in audio signals, improving the quality of the audio signal.

[0115] The audio signal processing method and apparatus provided in this application can be applied to various scenarios involving voice data processing, such as voice calls, video calls, and zoom photography. Voice calls and video calls are already widely used in interpersonal, family, and corporate meetings, where enhancing the speaker's voice is of great significance. The audio signal processing method provided in this application can improve the quality of voice calls and video calls, enhance the voice interaction experience, and has significant practical value.

[0116] Optionally, the audio signal processing method provided in this application embodiment can be applied to electronic devices with audio signal acquisition function (i.e., with microphones), such as mobile phones, tablets, laptops, smart speakers, televisions and other electronic devices equipped with microphones.

[0117] For example, taking a mobile phone as an example, Figure 2A schematic diagram of the structure of a mobile phone 200 is shown. The mobile phone 200 includes a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, a sensor module 280, buttons 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc. The sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an accelerometer sensor 280E, a distance sensor 280F, a proximity sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.

[0118] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the mobile phone 200. In other embodiments of this application, the mobile phone 200 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0119] Processor 210 may include one or more processing units, such as application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0120] The controller can serve as the central nervous system and command center of the mobile phone 200. Based on the instruction opcode and timing signals, the controller generates operation control signals to control the fetching and execution of instructions.

[0121] The processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can store instructions or data that the processor 210 has just used or that are used repeatedly. If the processor 210 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0122] In some embodiments, the processor 210 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0123] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 210 may include multiple I2C buses. The processor 210 can couple to the touch sensor 280K, charger, flash, camera 293, etc., through different I2C bus interfaces. For example, the processor 210 can couple to the touch sensor 280K through the I2C interface, enabling the processor 210 and the touch sensor 280K to communicate through the I2C bus interface, thus realizing the touch function of the mobile phone 200.

[0124] The I2S interface can be used for audio communication. In some embodiments, the processor 210 may include multiple I2S buses. The processor 210 can be coupled to the audio module 270 via the I2S bus to enable communication between the processor 210 and the audio module 270. In some embodiments, the audio module 270 can transmit audio signals to the wireless communication module 260 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.

[0125] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 270 and the wireless communication module 260 can be coupled via the PCM bus interface. In some embodiments, the audio module 270 can also transmit audio signals to the wireless communication module 260 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0126] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 210 and the wireless communication module 260. For example, the processor 210 communicates with the Bluetooth module in the wireless communication module 260 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 270 can transmit audio signals to the wireless communication module 260 via the UART interface to enable music playback through Bluetooth headphones.

[0127] The MIPI interface can be used to connect the processor 210 to peripheral devices such as the display screen 294 and the camera 293. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 210 and the camera 293 communicate via the CSI interface to enable the mobile phone 200 to take pictures. The processor 210 and the display screen 294 communicate via the DSI interface to enable the mobile phone 200 to display.

[0128] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 210 to a camera 293, a display screen 294, a wireless communication module 260, an audio module 270, a sensor module 280, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0129] USB port 230 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, or USB Type-C port. USB port 230 can be used to connect a charger to charge mobile phone 200, and can also be used for data transfer between mobile phone 200 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.

[0130] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the mobile phone 200. In other embodiments of this application, the mobile phone 200 may also adopt different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0131] The charging management module 240 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 240 receives charging input from the wired charger via a USB interface 230. In some wireless charging embodiments, the charging management module 240 receives wireless charging input via the wireless charging coil of the mobile phone 200. While charging the battery 242, the charging management module 240 can also supply power to the electronic device via the power management module 241.

[0132] The power management module 241 connects the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives input from the battery 242 and / or the charging management module 240, providing power to the processor 210, internal memory 221, external memory, display screen 294, camera 293, and wireless communication module 260. The power management module 241 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 241 may also be located within the processor 210. In other embodiments, the power management module 241 and the charging management module 240 may be housed in the same device.

[0133] The wireless communication function of mobile phone 200 can be realized through antenna 1, antenna 2, mobile communication module 250, wireless communication module 260, modem processor and baseband processor.

[0134] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in mobile phone 200 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.

[0135] The mobile communication module 250 can provide solutions for wireless communication applications including 2G / 3G / 4G / 5G on the mobile phone 200. The mobile communication module 250 may include at least one filter, switch, power amplifier, low-noise amplifier (LNA), etc. The mobile communication module 250 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 250 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 250 may be housed in the processor 210. In some embodiments, at least some functional modules of the mobile communication module 250 and at least some modules of the processor 210 may be housed in the same device.

[0136] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 270A, receiver 270B, etc.) or displays images or videos through the display screen 294. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 210 and may be housed in the same device as the mobile communication module 250 or other functional modules.

[0137] The wireless communication module 260 can provide solutions for wireless communication applications on the mobile phone 200, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 260 can be one or more devices integrating at least one communication processing module. The wireless communication module 260 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 210. The wireless communication module 260 can also receive signals to be transmitted from processor 210, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0138] In some embodiments, antenna 1 of mobile phone 200 is coupled to mobile communication module 250, and antenna 2 is coupled to wireless communication module 260, enabling mobile phone 200 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).

[0139] The mobile phone 200 implements its display function through a GPU, a display screen 294, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0140] Display screen 294 is used to display images, videos, etc. Display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, mobile phone 200 may include one or N displays 294, where N is a positive integer greater than 1.

[0141] The mobile phone 200 can achieve shooting functions through ISP, camera 293, video codec, GPU, display 294 and application processor.

[0142] The ISP (Image Signal Processor) is used to process data fed back from the camera 293. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 293.

[0143] Camera 293 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, mobile phone 200 may include one or N cameras 293, where N is a positive integer greater than 1.

[0144] A digital signal processor (DSP) is used to process digital signals. Besides digital image signals, it can also process other digital signals (such as audio signals). For example, when the mobile phone 200 is selecting a frequency, the DSP performs Fourier transforms on the frequency energy.

[0145] Video codecs are used to compress or decompress digital video. Mobile phone 200 can support one or more video codecs. Thus, mobile phone 200 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0146] NPU stands for Neural Network (NN) Computing Processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in mobile phones, such as image recognition, facial recognition, speech recognition, and text understanding.

[0147] The external storage interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the mobile phone 200. The external memory card communicates with the processor 210 through the external storage interface 220 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0148] Internal memory 221 can be used to store computer executable program code, which includes instructions. Processor 210 executes various functional applications and data processing of mobile phone 200 by running the instructions stored in internal memory 221. Internal memory 221 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of mobile phone 200 (such as audio data, phonebook, etc.). Furthermore, internal memory 221 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0149] The mobile phone 200 can perform audio functions, such as music playback and recording, through an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, and an application processor.

[0150] The audio module 270 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 270 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 270 may be located in the processor 210, or some functional modules of the audio module 270 may be located in the processor 210.

[0151] The speaker 270A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The mobile phone 200 can listen to music or make hands-free calls through the speaker 270A.

[0152] The receiver 270B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the mobile phone 200 answers a call or voice message, the receiver 270B can be brought close to the user's ear to listen to the voice.

[0153] Microphone 270C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 270C, inputting the sound signal into microphone 270C. Mobile phone 200 may have at least one microphone 270C. In some embodiments, mobile phone 200 may have two microphones 270C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, mobile phone 200 may also have three, four, or more microphones 270C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0154] The headphone jack 270D is used to connect wired headphones. The headphone jack 270D can be a USB 230 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0155] Pressure sensor 280A is used to sense pressure signals and convert them into electrical signals. Gyroscope sensor 280B is used to determine the motion posture of mobile phone 200. Barometric pressure sensor 280C is used to measure air pressure. Magnetic sensor 280D includes a Hall sensor. Accelerometer sensor 280E can detect the magnitude of acceleration of mobile phone 200 in various directions (generally three axes). Proximity sensor 280G may include, for example, a light-emitting diode (LED) and a photodetector, such as a photodiode. The LED can be an infrared LED. Ambient light sensor 280L is used to sense ambient light brightness. Fingerprint sensor 280H is used to collect fingerprints. Temperature sensor 280J is used to detect temperature. Touch sensor 280K, also called a "touch panel," can be located on display screen 294. Touch sensor 280K and display screen 294 together form a touch screen, also called a "touchscreen." Touch sensor 280K is used to detect touch operations applied to or near it.

[0156] The bone conduction sensor 280M can acquire vibration signals. In some embodiments, the bone conduction sensor 280M can acquire vibration signals from the vibrating bone segments of the human vocal cords. The bone conduction sensor 280M can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor 280M can also be incorporated into headphones to form bone conduction headphones. The audio module 270 can parse the voice signals from the vibrating bone segments of the vocal cords acquired by the bone conduction sensor 280M to realize voice functionality. The application processor can parse heart rate information from the blood pressure signals acquired by the bone conduction sensor 280M to realize heart rate detection functionality.

[0157] Keypad 290 includes a power button, volume buttons, etc. Keypad 290 can be a mechanical keypad or a touch-sensitive keypad. Mobile phone 200 can receive keypad input and generate key signal inputs related to user settings and function control of mobile phone 200.

[0158] Motor 291 can generate vibration alerts. Motor 291 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can be corresponding to touch operations applied to different applications (such as taking photos, playing audio, etc.). Motor 291 can also correspond to different vibration feedback effects for touch operations applied to different areas of the display screen 294. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.

[0159] Indicator 292 can be an indicator light, which can be used to indicate charging status, power changes, messages, missed calls, notifications, etc.

[0160] The SIM card interface 295 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 295 to make contact with and separate from the mobile phone 200. The mobile phone 200 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 295 can support Nano SIM cards, Micro SIM cards, and other SIM cards. Multiple cards can be inserted into the same SIM card interface 295 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 295 is also compatible with different types of SIM cards. The SIM card interface 295 is also compatible with external memory cards. The mobile phone 200 interacts with the network through the SIM card to achieve functions such as calls and data communication. In some embodiments, the mobile phone 200 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the mobile phone 200 and cannot be separated from the mobile phone 200.

[0161] It is understood that in the embodiments of this application, the electronic device (e.g., the mobile phone 200 described above) can perform some or all of the steps in the embodiments of this application. These steps or operations are merely examples, and the embodiments of this application can also perform other operations or variations thereof. Furthermore, the various steps can be performed in different orders as presented in the embodiments of this application, and it is not necessary to perform all the operations in the embodiments of this application. The embodiments of this application can be implemented individually or in any combination, and this application does not limit this.

[0162] Based on the above application scenarios, taking a linear microphone array as an example, the audio signal processing method provided in the embodiments of this application will be described in detail below, such as... Figure 3 As shown, the method for processing the audio signal includes steps 301 to 303.

[0163] Step 301: The electronic device uses N beams to perform beam processing on the audio signal to be processed, resulting in N beam signals.

[0164] Where N is an integer greater than or equal to 2.

[0165] It should be understood that the audio signal to be processed is an audio signal including at least one sound source collected by a microphone within the target pickup range, and the coverage area of ​​the N beams is within the target pickup range.

[0166] Optionally, in this embodiment, the target pickup range is 0°-180°. It should be noted that 0°-180° here is defined with reference to the direction parallel to the microphone array. For example, refer to the above... Figure 1 Assuming the microphone array is a linear array and parallel to the horizontal direction, the target pickup range can be... Figure 1The range shown is 0°-180°. Of course, the target pickup range can be 0°-360°, and this application embodiment does not limit it.

[0167] In this embodiment of the application, for the set target pickup range, N beams covering the target pickup range are first designed, and the beam coefficient corresponding to each beam is calculated. The specific process is as follows: S1 to S2.

[0168] S1. Set the width and center angle of each beam.

[0169] In this embodiment, the beam width is the beam pointing range. For example, assuming N is 3, the beam widths, center angles, azimuth angles of the left and right boundaries of the three beams are shown in Table 1 below. Referring to Table 1, the schematic diagrams of the three beams can be found... Figure 4 .

[0170] Table 1

[0171] Beam Central angle width Left boundary azimuth Right boundary azimuth Beam 1 20° 40° 40° 0° Beam 2 90° 100° 140° 40° Beam 3 160° 40° 180° 140°

[0172] Table 1 and Figure 4 Beam 1 in the table is a wide beam. For example, the N beams shown in Table 1 above can be applied to applications such as... Figure 5 The scene shown is a wide-angle, multi-person conversation.

[0173] For example, assuming N is 6, the widths of the six beams, their center directions, and the azimuth angles of their left and right boundaries are shown in Table 2 below. Referring to Table 2, the schematic diagrams of the six beams can be found... Figure 6 .

[0174] Table 2

[0175] Beam Central direction width Left boundary azimuth Right boundary azimuth Beam 1 15° 30° 30° 0° Beam 2 45° 30° 60° 30° Beam 3 75° 30° 90° 60° Beam 4 105° 30° 120° 90° Beam 5 135° 30° 150° 120° Beam 6 165° 30° 180° 150°

[0176] Table 2 and Figure 6 The beams shown are all narrow beams. For example, the N beams shown in Table 2 above can be applied to, for example... Figure 7 The scene shown is a narrow-angle, multi-person conversation.

[0177] Optionally, the center angle and beamwidth of the aforementioned beam can be preset or adjusted in real time according to the location of the sound source within the target pickup range (also known as the target area).

[0178] S2. Calculate the beam coefficient.

[0179] In this embodiment of the application, each beam corresponds to a set of beam coefficients, and the beam coefficient of the i-th beam among the above N beams satisfies the objective function shown in the following expression (1):

[0180]

[0181] The constraints of the objective function are expressions (2) and (2):

[0182] Constraint 1: ||η i (θ d ,f)Q(θ,f)||-1|<ε θ∈[θ i_l ,θ i_r (2)

[0183] Constraint 2: ||η i (θ d ,f)|| 2 <γ (3)

[0184] Where, η i (θ d Let Q(θ,f) be a set of beam coefficients corresponding to the i-th beam, and let Q(θ,f) be the array steering vector of the microphone array. d Let θ be the angle corresponding to the center direction of the i-th beam, and θ be the azimuth angle of the sound source. i_l Let θ be the azimuth angle of the left boundary of the i-th beam. i_r Let θ be the azimuth angle of the right boundary of the i-th beam, ε be the beam tolerance factor, γ be the beam robustness factor, and f be the frequency.

[0185] For the i-th beam, at frequency f, the beam coefficient η of the i-th beam... i (θ d f) is:

[0186] η i (θ d ,f)=[η i1 η i2 ... η im ]

[0187] Where m is the number of microphones included in the microphone array.

[0188] refer to Figure 8 The linear microphone array comprises m microphones, and the steering vector Q(θ,f) of the above microphone array can be calculated using the following expression (4):

[0189]

[0190] Where, d m This represents the distance of the m-th microphone from the reference position. Optionally, the position of the 1st microphone can be used as the reference position. The above d1 represents the distance of the 1st microphone from the 1st microphone, and d1 = 0.

[0191] In this embodiment of the application, the objective function means that the audio signal from outside the coverage area of ​​the i-th beam is attenuated to the greatest extent after being processed by the i-th beam.

[0192] The meaning of constraint 1 above is: after the audio signal from the coverage area of ​​the i-th beam is processed by the i-th beam, the audio signal is preserved to the greatest extent.

[0193] The meaning of constraint 2 above is: for the i-th beam, the attenuation coefficient of the white noise after processing by the i-th beam is less than γ. Generally, this value of γ is less than 1, and the smaller the value, the better.

[0194] Optionally, in this embodiment, the objective function can be solved using a convex optimization correlation method to obtain the beam coefficient of the i-th beam. Following this method, the beam coefficients of N beams are obtained and stored for beam processing of the acquired audio signals.

[0195] In the embodiments of this application, the beam width and center angle of the beam can be flexibly set according to actual needs in the above beam design method, that is, the viewing angle can be freely selected. Therefore, the beam design method can realize flexible beam design.

[0196] Optionally, step 301 can be implemented through step 3011.

[0197] Step 3011: The electronic device uses a set of beam coefficients corresponding to N beams to perform beam processing on the audio signal to be processed, and obtains N beam signals.

[0198] Assume the acquired audio signal to be processed is X, X = [x1 x2 ... x m ], where x m This represents the signal captured by the m-th microphone. Taking the i-th beam as an example again, the audio signal to be processed is processed using a set of filter coefficients corresponding to the i-th beam to obtain the i-th beam signal, which is Y. i Y i Satisfy the following expression (5):

[0199] Y i =η i (θ d ,f)·X (5)

[0200] Where, η i (θ d f) represents the beam coefficient of the i-th beam, and · represents the dot product operation.

[0201] Step 302: The electronic device determines the main beam signal and N-1 interfering beam signals among the N beam signals.

[0202] Among them, the main beam signal is the beam signal where the target sound source is located, that is, the main beam signal is the beam signal corresponding to the beam that covers the location of the target sound source, and the interference beam signal is the beam signal corresponding to the beam that does not cover the location of the target sound source.

[0203] In this embodiment, the target sound pickup range may include multiple sound sources (e.g., multiple participants within the target sound pickup range), one of which is the sound source the user is interested in (referred to as the target sound source, such as the speaker). This embodiment aims to enhance the audio signal corresponding to the target sound source. Optionally, the target sound source can be selected according to changes in the scenario; for example, in a multi-person conference scenario, the participant with the loudest voice can be selected as the target sound source.

[0204] Optionally, the location of the target sound source is the orientation of the target sound source. The orientation of the target sound source can be represented by an angle (then the location of the target sound source is the azimuth angle of the target sound source), or it can be identified by a vector or by coordinates. This application does not limit the specific orientation of the target sound source.

[0205] like Figure 9 As shown, before the electronic device determines the main beam signal and N-1 interfering beam signals among the N beam signals, the audio signal processing method provided in this application embodiment further includes step 304.

[0206] Step 304: The electronic device determines the location of the target sound source, and the location of the target sound source is used to determine the main beam signal.

[0207] In one implementation, if the electronic device has a camera, the camera can be used as an auxiliary device to determine the location of the target sound source. Specifically, in conjunction with... Figure 9 ,like Figure 10 As shown, step 304 can be achieved through steps 3041 to 3043.

[0208] Step 3041: The electronic device acquires image information of the target sound pickup range, which includes at least one person.

[0209] Step 3042: The electronic device performs lip movement detection on at least one person based on the image information to determine the lip movement state of at least one person.

[0210] In this embodiment, the lip movement state is either present or absent. If the electronic device determines that the mouth of a person in the image is closed based on the image information, then the electronic device determines that the person's lip movement state is absent; if the electronic device determines that the mouth of a person in the image is open based on the image information, then the electronic device determines that the person's lip movement state is present.

[0211] Optionally, artificial intelligence (AI) image processing techniques can be used to process one or more image frames to obtain the lip movement state of the person in the image. For example, machine learning algorithms or deep learning algorithms can be used to train a model for detecting lip movement state (e.g., a neural network model), and then the image to be detected can be input into the lip movement detection model to obtain the lip movement state of the person in the image.

[0212] Step 3043: The electronic device determines the position of the target sound source as the position of at least one person whose lip movement is present.

[0213] For example, refer to Figure 5 Electronic devices collect data including Figure 5 The electronic device uses images of three participants and determines that participant 2 has lip movements, thus identifying the position of participant 2 as the location of the target sound source.

[0214] In another implementation, the electronic device can employ sound source localization technology, such as direction of arrival (DOA) technology, to determine the location of the target sound source. Specifically, in conjunction with... Figure 9 ,like Figure 11 As shown, step 304 can be achieved through steps 3044 to 3045.

[0215] Step 3044: The electronic device performs sound source localization processing on the audio signal to be processed to determine the direction of arrival of the sound source in the audio signal to be processed.

[0216] Step 3045: The electronic device determines the location of the sound source in the direction of arrival as the location of the target sound source.

[0217] Optionally, the electronic device can determine the direction of arrival (DOA) of the audio signal to be processed based on the DOA of multiple audio frames within a preset time period. For example, the DOA with the highest statistical value within the preset time period can be determined as the DOA of the audio signal to be processed. For a description of DOA technology, please refer to the relevant descriptions in the prior art; details will not be elaborated here.

[0218] In another implementation, the electronic device can determine the location of the target sound source based on the energy of the N beam signals obtained from the waveform processing described above. Specifically, combined with... Figure 9 ,like Figure 12 As shown, step 304 above can be achieved through step 3046.

[0219] Step 3046: The electronic device determines the position of the beam corresponding to the first beam signal among the N beam signals as the position of the target sound source. The first beam signal is the beam signal with the largest energy among the N beam signals.

[0220] In this embodiment of the application, the beam signal corresponding to a certain beam has the greatest energy, which means that the sound source within the coverage area of ​​the beam has greater energy, and the sound source within the coverage area of ​​the beam is the target sound source.

[0221] Step 303: The electronic device performs adaptive filtering on the main beam signal and N-1 interfering beam signals to obtain the target sound source enhancement signal.

[0222] It should be understood that the above adaptive filtering can be implemented using a sidelobe cancellation algorithm with an adaptive filtering structure. For N beams, assuming the beam processing result of the beam covering the direction of the target sound source (hereinafter referred to as the main beam) is y0(t,f), and the beam processing result of the other beams (hereinafter referred to as interference beams) is y0(t,f). i (t,f), where i takes values ​​of 1, 2, ..., N-1.

[0223] y0(t,f) satisfies the following expression (6):

[0224]

[0225] Where x0(t,f) represents the audio signal corresponding to the target sound source, w i (t,f) represents the coefficients of the audio signal corresponding to the interfering sound source reaching the main beam, x i (t,f) represents the audio signal corresponding to the interference source, and n0(t,f) represents the additive noise signal.

[0226] y i (t,f) satisfies the following expression (7):

[0227] y i (t,f)=x i (t,f)+w(t,f)x0(t,f)+n i (t,f) (7)

[0228] Where, x i(t,f) represents the audio signal corresponding to the interfering sound source, w(t,f) represents the coefficient of the audio signal corresponding to the target sound source reaching the interfering beam, x0(t,f) represents the audio signal corresponding to the target sound source, and n i (t,f) represents an additive noise signal.

[0229] It should be understood that in the above expressions (6) and (7), t is the current time and f is the frequency.

[0230] Combination Figure 9 ,like Figure 13 As shown, based on the above expressions (6) and (7), the adaptive filtering in step 303 is a quadratic adaptive filtering. The specific step 303 can be implemented through steps 3031 to 3032.

[0231] Step 3031: The electronic device uses the main beam signal as a reference signal to adaptively filter the N-1 interference beam signals to obtain the filtered N-1 interference beam signals. The audio signal corresponding to the target sound source in the filtered N-1 interference beam signals is filtered out.

[0232] Step 3031 corresponds to the first adaptive filtering process. The filtered N-1 interference beam signals satisfy expression (8):

[0233]

[0234] in, Let y represent the filtered i-th interfering beam signal. i (t,f) represents the i-th interfering beam signal before filtering (i.e., the beam processing result). The filtering coefficients for adaptive filtering of N-1 interfering beam signals are shown. It can be seen that the coefficients of the audio signal corresponding to the target sound source in the above expression (7) that reach the interfering beam are the filtering coefficients for adaptive filtering of the interfering beam signals. y0(t,f) represents the main beam signal; where i takes the values ​​1, 2, ..., N-1, t is the current time, and f is the frequency.

[0235] In this embodiment of the application, according to expression (8), the main beam signal is used as the reference signal, and the interference beam signal is used as the signal to be filtered. Through the above-mentioned first adaptive filtering, the audio signal corresponding to the target sound source can be filtered out from N-1 interference beam signals. It can be understood that the filtered interference signal does not contain the audio signal corresponding to the target sound source, or contains very little of the signal corresponding to the target sound source.

[0236] By solving expression (8), the filter coefficients for adaptive filtering of N-1 interfering beam signals can be obtained. Satisfies expression (9):

[0237]

[0238] Where w(t-1,f) represents the filter coefficient of the previous time step, μ(t,f) represents the control coefficient update flag, and g1 represents the filter gain of the first adaptive filter.

[0239] In this embodiment, the control coefficient update identifier μ(t, f) can be obtained based on voice activity detection (VAD) technology according to a certain control strategy. In one implementation, the control coefficient update identifier μ(t, f) satisfies expression (10):

[0240]

[0241] in,

[0242] Where α is the energy threshold of the beam signal, θ is the azimuth angle of the current sound source, and θ l The azimuth angle of the left boundary of the main beam, θ r The azimuth angle of the right boundary of the main beam.

[0243] Optionally, the value of α can be set according to the actual situation. The principle for selecting the value of α is that when there is a sound source within the main beam range and it is the main sound source, the ratio between the maximum value of the energy of the main beam signal and the maximum value of the energy of the interference beam signal is greater than α.

[0244] It should be understood that the current sound source is the sound source currently detected during the adaptive filtering process. The azimuth angle of the current sound source is the angle between the direction vector of the current sound source and the aforementioned reference direction (e.g., the horizontal direction). The azimuth angle of the current sound source can be obtained by determining the position of the current sound source. In the embodiments of this application, it can be seen from expressions (9) and (10) that the azimuth angle of the current sound source is used to guide the filter update, that is, the azimuth angle of the current sound source is used to update the filter coefficients of the adaptive filter.

[0245] Determining the azimuth angle of the current sound source is equivalent to determining the location of the current sound source. Similar to determining the location of the target sound source in the above embodiment, the specific method for determining the location of the current sound source includes: acquiring current image information of the target sound pickup range, which includes at least one person; and performing lip movement detection on at least one person based on the acquired image information to determine the lip movement state of at least one person; the lip movement state is either present or absent; and determining the location of the person whose lip movement state is present as the location of the current sound source.

[0246] Alternatively, the specific method for determining the current location of the sound source may include: performing sound source localization processing on the audio signal to be processed to determine the direction of arrival of the sound source in the audio signal to be processed; and determining the location of the sound source located in the direction of arrival as the current location of the sound source.

[0247] Alternatively, the specific method for determining the current location of the sound source may include: determining the location of the beam corresponding to the first beam signal among the N beam signals as the current location of the sound source, wherein the first beam signal is the beam signal with the highest energy among the N beam signals.

[0248] For a detailed description of the method for determining the location of the current sound source, please refer to the description of the process for determining the location of the target sound source in the above embodiments, which will not be repeated here.

[0249] Optionally, the above adaptive filtering can be Kalman filtering or other adaptive filtering methods, which are not limited in the embodiments of this application.

[0250] Step 3032: The electronic device uses the filtered N-1 interference beam signals as reference signals to adaptively filter the main beam signal to obtain the target sound source enhancement signal. The audio signal corresponding to the interference sound source in the target sound source enhancement signal is filtered out.

[0251] Step 3032 corresponds to the second adaptive filtering process. The target sound source enhancement signal is the filtered main beam signal. Specifically, the target sound source enhancement signal satisfies expression (11):

[0252]

[0253] in, Let yt represent the target sound source enhancement signal, and y0(t,f) represent the main beam signal before filtering. Let represent the i-th filtering coefficient of the main beam signal adaptively filtered using the i-th filtered interference beam signal as the reference signal. It can be seen that the coefficient of the audio signal corresponding to the interference source in the above expression (6) that reaches the main beam is the i-th filtering coefficient of the main beam signal adaptively filtered. Let f represent the i-th interference beam signal after the first adaptive filtering; where i takes the values ​​1, 2, ..., N-1, t is the current time, and f is the frequency.

[0254] In this embodiment of the application, according to expression (11), the interference beam signal is used as the reference signal, and the main beam signal is used as the signal to be filtered. Through the above-mentioned second adaptive filtering, N-1 audio signals corresponding to interference beams can be filtered out from the main beam signal. It can be understood that there are no audio signals corresponding to interference beams in the filtered main beam signal, or there are very few audio signals corresponding to interference beams.

[0255] By solving expression (11), the filter coefficients for adaptive filtering of the main beam signal can be obtained. Satisfies expression (12):

[0256]

[0257] in, μ represents the filter coefficient at the previous time step. i (t, f) represents the control coefficient update flag when the i-th interfering beam is used as the reference signal, g 2i This represents the filter gain of the adaptive filter.

[0258] In this embodiment of the application, the aforementioned control coefficient update identifier μ i The values ​​of (t, f) satisfy the following expression (13):

[0259]

[0260] Where θ is the azimuth angle of the current sound source, θ il Let θ be the azimuth angle of the left boundary of the i-th interfering beam. ir Let be the azimuth angle of the right boundary of the i-th interfering beam.

[0261] In summary, the audio signal processing method provided in this application can be applied to electronic devices with microphones. Within the target pickup range, after the electronic device acquires an audio signal (referred to as the audio signal to be processed) including at least one sound source through the microphone, it performs beamforming processing (i.e., beamforming processing) on ​​the audio signal to be processed using N beams to obtain N beam signals. The coverage area of ​​these N beams is within the target pickup range, where N is an integer greater than or equal to 2. Then, the electronic device determines the main beam signal and N-1 interfering beam signals among the N beam signals. The main beam signal is the beam signal where the target sound source is located. Furthermore, the electronic device performs adaptive filtering on the main beam signal and the N-1 interfering beam signals to obtain the target sound source enhancement signal. In this application embodiment, after beamforming the audio signal to be processed, adaptive filtering is performed on the main beam signal and the N-1 interfering beam signals among the N beam signals. This can more thoroughly filter out the interfering signals in the audio signal corresponding to the target sound source, that is, it can effectively suppress the interfering signals in the audio signal and improve the quality of the audio signal.

[0262] Furthermore, the audio signal processing method provided in this application embodiment can be applied to interactive scenarios such as multi-person conversations (e.g., video conversations or voice conversations). Since the audio signal processing method provided in this application embodiment can effectively suppress interference signals in the audio signal and improve the quality of the audio signal, it can also improve the voice interaction experience.

[0263] Accordingly, this application provides an electronic device that can be divided into functional modules according to the above method examples. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0264] When dividing each function into modules according to its corresponding function. Figure 14 A schematic diagram of a possible structure of the electronic device involved in the above embodiments is shown. For example... Figure 14 As shown, the electronic device includes a beam processing module 1401, a determination module 1402, and a filtering module 1403.

[0265] The beam processing module 1401 is used to perform beam processing on the audio signal to be processed using N beams respectively, to obtain N beam signals, where N is an integer greater than or equal to 2; the audio signal to be processed is an audio signal including at least one sound source collected by a microphone within the target pickup range, and the coverage area of ​​the N beams is within the target pickup range, for example, by performing step 301 in the above method embodiment.

[0266] The determination module 1402 is used to determine the main beam signal and N-1 interfering beam signals among N beam signals. The main beam signal is the beam signal where the target sound source is located. For example, step 302 in the above method embodiment is executed.

[0267] The filtering module 1403 is used to adaptively filter the main beam signal and N-1 interfering beam signals to obtain the target sound source enhancement signal, for example, by performing step 303 in the above method embodiment.

[0268] Optionally, the filtering module 1403 is specifically used to use the main beam signal as a reference signal to adaptively filter N-1 interference beam signals to obtain filtered N-1 interference beam signals, in which the audio signal corresponding to the target sound source is filtered out; and to use the filtered N-1 interference beam signals as a reference signal to adaptively filter the main beam signal to obtain a target sound source enhancement signal, in which the audio signal corresponding to the interference sound source is filtered out, for example, by performing steps 3031 to 3032 in the above method embodiment.

[0269] Optionally, the electronic device provided in this application embodiment further includes a sound source localization module 1404, which is used to determine the location of the target sound source. The location of the target sound source is used to determine the main beam signal, for example, by performing step 304 in the above method embodiment to determine the location of the target sound source.

[0270] Optionally, the electronic device provided in this application embodiment further includes an image acquisition module 1405 and a detection module 1406. The image acquisition module 1405 is used to acquire image information of the target sound pickup range, for example, by executing step 3041 in the above method embodiment. The detection module 1406 is used to perform lip movement detection on at least one person based on the image information acquired by the image acquisition module 1405, and determine the lip movement state of at least one person, for example, by executing step 3042 in the above method embodiment. The sound source localization module 1404 is specifically used to determine the position of the person whose lip movement state is present among the at least one person as the position of the target sound source, for example, by executing step 3043 in the above method embodiment.

[0271] Optionally, the sound source localization module 1404 is specifically used to perform sound source localization processing on the audio signal to be processed, determine the direction of arrival of the sound source in the audio signal to be processed; and the position of the sound source in the direction of arrival is determined as the position of the target sound source, for example, by executing steps 3044 to 3045 in the above method embodiment.

[0272] Optionally, the sound source localization module 1404 is specifically used to determine the position of the beam corresponding to the first beam signal among the N beam signals as the position of the target sound source. For example, by executing step 3046 in the above method embodiment, the first beam signal is the beam signal with the largest energy among the N beam signals.

[0273] The various modules of the above-mentioned electronic device can also be used to perform other actions in the above-mentioned method embodiments. All relevant content of each step involved in the above-mentioned method embodiments can be referred to in the functional description of the corresponding functional module, and will not be repeated here.

[0274] When using integrated units, Figure 15 A schematic diagram of another possible structure of the electronic device involved in the above embodiments is shown. For example... Figure 15 As shown, the electronic device provided in this application embodiment may include a processing module 1501 and a communication module 1502. The processing module 1501 can be used to control and manage the operation of the electronic device. For example, the processing module 1501 can be used to support the electronic device in executing steps 301 to 303, step 304 in the above method embodiments, and / or other processes used in the technology described herein. The communication module 1502 can be used to support communication between the electronic device and other network entities. Optionally, as... Figure 15 As shown, the electronic device may also include a storage module 1503 for storing the device's program code and data.

[0275] The processing module 1501 may be a processor or a controller (e.g., as described above). Figure 2The processor 210 shown may be, for example, a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the embodiments of this invention. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The communication module 1502 may be a transceiver, transceiver circuitry, or a communication interface, etc. (e.g., as described above). Figure 2 The mobile communication module 250 or wireless communication module 260 shown. Storage module 1503 can be a memory (e.g., the one described above). Figure 2 The internal memory 221 shown.

[0276] When the processing module 16501 is a processor, the communication module 1502 is a transceiver, and the storage module 1503 is a memory, the processor, transceiver, and memory can be connected via a bus. The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc.

[0277] For more details on how the modules included in the above-mentioned electronic device achieve the above functions, please refer to the descriptions in the preceding method embodiments, which will not be repeated here.

[0278] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0279] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center integrating one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).

[0280] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0281] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0282] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0283] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0284] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.

[0285] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method of processing an audio signal, characterized by, The method comprises the following steps: performing beam processing on the audio signal to be processed by using N beams respectively to obtain N beam signals, N being an integer greater than or equal to 2; the audio signal to be processed is an audio signal including at least one sound source collected by a microphone in a target sound pickup range; the coverage of the N beams is in the target sound pickup range; determining a main beam signal and N-1 interference beam signals in the N beam signals; the main beam signal is a beam signal in which a target sound source is located; performing adaptive filtering on the N-1 interference beam signals by taking the main beam signal as a reference signal to obtain N-1 filtered interference beam signals; the audio signal corresponding to the target sound source in the N-1 filtered interference beam signals is filtered out; performing adaptive filtering on the main beam signal by taking the N-1 filtered interference beam signals as a reference signal to obtain a target sound source enhancement signal; the audio signal corresponding to an interference sound source in the target sound source enhancement signal is filtered out.

2. The method according to claim 1, wherein the N-1 filtered interference beam signals satisfy: or wherein, represents the filtered i-th interfering beam signal, y i (t,f) represents the i-th interfering beam signal before filtering, represents the filtering coefficient for adaptively filtering the N-1 interfering beam signals, y0(t,f) represents the main beam signal; wherein i is 1, 2, …, N-1, t is the current time, and f is the frequency. the target sound source enhancement signal satisfies:

3. The method according to claim 2, wherein w(t-1,f) represents a filtering coefficient of the previous moment, μ(t,f) represents a control coefficient update identifier, and g1 represents a filtering gain of the adaptive filtering; or wherein, denotes the target sound source enhancement signal, y0(t,f) denotes the main beam signal before filtering, denotes the i-th filter coefficient for adaptively filtering the main beam signal with the i-th filtered interference beam signal as the reference signal.

4. The method according to claim 3, wherein the control coefficient update identifier μ(t,f) satisfies: The satisfies: or one beam corresponds to a set of beam coefficients; a beam coefficient of an i-th beam in the N beams satisfies a following objective function: The satisfies: wherein, denotes the filter coefficient of the previous time, μ i (t, f) denotes a control coefficient update flag, g 2i denotes the filter gain of the adaptive filter. a constraint condition of the objective function is:

6. The method according to any one of claims 1 to 4, wherein the main beam signal is a beam signal corresponding to a beam covering a position of the target sound source, and the interference beam signal is a beam signal corresponding to a beam not covering the position of the target sound source. wherein a is an energy threshold of the beam signal, θ is an azimuth angle of the current sound source, θ l is an azimuth angle of the left boundary of the main beam, and θ r is an azimuth angle of the right boundary of the main beam. The method further comprises: The control coefficient update identifier μ i (t, f) satisfies: where θ is the azimuth angle of the current sound source, θ il is the azimuth angle of the left border of the i-th interfering beam, θ ir is the azimuth angle of the right border of the i-th interfering beam.

5. The method according to any one of claims 1 to 4, characterized in that, determining a position of the target sound source, the position of the target sound source being used to determine the main beam signal. min|||η i (θ d ,f)Q(θ,f) The determination of the position of the target sound source comprises: ||h i (i d ,f)Q(θ,f)||-1|<θ∈[θ i_l ,i i_r ],and||the i (i d ,f)|| 2 <c; wherein η i (θ d ,f) is a set of beam coefficients corresponding to the i-th beam, Q(θ,f) is an array steering vector of the microphone array, θ d is a center angle of the i-th beam, θ is an azimuth angle of the sound source, θ i_l is a left boundary azimuth angle of the i-th beam, θ i_r is a right boundary azimuth angle of the i-th beam, ε is a beam tolerance factor, γ is a beam robustness factor, and f is a frequency. collecting image information of the target sound pickup range, the target sound pickup range including at least one person; performing lip movement detection on the at least one person according to the image information to determine a lip movement state of the at least one person; the lip movement state is existence of lip movement or non-existence of lip movement; 7. The method according to any one of claims 1 to 4, characterized in that, determining a position of a person with existence of lip movement in the at least one person as the position of the target sound source; or 8. The method of claim 7, wherein, performing sound source positioning processing on the audio signal to be processed to determine a direction of arrival of a sound source in the audio signal to be processed; determining a position of the sound source located in the direction of arrival as the position of the target sound source; or determining a position of a beam corresponding to a first beam signal in the N beam signals as the position of the target sound source, the first beam signal being a beam signal with the largest energy in the N beam signals. The method comprises the following steps: a beam processing module, a determination module, and a filtering module. ​ ​ ​ 9. An electronic device, comprising: ​ ​ The beam processing module is configured to perform beam processing on the audio signal to be processed by using N beams respectively to obtain N beam signals, N is an integer greater than or equal to 2; the audio signal to be processed is an audio signal including at least one sound source collected by a microphone in a target sound pickup range, and coverage ranges of the N beams are in the target sound pickup range; The determination module is configured to determine a main beam signal and N-1 interference beam signals from the N beam signals, the main beam signal being a beam signal in which a target sound source is located; The filtering module is configured to perform adaptive filtering on the N-1 interference beam signals by taking the main beam signal as a reference signal to obtain filtered N-1 interference beam signals in which audio signals corresponding to the target sound source are filtered out, and perform adaptive filtering on the main beam signal by taking the filtered N-1 interference beam signals as a reference signal to obtain a target sound source enhancement signal in which audio signals corresponding to interfering sound sources are filtered out.

10. The electronic device of claim 9, wherein the filtered N-1 interference beam signals satisfy: or the target sound source enhancement signal satisfies:.

11. The electronic device of claim 10, wherein w (t-1, f) represents a filtering coefficient at a previous time, μ (t, f) represents a control coefficient update identifier, and g1 represents a filtering gain of the adaptive filtering; or 12. The electronic device of claim 11, wherein the control coefficient update identifier μ (t, f) satisfies: or one beam corresponds to a set of beam coefficients, and a beam coefficient of an i-th beam in the N beams satisfies a following objective function: a constraint condition of the objective function is: wherein, represents the filtered i-th interfering beam signal, y i (t,f) represents the i-th interfering beam signal before filtering, represents the filtering coefficient for adaptively filtering the N-1 interfering beam signals, y0(t,f) represents the main beam signal; wherein i is 1, 2, …, N-1, t is the current time, and f is the frequency.

14. The electronic device of any one of claims 9 to 12, wherein the main beam signal is a beam signal corresponding to a beam covering a position of the target sound source, and the interference beam signal is a beam signal corresponding to a beam not covering the position of the target sound source. The electronic device further includes a sound source positioning module. wherein, denotes the target sound source enhancement signal, y0(t,f) denotes the main beam signal before filtering, denotes the i-th filtering coefficient for adaptively filtering the main beam signal with the i-th filtered interference beam signal as the reference signal. The sound source positioning module is further configured to determine a position of the target sound source, and the position of the target sound source is used to determine the main beam signal. The satisfies: The electronic device further includes an image acquisition module and a detection module. The image acquisition module is configured to acquire image information of the target sound pickup range, and the target sound pickup range includes at least one person. The satisfies: wherein, denotes the filter coefficient of the previous time, μ i (t, f) denotes a control coefficient update flag, g 2i denotes the filter gain of the adaptive filter. The detection module is configured to perform lip movement detection on the at least one person according to the image information to determine a lip movement state of the at least one person, and the lip movement state is presence of lip movement or absence of lip movement. The sound source positioning module is specifically configured to determine a position of a person with presence of lip movement from the at least one person as the position of the target sound source; or wherein a is an energy threshold of the beam signal, θ is an azimuth angle of the current sound source, θ l is an azimuth angle of the left boundary of the main beam, and θ r is an azimuth angle of the right boundary of the main beam. ​ The control coefficient update identifier μ i (t, f) satisfies: where θ is the azimuth angle of the current sound source, θ il is the azimuth angle of the left boundary of the i-th interfering beam, θ ir is the azimuth angle of the right boundary of the i-th interfering beam.

13. The electronic device of any of claims 9-12, wherein, ​ min ||η i (θ d ,f)Q(θ,f)| ​ |||η i (θ d , f) Q(θ, f) || -1 | < ε θ ∈ [θ i_l , θ i_r ] and ||η i (θ d , f) || 2 < γ wherein η i (θ d ,f) is a set of beam coefficients corresponding to the i-th beam, Q(θ,f) is an array steering vector of the microphone array, θ d is a center angle of the i-th beam, θ is an azimuth angle of the sound source, θ i_l is a left boundary azimuth angle of the i-th beam, θ i_r is a right boundary azimuth angle of the i-th beam, ε is a beam tolerance factor, γ is a beam robustness factor, and f is a frequency. ​ ​ 15. The electronic device of any of claims 9-12, wherein, ​ ​ 16. The electronic device of claim 15, wherein, ​ ​ ​ ​ ​ The sound source positioning module is specifically configured to perform sound source positioning processing on the audio signal to be processed, determine a direction of arrival of a sound source in the audio signal to be processed, and determine a position of the sound source located in the direction of arrival as the position of the target sound source. Or The sound source positioning module is specifically configured to determine a position of a beam corresponding to a first beam signal in the N beam signals as the position of the target sound source, the first beam signal being a beam signal with the maximum energy in the N beam signals.

17. An electronic device, comprising: The computer program product comprises instructions for performing the method of any one of claims 1 to 8 when the computer program product is run on a computer.

18. A computer program product, characterised in that, The computer program product comprises instructions for performing the method of any one of claims 1 to 8 when the computer program product is run on a computer.

19. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program product comprises instructions for performing the method of any one of claims 1 to 8 when the computer program product is run on a computer.

20. A chip, characterized by The computer program product comprises instructions for performing the method of any one of claims 1 to 8 when the computer program product is run on a computer.

Citation Information

Patent Citations

  • Speech recognition method and system based on linear microphone array

    CN106710603A

  • Filtering method and device based on fixed beam forming

    CN109102822A