Audio signal processing based on microphone arrangement

By using multi-microphone arrangement and phase information analysis technology in the video endpoints, the arrival angle of the target sound source and the noise sound source is determined, and the gain of the audio signal is adjusted, the problem that video endpoints in the prior art are difficult to provide high-quality hands-free voice pickup in a compact integrated environment, achieving better noise suppression and echo cancellation effects.

CN114642003BActive Publication Date: 2025-06-10CISCO TECHNOLOGY INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080075286.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-17
Filing Date
2020-10-23
Publication Date
2025-06-10
Estimated Expiration
2040-10-23

AI Technical Summary

Technical Problem

Existing video endpoints have difficulty providing high-quality hands-free voice pickup in compact integrated environments, especially in poor performance in echo cancellation and noise suppression.

Method used

Using a multi-microphone arrangement including a vertical microphone array and a horizontal microphone array, the audio signals obtained by the vertical microphone array are processed by analyzing the phase information obtained by the horizontal microphone array, thereby determining the arrival angle of the target sound source and the horizontal displacement sound source, and adjusting the gain of the audio signal to optimize the audio quality.

Benefits of technology

It effectively reduces noise interference, improves the signal-to-noise ratio of the target sound source audio, and improves the hands-free voice pickup quality and echo cancellation performance of the video endpoint.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114642003B_ABST
    Figure CN114642003B_ABST
Patent Text Reader

Abstract

In one example, a video endpoint obtains a first audio signal from a vertical microphone array, the first audio signal including audio from a target sound source and audio from a horizontally displaced sound source. The video endpoint obtains a second audio signal and a third audio signal from a horizontal microphone array, both the second audio signal and the third audio signal including audio from the target sound source and audio from the horizontally displaced sound source. Based on the second audio signal and the third audio signal, the video endpoint determines at least one of a first arrival angle of the audio from the target sound source or a second arrival angle of the audio from the horizontally displaced sound source. Based on at least one of the first arrival angle or the second arrival angle, the video endpoint adjusts the gain of the first audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to U.S. Provisional Application No. 62 / 929,143, filed on November 1, 2019, the entire content of which is incorporated herein by reference. Technical Field

[0003] This disclosure relates to video endpoints. Background Art

[0004] A video endpoint is an electronic device that allows a user to often conduct a teleconference with one or more remote users via one or more teleconference servers and additional video endpoints. A video endpoint can include various functions to help facilitate a session or teleconference, such as one or more cameras, speakers, microphones, displays, etc. Video endpoints are often used in professional (e.g., enterprise) settings. Brief Description of the Drawings

[0005] Figure 1 A front view of a video endpoint according to an example embodiment is shown, the video endpoint being configured to process an audio signal based on a microphone arrangement of the video endpoint.

[0006] Figures 2A - 2D A corresponding use case scenario according to an example embodiment is shown, in which the video endpoint is configured to obtain audio from a vertical microphone array of the video endpoint in the presence of various vertically displayed sound sources.

[0007] Figures 3A - 3D A corresponding use case scenario according to an example embodiment is shown, in which the video endpoint is configured to process an audio signal based on a vertical microphone array and a horizontal microphone array of the video endpoint.

[0008] Figure 4 Is a block diagram showing an audio signal processing flow based on a microphone arrangement of a video endpoint according to an example embodiment.

[0009] Figure 5 A graph showing a mathematical directivity function that can be applied to one or more audio signals according to an example embodiment is shown.

[0010] Figure 6 A block diagram of a computing device according to an exemplary embodiment is shown, the computing device being configured to process an audio signal based on a microphone arrangement of the computing device.

[0011] Figure 7 A flowchart of a method for processing an audio signal based on a microphone arrangement of a video endpoint according to an exemplary embodiment is shown. Detailed Description

[0012] Overview

[0013] Aspects of the present invention are set forth in the independent claims, and the preferred features are set forth in the dependent claims. The features of one aspect can be applied to any aspect, either alone or in combination with the features of other aspects.

[0014] In one exemplary embodiment, a video endpoint is provided that includes a vertical microphone array and a horizontal microphone array. The video endpoint obtains a first audio signal from the vertical microphone array, the first audio signal including audio from a target sound source and audio from a horizontally displaced sound source. The video endpoint obtains a second audio signal and a third audio signal from the horizontal microphone array, both the second audio signal and the third audio signal including audio from the target sound source and audio from the horizontally displaced sound source. Based on the second audio signal and the third audio signal, the video endpoint determines at least one of a first angle of arrival of the audio from the target sound source or a second angle of arrival of the audio from the horizontally displaced sound source. Based on at least one of the first angle of arrival or the second angle of arrival, the video endpoint adjusts the gain of the first audio signal.

[0015] Exemplary Embodiment

[0016] Figure 1 An exemplary video endpoint 100 configured to process audio signals based on a particular microphone arrangement is shown. The video endpoint 100 includes a housing 110, a camera 120, a display panel or screen 130, a speaker 140, a vertical microphone array 150, and a horizontal microphone array 160. The housing 110 supports / protects / encloses one or more of the following: the camera 120, the screen 130, the speaker 140, the vertical microphone array 150, and the horizontal microphone array 160. The camera 120 is configured to capture video (e.g., video of a user of the video endpoint 100). The screen 130 is configured to present an image (e.g., an image of a second user of a remote video endpoint). The speaker 140 is configured to output audio (e.g., audio generated by the second user) and may include one or more speaker sub-components.

[0017] The vertical microphone array 150 can be positioned along the vertical side (e.g., the baffle) of the display screen 130. The vertical microphone array 150 includes, for example, microphones 150(1)-150(6). The microphones 150(1)-150(6) can be non-uniformly spaced microelectromechanical system (MEMS) microphones that are configured for fixed (non-adaptive) differential filter-and-sum beamforming. In one non-limiting example, microphone 150(1) is vertically displaced 217 mm in height from the bottom of the video endpoint 100; microphone 150(2) is vertically displaced 331 mm in height from the bottom of the video endpoint 100; microphone 150(3) is vertically displaced 369 mm in height from the bottom of the video endpoint 100; microphone 150(4) is vertically displaced 388 mm in height from the bottom of the video endpoint 100; microphone 150(5) is vertically displaced 407 mm in height from the bottom of the video endpoint 100; and microphone 150(6) is vertically displaced 445 mm in height from the bottom of the video endpoint 100. Generally, the vertical microphone array 150 can have any suitable configuration (e.g., any suitable number of microphones vertically displaced at any suitable corresponding height from the bottom of the video endpoint 100). Additionally, the microphones 150(1)-150(6) can be of any suitable type / size.

[0018] The horizontal microphone array 160 can be positioned along the top portion of the display screen 130 proximate to the camera 120. The horizontal microphone array 160 includes microphones 160(1) and 160(2). The microphones 160(1) and 160(2) can be configured for non-linear noise suppression. In one non-limiting example, microphones 160(1) and 160(2) are spaced 19 mm apart horizontally, but generally the horizontal microphone array 160 can have any suitable configuration (e.g., any suitable number of microphones spaced at any suitable corresponding horizontal distance). Additionally, the microphones 160(1) and 160(2) can be of any suitable type / size.

[0019] The video endpoint 100 can be configured to communicate wirelessly with one or more other video endpoints or via a wired technology such as Ethernet, allowing a user of the video endpoint 100 to communicate with one or more remote users of other video endpoints. The video endpoint 100 can be supported (e.g., placed / situated, adhered / fixed, etc.) by any suitable surface (e.g., a desktop). The video endpoint 100 can be implemented in any suitable environment (e.g., in an enterprise environment such as an open office environment). In one example, the video endpoint 100 can have a width 170 of 630 mm and a height 180 of 510 mm, but generally the video endpoint 100 can have any suitable dimensions.

[0020] The angle 190 can be the angle from a reference horizontal line to the line between the acoustic center of the vertical microphone array 150 at high frequencies and the speaker 140. In this example, the acoustic center of the vertical microphone array 150 at high frequencies is the microphone 150(4), and the angle 190 is approximately 47°, but generally the acoustic center of the vertical microphone array 150 at high frequencies (e.g., 1500 Hz to 20 kHz) can be configured / located at any suitable position, and the angle 190 can be any suitable angle.

[0021] Conventional video endpoint designs often fail to provide high-quality hands-free voice pickup in desktop and compact integrated video endpoints in crowded rooms. For example, in these compact integrated video endpoints, the distance between the speaker and the microphone is relatively short, which may reduce the performance of full-duplex communication with echo cancellation (AEC). In addition, destructive interference from desktop reflections may reduce the quality, and a user's laptop may shield the microphone from receiving the user's voice. The integrated video conferencing endpoint may also pick up noise (e.g., interference) from the keyboard and the desktop, or may also pick up noise from nearby colleagues if the endpoint is used in an open office environment.

[0022] Accordingly, the video endpoint 100 includes audio signal processing logic 195 that causes the video endpoint 100 to perform operations for improving hands-free microphone pickup. To this end, the video endpoint 100 can process the audio signal obtained from the vertical microphone array 150 based on the phase information obtained from the horizontal microphone array 160. This can enable the video endpoint 100 to minimize noise (e.g., interference) from noise sources while maximizing the audio from the target source (e.g., the user).

[0023] In short, in one example, the video endpoint 100 may obtain a first audio signal from the vertical microphone array 150, which includes audio from a target sound source and audio from a horizontally displaced sound source. The video endpoint 100 may also obtain a second audio signal and a third audio signal from the horizontal microphone array 160, both of which include audio from the target sound source and audio from the horizontally displaced sound source. Based on the second audio signal and the third audio signal, the video endpoint 100 may determine at least one of a first arrival angle (e.g., a first horizontal arrival angle) of the audio from the target sound source or a second arrival angle (e.g., a second horizontal arrival angle) of the audio from the horizontally displaced sound source. Based on at least one of the first arrival angle or the second arrival angle, the video endpoint 100 may adjust the gain of the first audio signal.

[0024] Now refer to Figures 2A - 2D , and continue to refer to Figure 1 . Figures 2A - 2D Use case scenarios 200A - 200D are shown, where the video endpoint 100 obtains an audio signal via the vertical microphone array 150 in the presence of various vertically displayed sound sources (e.g., vertically displayed noise sources). In each use case scenario 200A - 200D, there is a video endpoint 100 and a surface 210 (e.g., a desktop) configured to support the video endpoint 100. In these examples, the vertical microphone array 150 is configured to generate a vertically symmetric directivity pattern 220. The vertical microphone array 150 may use any suitable technique (e.g., beamforming technique) to generate the vertically symmetric directivity pattern 220. In this example, the vertically symmetric directivity pattern 220 is a torus symmetric about the vertical axis of the vertical microphone array 150, but any suitable vertically symmetric directivity pattern can be of any suitable shape (e.g., heart-shaped), and the vertically symmetric directivity pattern 220 may depend on the hardware geometry (e.g., transducer position).

[0025] The vertically symmetric directivity pattern 220 may have a maximum value (e.g., minimum suppression) along the predicted vertical displacement of the target sound source. In use case scenarios 200A - 200C, the target sound source is the user 230, and the audio 240 is generated by the user 230. The audio 240 may be the voice directed to a remote user in the video facilitated by the video endpoint 100. The vertically symmetric directivity pattern 220 may also have a zero value (e.g., maximum suppression) along the predicted vertical displacement of the vertically displaced sound source.

[0026] Figures 2A - 2D Different examples of vertically displaced sound sources are shown. In Figure 2AIn this case, the vertically displaced sound source is the speaker 140. The speaker 140 generates the audio 250, which can be the voice directed from a remote user to the user 230 in a session participated by the video endpoint 100. The audio 250 can reach the vertical microphone array 150 at an angle 190 (or a similar angle). Therefore, the vertically symmetric directivity pattern 220 can have a zero value at the angle 190 (or a similar angle). Since the zero value points to the speaker 140, the vertically symmetric directivity pattern 220 can improve the echo cancellation (AEC) performance by suppressing the audio 250, thereby improving the Echo-to-Near-end Ratio (ENR) of the video endpoint 100. In this example, the "near end" refers to the audio 240, and the "echo" refers to the audio 250. For example, a filter can be designed to maximize the suppression of the audio 250.

[0027] In Figure 2B this case, the vertically displaced sound source is the combination of the user 230 and the surface 210. That is, the surface 210 generates the reflected audio 260 from the user 230 towards the vertical microphone array 150. The reflected audio 260 can be the voice from the user 230, which is intended for a remote user in a session participated by the video endpoint 100. However, due to the longer sound path of the reflected audio 260 relative to the audio 240, the reflected audio 260 reaches the vertical microphone array 150 at a time later than the audio 240. Specifically, the reflected audio 260 can destructively interfere with the audio 240, thereby producing a comb filtering effect. The reflected audio 260 can reach the vertical microphone array 150 at an angle 190 (or a similar angle). Therefore, the vertically symmetric directivity pattern 220 can have a zero value at the angle 190 (or a similar angle). Since the zero value points to the point where the reflected audio 260 is reflected on the surface 210, the vertically symmetric directivity pattern 220 can attenuate the reflected audio 260 to improve the frequency response and audio quality.

[0028] In Figure 2C this case, the vertically displaced sound source is the user device 270 (e.g., a laptop). The user device 270 generates the audio 280, which may be key clicks and other similar noises caused by the interaction between the user 230 and the user device 270. The audio 280 can reach the vertical microphone array 150 at an angle 190 (or a similar angle). Therefore, the vertically symmetric directivity pattern 220 can have a zero value at the angle 190 (or a similar angle). Since the zero value points to the user device 270, the vertically symmetric directivity pattern 220 can attenuate the audio 280. In addition, the part of the vertical microphone array 150 that performs mid-high frequency pickup can be sufficiently enhanced to avoid being blocked by the user device 270.

[0029] In Figure 2DIn this case, the vertical displacement sound source is a heating, ventilation, and air conditioning (HVAC) unit / vent 290 located on the ceiling above the video endpoint 100. The HVAC unit / vent 290 generates audio 295, which may be noise generated by the opening or closing of the HVAC unit / vent 290, air flow, etc. The audio 295 may reach the vertical microphone array 150 with a zero value in the vertical symmetric directivity pattern 220 pointing to the HVAC unit / vent 290. Therefore, the vertical symmetric directivity pattern 220 can attenuate the noise from the HVAC unit / vent 290.

[0030] Now refer to Figures 3A - 3D and continue to refer to Figure 1 and Figures 2A - 2D . Figures 3A - 3D Use case scenarios 300A - 300D are shown, where the video endpoint 100 processes audio based on the vertical microphone array 150 and the horizontal microphone array 160. Based on the audio obtained via the horizontal microphone array 160, the video endpoint 100 can determine at least one of the audio arrival angle from the target sound source or the audio arrival angle from the horizontally displaced sound source. In use case scenarios 300A - 300C, the target sound source includes the user 230. The user 230 generates audio 240 towards the vertical microphone array 150 and audio 310 towards the horizontal microphone array 160. The audio 240 and 310 can be the voice directed to the remote user in the session facilitated by the video endpoint 100.

[0031] First, turning to Figure 3A , the video endpoint 100 can obtain the audio 310 and determine the arrival angle 320 of the audio 310 based on the phase information of the audio 310. The arrival angle 320 can depend on the difference between the time when the first of the microphones 160(1) and 160(2) obtains the audio 310 and the time when the second of the microphones 160(1) and 160(2) obtains the audio 310. In one example, if the microphones 160(1) and 160(2) obtain the audio 310 simultaneously (i.e., the difference is zero), then the video endpoint 100 can conclude that the arrival angle 320 is 90°.

[0032] In the example of use case scenario 300A, the horizontally displaced sound source is person 330, who can be a colleague of user 230 in an open office environment. Person 330 can generate audio 340 towards the vertical microphone array 150 and audio 350 towards the horizontal microphone array 160. Audio 340 and 350 can be the noise generated by person 330 in the open office environment (e.g., during a conversation with another colleague). The video endpoint 100 can obtain audio 350 and determine the angle of arrival 360 of audio 350 based on the phase information of audio 350. The angle of arrival 360 can depend on the difference between the time when the first of microphones 160(1) and 160(2) obtains audio 350 and the time when the second of microphones 160(1) and 160(2) obtains audio 350. In one example, microphone 160(1) can obtain audio 350 after microphone 160(2), so the video endpoint 100 can determine that the angle of arrival 360 is less than the angle of arrival 320.

[0033] In one example, the video endpoint 100 can determine that the angle of arrival 320 is within the range 370 (e.g., the effective beam width of the horizontal microphone array 160), and the angle of arrival 360 is outside the range 370. The range 370 can indicate whether a given audio is the target sound or noise. Thus, audio 310 is within the range 370 because user 230 stands at a suitable position for using the video endpoint 100, and audio 350 is outside the range 370 because person 330 is too far from the video endpoint 100 to actually use the video endpoint 100. The range 370 can be pre-configured (e.g., fixed) and / or dynamically adjusted.

[0034] Based on the angle of arrival 320 and / or 360, the video endpoint 100 can adjust the gain of the audio signal obtained from the vertical microphone array 150. The audio signal can include audio 240 and 340. For example, the video endpoint 100 can increase the audio level of audio 240 because the angle of arrival 320 is within the range 370, and / or can attenuate the gain of audio 340 because the angle of arrival 360 is outside the range 370. The video endpoint 100 can adjust the gains of audio 240 and 340 simultaneously, but adjusts the gains of audio 240 and 340 for different frequency bins.

[0035] Video endpoint 100 may process (e.g., adjust its gain) audio 240 and / or 340 (instead of audio 310 and 350) because, to provide a high-quality audio output signal, the vertical microphone array 150 may provide more favorable characteristics than the horizontal microphone array 160. For example, due to the vertical symmetric directivity pattern 220, the vertical microphone array 150 may have better frequency response, less noise, improved ENR, etc. compared to the horizontal microphone array 160. Thus, the horizontal microphone array 160 achieves time-varying spatial interference suppression by instructing the video endpoint 100 to attenuate the horizontal displacement noise sources in one or more audio signals obtained from the vertical microphone array 150. In other words, the video endpoint 100 may adjust the audio signals obtained from the vertical microphone array 150 based on specific characteristics (e.g., angle of arrival) of the audio signals determined by analyzing similar audio signals obtained from the horizontal microphone array 160.

[0036] In one example, the video endpoint 100 may obtain a video signal from the camera 120. Based on the video signal, the video endpoint 100 may adjust the range 370. The video endpoint may also determine that the angle of arrival 360 is outside the range 370 and, in response, adjust the gain of the audio signals obtained from the vertical microphone array 150. For example, the video endpoint 100 may increase the audio level of audio 240 and / or attenuate the gain of audio 340.

[0037] In one example, the range 370 may be pre-configured to match the field of view of the camera 120 (e.g., if the camera 120 has a fixed 70° field of view, the range 370 may also be fixed at 70°). In another example, the video endpoint 100 may adjust the range 370 by performing face recognition on the face of the user 230. For example, the video endpoint 100 may adjust the range 370 based on the position / size of the face of the user 230 within the field of view of the camera 120. In yet another example, the video endpoint 100 may adjust the range 370 based on the zoom level of the camera 120. For example, the zoom level of the camera 120 (e.g., changing the field of view of the camera 120) may be adjusted automatically or manually based on the position / size of the face of the user 230. Since the horizontal microphone array 160 is very close to the camera 120, the microphones 160(1) and 160(2) may be mapped to the coordinate system of the camera 120 to adapt the effective horizontal beam width to the zoom level / face detection.

[0038] Figure 3B and Figure 3C shows how, in addition to suppressing noise, the angle of arrival 320 may also be used to compensate for the asymmetric position and distance to the vertical microphone array 150. In Figure 3BIn this case, the video endpoint 100 can obtain the audio 310 and determine the angle of arrival 320 based on the phase information of the audio 310. The video endpoint 100 can also determine the horizontal displacement of the user 230 based on the angle of arrival 320. For example, if the angle of arrival 320 is 120°, the video endpoint 100 can determine that the user 230 has a horizontal displacement of 120°. The camera 120 can also assist the video endpoint 100 in determining the horizontal displacement of the user 230. The video endpoint 100 can also adjust the gain of the audio signal obtained from the vertical microphone array 150 based on the horizontal displacement of the user 230. For example, a 120° horizontal displacement of the user 230 can indicate that the user 230 is very close to the vertical microphone array 150, which can result in a relatively loud sound of the audio 240 at the remote video endpoint. Therefore, the video endpoint 100 can attenuate the gain of the audio 240 to handle the problem that the user 230 is very close to the vertical microphone array 150.

[0039] In Figure 3C this case, the video endpoint 100 can obtain the audio 310 and determine the angle of arrival 320 based on the phase information of the audio 310. The video endpoint 100 can also determine the horizontal displacement of the user 230 based on the angle of arrival 320. For example, if the angle of arrival 320 is 60°, the video endpoint 100 can determine that the user 230 has a horizontal displacement of 60°. The camera 120 can also assist the video endpoint 100 in determining the horizontal displacement of the user 230. The video endpoint 100 can also adjust the gain of the audio signal obtained from the vertical microphone array 150 based on the horizontal displacement of the user 230. For example, a horizontal displacement of 60° for the user 230 can indicate that the user 230 is far from the vertical microphone array 150, which may result in the audio 240 sounding relatively quiet at the far end. Therefore, the video endpoint 100 can increase the audio level of the audio 240 to handle the problem that the distance from the user 230 to the vertical microphone array 150 is relatively far.

[0040] In Figure 3DIn the example, both user 230 and person 330 are users participating in the session that video endpoint 100 is participating in, and audio 240, 310, 340, and 350 are directed to the users at the remote video endpoint during the session. In this example, video endpoint 100 can adjust the gain of the audio signal obtained from vertical microphone array 150 to equalize the audio levels of audio 240 and 340. For example, video endpoint 100 can determine the horizontal displacement of user 230 based on angle of arrival 320, and determine the horizontal displacement of person 330 based on angle of arrival 360. Camera 120 can also assist video endpoint 100 in determining the horizontal displacement of user 230. Based on angles of arrival 320 and 360, video endpoint 100 can determine that user 230 is close to vertical microphone array 150 and person 330 is relatively far from vertical microphone array 150. Thus, video endpoint 100 can attenuate the gain of audio 240 and / or increase the audio level of audio 340 to equalize audio 240 and 340. This may have the effect of making the audio levels of audio 240 and 340 substantially similar / equal at the far end.

[0041] According to the techniques described herein, linear beamforming can be performed in the vertical plane (e.g., via vertical microphone array 150), and non-linear suppression can be performed in the horizontal plane (e.g., via horizontal microphone array 160). This is because most of the unwanted sources in the vertical plane may be spatially reasonably time-invariant and / or stationary (e.g., speakers, HVAC noise, key clicks, tabletop reflections, etc.), while noise suppression in the horizontal plane may be both spatially and temporally different and may not be needed or desired at all in some cases. Although the angles of arrival described herein are angles relative to a reference horizontal line, any suitable (one or more) reference lines can be used to determine the angles of arrival. Additionally, the angles of arrival can be determined based on pre-configured settings that relate the phase difference to the angle of arrival and / or based on any suitable factors (e.g., room temperature, air pressure, frequency composition of the audio signal, etc.).

[0042] Figure 4 An example audio signal processing flow 400 based on the microphone arrangement of video endpoint 100 is shown. The operations of signal processing flow 400 can be performed by suitable hardware and / or software of video endpoint 100. For Figure 4 the description, reference is also made to Figure 1。At 405, the video endpoint 100 obtains an audio signal from the vertical microphone array 150 (e.g., six corresponding audio signals from microphones 150(1)-150(6)). The audio signal can include audio from a target sound source and audio from a horizontally displaced sound source. At 410, the audio signal passes through a beamformer and enters a filter bank at 415. At 420, the video endpoint 100 performs various operations on the audio signal (e.g., echo cancellation, automatic gain control (AGC), equalization, motion detection, diagnostic functions, etc.). Other operations can be performed in one channel for pipelining.

[0043] At 425, the video endpoint 100 obtains an audio signal from the horizontal microphone array 160 (e.g., one audio signal from microphone 160(1) and one audio signal from microphone 160(2)). These two audio signals can include audio from a target sound source and audio from a horizontally displaced sound source. At 430, the audio signal enters the filter bank. The filter bank can divide the audio signal into different frequency points. At 435, the angle-of-arrival finder determines at least one of a first angle of arrival of the audio from the target sound source or a second angle of arrival of the audio from the horizontally displaced sound source. The angle-of-arrival finder can calculate the phase difference for each time and frequency band block and convert the phase difference into an angle corresponding to the first angle of arrival and / or the second angle of arrival. The (one or more) angles of arrival for each frequency subband may be different. For example, the voices of colleagues may have different angles of arrival from the voice of the user, and these voices can be distinguished based on frequency.

[0044] At 440, the video endpoint 100 can adjust the gain of the audio signal obtained from the vertical microphone array 150. The video endpoint 100 can adjust the gain based on the first angle of arrival and / or the second angle of arrival. For example, for each time / frequency point, a mathematical directivity function can be applied based on the angle (e.g., using a smoothing vector, interband gain limiting, etc.). Thus, in one example, the video endpoint 100 determines the (one or more) angles of arrival in the horizontal plane and provides non-linear spatial interference suppression for the audio signal obtained from the vertical microphone array 150. At 445, the video endpoint 100 performs various operations on the audio signal (e.g., echo cancellation, automatic gain control (AGC), equalization, motion detection, diagnostic functions, etc.). Other operations can be performed in one channel for pipelining. At 450, the audio signal passes through an inverse filter bank, and at 455, the video endpoint 100 outputs the audio signal (e.g., towards a remote video endpoint).

[0045] Figure 5 An example graph 500 is shown, and graph 500 shows what can be applied to 440( Figure 4)'s mathematical directivity functions 540 and 550. Specifically, the y-axis of the curve graph 500 corresponds to the attenuation factor, and the x-axis of the curve graph 500 corresponds to the arrival angle of the audio. The curve graph 500 is divided into region 510, region 520, and region 530. Region 520 may represent the effective arrival angle range of the beam width of the horizontal microphone array. Therefore, any audio signal within region 520 can be regarded as a target audio signal, and any audio signal outside region 520 can be regarded as noise. In one example, the video endpoint 100 may apply the mathematical directivity functions 540 and / or 550 to any audio signal with any given arrival angle. In another example, the video endpoint 100 may apply the mathematical directivity functions 540 and / or 550 to any audio signal with an arrival angle corresponding to the incident angle within region 520, and may apply a zero attenuation factor to the audio signal with an arrival angle corresponding to the incident angle within region 510 and / or 530.

[0046] In one example, region 520 may be adjusted based on the camera. For example, if the camera zooms in, the width of region 520 may decrease along with the camera's field of view. In another example, the width of region 520 may be adjusted based on face recognition. For example, the width of region 520 may be increased or decreased to include the entire face of the user. Adjusting the width of region 520 may prompt the video endpoint 100 to modify the mathematical directivity functions 540 and / or 550, or switch from one of the mathematical directivity functions 540 and / or 550 to the other.

[0047] Figure 6 The hardware block diagram of an example device 600 (e.g., video endpoint 100) is shown. It should be realized that Figure 6 Only the description of one embodiment is provided, and it does not mean any limitation to the environment in which different embodiments can be implemented. Many modifications can be made to the depicted environment.

[0048] As shown in the figure, the device 600 includes a bus 612, and the bus 612 provides communication between (one or more) computer processors 614, a memory 616, a persistent storage device 618, a communication unit 620, and (one or more) input / output (I / O) interfaces 622. The bus 612 can be implemented with any architecture designed to transfer data and / or control information between processors (e.g., microprocessors, communication and network processors, etc.), system memories, peripheral devices, and any other hardware components within the system. For example, the bus 612 can be implemented with one or more buses.

[0049] The memory 616 and the persistent storage device 618 are computer-readable storage media. In the illustrated embodiment, the memory 616 includes a random access memory (RAM) 624 and a cache memory 626. Generally, the memory 616 can include any suitable volatile or non-volatile computer-readable storage media. Instructions for the audio signal processing logic 195 can be stored in the memory 616 or the persistent storage device 618 for execution by the computer processor(s) 614.

[0050] One or more programs can be stored in the persistent storage device 618 for execution by one or more corresponding computer processors 614 via one or more memories in the memory 616. The persistent storage device 618 can be a disk drive, a solid state drive, a semiconductor storage device, a read-only memory (ROM), an erasable programmable ROM (EPROM), a flash memory, or any other computer-readable storage media capable of storing program instructions or digital information.

[0051] The medium used by the persistent storage device 618 can also be removable. For example, a removable hard disk drive can be used for the persistent storage device 618. Other examples include optical discs and magnetic disks, thumb drives, and smart cards inserted into a drive for transfer to another computer-readable storage medium (which is also part of the persistent storage device 618).

[0052] In these examples, the communication unit 620 provides communication with other data processing systems or devices. In these examples, the communication unit 620 includes one or more network interface cards. The communication unit 620 can provide communication by using one or both of a physical communication link and a wireless communication link.

[0053] (One or more) I / O interfaces 622 allow for data input and output with other devices that can be connected to the device 600. For example, (one or more) I / O interfaces 622 can provide a connection to an external device 628, such as a keyboard, a camera, a keypad, a touch screen, and / or some other suitable input device. The external device 628 can also include a portable computer-readable storage medium, such as a database system, a thumb drive, a portable optical disc or magnetic disk, and a memory card.

[0054] Software and data for implementing the embodiments can be stored on such portable computer-readable storage media and can be loaded onto the persistent storage device 618 via (one or more) I / O interfaces 622. (One or more) I / O interfaces 622 can also be connected to a display 630. The display 630 provides a mechanism for displaying data to a user and can be, for example, a display screen of a video endpoint.

[0055] Figure 7It is a flowchart of an example method 700 for processing an audio signal based on a microphone arrangement of a video endpoint. At 710, the video endpoint obtains a first audio signal from a vertical microphone array, and the first audio signal includes audio from a target sound source and audio from a horizontally displaced sound source. At 720, the video endpoint obtains a second audio signal and a third audio signal from a horizontal microphone array, and both the second audio signal and the third audio signal include audio from the target sound source and audio from the horizontally displaced sound source. At 730, based on the second audio signal and the third audio signal, the video endpoint determines at least one of a first arrival angle of the audio from the target sound source or a second arrival angle of the audio from the horizontally displaced sound source. At 740, based on at least one of the first arrival angle or the second arrival angle, the video endpoint adjusts the gain of the first audio signal.

[0056] The programs described herein are identified based on the applications in which they are implemented in particular embodiments. However, it should be realized that any specific program nomenclature herein is for convenience only, and thus, embodiments should not be limited to use in any particular application identified and / or implied by such nomenclature.

[0057] Data related to the operations described herein can be stored in any conventional or other data structure (e.g., files, arrays, lists, stacks, queues, records, etc.), and can also be stored in any desired storage unit (e.g., databases, data or other repositories, queues, etc.). Data transmitted between entities can include any desired format and arrangement, and can include any number of fields of any type of any size to store data. The definition of any data set and the data model can indicate the overall structure in any desired manner (e.g., computer-related languages, graphical representations, lists, etc.).

[0058] This embodiment can use any number of any type of user interface (e.g., graphical user interface (GUI), command line, prompt, etc.) to obtain or provide information, where the interface can include any information arranged in any manner. The interface can include any number of any type of input or drive mechanisms (e.g., buttons, icons, fields, boxes, links, etc.), which are set at any position to input / display information and initiate desired actions via any suitable input device (e.g., mouse, keyboard, etc.). The interface screen can include any suitable actuators (e.g., links, labels, etc.) to navigate between screens in any manner.

[0059] The environment of this embodiment may include any number of computers or other processing systems (e.g., client or end-user systems, server systems, etc.) and databases or other repositories arranged in any desired manner, where this embodiment can be applied to any desired type of computing environment (e.g., cloud computing, client-server, network computing, mainframe, stand-alone systems, etc.). The computers or other processing systems used in this embodiment can be implemented by any number of any individual or other type of computer or processing system (e.g., desktop computer, laptop computer, personal digital assistant (PDA), mobile device, etc.), and can include any commercial operating system and any combination of commercial and custom software (e.g., machine learning software, etc.). These systems can include any type of monitor and input device for inputting and / or viewing information (e.g., keyboard, mouse, voice recognition, etc.).

[0060] It should be understood that the software of this embodiment can be implemented in any desired computer language and can be developed by those of ordinary skill in the computer art based on the functional descriptions contained in the specification and the flowcharts shown in the drawings. In addition, any reference herein to software performing various functions generally refers to a computer system or processor that performs these functions under the control of the software. The computer system of this embodiment can alternatively be implemented by any type of hardware and / or other processing circuitry.

[0061] The various functions of the computer or other processing systems can be distributed in any number of software and / or hardware modules or units, processing or computer systems, and / or circuits in any manner, where the computers or processing systems can be located locally or remotely from each other and communicate via any suitable communication medium (e.g., local area network (LAN), wide area network (WAN), intranet, Internet, hardwired, modem connection, wireless, etc.). For example, the functions of this embodiment can be distributed in any manner among various end-user / client and server systems and / or any other intermediate processing devices. The software and / or algorithms shown in the above and in the flowcharts can be modified in any way to implement the functions described herein. In addition, the functions in the flowcharts or descriptions can be executed in any order to achieve the desired operations.

[0062] The software of this embodiment can be used on a non-transitory computer-usable medium (e.g., magnetic or optical medium, magneto-optical medium, floppy disk, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory device, etc.) of a fixed or portable program product device or apparatus used in conjunction with a stand-alone system or a system connected via a network or other communication medium.

[0063] The communication network can be implemented by any number of any type of communication network (e.g., LAN, WAN, Internet, intranet, virtual private network (VPN), etc.). The computer or other processing system of this embodiment can include any conventional or other communication device that communicates via a network through any conventional or other protocol. The computer or other processing system can utilize any type of connection (e.g., wired, wireless, etc.) to access the network. The local communication medium can be implemented by any suitable communication medium (e.g., LAN, hardwired, wireless link, intranet, etc.).

[0064] Each element described herein can be coupled and / or interact with each other through an interface and / or through any other suitable connection (wired or wireless) that provides a viable communication path. The interconnections, interfaces, and their variations discussed herein can be used to provide connections between elements in a system, and / or can be used to provide communication, interaction, operation, etc. between elements directly or indirectly connected in a system. Any combination of interfaces can be provided for the elements described herein to facilitate the operations discussed for the various embodiments described herein.

[0065] The system can use any number of any conventional or other database, data storage, or storage structure (e.g., file, database, data structure, data or other repository, etc.) to store information. The database system can be implemented by any number of any conventional or other database, data storage, or storage structure to store information. The database system can be included within and / or coupled to the server and / or client system. The database system and / or storage structure can be remote from or local to the computer or other processing system and can store any desired data.

[0066] The presented embodiments can be in various forms, such as a system, a method, and / or a computer program product at any possible level of integrated technical details. The computer program product can include one or more computer-readable storage media having computer-readable program instructions thereon for causing a processor to execute the aspects described herein.

[0067] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer floppy disk, a hard disk, RAM, ROM, EPROM, flash memory, static RAM (SRAM), a portable CD-ROM, a DVD, a memory stick, a floppy disk, a mechanically encoded device, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse propagating through an optical fiber cable), or an electrical signal transmitted through a wire.

[0068] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a LAN, a WAN, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.

[0069] The computer-readable program instructions for performing the operations of this embodiment may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (e.g., Python, C++) and procedural programming languages (e.g., the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network connection including a LAN or WAN, or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit in order to perform the aspects described herein.

[0070] Aspects of this embodiment are described herein with reference to the flowchart and / or block diagram of a method, apparatus (system), and computer program product according to an embodiment. It should be understood that each block of the flowchart illustration and / or block diagram, and the combination of blocks in the flowchart illustration and / or block diagram, can be implemented by computer-readable program instructions.

[0071] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus for generating a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create a module for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium, which can direct a computer, a programmable data processing apparatus, and / or other devices to work in a particular manner, such that the computer-readable storage medium having instructions stored therein includes an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0072] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0073] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.

[0074] In summary, in one example, a video endpoint obtains a first audio signal from a vertical microphone array, the first audio signal including audio from a target sound source and audio from a horizontally displaced sound source. The video endpoint obtains a second audio signal and a third audio signal from a horizontal microphone array, both the second audio signal and the third audio signal including audio from the target sound source and audio from the horizontally displaced sound source. Based on the second audio signal and the third audio signal, the video endpoint determines at least one of a first arrival angle of the audio from the target sound source or a second arrival angle of the audio from the horizontally displaced sound source. Based on at least one of the first arrival angle or the second arrival angle, the video endpoint adjusts the gain of the first audio signal.

[0075] In one form, a device is provided. The device includes: a vertical microphone array; a horizontal microphone array; and a processor coupled to the vertical microphone array and the horizontal microphone array, wherein the processor is configured to perform the following operations: obtain a first audio signal from the vertical microphone array, the first audio signal including audio from a target sound source and audio from a horizontally displaced sound source; obtain a second audio signal and a third audio signal from the horizontal microphone array, both the second audio signal and the third audio signal including audio from the target sound source and audio from the horizontally displaced sound source; determine at least one of a first arrival angle of the audio from the target sound source or a second arrival angle of the audio from the horizontally displaced sound source based on the second audio signal and the third audio signal; and adjust a gain of the first audio signal based on at least one of the first arrival angle or the second arrival angle.

[0076] In one example, the device further includes a camera, and the processor is further configured to perform the following operations: obtain a video signal from the camera; adjust a range of arrival angles based on the video signal; determine that the second arrival angle is outside the range of arrival angles; and in response to determining that the second arrival angle is outside the range of arrival angles, adjust the gain of the first audio signal by increasing an audio level of the audio from the target sound source or attenuating a gain of the audio from the horizontally displaced sound source. In another example, the video signal includes a video feed of the target sound source, the target sound source including a user's face, and the processor is further configured to perform the following operations: adjust the range of arrival angles by performing face recognition on the user's face. In yet another example, the processor is further configured to perform the following operations: adjust the range of arrival angles based on a zoom level of the camera.

[0077] In one example, the processor is further configured to perform the following operations: determine a horizontal displacement of the target sound source relative to the vertical microphone array based on the first arrival angle; and adjust the gain of the first audio signal based on the horizontal displacement of the target sound source relative to the vertical microphone array.

[0078] In one example, the processor is further configured to perform the following operations: adjust the gain of the first audio signal based on at least one of the first arrival angle or the second arrival angle to equalize an audio level of the audio from the target sound source and an audio level of the audio from the horizontally displaced sound source.

[0079] In one example, the processor is further configured to perform the following operations: use beamforming techniques to generate a vertically symmetric directivity pattern via the vertical microphone array, the vertically symmetric directivity pattern having a maximum value along a predicted vertical displacement of the target sound source and a zero value along a predicted vertical displacement of a vertically displaced sound source. In another example, the device further includes a speaker, wherein the vertically displaced sound source includes at least one of the following: a speaker, a surface configured to support the device, or a user equipment on the surface.

[0080] In one example, the device is a video endpoint, the video endpoint including a housing that supports a display screen, and wherein the vertical microphone array is positioned along a vertical side of the display screen and the horizontal microphone array is positioned along a top portion of the display screen proximate a camera of the video endpoint.

[0081] In another form, a method is provided. The method includes: obtaining a first audio signal from a vertical microphone array, the first audio signal including audio from a target sound source and audio from a horizontally displaced sound source; obtaining a second audio signal and a third audio signal from a horizontal microphone array, both the second audio signal and the third audio signal including audio from the target sound source and audio from the horizontally displaced sound source; determining at least one of a first arrival angle of the audio from the target sound source or a second arrival angle of the audio from the horizontally displaced sound source based on the second audio signal and the third audio signal; and adjusting a gain of the first audio signal based on the at least one of the first arrival angle or the second arrival angle.

[0082] In another form, one or more non-transitory computer-readable storage media are provided. The non-transitory computer-readable storage media are encoded with instructions that, when executed by a processor, cause the processor to perform the following operations: obtaining a first audio signal from a vertical microphone array, the first audio signal including audio from a target sound source and audio from a horizontally displaced sound source; obtaining a second audio signal and a third audio signal from a horizontal microphone array, both the second audio signal and the third audio signal including audio from the target sound source and audio from the horizontally displaced sound source; determining at least one of a first arrival angle of the audio from the target sound source or a second arrival angle of the audio from the horizontally displaced sound source based on the second audio signal and the third audio signal; and adjusting a gain of the first audio signal based on the at least one of the first arrival angle or the second arrival angle.

[0083] For purposes of illustration, a description of various embodiments has been presented, but is not intended to be exhaustive or limiting to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein were chosen in order to best explain the principles of the embodiments, the practical application or technical improvement over technologies found in the marketplace, or to enable other ordinary skilled artisans in the art to understand the embodiments disclosed herein.

[0084] The foregoing description is only by way of example. Although these techniques have been illustrated and described herein as embodied in one or more specific examples, it is not intended to be limited to the details shown, since various modifications and structural changes may be made within the scope of equivalents of the claims.

Claims

1. An apparatus for processing an audio signal, comprising: a vertical microphone array; a horizontal microphone array; and a processor coupled to the vertical microphone array and the horizontal microphone array, wherein the processor is configured to perform the following operations: obtain a first audio signal from the vertical microphone array, the first audio signal including audio from a target sound source and audio from a horizontally displaced sound source; obtain a second audio signal and a third audio signal from the horizontal microphone array, both the second audio signal and the third audio signal including audio from the target sound source and audio from the horizontally displaced sound source; determine at least one of a first arrival angle of the audio from the target sound source or a second arrival angle of the audio from the horizontally displaced sound source based on the second audio signal and the third audio signal; and adjust a gain of the first audio signal based on at least one of the first arrival angle or the second arrival angle.

2. The apparatus according to claim 1, further comprising a camera, wherein the processor is further configured to perform the following operations: obtain a video signal from the camera; adjust a range of arrival angles based on the video signal; determine that the second arrival angle is outside the range of arrival angles; and in response to determining that the second arrival angle is outside the range of arrival angles, adjust the gain of the first audio signal by increasing an audio level of the audio from the target sound source or attenuating a gain of the audio from the horizontally displaced sound source.

3. The apparatus according to claim 2, wherein the video signal includes a video feed of the target sound source, wherein the target sound source includes a user's face, and wherein the processor is further configured to perform the following operation: adjust the range of arrival angles by performing face recognition on the user's face.

4. The apparatus according to claim 2 or 3, wherein the processor is further configured to perform the following operation: adjust the range of arrival angles based on a zoom level of the camera.

5. The apparatus according to claim 1, wherein the processor is further configured to perform the following operations: determine a horizontal displacement of the target sound source relative to the vertical microphone array based on the first arrival angle; and adjust the gain of the first audio signal based on the horizontal displacement of the target sound source relative to the vertical microphone array.

6. The apparatus according to claim 1, wherein the processor is further configured to perform the following operation: adjust the gain of the first audio signal based on at least one of the first arrival angle or the second arrival angle to equalize an audio level of the audio from the target sound source and an audio level of the audio from the horizontally displaced sound source.

7. The apparatus according to claim 1, wherein the processor is further configured to perform the following operation: Using beamforming techniques to generate a vertically symmetric directivity pattern via the vertical microphone array, the vertically symmetric directivity pattern having a maximum value along the predicted vertical displacement of the target sound source and a zero value along the predicted vertical displacement of the vertically displaced sound source.

8. The apparatus according to claim 7, further comprising a speaker, wherein, the vertically displaced sound source includes at least one of the following: the speaker, a surface configured to support the apparatus, or a user device on the surface.

9. The apparatus according to claim 1, wherein, the apparatus is a video endpoint, the video endpoint including a housing that supports a display screen, and wherein the vertical microphone array is positioned along a vertical side of the display screen and the horizontal microphone array is positioned along a top portion of the display screen proximate a camera of the video endpoint.

10. A method for processing an audio signal, comprising: obtaining a first audio signal from a vertical microphone array, the first audio signal including audio from a target sound source and audio from a horizontally displaced sound source; obtaining a second audio signal and a third audio signal from a horizontal microphone array, both the second audio signal and the third audio signal including audio from the target sound source and audio from the horizontally displaced sound source; determining at least one of a first arrival angle of the audio from the target sound source or a second arrival angle of the audio from the horizontally displaced sound source based on the second audio signal and the third audio signal; and adjusting a gain of the first audio signal based on at least one of the first arrival angle or the second arrival angle.

11. The method according to claim 10, further comprising: obtaining a video signal from a camera; adjusting a range of arrival angles based on the video signal; determining that the second arrival angle is outside the range of arrival angles; and in response to determining that the second arrival angle is outside the range of arrival angles, adjusting the gain of the first audio signal by increasing an audio level of the audio from the target sound source or attenuating a gain of the audio from the horizontally displaced sound source.

12. The method according to claim 11, wherein, the video signal includes a video feed of the target sound source and the target sound source includes a user's face, the method further comprising: adjusting the range of arrival angles by performing face recognition on the user's face.

13. The method according to claim 11 or 12, further comprising: adjusting the range of arrival angles based on a zoom level of the camera.

14. The method according to claim 10, further comprising: determining a horizontal displacement of the target sound source relative to the vertical microphone array based on the first arrival angle; and adjusting the gain of the first audio signal based on the horizontal displacement of the target sound source relative to the vertical microphone array.

15. The method according to claim 10, further comprising: Adjust the gain of the first audio signal based on at least one of the first angle of arrival or the second angle of arrival to equalize the audio level of the audio from the target sound source and the audio level of the audio from the horizontally displaced sound source.

16. The method according to claim 10, further comprising: Using beamforming techniques to generate a vertically symmetric directivity pattern via the vertical microphone array, the vertically symmetric directivity pattern having a maximum along the predicted vertical displacement of the target sound source and a zero along the predicted vertical displacement of the vertically displaced sound source.

17. One or more non-transitory computer-readable storage media encoded with instructions that, when executed by a processor, cause the processor to perform the following operations: Obtain a first audio signal from a vertical microphone array, the first audio signal including audio from a target sound source and audio from a horizontally displaced sound source; Obtain a second audio signal and a third audio signal from a horizontal microphone array, both the second audio signal and the third audio signal including audio from the target sound source and audio from the horizontally displaced sound source; Based on the second audio signal and the third audio signal, determine at least one of a first angle of arrival of the audio from the target sound source or a second angle of arrival of the audio from the horizontally displaced sound source; and Adjust the gain of the first audio signal based on at least one of the first angle of arrival or the second angle of arrival.

18. The one or more non-transitory computer-readable storage media according to claim 17, wherein the instructions further cause the processor to perform the following operations: Obtain a video signal from a camera; Based on the video signal, adjust the range of the angle of arrival; Determine that the second angle of arrival is outside the range of the angle of arrival; and In response to determining that the second angle of arrival is outside the range of the angle of arrival, adjust the gain of the first audio signal by increasing the audio level of the audio from the target sound source or attenuating the gain of the audio from the horizontally displaced sound source.

19. The one or more non-transitory computer-readable storage media according to claim 17 or 18, wherein the instructions further cause the processor to perform the following operations: Based on the first angle of arrival, determine the horizontal displacement of the target sound source relative to the vertical microphone array; and Based on the horizontal displacement of the target sound source relative to the vertical microphone array, adjust the gain of the first audio signal.

20. The one or more non-transitory computer-readable storage media according to claim 17, wherein the instructions further cause the processor to perform the following operations: Adjust the gain of the first audio signal based on at least one of the first angle of arrival or the second angle of arrival to equalize the audio level of the audio from the target sound source and the audio level of the audio from the horizontally displaced sound source.

21. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the steps of the method according to any one of claims 10 to 16.

22. A computer-readable medium comprising instructions that, when executed by a computer, cause the computer to perform the steps of the method according to any one of claims 10 to 16.

Citation Information

Patent Citations

  • Systems, methods, and apparatus for estimating direction of arrival

    CN104220896A

  • Systems and methods for surround sound echo reduction

    CN104429100A