Audio fence system and method

By deploying dedicated lobes outside the audio coverage area to capture unwanted sounds and using a source remover to generate a mask, the problem of off-axis noise leakage in existing technologies is solved, improving the audio quality of the audio system and the performance of the automatic mixer.

CN120981850APending Publication Date: 2025-11-18SHURE ACQUISITION HLDG INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202480024134.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-03
Filing Date
2024-03-03
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing audio systems struggle to effectively remove unwanted sounds from outside the audio coverage area, causing off-axis noise to leak into the desired audio mix and affecting the performance of automatic mixers.

Method used

By deploying dedicated audio pickup lobes outside the audio coverage area, unwanted sounds are captured, and a source remover is used to generate a mask to remove off-axis noise from the desired audio signal.

Benefits of technology

It effectively reduces or eliminates interference from sounds outside the coverage area on the desired audio output, improving audio quality and the performance of the automatic mixer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120981850A_ABST
    Figure CN120981850A_ABST
Patent Text Reader

Abstract

Systems and methods are provided herein for deploying a first microphone lobe toward a first location using at least one microphone, the first microphone lobe configured to capture one or more first audio signals from a first audio source located within a first audio pickup region; deploying a second microphone lobe toward a second position using the at least one microphone, the second microphone lobe configured to capture one or more second audio signals from a second audio source located outside the first audio pickup region; and removing, using at least one processor, off-axis noise from the one or more first audio signals by applying a mask determined based on the one or more second audio signals to the one or more first audio signals.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 449,845, filed March 3, 2023, the contents of which are incorporated by reference in their entirety. TECHNICAL FIELD

[0003] The present disclosure relates generally to removing unwanted sound from a desired audio signal in an audio caging scenario. In particular, the present disclosure relates to systems and methods for removing unwanted sound from a desired audio signal using audio signals captured by one or more lobes deployed outside a desired audio coverage area. BACKGROUND

[0004] Audio environments such as conference rooms, boardroom meeting and other meeting places, video conferencing setups, and the like can involve the use of multiple microphones or microphone array lobes to capture sound from various audio sources. For example, the audio sources can include human speakers. The captured sound can be disseminated through speakers (for amplification) to local observers in the environment, and / or to others remotely (such as via television broadcast, webcast, and the like). For example, personnel in a conference room can be on a teleconference with personnel at a remote location. Each of the microphones or array lobes can form a channel. The captured sound can be input as multi-channel audio and provided or output as a single mixed audio channel.

[0005] Generally, audio capture devices such as, for example, conferencing devices, are available in various sizes, form factors, mounting options, and wiring options to meet the needs of a particular environment. The types of conferencing devices, their operating characteristics (e.g., lobe direction, gain, and the like), and their placement in a particular audio environment can depend on many factors including, for example, the location of the audio sources, the location of the listeners, physical space requirements, aesthetics, room layout, and / or other considerations. For example, in some environments, conferencing devices can be placed on a table or podium close to the audio sources and / or listeners. For example, in other environments, conferencing devices can be mounted overhead or on a wall to capture sound from or project sound toward the entire room.

[0006] Some existing audio systems ensure optimal audio coverage of a given environment by delineating "audio coverage areas," which represent zones in the environment designated for capturing audio signals such as, for example, speech produced by human speakers. For example, an audio coverage area defines a space in which a microphone can deploy a beamformed audio lobe. A given environment or room can include one or more audio coverage areas, depending on the size, shape, and type of the environment. For example, an audio coverage area for a typical conference room can include the seating area around a conference table, while an audio coverage area for a typical classroom can include the space around a chalkboard and / or podium at the front of the room. Some audio systems have fixed audio coverage areas, while other audio systems are configured to dynamically create audio coverage areas for a given environment.

[0007] In some cases, sound captured within a given audio coverage area includes speech from human speakers, as well as unwanted audio such as errant non-speech or non-human noise in the environment (such as, for example, sudden, impulsive, or recurring sounds such as paper being shuffled, bags and containers being opened, chewing, sneezing, coughing, typing, etc.), errant speech noise (such as a bystander's commentary, overheard conversations between other people in the environment, etc.), or other noise disturbances. Noise reduction techniques can be used to reduce certain background, static, or fixed noise such as fan and HVAC system noise. However, such noise reduction techniques are not ideal for reducing or rejecting errant noise, unwanted speech, and other stray noise disturbances. Voice activity detection (VAD) algorithms that detect the presence or absence of human speech or voice in an audio stream can also be applied to one or more channels of a microphone to minimize unwanted audio in the captured sound. However, VAD techniques can not be effective at removing errant human speech from the desired audio stream. When a microphone does not capture human speech or voice, an automatic mixer can automatically reduce the strength of the audio input signal for a particular microphone to mitigate the effects of background, static, or fixed noise. However, completely or near-completely rejecting unwanted audio can compromise the performance of existing automatic mixers, as automatic mixers typically rely on relatively simple rules to select which channel to "gate" open, such as, for example, the first arrival time or highest amplitude at a given instant in time. SUMMARY

[0008] The techniques of this disclosure provide systems and methods designed to, among other things: (1) deploy a beamformed audio lobe toward an audio source located outside a designated audio coverage area; and (2) remove unwanted sound produced by the out-of-coverage audio source from a desired audio mix using audio captured by the "out-of-coverage" lobe.

[0009] One example embodiment includes a method using at least one processor in communication with at least one microphone, the method comprising: deploying, using the at least one microphone, a first microphone lobe toward a first location, the first microphone lobe configured to capture one or more first audio signals from a first audio source located within a first audio pickup zone; deploying, using the at least one microphone, a second microphone lobe toward a second location, the second microphone lobe configured to capture one or more second audio signals from a second audio source located outside the first audio pickup zone; and removing, using the at least one processor, off-axis noise from the one or more first audio signals by applying a mask determined based on the one or more second audio signals to the one or more first audio signals.

[0010] Another example embodiment includes a system comprising: at least one microphone configured to deploy a first microphone lobe toward a first location to capture one or more first audio signals from a first audio source located within a first audio pickup zone, and to deploy a second microphone lobe toward a second location to capture one or more second audio signals from a second audio source located outside the first audio pickup zone; and at least one processor communicatively coupled to the at least one microphone and configured to remove off-axis noise from the one or more first audio signals by applying a mask determined based on the one or more second audio signals to the one or more first audio signals.

[0011] Another example embodiment includes a non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to: deploy, using the at least one microphone, a first microphone lobe toward a first location to capture one or more first audio signals from a first audio source located within a first audio pickup zone; deploy, using the at least one microphone, a second microphone lobe toward a second location to capture one or more second audio signals from a second audio source located outside the first audio pickup zone; and remove off-axis noise from the one or more first audio signals by applying a mask determined based on the one or more second audio signals to the one or more first audio signals.

[0012] Another example embodiment includes a digital signal processing (DSP) component configured to: receive one or more first audio signals associated with a first location; receive one or more second audio signals associated with a second location; identify the first location as being within a first audio pickup zone; identify the second location as being outside the first audio pickup zone; generate a first audio mix using the one or more first audio signals; generate a second audio mix using the one or more second audio signals; and remove off-axis noise from the first audio mix by applying a mask determined based on the second audio mix to the first audio mix. According to some aspects, the DSP component is further configured to deploy a first microphone lobe toward the first location to pick up the one or more first audio signals; and deploy a second microphone lobe toward the second location to pick up the one or more second audio signals.

[0013] Another example embodiment includes a non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to: receive one or more first audio signals associated with a first location; receive one or more second audio signals associated with a second location; identify the first location as being within a first audio pickup zone; identify the second location as being outside the first audio pickup zone; generate a first audio mix using the one or more first audio signals; generate a second audio mix using the one or more second audio signals; and remove off-axis noise from the first audio mix by applying a mask determined based on the second audio mix to the first audio mix.

[0014] Another example embodiment includes a method using at least one processor in communication with at least one microphone, the method comprising: receiving one or more first audio signals associated with a first location; receiving one or more second audio signals associated with a second location; identifying the first location as being within a first audio pickup zone; identifying the second location as being outside the first audio pickup zone; generating a first audio mix using the one or more first audio signals; generating a second audio mix using the one or more second audio signals; and removing off-axis noise from the first audio mix by applying a mask determined based on the second audio mix to the first audio mix. According to some aspects, the method further comprises deploying a first microphone lobe toward the first location to pick up the one or more first audio signals; and deploying a second microphone lobe toward the second location to pick up the one or more second audio signals.

[0015] These and other embodiments, along with various arrangements and aspects, will become readily apparent from the following detailed description, the accompanying drawings, and the claims, where reference numerals have been used to describe the same or similar elements and functions in the several views. The detailed description particularly illustrates embodiments by describing exemplar}' modes of practicing the application. These are, however, merely exemplary and are not intended to limit the principles of the application. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 is a schematic diagram of an exemplary environment having audio coverage areas specified around microphones in accordance with one or more embodiments.

[0017] Figure 2 is a schematic diagram of an exemplary audio system in accordance with one or more embodiments.

[0018] Figure 3 is a schematic diagram of an exemplary audio processor that can be included in the system of Figure 2

[0019] Figure 4 is a schematic diagram of an exemplary environment having lobes deployed by the microphones of the system of Figure 2

[0020] Figure 5 is a schematic diagram of an exemplary source remover that can be included in the audio processor of Figure 3

[0021] Figure 6 is a schematic diagram of another exemplary source remover that can be included in the audio processor of Figure 3

[0022] Figure 7 is a flowchart illustrating exemplary operations for removing unwanted audio from a desired audio signal using the system of Figure 2

[0023] Figure 8 is a schematic diagram of an exemplary audio processor that can be included in the system of Figure 2

[0024] Figure 9 is a schematic diagram of another exemplary audio processor that can be included in the system of Figure 2 DETAILED DESCRIPTION

[0025] ​​​​​​​Generally, audio systems use an audio coverage area to focus one or more beamformed audio pickup lobes on sound produced by an audio source located within a predefined zone or acceptance zone of a given environment (e.g., a room), and audio signals captured by the audio pickup lobes are provided to corresponding channels of an automatic mixer to generate a desired audio mix. When a detected audio source falls outside the audio coverage area, existing audio systems simply refrain from deploying a lobe toward the source location, and rely on natural attenuation of the audio signals to prevent or at least minimize detection of such "out-of-coverage" sound by nearby active lobes. However, in some cases, out-of-coverage sound (which can include human speech and / or noise) can bleed or leak into audio captured by an "in-coverage" lobe (also referred to as "acoustic bleed-in"), and thus can be present in the desired audio mix as "off-axis" noise. For example, during a double-talk scenario, or when one person inside the audio coverage area and another person just outside the audio coverage area speak at the same time, a lobe focused within the audio coverage area can capture both in-coverage speech and out-of-coverage speech. The latter unwanted audio can also cause problems for proper automatic mixer channel selection, which attempts to select channels containing speech while also avoiding false noise.

[0026] Provided herein are systems and methods for actively locating sound from an audio source located outside an audio coverage area and using "out-of-coverage" sound to remove off-axis noise from a desired audio mix captured within the audio coverage area. For example, embodiments include using at least one microphone of an audio system to detect out-of-coverage audio, or unwanted sound produced by an audio source located outside the audio coverage area, and deploying a dedicated audio pickup lobe toward the out-of-coverage audio source in order to track or capture sound from the audio source. The at least one microphone can also deploy one or more audio pickup lobes toward audio sources located within the audio coverage area in order to capture sound from "in-coverage" audio sources. Audio signals captured by the dedicated lobe (or "out-of-coverage audio signals") are provided to one input of a source remover, while desired audio signals captured by active lobes deployed within the audio coverage area (or "in-coverage audio signals") are provided to another input of the source remover. The source remover generates a mask based on the out-of-coverage audio signals and the in-coverage audio signals, and applies the mask to the in-coverage audio signals to remove off-axis noise due to the out-of-coverage audio signals.

[0027] As used herein, the terms "lobe" and "microphone lobe" refer to a beamformed audio beam generated by a given microphone array (or array microphone) to pick up audio signals at selected locations such as where the lobe is pointed. While the technology disclosed herein is described with reference to microphone lobes generated by array microphones, the same or similar technology can also be used for other forms or types of microphone coverage (e.g., cardioid patterns, etc.) and / or for microphones that are not array microphones (e.g., handheld microphones, boundary microphones, lapel microphones, etc.). Thus, the term "lobe" is intended to encompass any type of audio beam or coverage.

[0028] Figure 1 An exemplary conference or other audio environment 100 including one microphone 102 and multiple audio sources 104 is shown in accordance with an embodiment. The audio environment 100 can be a conference room, boardroom, classroom, or other meeting space; a theater, stadium, auditorium, or other performance or event venue; or any other space. The audio sources 104 can be human speakers or talkers participating in a teleconference, television broadcast, webcast, class, seminar, performance, sporting event, or any other activity, and can be located at different locations around the environment 100. For example, the audio sources 104 can be teleconference local participants seated in respective chairs arranged around a table, or local audience members seated in chairs arranged in front of a lectern or other presentation space.

[0029] The microphone 102 can be configured to detect sound from the audio sources 104, such as human speech or voices emitted by the audio sources 104 and / or music, applause, or other sounds produced by the audio sources 104, and convert the detected sound into one or more audio signals. The microphone 102 can also capture other sounds present in the environment 100, including unwanted or undesirable sounds such as background noise (e.g., from a fan, vent, heating, ventilation, and air conditioning (HVAC) system, etc.), stray noise (e.g., typing, rustling of paper, opening of a bag of potato chips or other food container, etc.), or other non-human noise (e.g., sounds from audiovisual equipment, electronic equipment, etc.); non-speech human noise (e.g., sneezing, coughing, chewing, etc.); and human speech noise (e.g., comments by a bystander or conversation from a non-participant or other person present in the environment 100, audio from a remote participant played over an audio speaker in the environment 100, etc.). These undesirable sounds can also be captured in the audio signals produced by the microphone 102.

[0030] While Figure 1Only one microphone 102 is shown, but microphone 102 may include one or more of an array microphone, a non-array microphone (e.g., a directional microphone, such as a lavalier, boundary, etc.), or any other type of audio input device capable of capturing speech and other sounds. As examples, microphone 102 may include, but is not limited to, SHURE MXA310, MX690, MXA910, etc. Microphone 102 may be placed in any suitable location, including walls, ceilings, tables, podiums, and / or any other surface in environment 100, and may conform to various sizes, form factors, mounting options, and wiring options to meet the needs of a particular environment. The exact type, number, and placement of microphones in a particular environment may depend on the location of the audio source, the audience, physical space requirements, aesthetics, room layout, stage layout, and / or other considerations.

[0031] As shown in the figure, the audio environment 100 also includes an audio coverage area 106 (also referred to herein as the "audio pickup area") representing an acceptable audio pickup band for the microphone 102. Specifically, the audio coverage area 106 defines a zone or space within which the microphone 102 can deploy or focus beamforming audio lobes 108 to capture or detect desired audio signals, such as sound generated by an audio source 104 located within the audio coverage area 106. For example, as... Figure 1 As shown, the first audio pickup lobe 108 (e.g., lobe 1) can be deployed or pointed toward the first audio source 104, and the second audio pickup lobe 108 (e.g., lobe 2) can be deployed or pointed toward the second audio source 104 located on the opposite side of the microphone 102. In embodiments, the microphone 102 can be an audio system (such as, for example...) Figure 2 As part of the audio system 200 shown, the audio system is configured to define an audio coverage area 106 based on, for example, the known or calculated location of the microphone 102, the known or expected location of the audio source 104, and / or the real-time location of the audio source 104.

[0032] Although Figure 1 The specific configuration of environment 100 is shown, but it should be understood that other configurations are also conceivable and possible, including, for example, different arrangements of audio sources 104, audio sources that move within the room, different arrangements of audio coverage area 106, different locations of microphones 102, different numbers of audio sources, microphones and / or audio coverage areas, etc.

[0033] Environment 100 may also include one or more other audio sources 110 located outside the audio coverage area 106, such as... Figure 1The other audio sources 110 (also referred to herein as "out-of-coverage audio sources") can be human speakers or talkers located at or near the periphery of the audio coverage area 106 or are close enough in proximity that sound produced by the other audio sources 110 can also be detected by the lobes 108 deployed within the audio coverage area 106. The sound produced by the out-of-coverage audio sources 110 can be unwanted speech audio or other human voice noise, such as comments, conversations, or other speech by bystanders that does not belong to the teleconference or other activity being captured by the audio coverage area 106.

[0034] In some cases, an audio fence 112 can be formed around the audio coverage area 106 in order to prevent or block unwanted sound produced by the out-of-coverage audio sources 110 from entering the desired audio output (e.g., the mix of audio signals captured inside the audio coverage area 106). For example, the audio fence 112 can be formed by defining an additional audio coverage area 114 (or "outer coverage area") around the periphery of the preferred audio coverage area 106 and muting the lobes deployed in the outer coverage area 114 so that audio signals captured outside the audio coverage area 106 are not included in the desired audio output. Even so, unwanted sound produced by the out-of-coverage audio sources 110 can seep or leak into the audio signals captured by the in-coverage lobes 108, e.g., as off-axis noise, and thus can still be present or audible in the desired audio output. For example, during a doubletalk scenario where one of the desired audio sources 104 (e.g., talker A) is speaking simultaneously with one of the out-of-coverage audio sources 110 (e.g., talker B), sound produced by talker B can be inadvertently picked up by the same lobe 108 deployed towards talker A due to acoustic leakage.

[0035] In various embodiments, the microphones 102 can be configured to further minimize or remove unwanted acoustic leakage by deploying additional audio pickup lobes 116 configured to capture unwanted sound produced outside the audio coverage area 106. In particular, the additional lobes 116 (or "outer lobes") can be pointed towards the out-of-coverage audio sources 110 or other locations outside the audio coverage area 106. For example, as Figure 1 As shown, a first additional lobe 116 (e.g., lobe 4) can be deployed towards a first out-of-coverage audio source 110 in order to capture sound produced by talker B, and a second additional lobe 116 (e.g., lobe 5) can be deployed towards a second out-of-coverage audio source 110 in order to capture sound produced by the second source 110 (e.g., clapping). The unwanted audio signals captured by the additional lobes 116 can be provided to a source remover (e.g., a beamformer) included in the audio system to remove the unwanted sound from the audio signals captured by the in-coverage lobes 108. Figure 3the source remover 306) to remove off-axis noise caused by unwanted sound leaking or bleeding into the desired audio output. As described herein and Figures 3 to 5 As shown, the source remover can compute a mask based on the unwanted audio signals captured outside the audio coverage area 106 and the desired audio signals captured inside the audio coverage area 106, and apply the mask to the desired audio signals in order to minimize or remove the effects of unwanted acoustic leakage in the desired audio output.

[0036] With additional reference to Figure 2 , an example audio system 200 according to an embodiment is shown that is configured to remove off-axis noise from a desired audio output by implementing one or more of the techniques described herein. As shown, the audio system 200 (also referred to herein as a “system”) includes a microphone 202, a beamformer 204 communicatively coupled to the microphone 202, and an audio processor 206 communicatively coupled to the beamformer 204 and including a source remover 208. The audio system 200 can be used in an environment such as the environment 100, for example, in order to facilitate communication with personnel at a remote location and / or audio augmentation at the same location in addition to improving the signal quality of the audio output. For example, the microphone 202 can be the same as or substantially similar to the microphone 102 of Figure 1 , and can be used to capture sound from one or more of the audio sources 104 and / or 110 shown. Figure 1

[0037] Generally, the microphone 202 is configured to detect sound from audio sources in the environment and convert the sound into audio signals. While Figure 2 only one microphone is shown in the environment 100, the audio system 200 can work with any type and any number of microphones 202, including one or more microphone transducers (or elements), one or more microphone arrays, one or more directional microphones, or any combination thereof. In an embodiment, the microphone 202 generates a plurality of audio signals 210 based on the captured sound, and provides the audio signals 210 (also referred to herein as “detected audio signals”) to the beamformer 204.

[0038] The beamformer 204 can be configured to process the audio signals 210 and generate one or more beamformed audio signals 212 based thereon, or otherwise direct an audio pickup beam or microphone lobe to a particular location in the environment (e.g., as shown in Figure 1 ​in a different angle relative to the microphone 202. The beamformer 204 can be further configured to provide one or more beamformed audio signals 212 (or "lobe signals") to the audio processor 206 for further processing and mixing, as described herein. The beamformer 204 can include any type of beamforming algorithm or other beamforming technique configured to deploy or place microphone lobes, including, for example, a delay-and-sum beamforming algorithm, a minimum variance distortionless response ("MVDR") beamforming algorithm, etc. Although Figure 2 The beamformer 204 is shown as a separate or independent device communicatively coupled to the microphone 202, but in other embodiments, the beamformer 204 can be included in the microphone 202, the audio processor 206, or other component of the audio system 200.

[0039] When multiple microphone lobes are formed, the beamformer 204 can include multiple audio channels (not shown), and each channel can be assigned to a respective lobe for separately receiving and processing audio signals corresponding to that lobe. For example, the microphone 202 can be configured to provide each of the multiple audio signals 210 to a respective one of the multiple audio channels at the beamformer 204. Likewise, the audio processor 206 and / or each component thereof can be configured to include multiple audio channels for respectively receiving the lobe signals 212 output by the beamformer 204, e.g., as shown. Figure 3 As will be appreciated, other components of the audio system 200 can also include multiple channels respectively assigned to the multiple audio channels of the microphone 202 in order to allow for separate processing and / or handling of the audio signals 210 and / or lobe signals 212.

[0040] For ease of explanation, the techniques described herein can refer to using the multiple audio signals 210 captured by the microphone 202, even though these techniques can utilize any type of acoustic source, including the beamformed audio signals 212 generated by the beamformer 204. Additionally or alternatively, the multiple audio signals 210 captured by the microphone 202 can be converted into the frequency domain, in which case certain components of the audio system 200 can operate in the frequency domain.

[0041] In some embodiments, the microphones 202 and / or the beamformer 204 are configured to generate up to eight microphone lobes, and thus have at least eight audio channels. As will be appreciated, other numbers of channels / lobes (e.g., twelve, six, four, etc.) are also contemplated. In some embodiments, the total number of lobes can be fixed (e.g., eight). In other embodiments, the number of lobes can be user-selectable and / or automatically determined based on the locations of various audio sources detected by the microphones 202. Similarly, in some embodiments, the directionality and / or location of each lobe can be fixed such that the lobes always form a particular configuration. In other embodiments, the directionality and / or location of each lobe can be steerable or selectable based on user input and / or in response to, for example, detecting a new audio source, a known audio source moving to a new location, or repositioning or resetting an existing lobe.

[0042] In various embodiments, the microphones 202 and / or the beamformer 204 can include an automatic lobe deployer (“ALD”) configured to automatically deploy or position microphone lobes in the direction of detected audio sources, for example, this can include deploying a new lobe based on newly detected audio activity, repositioning an existing lobe to a newly detected location of a known audio source, or resetting an existing lobe to an initial lobe position. Exemplary embodiments of an audio system configured to use automatic lobe deployment techniques are disclosed in commonly-assigned U.S. Patent No. 11,438,691, the contents of which are incorporated by reference herein in their entirety.

[0043] In some embodiments, the microphones 202 can be configured to detect audio using general or non-directional lobes, and upon detecting an audio signal at a given location, the microphones 202 and / or the beamformer 204 can deploy a directional lobe toward the given location for capturing the detected audio signal. In other embodiments, the audio system 200 can not include a beamformer, in which case each of the audio signals 210 captured by the microphones 202 can be provided directly to the audio processor 206. For example, the microphones 202 can include a plurality of omnidirectional microphones, each configured to capture audio signals 210 using omnidirectional lobes. In this case, the plurality of audio signals 210 can still be provided to respective audio channels associated with the audio system 200.

[0044] In various embodiments, the beamformer 204 uses location data 214 obtained from the microphones 202 to determine appropriate lobe placement for optimally capturing audio sources detected by the microphones 202. The location data 214 (also referred to as “sound localization data”) can indicate the location of a detected audio source relative to the microphones 202. The microphones 202 can be configured to use, for example, time-of-arrival (TOA) techniques to determine the location of a detected audio source. In some embodiments, the location data 214 can be obtained from the microphones 202 using a sound localization algorithm, such as the one disclosed in commonly-assigned U.S. Patent No. 10, 1 13, 1 17, the contents of which are incorporated by reference herein in their entirety. Figure 2The positioning module 216, included in the microphone 202 or one or more other components of the audio system 200, generates location data 214. In embodiments, the positioning module 216 includes an algorithm or other software configured to generate the location of a detected sound or audio source and determine coordinates (also referred to herein as “position coordinates”) representing the position or orientation of the detected audio source relative to the microphone 202. Various methods for generating sound localization are known in the art, including, for example, generalized cross-correlation (“GCC”) and other methods.

[0045] The positioning coordinates generated by positioning module 216 can be included in the position data 214 provided to beamformer 204. Positioning coordinates can be Cartesian or rectangular coordinates representing a position point in three dimensions or x, y, and z values. In some cases, positioning coordinates can be converted to polar or spherical coordinates, i.e., azimuth (phi), elevation (theta), and radius (r), as known in the art, for example, using transformation formulas. Spherical coordinates can be used in various embodiments to determine additional information about audio system 200, such as, for example, the angular spacing or distance between the audio source and microphone 202, which can be used to configure one or more aspects of a given audio coverage area and / or by a source remover to optimize noise removal, as described herein.

[0046] In various embodiments, the audio system 200 includes an audio coverage area module 218 configured to set and / or configure one or more audio coverage areas for capturing sound from a desired audio source. The audio coverage area module 218 may be included in, for example... Figure 2 In the microphone 202 or one or more other components of the audio system 200 shown. With Figure 1 Similar to the illustrated audio coverage area 106, each audio coverage area represents an audio pickup band, or microphone 202 may deploy lobes within it for detecting other areas or spaces of audio. Audio coverage area module 218 can be configured to determine the size, shape, and location of each audio coverage area based on the location or intended location of the audio source in the environment, whether the audio source is sitting, standing, or moving within the environment, and / or other relevant information about the environment itself (e.g., the location of other audio equipment; vents, HVAC systems, or other noise sources; furniture or other objects; etc.). For example, in a typical conference room with multiple chairs arranged around a table, audio coverage area module 218 can position the audio coverage area above the chairs and / or table to create a sound band that focuses audio pickup onto any human speaker sitting at the table.

[0047] In some embodiments, the audio coverage area module 218 can be configured to automatically configure the audio coverage area based on the actual or real-time location of the audio source, e.g., as provided by or determined based on the location data 214 generated by the localization module 216. In such cases, the audio coverage area module 218 can be configured to select a set of initial boundaries for a given audio coverage area based on the initial audio source location obtained from the localization module 216. Upon receiving new location data 214 indicating that the audio source location has moved or changed, the audio coverage area module 218 can be configured to dynamically adjust the boundaries of the audio coverage area to encompass or cover the new audio source location. An exemplary embodiment of an audio system configured to automatically define or configure an audio coverage area based on real-time audio source location is disclosed in commonly-assigned U.S. Patent Application No. 18 / 151,346, filed January 6, 2023, and entitled “System and Method For Automatic Setup of Audio Coverage Area,” the contents of which are incorporated herein by reference in their entirety.

[0048] Once the audio coverage area module 218 has completed setup of the audio coverage area, the audio coverage area module 218 can output information about the audio coverage area to the microphones 202 and / or the beamformer 204, such as, for example, location coordinates or other information used to outline the area being covered or otherwise define the boundaries of the coverage area. Based on the received coverage information, the beamformer 204 can be configured to implement the audio coverage area by deploying or pointing the audio pickup lobe toward the area defined by the coverage information or, more generally, toward the specified location coordinates.

[0049] The audio processor 206 can be any type of processor capable of combining the desired audio signals received from the beamformer 204 to generate a mixed audio output (or “desired audio mix”) and removing off-axis noise from the mixed audio output using the source remover 208 or otherwise implementing the techniques described herein. In various embodiments, the audio processor 206 can be an audio signal processor, a digital signal processor (“DSP”), a software-implemented digital signal processing component, or any combination thereof. In some embodiments, the audio processor 206 can be or can be included in an aggregator configured to aggregate or collect data and / or audio from various components of the audio system 200 and apply appropriate processing techniques to the collected data and / or audio in accordance with the techniques described herein.

[0050] In various embodiments, the audio processor 206 can be configured to receive, at respective audio channels, beamformed audio signals 212 (or "lobe signals") corresponding to each of the lobes deployed by the beamformer 204. As described herein, audio pickup lobes deployed in a given environment can include in-coverage lobes, configured to pick up sound (or "desired audio") produced by audio sources located within a selected audio coverage area (such as, for example, the in-lobe 108 deployed in the audio coverage area 106 in Figure 1 Figure 1 out-of-coverage lobes, configured to pick up sound (or "undesired audio") produced by audio sources located outside the selected audio coverage area (such as, for example, the out-lobe 116 deployed outside the audio coverage area 106 in Figure 1 Thus, the audio processor 206 can include a plurality of input audio channels for receiving desired audio signals captured by or otherwise associated with in-coverage lobes (e.g., speech from the audio source 104 in Figure 1 Figure 1 and one or more reference audio channels for receiving undesired audio signals captured by or otherwise associated with out-of-coverage lobes (e.g., sound from the audio source 110 in Figure 1 By separating the lobe signals 212 into separate channels, the audio processor 206 can be configured to create a desired audio mix consisting of in-coverage audio signals and a noise mix consisting of out-of-coverage audio signals, and use the noise mix via the source remover 208 to remove off-axis noise from the desired audio mix, as described herein.

[0051] With additional reference to Figure 3 , an example audio processor 300 that can be used to implement the audio processor 206 shown in Figure 2 is shown, in accordance with an embodiment. The audio processor 300 includes a first audio mixer 302 (e.g., mixer A) for generating a desired audio mix (e.g., mix A) using in-coverage audio signals 303 or audio signals captured by desired or in-lobes (e.g., lobes 1 through n) deployed within an effective audio coverage area. The audio processor 300 further includes a second audio mixer 304 (e.g., mixer B) for generating a noise mix (e.g., mix B) using out-of-coverage audio signals 305 or audio signals captured by additional or out-lobes (e.g., lobes n+1 through n+m) deployed outside the effective audio coverage area. The audio processor 300 further includes a source remover 306 for removing off-axis noise from the desired audio mix by applying a mask determined based on the noise mix to the desired audio mix. The source remover 306 can be the same as or similar to the source remover 208 shown in Figure 2

[0052] Although Figure 3 While shown as separate components, any of the first audio mixer 302, the second audio mixer 304, and / or the source remover 306 can be combined into a single component of the audio processor 300 in other embodiments. In yet other embodiments, certain components of the audio processor 300 can be included separately in other devices, such as, for example, the microphone 202 or a computing device of the audio system 200. Figure 2

[0053] The first audio mixer 302 (also referred to herein as the “first mixer”) can be an automatic mixer, an audio mixing module, or any other type of mixer configured to generate a mixed audio signal that conforms to a desired mix of the in-coverage audio signals 303 obtained by the inner lobes (e.g., lobes 1 through n). For example, the desired mix can be obtained by emphasizing the audio signals 303 from certain lobes deployed within the audio coverage area and / or de-emphasizing or suppressing the audio signals 303 from other lobes deployed in the audio coverage area. Exemplary embodiments of audio mixers are disclosed in commonly-assigned U.S. Patent Nos. 4,658,425, 5,297,210, and 11,302,347, each of which is incorporated by reference herein in its entirety. As shown, the first audio mixer 302 can include a plurality of input audio channels for receiving the plurality of in-coverage audio signals 303, respectively, and an output audio channel for providing the mixed audio signal (or “desired audio mix”) to the source remover 306. Generally, each of the input audio channels can be gated open (e.g., allowed with little or no suppression) or gated closed (e.g., suppressed or attenuated), depending on whether the contribution of that channel or audio signal 303 captured by the corresponding lobe has been selected for inclusion in the desired audio mix. For example, an input audio channel can be gated closed if the corresponding audio signal 303 contains noise audio or does not contain speech audio. In this way, the first audio mixer 302 can be configured to generate the desired audio mix using only the contributions of the desired input audio channels, while excluding all other channels. Figure 3

[0054] The second audio mixer 304 (also referred to herein as the “second mixer” or “noise mixer”) can be any type of summer or other mixer for combining or summing the out-of-coverage audio signals 305 or otherwise generating a mix of sounds captured by the outer lobes (e.g., lobes n+1 through n+m) deployed outside the effective audio coverage area. As shown, the second audio mixer 304 can include a plurality of input audio channels for receiving the plurality of out-of-coverage audio signals 305 and an output audio channel for providing the mixed audio signal (or “noise mix”) to the source remover 306. Generally, each of the input audio channels can be gated open or gated closed, depending on whether the contribution of that channel or audio signal 305 captured by the corresponding lobe has been selected for inclusion in the noise mix. For example, an input audio channel can be gated closed if the corresponding audio signal 305 contains noise audio or does not contain speech audio. In this way, the second audio mixer 304 can be configured to generate the noise mix using only the contributions of the desired input audio channels, while excluding all other channels. Figure 3 ​​As shown, the second audio mixer 304 can include a plurality of input audio channels for receiving the out-of-coverage audio signals 305 and an output audio channel for providing the resulting noise mix (or "out-of-coverage mix") to the source remover 306, respectively.

[0055] With additional reference to Figure 4 , an example of acoustic leakage or spillage in an environment 400 is shown that includes a microphone 402 (e.g., similar to the microphone 202), a first sound source (e.g., source A) located within an audio coverage area 404, and a second audio source (e.g., source B) located outside the audio coverage area 404. As demonstrated, acoustic leakage can occur when sound produced by the source B (e.g., speech or other human noise) is audible in the audio signals captured by a first microphone lobe (e.g., lobe A) directed toward the source A, even though the source B falls outside the audio coverage area 404. For example, using only the audio signals captured by the lobe A or otherwise generating a desired audio mix inside the audio coverage area 404 can still include sound from the source B as off-axis noise.

[0056] According to embodiments, the source remover 306 can leverage the directionality of the microphone 402 (or its microphone lobes) to remove off-axis noise from the desired audio mix. In particular, the microphone 402 can be configured to deploy a second microphone lobe (e.g., lobe B) toward the source B or other audio sources located outside the audio coverage area 404 and provide sound captured by the lobe B (e.g., the out-of-coverage signals 305) to the source remover 306 for removing off-axis noise from the desired or in-coverage audio signals 303. The source remover 306 can be configured to generate a mask based on the audio signals 305 captured by the out-of-coverage lobe or the noise mix generated by the second audio mixer 304 and apply the mask (or "noise mask") to the desired mix generated based on the audio signals 303 captured by the in-coverage lobe. In this way, off-axis noise originating from the out-of-coverage signals 305 can be removed from the desired audio mix.

[0057] In various embodiments, source remover 306 can be configured to use a ratio of the desired audio mix to the noise mix to compute a mask (or mask value) and multiply the desired audio mix by the mask value to obtain a modified audio output that does not have off-axis noise. The mask can have any value in a range from about zero (i.e., apply a full mask) to about one (i.e., do not apply a mask). Source remover 306 can also be configured to adjust the aggressiveness of the mask, or the extent to which sources (e.g., noise) are removed from the desired audio mix. Furthermore, in some cases, the amount of removal applied to certain frequency bands of the desired audio mix can be adjusted according to known beamforming suppression in those frequency bands. In some embodiments, source remover 306 can be used to implement source separation in addition to or instead of removing in-coverage audio sources from the output of microphones 202. These and other aspects of the mask will be described in greater detail below according to exemplary embodiments. However, it will be appreciated that other embodiments can use other types of masks and / or any other combination of the techniques described herein to remove off-axis noise from the microphone output.

[0058] Reference is now made to Figure 5 FIG. 5, which shows an exemplary source remover 500 according to an embodiment that can be used to implement any of the source removers described herein, including source remover 208 of Figure 2 and / or source remover 306 of Figure 3 Source remover 500 can be configured to remove off-axis noise from a desired signal d that is contaminated by an interfering signal r. The desired signal can be, for example, mix A of Figure 3 or any other mix of audio signals captured using lobes deployed inside an audio coverage area (e.g., area 404 of Figure 4 and the interfering signal can be, for example, mix B of Figure 3 or any other mix of audio signals captured using lobes deployed outside the audio coverage area. As shown, source remover 500 includes a first input for receiving the desired signal, a second input for receiving the interfering signal, and an output for providing a corrected or modified version dr of the desired signal, such as, for example, mix A with off-axis noise removed due to mix B.

[0059] In various embodiments, source remover 500 can be configured to take a ratio of d to r and apply the ratio d / r as a gain or “mask” to the desired signal d to obtain a corrected signal dr. In other words, at a given time n, source remover 500 can remove off-axis noise in the desired signal due to acoustic leakage of the interfering signal by using a noise removal formula such as Equation 1:

[0060] dr[n] = d[n] * (d[n] / r[n]) = d[n] * m[n], (1)

[0061] where m is a mask value equal to d / r. According to embodiments, the mask value m can be upper bounded by one (i.e., no mask is applied) and lower bounded by zero (i.e., a full mask is applied). Further, the source remover 500 can obtain the squared norms of the desired signal and the interference signal to use as the d value and the r value, respectively, in the noise removal formula.

[0062] In some embodiments, the source remover 500 can be configured to operate in the frequency domain, e.g., like other components of the audio system 200. In this case, the noise removal formula can be applied to individual bins of a fast Fourier transform (FFT) of the audio signal. For example, the source remover 500 can be configured to compute a plurality of mask values or gains based on selected frequency bands (or "subbands") of the audio signal, and apply the mask values to corresponding bins of the FFT, respectively. In some cases, the source remover 500 can be configured to apply a mask to every bin of the FFT, e.g., by computing N mask values for an FFT having N bins in total. In other cases, as will be appreciated, the source remover 500 can be configured to apply a mask to only positive frequency bins of the FFT, e.g., by computing N / 2+1 mask values for N / 2+1 positive frequency bins of the FFT.

[0063] When operating in the frequency domain, each bin used by the source remover 500 has an associated "cross-over" threshold c that defines the point at which the mask switches between positive and negative gains. For example, in a given subband, if d is equal to g*r, then the ratio of the desired to the reference (i.e., d / r) is equal to g, and pre-multiplying the mask by 1 / g ensures that the mask value is 1, or 0 decibels (dB). In this case, since g is the point at which the mask switches between positive and negative gains, the cross-over threshold c can be set to 1 / g, and the mask value m can be set to c*(d / r). Thus, the noise removal formula dr[n] = d[n] * m[n] (or equation 1) described above becomes as shown by equation 2:

[0064] dr[n] = d[n] * c * (d[n] / r[n]) = d[n] * (1 / g) * (d[n] / r[n]). (2)

[0065] In embodiments, the cross-over threshold c and / or its denominator g can be predetermined or set during tuning or setup, e.g., by an operator of the audio system 200. In some cases, the cross-over threshold can be adaptive depending on one or more criteria, such as, for example, room size, amount of desired gain, reverberation in the environment, relative sound level at quiet times, and so on.

[0066] Generally, when the mask value m is less than one, e.g., because the desired signal d has fallen below g times the interference signal r, the noise removal formula causes the source remover 500 to output a corrected microphone signal dr that is attenuated compared to the desired signal d (or the desired audio mix). In some embodiments, the cross-over threshold c can be configured to have a more significant and / or customized impact on the performance of the source remover 500, e.g., in order to adjust the mask based on known beamforming rejection criteria. In one embodiment, the source remover 500 is configured to set the denominator of g or the cross-over threshold c to a value of one for the lowest frequency band, and a value of thirty-two or higher for higher frequency bands, with a gradient of values in between. For example, the g value can be configured to smoothly or uniformly transition from one to thirty-two for frequencies within the 0 to 9 kilohertz (kHz) range (within which human speech is likely to exist), and from thirty-two to one thousand for frequencies within the 9 kHz to 24 kHz range (within which speech is unlikely to exist, but noise can still exist). Thus, the mask can be customized to be more aggressive, or provide greater attenuation, in bandwidths that are unlikely to contain speech audio. In other embodiments, the source remover 500 can be customized according to other frequency bands and / or other ranges of the desired audio mix.

[0067] In some embodiments, the source remover 500 is configured to scale the aggressiveness of the noise removal from 0% (i.e., no removal) to 100% (i.e., full removal) by applying an aggressiveness scalar x to the mask value m. For example, the scalar x can be configured to have a value selected from the range of zero to one, and the modified mask value y can be calculated using Equation 3:

[0068] y = (1 - m) * x + m. (3)

[0069] In this case, the modified mask value can be applied to the desired signal d to obtain the corrected microphone output dr, i.e., using Equation 4:

[0070] dr = d * y = d * ((1 - m) * x + m). (4)

[0071] As will be appreciated, when the scalar x is equal to zero, the modified mask value y becomes equal to the mask value m, and thus, a full mask is applied to the desired signal d (i.e., dr= d*m). And when the scalar x is equal to one, the modified mask value y becomes one, which means that the gain is set to one and no mask is applied to the desired signal d (i.e., dr= d). In some cases, the scalar x can be automatically selected by the source remover 500. In other cases, the scalar x can be a user-selected value that is provided to the source remover 500 via a user interface of the audio processor 206 or other data input device of the audio system 200. In one example embodiment, the user inputs a value v between zero and one, and the source remover 500 is configured to invert the value v using the formula x = 1 - v, such that when applied to the mask value m, a user input of “0” means no removal, and a user input of “1” means full removal.

[0072] In other embodiments, the degree of aggressiveness of the mask can be scaled by raising the mask value m to an exponent. In this case, the exponent can be another degree of aggressiveness scalar s, and the modified mask value can be equal to m s Since the mask value m is a value less than or equal to one, the degree of aggressiveness of the mask can be increased by setting the scalar s to a value higher than one, and can be decreased by setting the scalar s to a value lower than one.

[0073] In some embodiments, the degree of aggressiveness of the source removal techniques described herein, or the fullness of the removal of a given source from an audio output, can be tuned by adjusting the signal power of an interfering signal (e.g., r) before it is received by the source remover. For example, the audio processor 300 can further include an amplifier or the like (not shown) that is external to the source remover 306 and is configured to amplify or attenuate the interfering signal (e.g., mix B) based on a desired degree of aggressiveness of the source remover. The desired degree of aggressiveness can be set by a user, e.g., via a user input, or automatically determined by the audio processor 300 based on gating decisions and the geometry of the sound sources. In some embodiments, the desired degree of aggressiveness can be an input value that adjusts the gain setting of the amplifier based on the amount of amplification or attenuation needed to reach the desired degree of aggressiveness level of the source remover.

[0074] While the source remover is primarily described herein as operating in the frequency domain, in other embodiments, the source remover and the rest of the audio system can be configured to operate in the time domain. In this case, as will be appreciated, the noise removal formula can be applied to the entire frequency band of the audio signal, e.g., by calculating a single mask value for the full bandwidth (e.g., 0 to 24 kilohertz (kHz) for a 48 kHz sampling rate), and the other techniques described herein can be modified accordingly.

[0075] In some embodiments, noise removal by source remover 500 can be most or more effective when the angular separation between the interference source and the speech source is within a predetermined range, such as, for example, 90 to 180 degrees, or 120 to 180 degrees, etc. If the interference source and the speech source are too close together, for example, when the angular separation is significantly less than the predetermined range (e.g., less than 90 degrees, less than 45 degrees, less than 30 degrees, etc.), source remover 500 can have difficulty distinguishing one source from the other. In such cases, source remover 500 can be configured to, for example, increase the aggressiveness of source remover 500 via the scalar x to compensate for the minimum separation.

[0076] The source removal techniques described herein can be used to remove out-of- coverage sounds from audio signals captured within an audio coverage area of a conference room or other environment (with multiple participants located at multiple microphones) or in any other noisy environment. In some embodiments, source remover 500 can also be configured to implement source separation or isolation for audio signals received from microphones 202. For example, using the techniques described herein, source remover 500 can separate a first audio signal corresponding to a first audio source from a second audio signal corresponding to a second audio source, i.e., remove the second audio signal from the first audio signal, and vice versa.

[0077] Figure 6 Another exemplary source remover 600 that can be used to implement any of the source removers described herein is shown in accordance with an embodiment. Like source remover 500, source remover 600 includes a first input for receiving a desired signal d, such as, for example, a mix A or other desired audio mix that includes sounds captured by lobes deployed within an audio coverage area; a second input for receiving an interference signal r, such as, for example, a mix B or other noise mix that includes sounds captured by lobes deployed outside the audio coverage area; and an output for providing a corrected or modified version of the desired signal dr, such as, for example, mix A with off-axis noise due to mix B removed. Also like source remover 500, source remover 600 can be configured to remove off-axis noise from the desired signal d that results from interference signal r bleeding into the desired signal by computing a mask m based on the interference signal and applying the mask to the desired signal. Figure 5 Figure 3 Figure 3

[0078] ​​​However, unlike source remover 500, source remover 600 includes a neural network 602 for computing a mask m based on the input desired signal and the input interfering signal. Source remover 600 also includes a multiplier 604 or other similar component for applying the mask to the desired signal (or mix A), as shown. When operating in the frequency domain, the input desired signal can include subband energy from the inner lobes, the input interfering signal can include subband energy from the outer lobes, and the neural network 602 can output the mask as a corresponding plurality of subband gains (e.g., N gains for N frequency partitions, etc.) to apply to the respective subband energy of the input desired signal, as described herein. According to embodiments, the neural network 602 can be implemented using any type of neural network including, for example, a convolutional neural network (“CNN”), a recurrent neural network (“RNN”), or any combination thereof.

[0079] In various embodiments, the neural network 602 can be configured to compute appropriate mask values for the input signal based at least in part on one or more coefficients obtained by the neural network 602 during a training phase. For example, the neural network 602 can be trained using sample audio signals captured by pairs of microphone lobes, each pair including an inner lobe pointing inside an audio coverage area in an environment and an outer lobe pointing outside the audio coverage area. During the training phase, the neural network 602 can be configured to test different values for one or more of the parameters used to compute the mask until it identifies a set of parameter values that correspond to a mask that, when applied to the desired signal, produces an output that closely resembles the original desired signal. The set of parameter values can be saved as one or more coefficients in a memory of the neural network 602 and / or source remover 600 and can be retrieved by the neural network 602 during normal or real-time use of the source remover 600.

[0080] In some embodiments, Figure 3 The illustrated source remover 306 can be implemented using a combination of the source remover 500 of Figure 5 and the source remover 600 of Figure 6 For example, the two source removers 500 and 600 can be operated independently to generate respective masks (e.g., ml and m2), and the final mask m can be determined based on the two masks ml and m2, for example, by selecting the minimum of the two, taking the average of the two, or any other suitable technique. Other techniques for combining the two source removers 500 and 600 are also contemplated.

[0081] Reference is now made to Figure 7An exemplary method or process 700 according to an embodiment is illustrated, which includes operations for removing off-axis noise caused by an audio source located outside an audio coverage area from a desired audio signal. Process 700 can be implemented using at least one processor communicating with at least one microphone or otherwise using an audio system. For ease of explanation, reference will be made below. Figure 2 The audio system 200 (including microphone 202) and / or Figure 3 The process 700 is described using an audio processor 300 (including a source remover 306), but it should be understood that the process 700 can also be implemented using other audio systems, processors, or devices. In embodiments, one or more processors and / or other processing components within the audio system 200 can perform any, some, or all of the steps of process 700. For example, process 700 can be implemented using a digital signal processing (“DSP”) component having multiple audio channels for receiving multiple audio signals captured by one or more microphones, respectively. The DSP component can be included in one or more microphones of the audio system (e.g., [microphone 300]). Figure 2 Microphone 202), processor (e.g., Figure 2 The audio processor 206 and / or one or more other components, or integrated therewith, are used in the audio system. In some embodiments, process 700 may be executed by a computing device included in the audio system (or more specifically, a processor of the computing device that executes software stored in memory). In some cases, the computing device may further perform the operation of process 700 by interacting with or interface-connecting with one or more other devices internal or external to the audio system 200 and communicatively coupled to the computing device. Process 700 may also be performed in conjunction with the processor and / or other processing components by utilizing one or more other types of components (e.g., memory, input and / or output devices, transmitters, receivers, buffers, drivers, discrete components, etc.) to perform any, some, or all of the steps of process 700.

[0082] like Figure 7 As shown, process 700 may begin at step 702 by using at least one microphone to transmit a first audio source (e.g., Figure 1 Speaker A) is identified as being located in the audio pickup area (e.g., Figure 1 The first location within the audio coverage area 106). In various embodiments, step 702 further includes: using at least one microphone to detect an audio signal (e.g., Figure 1 The audio signal 210); for example, based on the positioning coordinates or other location data associated with the audio signal (e.g., Figure 1identify or locate an origin or source of the audio signal; and compare the location of the audio signal to one or more boundaries of the audio pickup zone to determine that the source of the audio signal is located within the audio pickup zone.

[0083] At step 704, the process 700 can also include using the at least one microphone to identify a second audio source (e.g., speaker B in Figure 1 as being located at a second location outside of the audio pickup zone. In various embodiments, step 704 further includes using the at least one microphone to detect another audio signal (e.g., audio signal 210 of Figure 1 ); identifying or locating an origin or source of the other audio signal, e.g., based on positioning coordinates or other location data (e.g., location data 214 of Figure 1 associated with the other audio signal; and comparing the location of the other audio signal to one or more boundaries of the audio pickup zone to determine that the source of the other audio signal is located outside of the audio pickup zone.

[0084] At step 706, the process 700 includes using the at least one microphone to deploy a first microphone lobe (e.g., lobe 1 of Figure 1 toward the first location to capture one or more first audio signals (e.g., in-coverage audio signals 303 of Figure 1 from or generated by the first audio source inside the audio coverage area. As an example, the one or more first audio signals can be speech or other desired audio produced by the first audio source inside the audio coverage area.

[0085] The process 700 also includes, at step 708, using the at least one microphone to deploy a second microphone lobe (e.g., lobe 4 of Figure 3 toward the second location to capture one or more second audio signals (e.g., out-of-coverage audio signals 305 of Figure 1 from or generated by the second audio source outside the audio coverage area. As an example, the one or more second audio signals can include unwanted speech and / or noise produced by the second audio source outside the audio coverage area.

[0086] In some embodiments, as will be appreciated, the process 700 also includes receiving, at the at least one processor, the one or more first audio signals associated with the first microphone lobe and the one or more second audio signals associated with the second microphone lobe, and providing the audio signals to the at least one processor and / or a DSP component of the audio system for processing. In some embodiments, the process 700 can include deploying additional in-lobe (i.e., within the audio pickup zone) to capture sound from other desired audio sources located within the audio pickup zone, e.g., as by Figure 1The inner lobe 108 is shown in the diagram. Similarly, in some embodiments, process 700 may also include deploying additional outer lobes (i.e., outside the audio pickup area) to capture sound from other unwanted audio sources located outside the audio pickup area, such as those from... Figure 3 The outer lobe 116 is shown in the figure.

[0087] In various embodiments, process 700 further includes, at step 710, using at least one processor, generating a first audio mix based on or using one or more first audio signals captured within the audio pickup area. For example, the first audio mix (e.g., Figure 3 Mix A) can be the desired mix of audio signals captured within an audio coverage area, and can be achieved using an automatic mixer or other audio mixer (e.g., included in or otherwise communicatively coupled to at least one processor) Figure 3 The first audio mixer (302) in the middle is used to generate it.

[0088] Process 700 may further include, at step 712, using at least one processor, generating a second audio mix based on or using one or more second audio signals captured outside the audio pickup area. For example, the second audio mix (e.g., Figure 3 Mix B) can be a noise mix consisting of all audio signals captured outside the audio coverage area, and can be made using at least one adder, audio mixer, or other suitable device included in or otherwise communicatively coupled to it (e.g., Figure 7 The second audio mixer (304) in the middle is used to generate it.

[0089] like Figure 3 As shown, process 700 may further include calculating a mask based on the second audio mix at step 714. For example, the mask can be calculated using the ratio of the first audio mix to the second audio mix. Process 700 further includes removing off-axis noise from the first audio mix at step 716 by applying the mask determined based on the second audio mix to the first audio mix (i.e., at step 714). In embodiments, the mask may be provided by a source remover included in the audio system (e.g., Figure 8The source remover (306) calculates and applies a mask to the first audio mix to remove any off-axis noise from the first audio mix. In some embodiments, step 716 further includes adjusting the aggressiveness of the mask, for example, by applying a scaling factor to the ratio of the first audio mix to the second audio mix, as described herein. In other embodiments, step 716 further includes adjusting the aggressiveness of removing off-axis noise from the first audio mix by attenuating the second audio mix before calculating the mask. In some embodiments, the mask has a value ranging from about zero to about one, where a mask value of zero indicates that a full mask is applied, and a mask value of one indicates that no mask is applied. In some embodiments, the source remover may be configured to operate in the frequency domain such that a mask is applied to each subband of the audio signal. Once step 716 is completed, process 700 may end.

[0090] Figure 2 This demonstrates what can be used for implementation according to some embodiments. Figure 3 Another exemplary audio processor 800 is shown as the audio processor 206. As illustrated, the audio processor 800 includes an audio mixer 802, a processing module 804, and a plurality of source removers 806. Each of the source removers 806 (e.g., “SR 1”, “SR 2”, etc.) may be similar to... Figure 5 Source Remover 306 Figure 6 Source Remover 500 Figure 5 The source remover 600 or any combination thereof, or otherwise configured to implement one or more of the source removal techniques described herein.

[0091] According to an embodiment, the audio processor 800 can be configured to pair each inner lobe with an appropriate outer lobe based on the proximity between the lobes and apply a source removal technique to each pairing, and then generate the desired audio mix at the audio mixer 802. (See reference...) Figure 8 As explained, for example, source removal may be more effective or most effective when the speech source and interference source are expected to be sufficiently separated angularly (e.g., at approximately 180 degrees, at least 90 degrees, etc.) and / or physically (e.g., on opposite sides relative to the audio coverage area). Therefore, the audio processor 800 can improve the overall effectiveness of the source remover 806 by creating lobe pairs with sufficient or maximum spacing and providing each lobe pair to an individual source remover 806 for individualized source removal.

[0092] More specifically, such as Figure 1 As shown, the processing module 804 includes a first plurality of input terminals or input audio channels for receiving audio signals 803 within the coverage area or from signals deployed in the effective audio coverage area (e.g., Figure 2the inner lobes (e.g., lobes 1 through n) that capture audio signals within the audio coverage area 106. In addition, the processing module 804 includes a second plurality of inputs to receive out-of-coverage audio signals 805 or audio signals captured by the outer lobes (e.g., lobes n+1 through n+m) that are disposed outside the active audio coverage area, respectively.

[0093] The processing module 804 also receives location information (not shown) for each of the audio signals 803 and 805 and / or associated with each of the inner lobes and each of the outer lobes. The location information can include an audio source location, or a location of an audio source that produces the audio signal; a lobe location, or a location to which a lobe that picks up the audio signal is pointed; or any other location associated with the audio signal. In some embodiments, the location information can include positioning coordinates and / or other data included in the location data 214 received from the microphones 202. Figure 1

[0094] In various embodiments, the processing module 804 can be configured (e.g., using an algorithm or other software) to determine or calculate a physical distance and / or angular separation between a given inner lobe and each of the outer lobes, for example, based on the received location information associated with the lobes, and pair or associate the given inner lobe with the outer lobe that is located the farthest away and / or has the largest angular separation from the inner lobe. For example, in the example of FIG. 8, lobe 1 can be paired with lobe 5 because the angular and / or physical distance between lobe 1 and lobe 5 is greater than the angular and / or physical distance between lobe 1 and lobe 4. For similar reasons, for example, each of lobes 2 and 3 can be paired with lobe 4 rather than lobe 5. Figure 1

[0095] In cases where multiple out-of-coverage audio signals 805 result in off-axis noise in a given in-coverage audio signal 803 or multiple outer lobes are sufficiently separated from a given inner lobe, the processing module 804 can be configured to generate an appropriate mix of the out-of-coverage audio signals 805 and pair the out-of-coverage mix (e.g., “mix 1,” “mix 2,” etc.) with the corresponding in-coverage audio signal 803, as shown. In such cases, the processing module 804 can include an audio mixer, an audio mixing module, or other component for mixing audio signals.

[0096] In some embodiments, the processing module 804 can be further configured to estimate or measure the energy level of the outer lobe signals and determine which of the outer lobes to use for pairing purposes and / or include in the out-of-coverage mix based thereon. For example, when there is no active audio source (e.g., a speaker is not speaking or has moved away) at the lobe location, the outer lobe will typically produce low signal energy. As will be appreciated, such a lobe need not be included in the out-of-coverage mix or paired with an inner lobe.​​

[0097] The processing module 804 can be further configured to provide the audio signals associated with a given lobe pairing to the same source-remover 806 in order to maximize removal of the corresponding out-of-coverage audio signals 805 (or mix) from the in-coverage audio signals 803 within its pairing's coverage. As shown, the processing module 804 can include multiple outputs or output audio channels for providing the in-coverage signals 803 and the out-of-coverage mix to the appropriate inputs of the source-remover 806. For example, in Figure 1 coverage audio signals 803 associated with lobe 1 can be provided as a "desired" input signal to the first source-remover 806 (or "SR 1"), and an audio mix (e.g., "mix 1") including one or more out-of-coverage audio signals 805 associated with lobe 5 can be provided as an "interfering" input signal to the first source-remover 806. Using the techniques described herein, the first source-remover 806 can be configured to compute a mask based on mix 1 from lobe 5 (e.g., using a ratio of the desired signal to the interfering signal) and apply that mask to the in-coverage audio signals 803 from lobe 1, thus removing off-axis noise from lobe 1 due to acoustic leakage from lobe 5.

[0098] In some embodiments, each of the source-removers 806 further includes a data input (not shown) for receiving control data from the processing module 804 to control the aggressiveness of the mask applied by the given source-remover 806. As described herein, a scalar having any value on a scale from 0 (or no removal) to 1 (or full removal) can be used to control the aggressiveness of the mask. In various embodiments, the processing module 804 can be configured to compute the aggressiveness of each pair of lobes (e.g., lobe 1 and lobe 5) based on or as a function of the physical and / or angular separation between the corresponding lobes. For example, the processing module 804 can be configured to compute the aggressiveness of each pair of lobes based on the following equation: Figure 3The scalar values ​​of lobe 1 and lobe 5 in the audio source are calculated and provided as an aggressiveness scalar to the corresponding source remover 806 (e.g., via its data input). For example, the aggressiveness scalar can have a high value if the paired lobes are close to each other or have a physical and / or angular spacing below a preset threshold (e.g., about twenty degrees). Similarly, the aggressiveness scalar can have a low value if the paired lobes are sufficiently separated or have a physical and / or angular spacing exceeding a preset threshold. In this way, each source remover 806 can be configured to customize or individually scale the aggressiveness of its own mask based on the specific audio source assigned to that source remover 806, rather than applying the same or universal scalar to all masks. In other embodiments, the audio processor 800 can be configured to adjust the aggressiveness of source removal by adjusting the signal level of the out-of-coverage mix before receiving it at the source remover 806 or before calculating the mask, or to remove off-axis noise from each desired audio mix. In this case, each mix 1, 2, or n can be attenuated or amplified depending on the desired radicality of the corresponding source remover 806, as described herein.

[0099] In an embodiment, the modified or noise-removed output of each source remover 806 can be provided to the audio mixer 802 to generate the desired audio mix. The audio mixer 802 can be an automatic mixer, an automatic mixing module, or any other type of mixer configured to generate a mixed audio signal conforming to the desired mix of audio signals received at one or more inputs, such as those similar to... Figure 8 The first audio mixer 302. As shown, the audio mixer 802 includes multiple input channels for receiving modified audio signals from each of a plurality of source removers 806, and an output for providing a mixed audio output or a desired audio mix.

[0100] According to various embodiments, the processing module 804 may be implemented in hardware and / or in software stored in the memory of the audio processor 800 or in other components of the audio system 200. Although Figure 2 These are shown as separate components, but in other embodiments, any of the audio mixer 802, processing module 804, and / or multiple source removers 806 may be combined into a single component of the audio processor 800. In yet another embodiment, certain components of the audio processor 800 may be individually included, such as... Figure 9 In the microphone 202 or audio system 200, or in other devices such as computing devices.

[0101] Figure 2 This demonstrates what can be used for implementation according to some embodiments. Figure 3Another exemplary audio processor 900 is shown as the audio processor 206. As shown, the audio processor 900 includes an audio mixer 902 for generating a desired audio mix based on an audio signal 903 within coverage area, a processing module 904 for generating a modified mix of an audio signal 905 outside coverage area based on the proximity between the inner and / or outer lobes, and a source remover 906 for removing off-axis noise from the desired audio mix using a mask calculated based on the modified outside coverage area mix. The source remover 906 may be similar to... Figure 5 Source Remover 306 Figure 6 Source Remover 500 Figure 9 The source remover 600 or any combination thereof, or otherwise configured to implement one or more of the source removal techniques described herein. In embodiments, the processing module 904 may be configured to generate a modified out-of-coverage mix that emphasizes or weakens certain out-of-coverage audio signals 905, such that the source remover 906 can remove audio signals 905 that contribute more noise to the on-coverage mix to a greater extent, or otherwise improve source removal in the desired audio mix.

[0102] like Figure 1 As shown, the audio mixer 902 includes multiple input audio channels for receiving audio signals from sources deployed in the audio coverage area (e.g., [insert audio channel name here]). Figure 3 The audio mixer 902 captures multiple audio signals 903 within the coverage area (e.g., lobe 1, lobe 2, etc.) within the inner lobe of the audio coverage area 106. The audio mixer 902 also includes an output audio channel for providing a desired mix (or “desired coverage mix”) of the audio signals 903 within the coverage area to the “desired” input of the source remover 906. In an embodiment, the audio mixer 902 can be coupled with… Figure 8 The first audio mixer 302 is basically similar.

[0103] As shown, processing module 904 includes multiple input terminals or input audio channels for receiving multiple out-of-coverage audio signals 905 captured by outer lobes (e.g., lobe n+1, lobe n+2, etc.) deployed outside the audio coverage area. Processing module 904 also includes an output terminal or output audio channel for providing a modified mix of the out-of-coverage audio signals 905 to the "interference" input of source remover 906. Although not shown, it is related to... Figure 1 Like the processing module 804, the processing module 904 also receives position information associated with each of the audio signals 903 and 905 and / or with each of the inner lobes and the outer lobes.

[0104] In embodiments, the processing module 904 can be configured (e.g., using an algorithm or other software) to generate a modified mix of the out-of-coverage audio signals 905 that is designed to cause the source remover 906 to more aggressively remove the out-of-coverage audio signals 905 that have a louder presence in the in-coverage audio signals 903. In particular, the processing module 904 can be configured to apply a gain or frequency weight to one or more of the out-of-coverage audio signals 905 based on or as a function of the physical and / or angular spacing of the in-lobe with respect to each other and / or the outer lobes. The processing module 904 can determine or calculate the physical and / or angular spacing between a given in-lobe and each of the other in-lobes and each of the outer lobes based on the location information received for each of the lobes. Based on the spacing information, the processing module 904 can determine or identify which of the outer lobes is likely to make a greater contribution to or have a louder presence in the in-lobes. As an example, if two or more in-lobes (e.g., lobes 2 and 3 in Figure 1 are in close physical and / or angular proximity to the same outer lobe (e.g., lobe 5 in Figure 9 ), then sound from that outer lobe can be picked up or bleed into both in-lobes and, thus, can be included in at least two of the audio signals 903 mixed together by the audio mixer 902 to generate the in-coverage mix.

[0105] In some embodiments, for each outer lobe that is identified as being in close proximity to multiple in-lobes, the processing module 904 can calculate a gain or frequency weight based on the number of in-lobes in close physical and / or angular proximity to the outer lobe, the actual distance and / or angular spacing between the outer lobe and each nearby in-lobe, and / or any other relevant spacing information that results in a higher gain being calculated for the out-of-coverage audio signal 905 corresponding to that outer lobe. For example, the processing module 904 can apply a gain greater than one to the noisier out-of-coverage audio signals 905 in order to emphasize these audio signals and, thus, improve their removal capability when using the out-of-coverage mix to calculate the source removal mask. In other embodiments, the processing module 904 can instead apply a lower gain to other out-of-coverage audio signals 905 (i.e., signals captured by outer lobes that are further away from the identified outer lobe) in order to de-emphasize the other out-of-coverage audio signals 905 as compared to the noisier signals 905. For example, the processing module 904 can apply a gain less than one to the other signals 905. As will be appreciated, other techniques can be used to adjust or modify the frequency shape of the out-of-coverage mix in order to improve the effectiveness of the source remover 906.

[0106] In various embodiments, processing module 904 applies the computed gains to the corresponding out-of-coverage audio signals 905, then combines or mixes the out-of-coverage audio signals 905 together and outputs the modified out-of-coverage mix to source remover 906. Source remover 906 then computes a mask based on the ratio of the in-coverage mix to the modified out-of-coverage mix using the techniques described herein and applies the mask to the in-coverage mix (e.g., by multiplying the in-coverage mix by the mask) to remove off-axis noise. Thus, audio processor 900 can be configured to more aggressively or to a greater extent remove out-of-coverage audio sources from the desired audio output that are in close proximity to two or more of the in-coverage lobes and thus responsible for or contribute to a higher proportion of acoustic bleed-through in the in-coverage mix.

[0107] According to various embodiments, processing module 904 can be implemented in hardware and / or in software stored in a memory of audio processor 900 or other components of audio system 200. While shown as a single component, in other embodiments, any of audio mixer 902, processing module 904, and / or source remover 906 can be combined into a single component of audio processor 900. In yet other embodiments, certain components of audio processor 900 can be included separately in other devices such as, for example, microphone 202 or a computing device of audio system 200. Figure 2 While shown as separate components, in other embodiments, any of audio mixer 902, processing module 904, and / or source remover 906 can be combined into a single component of audio processor 900. In yet other embodiments, certain components of audio processor 900 can be included separately in other devices such as, for example, microphone 202 or a computing device of audio system 200. Figure 2

[0108] ​In various embodiments, any of the beamformers and / or microphones described herein can be further configured to provide audio fence coverage for a desired source and an interfering source that are located in the same direction, and thus have lobes that at least partially overlap due to the physical characteristics of the beamformer. More specifically, when a desired audio source and an interfering or noise source are located in the same azimuthal direction or otherwise have directionally similar locations, there can be significant overlap between the in-coverage lobe and the out-of-coverage lobe that are steered toward the two sources, respectively. Large overlap in lobe placement can result in various problems, including suboptimal noise removal and / or audio mixing. In some embodiments, the beamformer and / or microphone can be configured to use cosine angle information to determine the best coverage options for a desired source and an interfering source that are located in the same direction. In some cases, the beamformer and / or microphone can be configured to measure the lobe placement overlap for directionally similar sources, and strategically place the lobes to cover both sources. For example, the beamformer can be configured to adjust the position of the lobes so that the desired audio source is covered when the desired speaker is speaking, and the interfering source is covered when the desired speaker is not speaking. In such cases, the audio signals captured at the different lobe positions can be used as the desired input and the reference input to a source remover in order to remove noise from the desired audio, as described herein.

[0109] In various embodiments, any of the audio processors and / or microphones described herein can be further configured to account for jitter in the positioning information received for a given speaker or other desired audio source. More specifically, due to the stochastic nature of acoustics, speaker locations are rarely static and often contain jitter that makes the exact source location uncertain. This can be particularly problematic for audio sources that are located at or near the boundary between an in-coverage region and an out-of-coverage region of an acoustic fence or coverage area, as in such cases the source location can be interpreted as being both inside and outside the coverage area depending on the direction of the jitter. When it comes to source removal, jitter in the source location can result in jitter in the final audio output if a given audio source is constantly mixed in and out based on its fluctuating location. In some embodiments, the audio processor and / or microphone can be configured to use a spatio-temporal lag to account for the spatial characteristics of a boundary source or an audio source that is located at or near the boundary between an in-coverage region and an out-of-coverage region. In some cases, the audio processor and / or microphone can use the Euclidean distance between the lobe position and the detected audio activity location to determine the most likely placement of the lobe within the spatial characteristics of the audio source. The audio processor and / or microphone can also use a spatio-temporal timer to measure the active and inactive frequencies of the audio source, as well as the amount of spatial change that occurs at its location over a particular time period. These temporal results can be used to slow down the speed at which the audio source is classified as an in-coverage source or an out-of-coverage source, and thus reduce jitter in the source removal application of the final audio output.

[0110] In various embodiments, any of the audio processors and / or microphones described herein can be further configured to provide a gradual source removal for audio sources located at or near a boundary between an acoustically fenced or coverage area and an out-of-coverage area to help minimize the effects of audio source flutter at the boundary. Specifically, the attenuation gain or volume of lobes positioned near the boundary can be adjusted using an appropriate roll-off gain staging such that audio sources located at or near the boundary have less relevance or weight in the final audio output. For example, the area surrounding the edge of a given in-coverage area or portion adjacent to an out-of-coverage area can be designated as a “roll-off zone,” or an area where roll-off gain staging is applied. The gain gradient of such a roll-off zone can range from 0 decibels (dB) attenuation for audio sources located at or near the inner edge of the roll-off zone (i.e., adjacent to the in-coverage area) to -12 dB attenuation for audio sources located at or near the outer edge of the roll-off zone (i.e., adjacent to the out-of-coverage area). According to some embodiments, the audio processor and / or microphone can be configured to determine the roll-off gain for a given lobe based on a Euclidean or other relevant distance measurement, such as the distance of the lobe from the center of the in-coverage area. This makes the linear slope of the roll-off zone more gradual or less steep, which enables more gradual source removal near the acoustic fence.

[0111] While specific techniques for computing a mask for removing off-axis noise or otherwise isolating one audio source from another are described herein, it should be understood that other techniques can alternatively be used. For example, in some embodiments, the source remover can be configured to use Equation 5 to compute a mask for removing off-axis noise from the desired audio:

[0112] mask = mic_power / (mic_power + ref_power + EPS), (5)

[0113] Where mic_power is the desired audio signal (e.g., d), ref_power is the interfering audio signal or noise (e.g., r), and EPS is a very small number included to prevent division by zero. In this case, the mask is still calculated based on the desired audio signal and the interfering or noise audio signal, although using a different formula, and can still be used as described herein. For example, the mask value can still have any value in the range of approximately zero (i.e., a full mask is applied) to approximately one (i.e., no mask is applied), and the desired audio signal can be multiplied by the mask value to obtain a modified audio output without off-axis noise (e.g., dr = d * m), as described herein. Furthermore, the mask value can be applied to each frequency partition, for example, when the source remover operates in the frequency domain, and can be modified to scale for aggression, also as described herein. In various embodiments, the mask can be calculated using any of a plurality of equations, provided that the mask value (e.g., m) approaches one as the desired signal power (e.g., d) increases relative to the interfering signal power (e.g., r) and the mask value approaches 0 as the interfering signal power increases relative to the desired signal power.

[0114] Therefore, the techniques described herein can be used to remove off-axis noise from a mix of desired audio signals captured within an audio coverage area, caused by unwanted sounds from outside the audio coverage area infiltrating the desired audio signal. Specifically, the source remover can be configured to remove off-axis noise by applying a mask to the desired audio mix, the mask being calculated based on the audio signal outside the coverage area and the desired audio signal.

[0115] Return to reference Figure 2 In various embodiments, the audio system 200 may also include Figure 2 Various components not shown in the text include, for example, one or more speakers, displays, computing devices, and / or cameras. Furthermore, one or more of the components in system 200 may include one or more digital signal processors or other processing components, controllers, wireless receivers, wireless transceivers, etc., although not shown or mentioned above. It should be understood that... Figure 7 The components shown are merely exemplary, and any number, type, and placement of various components in system 200 are contemplated and possible.

[0116] One or more components of audio system 200 can be in wired or wireless communication with one or more other components of system 200. For example, microphone 202 can transmit a plurality of audio signals 210 to beamformer 204, audio processor 206, or a computing device including one or more of the aforementioned components using a wired or wireless connection. In some embodiments, one or more components of audio system 200 can communicate with one or more other components of system 200 via a suitable application programming interface (API). For example, one or more APIs can enable components of audio processor 206 to transmit audio and / or data signals between one another.

[0117] In some embodiments, one or more components of audio system 200 can be combined into or reside in a single unit or device. For example, all components of audio system 200 can be included in a same device such as microphone 202 or a computing device including microphone 202. As another example, audio processor 206 can be included in or combined with microphone 202 in addition to or instead of beamformer 204. In some embodiments, audio system 200 can take the form of a cloud-based system or other distributed system such that components of system 200 can or can not be in physical proximity to one another.

[0118] Components of audio system 200 can be implemented in hardware (e.g., discrete logic circuits, an application specific integrated circuit (ASIC), a programmable gate array (PGA), a field programmable gate array (FPGA), a digital signal processor (DSP), a microprocessor, etc.), using software executable by one or more servers or computers, or other computing devices having a processor and memory (e.g., a personal computer (PC), a laptop, a tablet, a mobile device, a smart device, a thin client, etc.), or through a combination of both hardware and software. For example, some or all components of microphone 202, beamformer 204, and / or audio processor 206 can be implemented using discrete circuit devices and / or using one or more processors (e.g., audio processors and / or digital signal processors) executing program code stored in memory (not shown) that is configured to perform one or more processes or operations described herein, such as, for example Figure 7 The illustrated method 700. Thus, in embodiments, one or more of the components of audio system 200 can include one or more processors, memory devices, computing devices, and / or other hardware components not shown in the figures.

[0119] All or part of the processes described herein, including the method 700, can be performed by one or more processors that execute software instructions Figure 2 The method 700, can be performed by one or more processors that execute software instructions Figure 7by one or more processing devices or processors (e.g., analog-to-digital converters, encryption chips, etc.) internal or external to the audio system 200. Moreover, one or more other types of components (e.g., memory, input and / or output devices, transmitters, receivers, buffers, drivers, discrete components, logic circuits, etc.) can also be used in conjunction with the processors and / or other processing components to perform any, some or all of the steps of the method 700. As an example, in some embodiments, each of the methods described herein can be performed by a processor executing software stored in a memory. The software can include, for example, program code or computer program modules including software instructions executable by the processor. In some embodiments, the program code can be a computer program stored on a non-transitory computer-readable medium executable by a processor of the relevant device.

[0120] Any of the processors described herein can include a general purpose processor (e.g., a microprocessor) and / or a special purpose processor (e.g., an audio processor, a digital signal processor, etc.). In some instances, the processors described herein can be any suitable processing device or set of processing devices, such as, but not limited to, a microprocessor, a microcontroller-based platform, an integrated circuit, one or more field programmable gate arrays (FPGAs), and / or one or more application specific integrated circuits (ASICs).

[0121] Any of the memories or memory devices described herein can be volatile memory (e.g., RAM including non-volatile RAM, magnetic RAM, ferroelectric RAM, etc.), non-volatile memory (e.g., disk storage, flash memory, EPROM, EEPROM, memristor-based non-volatile solid state memory, etc.), unalterable memory (e.g., EPROM), read-only memory, and / or a mass-based storage device (e.g., a hard drive, a solid state drive, etc.). In some instances, the memories described herein include multiple kinds of memory, particularly volatile memory and non-volatile memory.

[0122] Moreover, any of the memories described herein can be a computer-readable medium on which one or more sets of instructions can be embedded. The set of instructions can completely or at least partially reside on any one or more of the memories, computer-readable media, and / or within one or more processors during execution thereof. In some embodiments, the memories described herein can include one or more data storage devices configured for the persistent storage of data required to be stored and called upon by an end user. In this case, the data storage devices can save the data in flash memory or other memory devices. In some embodiments, the data storage devices can be implemented using, for example, a SQLite database, UnQLite, Berkeley DB, BangDB, etc.

[0123] Any of the computing devices described herein can be any general-purpose computing device including at least one processor and one memory device. In some embodiments, the computing device can be a standalone computing device included in the audio system 200, or can reside in another component of the audio system 200 such as, for example, the microphone 202, the audio processor 206, or the beamformer 204. In such embodiments, the computing device can be physically located in a given environment or room such as, for example, the same environment as the microphone 202, and / or be dedicated to that given environment or room. In other embodiments, the computing device can not be physically proximate to the microphone 202, but can reside in an external network such as a cloud computing network, or can otherwise be distributed in a cloud-based environment. Further, in some embodiments, the computing device can be implemented in firmware or be entirely software-based as part of a network that can be accessed or otherwise in communication with via another device including other computing devices such as, for example, desktops, laptops, mobile devices, tablets, smart devices, and the like. Thus, the term “computing device” should be understood to include distributed systems and devices such as cloud-based systems and devices, as well as software, firmware, and other components configured to perform one or more of the functions described herein. Further, one or more features of the computing device can be physically remote and can be communicatively coupled to the computing device.

[0124] In some embodiments, any of the computing devices described herein can include one or more components configured to facilitate a teleconference, meeting, class, or other activity, and / or to process audio signals associated therewith to improve the audio quality of the activity. For example, in various embodiments, any of the computing devices described herein can include a digital signal processor (“DSP”) configured to process audio signals received from various microphones or other audio sources using, for example, automatic mixing, matrix mixing, delay, compressor, parametric equalizer (“PEQ”) functions, acoustic echo cancellation, etc. In other embodiments, the DSP can be a standalone device that is operatively coupled or connected to the computing device using a wired or wireless connection. One example embodiment of a DSP when implemented in hardware is the P300 IntelliMix Audio Conference Processor from SHURE, the user manual of which is incorporated herein by reference in its entirety. As further explained in the P300 manual, this audio conference processor includes algorithms optimized for audio / video conferencing applications and to provide a high-quality audio experience including eight channels of acoustic echo cancellation, noise reduction, and automatic gain control. Another example embodiment of a DSP when implemented in software is IntelliMix Room from SHURE, the user guide of which is incorporated herein by reference in its entirety. As further explained in the IntelliMix Room user guide, this DSP software is configured to optimize the performance of networked microphones with audio and video conferencing software and is designed to run on the same computer as the meeting software. In other embodiments, as will be appreciated, other types of audio processors, digital signal processors, and / or DSP software components can be used to perform one or more of the audio processing techniques described herein.

[0125] Further, any of the computing devices described herein can also include various other software modules or applications (not shown) configured to facilitate and / or control conferencing activities, such as, for example, internal or proprietary conferencing software and / or third-party conferencing software (e.g., Microsoft Skype, Microsoft Teams, Bluejeans, Cisco WebEx, GoToMeeting, Zoom, Join.me, etc.). Such software applications can be stored in the memory of the computing device and / or can be stored on a remote server (e.g., local or as part of a cloud computing network) and accessed by the computing device via a network connection. Some software applications can be configured as distributed, cloud-based software, with one or more portions of the application residing in the computing device and one or more other portions residing in the cloud computing network. One or more of the software applications can reside in an external network, such as a cloud computing network. In some embodiments, one or more of the software applications can be accessed via a web portal architecture, or otherwise provided as a software as a service (SaaS).

[0126] Generally, a computer program product in accordance with the embodiments described herein includes a computer usable storage medium (e.g., standard random access memory (RAM), an optical disc, a universal serial bus (USB) drive, etc.) having a computer readable program code embodied therein, wherein the computer readable program code is adapted to be executed (e.g., in conjunction with an operating system) by a processor to implement methods described herein. In this regard, the program code can be implemented in any desired language, and can be implemented as machine code, assembly code, byte code, interpretable source code, etc. (e.g., via C, C++, Java, ActionScript, Python, Objective-C, JavaScript, CSS, XML, and / or others). In some embodiments, the program code can be a computer program stored on a non-transitory computer readable medium that is executable by a processor of the relevant device.

[0127] The terms“non-transitory computer-readable medium” and“computer-readable medium” include a single medium or multiple media, such as a centralized or distributed database, and / or associated caches and servers that store one or more sets of instructions. Further, the terms“non-transitory computer- readable medium” and“computer-readable medium” include any medium that is capable of storing, encoding or carrying a set of instructions for execution by a processor or that cause a system to perform any one or more of the methodologies of the present disclosure. The term“computer- readable medium” as used herein expressly excludes propagating signals per se.

[0128] Any process descriptions or blocks in the figures can represent modules, segments, or portions of code which include one or more executable instructions for implementing specific logic functions (or steps) in the process, and alternate implementations are included within the scope of embodiments described herein in which the functions can be performed out of the same order as shown or discussed, including substantially concurrently or in reverse order, depending upon the functionality involved, as will be understood by those having ordinary skill in the art. ​ It is noted that in the description and drawings, like or substantially similar elements can be labeled with the same reference numerals. However, these elements can sometimes be labeled with different numbers, such as, for example, where such labeling aids in a clearer description. Further, system components can be arranged in various ways, as is known in the art. Moreover, the drawings set forth herein are not necessarily drawn to scale, and in some instances, the proportions can be exaggerated, for greater clarity, and / or related elements can be omitted, to emphasize and highlight the novel features described herein. Such labeling and drawing practices do not necessarily imply a substantiality of purpose. The above description is intended to be taken as a whole and interpreted and understood in accordance with the principles taught herein by those of ordinary skill in the art.

[0129] It is noted that in the description and drawings, like or substantially similar elements can be labeled with the same reference numerals. However, these elements can sometimes be labeled with different numbers, such as, for example, where such labeling aids in a clearer description. Further, system components can be arranged in various ways, as is known in the art. Moreover, the drawings set forth herein are not necessarily drawn to scale, and in some instances, the proportions can be exaggerated, for greater clarity, and / or related elements can be omitted, to emphasize and highlight the novel features described herein. Such labeling and drawing practices do not necessarily imply a substantiality of purpose. The above description is intended to be taken as a whole and interpreted and understood in accordance with the principles taught herein by those of ordinary skill in the art.

[0130] In this disclosure, the use of contractions is intended to include conjunctions. The use of definite or indefinite articles is not intended to indicate a cardinal number. Specifically, reference to“the” object or“a” or“an” object is intended to also mean one of possibly multiple such objects.

[0131] The present disclosure describes, illustrates, and exemplifies one or more specific embodiments of the application according to its principles. The present disclosure is intended to explain the principles of the application and to enable others skilled in the art to utilize the application in various embodiments, the true scope of which is not limited to the precise forms disclosed. That is, the foregoing description is intended to be illustrative rather than limiting of the true scope of the application, which is set forth in the claims. It is contemplated that the application can be practiced with the exact features shown and / or described or with other equivalent arrangements and / or components. The present disclosure is intended to cover and include all changes, substitutions, and alterations that come within the true scope and spirit of the application.

Claims

1. A method of using at least one processor in communication with at least one microphone, the method comprising: The first microphone lobe is deployed toward the first position using the at least one microphone, the first microphone lobe being configured to capture one or more first audio signals from a first audio source located within the first audio pickup area; The second microphone lobe is deployed toward the second position using the at least one microphone, and the second microphone lobe is configured to capture one or more second audio signals from a second audio source located outside the first audio pickup area; as well as The at least one processor removes off-axis noise from the one or more first audio signals by applying a mask determined based on the one or more second audio signals to the one or more first audio signals.

2. The method according to claim 1, further comprising: The mask is calculated using the at least one processor based on the one or more second audio signals and the one or more first audio signals.

3. The method of claim 1, wherein the mask has a value ranging from about zero to about one.

4. The method of claim 1, wherein the off-axis noise comprises audio from the second audio source picked up by the first microphone lobe.

5. The method of claim 1, further comprising: A first audio mix is ​​generated using the one or more first audio signals; A second audio mix is ​​generated using the one or more second audio signals; And the mask is calculated based on the first audio mix and the second audio mix.

6. The method of claim 1, further comprising: Using the at least one processor, the mask is calculated using a neural network based on the one or more first audio signals and the one or more second audio signals.

7. The method of claim 1, further comprising: The first audio source is identified as being located at the first position within the audio pickup area using the at least one microphone; as well as The second audio source is identified as being located at the second location outside the audio pickup area using the at least one microphone.

8. The method of claim 7, further comprising: Receive positioning data for the first audio source and the second audio source from the at least one microphone; as well as The location data is used to determine whether each of the first audio source and the second audio source is within the audio pickup area.

9. A system comprising: At least one microphone, said at least one microphone being configured to: The first microphone lobe is deployed toward the first position to capture one or more first audio signals from a first audio source located within the first audio pickup area, and The second microphone lobe is deployed toward the second position to capture one or more second audio signals from a second audio source located outside the first audio pickup area; as well as At least one processor, communicatively coupled to the at least one microphone, and configured to remove off-axis noise from the one or more first audio signals by applying a mask determined based on the one or more second audio signals to the one or more first audio signals.

10. The system of claim 9, wherein the at least one processor is further configured to calculate the mask based on the one or more second audio signals and the one or more first audio signals.

11. The system of claim 9, wherein the mask has a value ranging from about zero to about one.

12. The system of claim 9, wherein the off-axis noise comprises audio from the second audio source picked up by the first microphone lobe.

13. The system of claim 9, wherein the at least one processor comprises: A first audio mixer, configured to generate a first audio mix using the one or more first audio signals; A second audio mixer, configured to generate a second audio mix using the one or more second audio signals; as well as A source remover, configured to calculate the mask based on the second audio mix and apply the mask to the first audio mix.

14. The system of claim 9, wherein the at least one processor is further configured to use a neural network and to calculate the mask based on the one or more first audio signals and the one or more second audio signals.

15. The system of claim 9, wherein the at least one microphone is further configured to: Based on the positioning data used for the first audio source, the first audio source is identified as being located at the first position within the audio pickup area; and The second audio source is identified as being located at a second location outside the audio pickup area based on the positioning data used for the second audio source.

16. The system of claim 9, further comprising a beamformer configured to deploy the first microphone lobe and the second microphone lobe for the at least one microphone.

17. The system of claim 16, wherein the beamformer is included in the at least one microphone.

18. The system of claim 9, wherein the at least one processor is included in the at least one microphone.

19. A non-transitory computer-readable storage medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform the following operations: A first microphone lobe is deployed toward a first location using at least one microphone, the first microphone lobe being configured to capture one or more first audio signals from a first audio source located within a first audio pickup area; The second microphone lobe is deployed toward the second position using the at least one microphone, and the second microphone lobe is configured to capture one or more second audio signals from a second audio source located outside the first audio pickup area; as well as Off-axis noise is removed from the one or more first audio signals by applying a mask determined based on the one or more second audio signals.

Citation Information

Patent Citations

  • Low latency automixer integrated with voice and noise activity detection

    US11302347B2

  • Auto focus, auto focus within regions, and auto placement of beamformed microphone lobes with inhibition functionality

    US11438691B2

  • System and method for automatic setup of audio coverage area

    US20230224636A1

  • Microphone actuation control system suitable for teleconference systems

    US4658425A

  • Microphone actuation control system

    US5297210A