Autofocus, autofocus within area, and auto configuration of beamforming microphone lobes with suppression
By automatically focusing and configuring the beamforming lobes in response to the detection of sound activity through the array microphone system, the problem of traditional microphones picking up sound when the environment changes is solved, the coverage and quality of the audio source are improved, and the capture of unwanted sounds is reduced.
Patent Information
- Application Number
- CN202410766380.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-07
- Filing Date
- 2020-03-20
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2040-03-20
AI Technical Summary
Traditional array microphones have difficulty automatically adjusting their beamforming lobes to optimally pick up audio sources as the environment changes, and are prone to capturing unwanted audio such as room noise and echoes.
The array microphone system responds to the detection of sound activity to achieve automatic focus and configuration of the beamforming lobe, and suppresses or limits the adjustment of the lobe based on the far-end audio signal to ensure that the lobe picks up the audio source in the best position.
It improves the coverage and quality of audio sources in the environment, reduces the capture of unwanted sounds, and enhances the adaptability and capture accuracy of the microphone system.
Smart Images

Figure CN118803494B_ABST
Abstract
Description
[0001] Information about divisional applications
[0002] This application is a divisional application. The parent application is an invention patent application filed on March 20, 2020, entitled “Autofocus, Autofocus within a Region, and Autoconfiguration of Beamforming Microphone Lobes with Suppression Function,” with application number 202080036963.0.
[0003] CROSS-REFERENCE TO RELATED APPLICATIONS
[0004] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 821,800, filed on March 21, 2019, U.S. Provisional Patent Application No. 62 / 855,187, filed on May 31, 2019, and U.S. Provisional Patent Application No. 62 / 971,648, filed on February 7, 2020. The contents of each application are fully incorporated herein by reference in their entirety. Technical Field
[0005] The present application generally relates to an array microphone with automatic focus and configuration of beamforming microphone lobes. In particular, the present application relates to an array microphone that adjusts the focus and configuration of the beamforming microphone lobes based on detection of voice activity after the lobes have been initially configured, and allows the adjustment of the focus and configuration of the beamforming microphone lobes to be suppressed based on a remote far-end audio signal. Background Art
[0006] Meeting environments, such as conference rooms, boardrooms, video conferencing applications, and the like, may involve the use of microphones to capture sound from various audio sources active in such environments. For example, such audio sources may include a person speaking. The captured sound can be transmitted to a local audience in the environment via amplified speakers (for sound reinforcement), and / or to other people farther away from the environment (e.g., via a television broadcast and / or webcast). The type of microphone and its configuration in a particular environment may depend on the location of the audio source, physical space requirements, aesthetics, room layout, and / or other considerations. For example, in some environments, microphones may be configured on a table or podium near the audio source. In other environments, for example, microphones may be mounted overhead to capture sound from the entire room. Therefore, microphones of various sizes, form factors, mounting options, and wiring options may be used to meet the needs of a particular environment.
[0007] Traditional microphones typically have fixed polar patterns and few manually selectable settings. To capture sound in a conference setting, multiple traditional microphones are used simultaneously to capture audio sources within the environment. However, traditional microphones also tend to capture unwanted audio, such as room noise, echo, and other undesirable audio elements. Using multiple microphones exacerbates the capture of these unwanted noises.
[0008] Array microphones with multiple microphone elements offer benefits such as a steerable coverage or pickup pattern (with one or more lobes), which allows the microphone to focus on desired audio sources and reject undesirable sounds, such as room noise. The ability to manipulate the audio pickup pattern offers the following benefits: the accuracy of the microphone configuration can be reduced, making the array microphone more forgiving. Furthermore, array microphones offer the ability to pick up multiple audio sources with one array microphone or element, again due to the ability to manipulate the pickup pattern.
[0009] However, in certain environments and situations, the position of the lobes of the pickup pattern of the array microphone may not be optimal. For example, the audio source initially detected by the lobes may move and change position. In this case, the lobes may not be able to optimally pick up the audio source in its new position.
[0010] Therefore, there is an opportunity for array microphones to address these issues. More specifically, there is an opportunity for array microphones to automatically focus and / or configure the beamforming microphone lobes based on detection of sound activity after they have been initially configured, while also being able to suppress the focus and / or configuration of the beamforming microphone lobes based on distant far-end audio signals, which can result in higher quality sound capture and better coverage of the environment. Summary of the Invention
[0011] The present invention is directed to solving the above-mentioned problems by providing an array microphone system and method, which are designed, among other things, to: (1) automatically focus the beamforming lobes of the array microphone in response to the detection of sound activity after the lobes have been initially configured; (2) automatically configure the beamforming lobes of the array microphone in response to the detection of sound activity; (3) automatically focus the lobes within the lobe area in response to the detection of sound activity after the beamforming lobes of the array microphone have been initially configured; and (4) suppress or limit the automatic focusing or automatic configuration of the beamforming lobes of the array microphone based on the activity of a far-end audio signal.
[0012] In one embodiment, when new sound activity is detected at new coordinates substantially near the initial coordinates, the beamforming lobe that was positioned at the initial coordinates may be focused by moving the lobe to the new coordinates.
[0013] In another embodiment, the beamforming lobe may be configured or moved to new coordinates when new voice activity is detected at the new coordinates.
[0014] In yet another embodiment, when new sound activity is detected at the new coordinates, the beamforming lobe that has been positioned at the initial position may be focused by moving the lobe, but confined to the lobe area.
[0015] In another embodiment, the movement or configuration of the beamforming lobe may be suppressed or limited when the activity of the far-end audio signal exceeds a predetermined threshold.
[0016] These and other embodiments, as well as various arrangements and aspects, will become apparent and more fully understood from the following detailed description and accompanying drawings, which set forth illustrative embodiments that are indicative of the various ways in which the principles of the invention may be employed. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Schematic diagram of an array microphone with auto-focusing beamforming lobes in response to detection of voice activity, according to some embodiments.
[0018] Figure 2 is a flow chart illustrating operations for auto-focusing a beamforming lobe, according to some embodiments.
[0019] Figure 3 A flow chart illustrating operations for autofocusing of beamforming lobes utilizing a cost functional, according to some embodiments.
[0020] Figure 4 Schematic diagram of a beamforming lobe with an array microphone that automatically configures in response to detection of voice activity, according to some embodiments.
[0021] Figure 5 A flow chart illustrating operations for automatically configuring beamforming lobes according to some embodiments.
[0022] Figure 6 A flow chart illustrating operations for finding a lobe near detected sound activity, according to some embodiments.
[0023] Figure 7 is an exemplary depiction of a microphone having beamforming lobes within a lobe region in accordance with some embodiments.
[0024] Figure 8 is a flow chart illustrating operations for auto-focusing a beamforming lobe within a lobe region, according to some embodiments.
[0025] Figure 9A flow chart illustrating operations for determining whether detected acoustic activity is within the apparent radius of a lobe, according to some embodiments.
[0026] Figure 10 is an exemplary depiction of an array microphone having beamforming lobes within a lobe region and illustrating the apparent radius of the lobes in accordance with some embodiments.
[0027] Figure 11 A flow chart illustrating operations for determining movement of a petal within a movement radius of the petal, according to some embodiments.
[0028] Figure 12 is an exemplary depiction of an array microphone having beamforming lobes within a lobe region and illustrating the movement radius of the lobes in accordance with some embodiments.
[0029] Figure 13 is an exemplary depiction of an array microphone having beamforming lobes within lobe regions and showing boundary pads between lobe regions in accordance with some embodiments.
[0030] Figure 14 A flow chart illustrating operations for limiting lobe movement based on boundary pads between lobe regions, according to some embodiments.
[0031] Figure 15 is an exemplary depiction of an array microphone having beamforming lobes within regions and demonstrating movement of the lobes based on boundary pads between regions, in accordance with some embodiments.
[0032] Figure 16 Schematic diagram of an array microphone with auto-focusing of beamforming lobes in response to detection of voice activity and suppression of auto-focus based on a far-end audio signal, according to some embodiments.
[0033] Figure 17 Schematic diagram of an array microphone with beamforming lobes that automatically configure the array microphone in response to detection of voice activity and suppress the automatically configured array microphone based on a far-end audio signal, according to some embodiments.
[0034] Figure 18 A flow chart illustrating operations for automatically adjusting beamforming lobes of an array microphone based on far-end audio signal suppression, according to some embodiments.
[0035] Figure 19 Schematic diagram of an array microphone with automatic configuration of beamforming lobes of the array microphone in response to detection of voice activity and activity detection of the voice activity, according to some embodiments.
[0036] Figure 20A flow diagram illustrating operations for automatically configuring beamforming lobes including activity detection of voice activity, according to some embodiments. DETAILED DESCRIPTION
[0037] The following description describes, illustrates, and exemplifies one or more specific embodiments of the present invention based on the principles of the present invention. This description is not provided to limit the present invention to the embodiments described herein, but rather to disclose and teach the principles of the present invention in a manner that enables those skilled in the art to understand these principles and, with such understanding, to apply them to practice not only the embodiments described herein but also other embodiments conceivable in accordance with these principles. The scope of the present invention is intended to encompass all such embodiments that may fall within the scope of the appended claims, both literally and under the doctrine of equivalents.
[0038] It should be noted that in the specification and drawings, similar or substantially similar elements may be labeled with the same reference numerals. However, sometimes these elements may be labeled with different numbers, such as when such labeling helps to more clearly describe the situation. In addition, the drawings described herein are not necessarily drawn to scale, and in some cases, the proportions may be exaggerated to more clearly depict certain features. Such labeling and drawing conventions do not necessarily imply an underlying substantive purpose. As stated above, this specification is intended to be interpreted as a whole and in accordance with the principles of the present invention as taught herein and understood by those skilled in the art.
[0039] The array microphone systems and methods described herein can achieve automatic focusing and configuration of beamforming lobes in response to detection of sound activity, as well as allow the focus and configuration of beamforming lobes to be suppressed based on a remote far-end audio signal. In an embodiment, the array microphone may include a plurality of microphone elements, an audio activity locator, a lobe autofocuser, a database, and a beamformer. The audio activity locator can detect the coordinates and confidence score of new sound activity, and the lobe autofocuser can determine whether there is a previously configured lobe near the new sound activity. If such a lobe exists and the confidence score of the new sound activity is greater than the confidence score of the lobe, the lobe autofocuser can transmit the new coordinates to the beamformer so that the lobe moves to the new coordinates. In these embodiments, the position of the lobe can be improved and automatically focused on the latest position of the audio source inside and near the lobe, while also preventing the lobe from overlapping, pointing in an undesirable direction (e.g., towards unwanted noise), and / or moving too suddenly.
[0040] In other embodiments, the array microphone may include multiple microphone elements, an audio activity locator, a lobe autoconfigurator, a database, and a beamformer. The audio activity locator may detect the coordinates of new sound activity, and the lobe autoconfigurator may determine whether a lobe exists near the new sound activity. If no such lobe exists, the lobe autoconfigurator may transmit the new coordinates to the beamformer to cause an inactive lobe to be configured at the new coordinates, or to cause an existing lobe to be moved to the new coordinates. In these embodiments, the set of active lobes of the array microphone may point to the most recent sound activity in the coverage area of the array microphone.
[0041] In other embodiments, the audio activity locator may detect the coordinates and confidence score of new sound activity, and if the confidence score of the new sound activity is greater than a threshold, the lobe autofocuser may identify the lobe region to which the new sound activity belongs. Within the identified lobe region, if the coordinates are within the appearance radius of the current coordinates of the lobe (i.e., the three-dimensional region of space around the current coordinates of the lobe in which the new sound activity may be considered), then the previously configured lobe may be moved. The movement of the lobe within the lobe region may be limited to the movement radius of the current coordinates of the lobe, i.e., the maximum distance the lobe is allowed to move in three-dimensional space, and / or limited to outside the boundary pad between lobe regions, i.e., how close the lobe can move to the boundary between lobe regions. In these embodiments, the position of the lobe can be improved, and autofocus can be maintained on the latest position of the audio source within the lobe region associated with the lobe, while also preventing the lobe from overlapping, pointing in undesirable directions (e.g., toward unwanted noise), and / or moving too suddenly.
[0042] In other embodiments, the activity detector may, for example, receive a remote audio signal from a remote location. The sound of the remote audio signal may be played back in a local environment, such as on a speaker in a conference room. If the activity of the remote audio signal exceeds a predetermined threshold, automatic adjustment of the beamforming lobe (i.e., focus and / or configuration) may be suppressed. For example, the activity of the remote audio signal may be measured by the energy level of the remote audio signal. In this example, the energy level of the remote audio signal may exceed the predetermined threshold when there is a certain level of voice or speech contained in the remote audio signal. In this case, it may be desirable to prevent automatic adjustment of the beamforming lobe so that the lobe is not oriented to pick up sound from the remote audio signal, such as that played back in the local environment. However, if the energy level of the remote audio signal does not exceed the predetermined threshold, automatic adjustment of the beamforming lobe may be performed. Automatic adjustment of the beamforming lobe may include, for example, automatic focus and / or configuration of the lobe as described herein. In these embodiments, the position of the lobe may be refined and automatically focused and / or configured when the activity of the remote audio signal does not exceed a predetermined threshold, and automatically focused and / or configured may be suppressed or limited when the activity of the remote audio signal exceeds a predetermined threshold.
[0043] By using the systems and methods herein, the quality of coverage of an audio source in an environment can be improved by, for example, ensuring that the beamforming lobes optimally pick up the audio source even if the audio source has moved from its initial location and changed position. For example, the quality of coverage of an audio source in an environment can also be improved by reducing the likelihood that the beamforming lobes will be deployed (e.g., focused or configured) to pick up unwanted sounds (such as speech, voice, or other noise from the far end).
[0044] Figure 1 and 4 Schematic diagrams of array microphones 100 and 400 that can detect sounds from audio sources of various frequencies. Array microphones 100 and 400 can be used in a conference room or boardroom, for example, where the audio source may be one or more human speakers. Other sounds that may be undesirable may be present in the environment, such as noise from ventilation equipment, other people, audio / visual equipment, electronic devices, etc. In a typical scenario, the audio source may be seated in a chair at a table, although other configurations and arrangements of the audio source are contemplated and possible.
[0045] Array microphones 100 and 400 can be placed on or in a table, podium, tabletop, wall, ceiling, or the like to detect and capture sounds from an audio source, such as speech from a human speaker. Array microphones 100 and 400 can include, for example, any number of microphone elements 102a, 102b, ..., 102zz, 402a, 402b, ..., 402zz, and can form multiple pickup patterns with lobes to detect and capture sounds from an audio source. Any suitable number of microphone elements 102 and 402 is possible and contemplated.
[0046] Each of the microphone elements 102, 402 in the array microphone 100, 400 can detect sound and convert the sound into an analog audio signal. Components in the array microphone 100, 400, such as an analog-to-digital converter, a processor, and / or other components, can process the analog audio signals and ultimately generate one or more digital audio output signals. In some embodiments, the digital audio output signals may conform to the Dante standard for transmitting audio over Ethernet, or may conform to another standard and / or transmission protocol. In embodiments, each of the microphone elements 102, 402 in the array microphone 100, 400 can detect sound and convert the sound into a digital audio signal.
[0047] The beamformers 170 and 470 in the array microphones 100 and 400 can form one or more pickup patterns based on the audio signals from the microphone elements 102 and 402. The beamformers 170 and 470 can generate digital output signals 190a, 190b, 190c, ..., 190z, 490a, 490b, 490c, ..., 490z corresponding to each pickup pattern. A pickup pattern can be composed of one or more lobes (e.g., a main lobe, a side lobe, and a rear lobe). In other embodiments, the microphone elements 102 and 402 in the array microphones 100 and 400 can output analog audio signals so that other components and devices outside the array microphones 100 and 400 (e.g., a processor, a mixer, a recorder, an amplifier, etc.) can process the analog audio signals.
[0048] Automatic focusing of beamforming lobes in response to detection of voice activity Figure 1 The array microphone 100 may include microphone elements 102; an audio activity locator 150 in wired or wireless communication with the microphone elements 102; a flap autofocus unit 160 in wired or wireless communication with the audio activity locator 150; a beamformer 170 in wired or wireless communication with the microphone elements 102 and the flap autofocus unit 160; and a database 180 in wired or wireless communication with the flap autofocus unit 160. These components will be described in more detail below.
[0049] Automatic configuration of beamforming lobes in response to detection of voice activity Figure 4 The array microphone 400 may include microphone elements 402; an audio activity locator 450 in wired or wireless communication with the microphone elements 402; a lobe autoconfigurator 460 in wired or wireless communication with the audio activity locator 450; a beamformer 470 in wired or wireless communication with the microphone elements 402 and the lobe autoconfigurator 460; and a database 480 in wired or wireless communication with the lobe autoconfigurator 460. These components will be described in more detail below.
[0050] In embodiments, the array microphones 100, 400 may include other components that work with the audio activity locators 150, 450 and / or the beamformers 170, 470, such as an echo canceller or an automatic mixer. For example, as described herein, when a lobe is moved to new coordinates in response to detecting new voice activity, information from the lobe movement can be used by the echo canceller to minimize echo during the movement and / or by the automatic mixer to improve its decision-making capabilities. As another example, the lobe movement can be influenced by the decision of the automatic mixer, such as allowing lobe movement that the automatic mixer has identified as having associated voice activity. The beamformers 170, 470 can be any suitable beamformer, such as a delay-sum beamformer or a minimum variation distortionless response (MVDR) beamformer.
[0051] The various components included in the array microphones 100, 400 may be implemented using software that may be executed by one or more servers or computers, such as a computing device having a processor and memory, a graphics processing unit (GPU), and / or by hardware (e.g., discrete logic circuits, application specific integrated circuits (ASICs), programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.).
[0052] In some embodiments, the microphone elements 102 and 402 may be arranged in concentric rings and / or harmonically nested. In some embodiments, the microphone elements 102 and 402 may be arranged substantially symmetrically. In other embodiments, the microphone elements 102 and 402 may be arranged asymmetrically or in another arrangement. In other embodiments, for example, the microphone elements 102 and 402 may be arranged on a substrate, configured in a frame, or individually suspended. An embodiment of an array microphone is described in commonly assigned U.S. Patent No. 9,565,493, which is hereby incorporated by reference in its entirety. In an embodiment, the microphone elements 102 and 402 may be unidirectional microphones that are primarily sensitive in one direction. In other embodiments, the microphone elements 102 and 402 may have other directional or polar patterns, such as cardioid, subcardioid, or omnidirectional, as desired. The microphone elements 102 and 402 may be any suitable type of sensor that can detect sound from an audio source and convert the sound into an electrical audio signal. In one embodiment, the microphone elements 102, 402 may be microelectromechanical system (MEMS) microphones. In other embodiments, the microphone elements 102, 402 may be condenser microphones, balanced armature microphones, electret microphones, dynamic microphones, and / or other types of microphones. In embodiments, the microphone elements 102, 402 may be arranged in one or two dimensions. The array microphones 100, 400 may be arranged or mounted on a table, wall, ceiling, etc., and may be located, for example, next to, below, or above a video monitor.
[0053] Figure 2, an embodiment of a process 200 for autofocusing previously configured beamforming lobes of array microphone 100 is shown in FIG. Process 200 may be performed by lobe autofocuser 160, causing array microphone 100 to output one or more audio signals 180 from array microphone 100, where audio signals 180 may include sound picked up by beamforming lobes focused on new sound activity from an audio source. One or more processors and / or other processing components (e.g., analog-to-digital converters, encryption chips, etc.) internal or external to array microphone 100 may perform any, some, or all steps of process 200. One or more other types of components (e.g., memory, input and / or output devices, transmitters, receivers, buffers, drivers, discrete components, etc.) may also be used in conjunction with processors and / or other processing components to perform any, some, or all steps of process 200.
[0054] At step 202, coordinates and confidence scores corresponding to new sound activity may be received at the flap autofocuser 160 from the audio activity locator 150. The audio activity locator 150 may continuously scan the environment of the array microphone 100 to find new sound activity. The new sound activity discovered by the audio activity locator 150 may include a suitable audio source, such as a non-stationary human speaker. The coordinates of the new sound activity may be specific three-dimensional coordinates relative to the position of the array microphone 100, such as in Cartesian coordinates (i.e., x, y, z) or in spherical coordinates (i.e., radial distance / magnitude r, elevation angle θ (theta), azimuth angle θ (theta), and so on). ). For example, the confidence score of the new voice activity may represent the certainty of the coordinates and / or the quality of the voice activity. In embodiments, other suitable metrics related to the new voice activity may be received and utilized at step 202. It should be noted that Cartesian coordinates can be easily converted to spherical coordinates, and vice versa, as needed.
[0055] At step 204, the lobe autofocuser 160 may determine whether the coordinates of the new sound activity are near (i.e., in the vicinity of) an existing lobe. Whether the new sound activity is near an existing lobe may be based on the difference in azimuth and / or elevation of (1) the coordinates of the new sound activity and (2) the coordinates of the existing lobe relative to a predetermined threshold. The distance of the new sound activity from the microphone 100 may also affect the determination of whether the coordinates of the new sound activity are near an existing lobe. In some embodiments, the lobe autofocuser 160 may retrieve the coordinates of the existing lobe from the database 180 for use in step 204. Figure 6 An embodiment of determining whether the coordinates of new sound activity are near an existing lobe is described in more detail.
[0056] If the lobe autofocuser 160 determines at step 204 that the coordinates of the new sound activity are not near an existing lobe, then process 200 may end at step 210 and the position of the lobe of array microphone 100 is not updated. In this case, the coordinates of the new sound activity may be considered outside the coverage area of array microphone 100 and, therefore, the new sound activity may be ignored. However, if at step 204 the lobe autofocuser 160 determines that the coordinates of the new sound activity are near an existing lobe, then process 200 continues to step 206. In this case, the coordinates of the new sound activity may be considered an improved (i.e., more focused) position of the existing lobe.
[0057] At step 206, the petal autofocuser 160 may compare the confidence score of the new sound activity with the confidence score of the existing petal. In some embodiments, the petal autofocuser 160 may retrieve the confidence score of the existing petal from the database 180. If the petal autofocuser 160 determines at step 206 that the confidence score of the new sound activity is less than (i.e., inferior to) the confidence score of the existing petal, then the process 200 may end at step 210 and the position in the petal of the array microphone 100 is not updated. However, if the petal autofocuser 160 determines at step 206 that the confidence score of the new sound activity is greater than or equal to (i.e., better than or more favorable to) the confidence score of the existing petal, then the process 200 may continue to step 208. At step 208, the petal autofocuser 160 may transmit the coordinates of the new sound activity to the beamformer 170 so that the beamformer 170 may update the position of the existing petal to the new coordinates. In addition, the petal autofocuser 160 may store the new coordinates of the petal in the database 180.
[0058] In some embodiments, at step 208, petal autofocus 160 may limit the movement of an existing petal to prevent and / or minimize abrupt changes in the petal's position. For example, if a particular petal has recently moved within a certain recent period of time, petal autofocus 160 may not move the petal to the new coordinates. As another example, petal autofocus 160 may not move a particular petal to the new coordinates if the new coordinates are too close to the petal's current coordinates, too close to another petal, overlap with another petal, and / or are considered too far from the petal's existing position.
[0059] The process 200 may be continuously performed by the array microphone 100 as the audio activity locator 150 discovers new sound activity and provides the coordinates and confidence scores of the new sound activity to the lobe autofocuser 160. For example, the process 200 may be performed as an audio source (e.g., a human speaker) moves around a conference room so that one or more lobes can focus on the audio source to optimally pick up its voice.
[0060] Figure 3, an embodiment of a process 300 for autofocusing previously configured beamforming lobes of an array microphone 100 using a cost functional is shown. Process 300 may be performed by lobe autofocuser 160 so that array microphone 100 may output one or more audio signals 180, where audio signals 180 may include sound picked up by beamforming lobes focused on new sound activity from an audio source. One or more processors and / or other processing components (e.g., analog-to-digital converters, encryption chips, etc.) internal or external to microphone array 100 may perform any, some, or all of the steps of process 300. One or more other types of components (e.g., memory, input and / or output devices, transmitters, receivers, buffers, drivers, discrete components, etc.) may also be used in conjunction with processors and / or other processing components to perform any, some, or all of the steps of process 300.
[0061] Steps 302, 304, and 306 of process 300 of flap autofocuser 160 may be the same as those described above. Figure 2 Steps 202, 204, and 206 of process 200 are substantially the same. Specifically, coordinates and confidence scores corresponding to new sound activity may be received at the lobe autofocuser 160 from the audio activity locator 150. The lobe autofocuser 160 may determine whether the coordinates of the new sound activity are near (i.e., in the vicinity of) an existing lobe. If the coordinates of the new sound activity are not near (i.e., in the vicinity of) an existing lobe (or if the confidence score of the new sound activity is less than the confidence score of the existing lobe), the process 300 may proceed to step 324, and the position of the lobe of the array microphone 100 is not updated. However, if at step 306, the lobe autofocuser 160 determines that the confidence score of the new sound activity is greater than (i.e., better than or more favorable than) the confidence score of the existing lobe, the process 300 may continue to step 308. In this case, the coordinates of the new sound activity may be considered as candidate positions to which to move the existing lobe, and the cost functional of the existing lobe may be evaluated and maximized, as described below.
[0062] The cost functional of a petal may take into account the spatial aspects of the petal and the audio quality of the new sound activity. As used herein, cost functional and cost function have the same meaning. Specifically, in some embodiments, the cost functional of a petal i may be defined as the coordinate (LC) of the new sound activity. i ), the signal-to-noise ratio (SNR) of the lobe i ), the gain value of the petal (Gain i ), voice activity detection information (VAD) related to new sound activity i ) and the distance from the coordinates of the existing petal (distance (LO i )). In other embodiments, the cost functional for a petal may be a function of other information. For example, the cost functional for petal i may be expressed as J with Cartesian coordinatesi (x,y,z) or J with spherical coordinates i (azimuth, elevation, magnitude). Taking the cost functional with Cartesian coordinates as an example, the cost functional J i (x,y,z)=f(LC i ,distance(LO i ),Gain i ,SNR i ,VAD i ). Therefore, the flap can be obtained by evaluating and maximizing the cost functional J i The lobe is moved over the spatial grid of coordinates such that the movement of the lobe is in the direction of the gradient (i.e., steepest ascent) of the cost functional. In some cases, the maximum of the cost functional may be the same as the coordinates of the new sound activity (i.e., the candidate location) received by lobe autofocuser 160 at step 302. In other cases, the maximum of the cost functional may move the lobe to a location different from the coordinates of the new sound activity, taking into account the other parameters described above.
[0063] At step 308, the cost functional of the flap may be evaluated by the flap autofocus 160 at the coordinates of the new sound activity. In some embodiments, the evaluated cost functional may be stored by the flap autofocus 160 in the database 180. At step 310, the flap autofocus 160 may move the flap by each of the amounts Δx, Δy, and Δz in the x, y, and z directions, respectively, from the coordinates of the new sound activity. After each movement, the cost functional may be evaluated by the flap autofocus 160 at each of these positions. For example, the flap may be moved to position (x+Δx, y, z), and the cost functional may be evaluated at that position; then moved to position (x, y+Δy, z), and the cost functional may be evaluated at that position; and then moved to position (x, y, z+Δz), and the cost functional may be evaluated at that position. At step 310, the flap may be moved by the amounts Δx, Δy, and Δz in any order. In some embodiments, each of the evaluated cost functionals at these locations may be stored in database 180 by flap autofocus 160. As described below, the evaluation of the cost functional is performed by flap autofocus 160 at step 310 to compute estimates of the partial derivatives and the gradient of the cost functional. It should be noted that while the above description relates to Cartesian coordinates, similar operations may be performed for spherical coordinates (e.g., Δ azimuth, Δ elevation, Δ magnitude).
[0064] At step 312, the gradient of the cost functional may be computed by the flap autofocuser 160 based on the estimated set of partial derivatives. It can be calculated as follows:
[0065]
[0066] At step 314, the petal autofocuser 160 may focus the petal along the gradient calculated at step 312. Specifically, the petal can be moved to a new position: (x i +μgx i ,y i +μgy i ,z i +μgz i ). At step 314, the flap autofocuser 160 may also evaluate a cost functional for the flap at this new position. In some embodiments, this cost functional may be stored by the flap autofocuser 160 in the database 180.
[0067] At step 316, the lobe autofocuser 160 may compare the cost functional of the lobe at the new position (evaluated at step 314) with the cost functional of the lobe at the coordinates of the new sound activity (evaluated at step 308). If, at step 316, the cost functional of the lobe at the new position is less than the cost functional of the lobe at the coordinates of the new sound activity, then the step size μ at step 314 may be considered too large, and the process 300 may continue to step 322. At step 322, the step size may be adjusted, and the process may return to step 314.
[0068] However, if the cost functional of the lobe at the new position is not less than the cost functional of the lobe at the coordinates of the new sound activity at step 316, then process 300 may continue to step 318. At step 318, lobe autofocuser 160 may determine whether the difference between (1) the cost functional of the lobe at the new position (evaluated at step 314) and (2) the cost functional of the lobe at the coordinates of the new sound activity (evaluated at step 308) is close, that is, whether the absolute value of the difference is within a small amount ε. If the condition is not met at step 318, then it can be considered that the local maximum of the cost functional has not been reached. Process 300 may proceed to step 324, and the position of the lobe of array microphone 100 is not updated.
[0069] However, if the condition is met at step 318, then it can be considered that the local maximum of the cost functional has been reached and the lobe has been auto-focused, and process 300 continues to step 320. At step 320, lobe auto-focuser 160 can transmit the coordinates of the new sound activity to beamformer 170 so that beamformer 170 can update the position of the lobe to the new coordinates. In addition, lobe auto-focuser 160 can store the new coordinates of the lobe in database 180.
[0070] In some embodiments, at step 320, the petal autofocuser 160 may apply an annealing / dithering move of the petal. The annealing / dithering move may be applied to nudge the petal out of a local maximum of the cost functional in an attempt to find a better local maximum (and therefore a better position for the petal). The annealing / dithering position may be given by (xi +rx i ,y i +ry i ,z i +rz i ) definition, where (rx i ,ry i ,rz i ) is a small random value.
[0071] The process 300 may be continuously performed by the array microphone 100 as the audio activity locator 150 discovers new sound activity and provides the coordinates and confidence scores of the new sound activity to the lobe autofocuser 160. For example, the process 300 may be performed as an audio source (e.g., a human speaker) moves around a conference room so that one or more lobes can focus on the audio source to optimally pick up its voice.
[0072] In an embodiment, for example, the cost functional can be re-evaluated and updated in steps 308 to 318 and 322, and the coordinates of the lobes can be adjusted, for example, without receiving a set of coordinates of new sound activity at step 302. For example, the algorithm can detect which lobe of the array microphone 100 has the greatest sound activity without providing a set of coordinates of new sound activity. Based on the sound activity information from this algorithm, the cost functional can be re-evaluated and updated.
[0073] Figure 5 An embodiment of a process 500 for automatically configuring or deploying beamforming lobes for an array microphone 400 is shown in FIG. The process 500 may be performed by the lobe autoconfigurator 460 so that the array microphone 400 may automatically configure or deploy beamforming lobes for an array microphone 400. Figure 4 Array microphone 400 shown in FIG outputs one or more audio signals 480, where audio signal 480 may include sound of new sound activity from an audio source picked up by configured beamforming lobes. One or more processors and / or other processing components (e.g., analog-to-digital converters, encryption chips, etc.) internal or external to microphone array 400 may perform any, some, or all steps of process 500. One or more other types of components (e.g., memory, input and / or output devices, transmitters, receivers, buffers, drivers, discrete components, etc.) may also be used in conjunction with processors and / or other processing components to perform any, some, or all steps of process 500.
[0074] At step 502, coordinates corresponding to new sound activity may be received at the lobe autoconfigurator 460 from the audio activity locator 450. The audio activity locator 450 may continuously scan the environment of the array microphone 400 to find new sound activity. The new sound activity discovered by the audio activity locator 450 may include a suitable audio source, such as a non-stationary human speaker. The coordinates of the new sound activity may be specific three-dimensional coordinates relative to the position of the array microphone 400, such as in Cartesian coordinates (i.e., x, y, z) or in spherical coordinates (i.e., radial distance / magnitude r, elevation angle θ (theta), azimuth angle θ (θ)). ).
[0075] In an embodiment, configuration of the beamforming lobes may occur based on whether the amount of new sound activity exceeds a predetermined threshold. Figure 19 FIG1 is a schematic diagram of an array microphone 1900 that can detect sounds from audio sources of various frequencies and automatically configure beamforming lobes in response to detection of sound activity, taking into account the amount of activity of the new sound activity. In an embodiment, array microphone 1900 may include some or all of the same components as array microphone 400 described above, such as microphone 402, audio activity locator 450, lobe autoconfigurator 460, beamformer 470, and / or database 480. Array microphone 1900 may also include an activity detector 1904 in communication with lobe autoconfigurator 460 and beamformer 470.
[0076] Activity detector 1904 can detect the amount of activity in the new sound activity. In some embodiments, the amount of activity can be measured as the energy level of the new sound activity. In other embodiments, the amount of activity can be measured using methods in the time domain and / or frequency domain, such as by applying machine learning (e.g., using cepstral coefficients), measuring signal non-stationarity in one or more frequency bands, and / or searching for characteristics of the desired sound or voice.
[0077] In an embodiment, the activity detector 1904 may be a voice activity detector (VAD) that can determine the presence of speech and / or noise in the remote audio signal. For example, the activity detector 1904 can be implemented by analyzing the spectral variation of the remote audio signal, using linear predictive coding, applying machine learning or deep learning techniques to detect speech and / or noise, and / or using techniques such as ITU G.729 VAD, ETSI standards for VAD calculation as included in the GSM specification, or long-term pitch prediction.
[0078] Based on the amount of activity detected, automatic petal configuration may or may not be performed. When the detected activity of new sound activity meets predetermined criteria, automatic petal configuration may be performed. Conversely, when the detected activity of new sound activity does not meet the predetermined criteria, automatic petal configuration may not be performed. For example, meeting the predetermined criteria may indicate that the new sound activity includes speech, voice, or other sounds that are preferably picked up by a petal. As another example, not meeting the predetermined criteria may indicate that the new sound activity does not include speech, voice, or other sounds that are preferably picked up by a petal. By suppressing automatic petal configuration in this latter case, the petals will not be configured to avoid picking up sounds from the new sound activity.
[0079] As in Figure 20 As shown in process 2000 of FIG. 5 , after step 502 , at step 2003 , it may be determined whether the activity level of the new sound activity meets a predetermined criterion. For example, activity detector 1904 may receive the new sound activity from beamformer 470 . The detected activity level may correspond to the amount of speech, voice, noise, etc. in the new sound activity. In embodiments, the activity level may be measured as the energy level of the new sound activity or as the amount of speech in the new sound activity. In embodiments, the detected activity level may specifically indicate the volume of speech or voice in the new sound activity. In other embodiments, the detected activity level may be a speech-to-noise ratio or indicate the amount of noise in the new sound activity.
[0080] If the amount of activity does not meet the predetermined criteria at step 2003, process 2000 may end at step 522 and the position of the lobes of array microphone 1900 is not updated. When the volume of voice or speech in the new sound activity is relatively low and / or the speech-to-noise ratio is relatively low, the amount of activity detected for the new sound activity may not meet the predetermined criteria. Similarly, when a relatively high amount of noise is present in the new sound activity, the amount of activity detected for the new sound activity may not meet the predetermined criteria. Therefore, not automatically configuring the lobes to detect new sound activity can help ensure that undesirable sounds are not picked up.
[0081] If the amount of activity meets the predetermined criteria at step 2003, process 2000 may continue to step 504 as described below. When the amount of voice or speech in the new sound activity is relatively high and / or the speech-to-noise ratio is relatively high, the amount of activity detected for the new sound activity may meet the predetermined criteria. Similarly, when a relatively low amount of noise is present in the new sound activity, the amount of activity detected for the new sound activity may meet the predetermined criteria. Therefore, in this case, it may be desirable to automatically configure a lobe to detect new sound activity.
[0082] Returning to process 500, at step 504, the flap autoconfigurator 460 may update the timestamp to the current value of, for example, a clock. In some embodiments, the timestamp may be stored in the database 480. In some embodiments, the timestamp and / or clock may be a real-time value, such as hours, minutes, seconds, etc. In other embodiments, the timestamp and / or clock may be based on an increasing integer value, which may enable tracking of the chronological order of events.
[0083] The lobe autoconfigurator 460 may determine in step 506 whether the coordinates of the new sound activity are near (i.e., in the vicinity of) an existing active lobe. Whether the new sound activity is near an existing lobe may be based on the difference in azimuth and / or elevation of (1) the coordinates of the new sound activity and (2) the coordinates of the existing lobe relative to a predetermined threshold. The distance of the new sound activity from the microphone 400 may also affect the determination of whether the coordinates of the new sound activity are near an existing lobe. In some embodiments, the lobe autoconfigurator 460 may retrieve the coordinates of the existing lobe from the database 480 for use in step 506. Figure 6 An embodiment of determining whether the coordinates of new sound activity are near an existing lobe is described in more detail.
[0084] However, if at step 506, the lobe autoconfigurator 460 determines that the coordinates of the new sound activity are near an existing lobe, then the process 500 continues to step 520. At step 520, the timestamp of the existing lobe is updated from step 504 to the current timestamp. In this case, the existing lobe is considered to be able to cover (i.e., pick up) the new sound activity. The process 500 can end at step 522, and the position of the lobe of the array microphone 400 is not updated.
[0085] However, if at step 506, the flap autoconfigurator 460 determines that the coordinates of the new sound activity are near an existing flap, then the process 500 continues to step 508. In this case, the coordinates of the new sound activity can be considered to be outside the current coverage area of the array microphone 400, and therefore the new sound activity needs to be covered. At step 508, the flap autoconfigurator 460 can determine whether an inactive flap of the array microphone 400 is available. In some embodiments, if the flap is not pointing to a specific set of coordinates or if the flap is not deployed (i.e., does not exist), then the flap can be considered inactive. In other embodiments, based on whether a metric (e.g., time, age, etc.) of the deployed flap meets a certain criterion, the deployed flap can be considered inactive. If the flap autoconfigurator 460 determines at step 508 that there is an available inactive flap, then at step 510, the inactive flap is selected, and at step 514, the timestamp of the newly selected flap is updated to the current timestamp (from step 504).
[0086] However, if at step 508 the flap autoconfigurator 460 determines that there are no available inactive flaps, then the process 500 may continue to step 512. At step 512, the flap autoconfigurator 460 may select the currently active flap for recycling to point at the new sound activity coordinates. In some embodiments, the flap selected for recycling may be the active flap with the lowest confidence score and / or the oldest timestamp. The confidence score of the flap may represent, for example, the certainty of the coordinates and / or the quality of the sound activity. In embodiments, other suitable metrics related to the flap may be utilized. The oldest timestamp of the active flap may indicate that no sound activity has been detected recently for that flap, and may indicate that an audio source is no longer present in that flap. The flap selected for recycling at step 512 may have its timestamp updated to the current timestamp (from step 504) at step 514.
[0087] At step 516, a new confidence score may be assigned to the flap, whether it is the selected inactive flap from step 510 or the selected recirculation flap from step 512. At step 518, the flap autoconfigurator 460 may transmit the coordinates of the new acoustic activity to the beamformer 470 so that the beamformer 470 may update the position of the flap to the new coordinates. Additionally, the flap autoconfigurator 460 may store the new coordinates of the flap in the database 480.
[0088] Process 500 may be continuously performed by array microphone 400 as audio activity locator 450 discovers new sound activity and provides the coordinates of the new sound activity to lobe autoconfigurator 460. For example, process 500 may be performed as an audio source (e.g., a human speaker) moves around a conference room so that one or more lobes can be configured to optimally pick up the audio source's sound.
[0089] exist Figure 6 An embodiment of a process 600 for finding a previously configured lobe near a sound activity is shown in FIG. Process 600 may be used by lobe autofocuser 160 at step 204 of process 200, at step 304 of process 300, and / or at step 806 of process 800, and / or by autoconfigurator 460 at step 506 of process 500. Specifically, process 600 may determine whether the coordinates of a new sound activity are near an existing lobe of array microphone 100, 400. Whether the new sound activity is near an existing lobe may be based on a difference in azimuth and / or elevation of (1) the coordinates of the new sound activity and (2) the coordinates of the existing lobe relative to a predetermined threshold. The distance of the new sound activity from the array microphone 100, 400 may also affect the determination of whether the coordinates of the new sound activity are near an existing lobe.
[0090] At step 602, coordinates corresponding to new sound activity may be received from the audio activity locator 150, 450 at the flap autofocuser 160 or the flap autoconfigurer 460, respectively. The coordinates of the new sound activity may be specific three-dimensional coordinates relative to the position of the array microphone 100, 400, such as in Cartesian coordinates (i.e., x, y, z) or in spherical coordinates (i.e., radial distance / magnitude r, elevation angle θ (theta), azimuth angle θ (θ)). ). It should be noted that Cartesian coordinates can be easily converted to spherical coordinates, and vice versa, as needed.
[0091] At step 604, the flap autofocuser 160 or flap autoconfigurer 460 may determine whether the new sound activity is relatively far from the array microphone 100, 400 by evaluating whether the distance of the new sound activity is greater than a determined threshold. The distance of the new sound activity may be determined by the magnitude of a vector representing the coordinates of the new sound activity. If it is determined at step 604 that the new sound activity is relatively far from the array microphone 100, 400 (i.e., greater than the threshold), then at step 606, a lower azimuth threshold may be set for later use in process 600. If it is determined at step 604 that the new sound activity is not relatively far from the array microphone 100, 400 (i.e., less than or equal to the threshold), then at step 608, a higher azimuth threshold may be set for later use in process 600.
[0092] After setting the azimuth threshold at step 606 or step 608, process 600 may continue to step 610. At step 610, lobe autofocuser 160 or lobe autoconfigurer 460 may determine whether there are any lobes to be checked for proximity to new sound activity. If no lobes of array microphone 100, 400 are to be checked at step 610, process 600 may end at step 616, indicating that no lobes are near array microphone 100, 400.
[0093] However, if there is a lobe of the array microphone 100, 400 to be checked at step 610, then the process 600 may continue to step 612 and check one of the existing lobes. At step 612, the lobe autofocuser 160 or the lobe autoconfigurer 460 may determine whether the absolute value of the difference between (1) the azimuth of the existing lobe and (2) the azimuth of the new sound activity is greater than the azimuth threshold (the azimuth threshold set at step 606 or step 608). If the condition is met at step 612, then the lobe being checked may be considered not to be near the new sound activity. The process 600 may return to step 610 to determine whether there are other lobes to be checked.
[0094] However, if the condition is not met at step 612, then process 600 may proceed to step 614. At step 614, the lobe autofocuser 160 or the lobe autoconfigurer 460 may determine whether the absolute value of the difference between (1) the elevation angle of the existing lobe and (2) the elevation angle of the new sound activity is greater than a predetermined elevation angle threshold. If the condition is met at step 614, then the lobe under examination may be considered not to be in the vicinity of the new sound activity. Process 600 may return to step 610 to determine whether there are other lobes to be examined. However, if the condition is not met in step 614, then process 600 may end at step 618, indicating that the lobe under examination is in the vicinity of the new sound activity.
[0095] Figure 7 is an exemplary depiction of an array microphone 700 that can automatically focus previously configured beamforming lobes within associated lobe regions in response to detecting new sound activity. In embodiments, the array microphone 700 may include some or all of the same components as the array microphone 100 described above, such as the audio activity locator 150, the lobe autofocuser 160, the beamformer 170, and / or the database 180. Each lobe of the array microphone 700 can move within its associated lobe region, and the lobe may not cross the boundaries between lobe regions. It should be noted that although Figure 7 Eight lobes with eight associated lobe areas are depicted, but any number of lobes and associated lobe areas is possible and contemplated, such as Figure 10 、 12 , 13 and 15 depict four lobes with four associated lobe regions. It should also be noted that Figure 7 、 10 , 12, 13 and 15 are depicted as two-dimensional representations of the three-dimensional space around the array microphone.
[0096] At least two sets of coordinates may be associated with each lobe of the array microphone 700: (1) original or initial coordinates LOi (e.g., coordinates automatically or manually constructed when the array microphone 700 is set up), and (2) current coordinates to which the current lobe is pointing at a given time. In some embodiments, the set of coordinates may indicate the location of the center of the flap. In some embodiments, the set of coordinates may be stored in database 180.
[0097] Additionally, each lobe of the array microphone 700 may be associated with a lobe region in three-dimensional space surrounding it. In an embodiment, a lobe region may be defined as a set of points in space that are closer to the initial coordinates LO of the lobe than to the coordinates of any other lobe of the array microphone. i In other words, if p is defined as a point in space, then the point p and lobe i (LO i) is smallest compared to any other lobe, a point p may belong to a specific lobe region LR i , as in the following formula: A region defined in this way is called a Voronoi region or a Voronoi cell. For example, Figure 7 As can be seen in FIG, there are eight lobes with associated lobe regions with boundaries drawn between each of the lobe regions. A boundary between lobe regions is a set of points in space that are equidistant from two or more adjacent lobes. Some edges of the lobe regions may also be unbounded. In an embodiment, the distance D may be the distance between point p and LO. i (For example, In some embodiments, the lobe area can be recalculated as a particular lobe moves.
[0098] In embodiments, the lobe regions may be calculated and / or updated based on sensing the environment (e.g., objects, walls, people, etc.) in which the array microphone 700 is located using infrared sensors, visual sensors, and / or other suitable sensors. For example, the array microphone 700 may use information from the sensors to set approximate boundaries for the lobe regions, which in turn may be used to configure the associated lobe. In other embodiments, the lobe regions may be calculated and / or updated based on user-defined lobe regions, such as through a graphical user interface of the array microphone 700.
[0099] like Figure 7 As further shown in FIG, and described below, there may be various parameters associated with each lobe that may constrain its movement during the autofocus process. One parameter is the apparent radius of the lobe, which is the apparent radius at the initial coordinate LO of the lobe. i The three-dimensional region of space surrounding a lobe in which new sound activity may be considered. In other words, if new sound activity is detected in a lobe region, but outside the apparent radius of the lobe, there will not be any movement or autofocus of the lobe in response to the detection of the new sound activity. Thus, points outside the apparent radius of a lobe may be considered to be ignored or "irrelevant" parts of the associated lobe region. For example, in Figure 7 , point A is outside the apparent radius of lobe 5 and its associated lobe region 5, so any new sound activity at point A will not cause the lobe to move. Conversely, if new sound activity is detected in a particular lobe region and is within the apparent radius of its lobe, the lobe can be automatically moved and focused in response to the detection of the new sound activity.
[0100] Another parameter is the petal movement radius, which is the maximum distance in space that the petal is allowed to move. The petal movement radius is usually smaller than the petal's apparent radius and can be set to prevent the petal from moving too far from the array microphone or from the petal's initial coordinate LO. i Too far. For example, Figure 7In FIG, the point denoted as B is within both the apparent radius and the moving radius of the lobe 5 and its associated lobe area 5. If new sound activity is detected at point B, the lobe 5 can be moved to point B. As another example, in Figure 7 , the point denoted as C is within the appearance radius of the petal 5, but outside the movement radius of the petal 5 and its associated petal area 5. If new sound activity is detected at point C, the maximum distance the petal 5 can move is limited to the movement radius.
[0101] Another parameter is the petal boundary pad, which is the maximum distance in space that the petal is allowed to move toward adjacent petal areas and toward the boundaries between petal areas. Figure 7 In , the point indicated as D is outside the boundary pad of flap 8 and its associated flap area 8 (adjacent to flap area 7). The boundary pads of the flaps can be arranged to minimize the overlap of adjacent flaps. Figure 7 、 10 In Figures 1, 12, 13, and 15, the boundaries between the lobe regions are represented by dotted lines, and the boundary pad of each lobe region is represented by a dot-dash line parallel to the boundary.
[0102] Figure 8 , an embodiment of a process 800 for autofocusing previously configured beamforming lobes of array microphone 700 within associated lobe regions is shown. Process 800 may be performed by lobe autofocuser 160, causing array microphone 700 to output one or more audio signals 180 from array microphone 700, where audio signals 180 may include sound picked up by beamforming lobes focused on new sound activity from an audio source. One or more processors and / or other processing components (e.g., analog-to-digital converters, encryption chips, etc.) internal or external to array microphone 700 may perform any, some, or all of the steps of process 800. One or more other types of components (e.g., memory, input and / or output devices, transmitters, receivers, buffers, drivers, discrete components, etc.) may also be used in conjunction with processors and / or other processing components to perform any, some, or all of the steps of process 800.
[0103] Step 802 of process 800 for the flap autofocuser 160 may be the same as described above. Figure 2800 , the position of the lobes of the array microphone 700 is substantially the same as step 202 of process 200. Specifically, at step 802, coordinates and a confidence score corresponding to the new sound activity may be received from the audio activity locator 150 at the lobe autofocus 160. In an embodiment, other suitable metrics related to the new sound activity may be received and utilized at step 802. At step 804, the lobe autofocus 160 may compare the confidence score of the new sound activity to a predetermined threshold to determine whether the new confidence score is satisfactory. If the lobe autofocus 160 determines at step 804 that the confidence score of the new sound activity is less than the predetermined threshold (i.e., the confidence score is unsatisfactory), then the process 800 may end at step 820 and the position of the lobes of the array microphone 700 is not updated. However, if the lobe autofocus 160 determines at step 804 that the confidence score of the new sound activity is greater than or equal to the predetermined threshold (i.e., the confidence score is satisfactory), then the process 800 may continue to step 806.
[0104] At step 806, the petal autofocuser 160 may identify the petal region where the new sound activity is located, that is, the petal region to which the new sound activity belongs. In an embodiment, at step 806, the petal autofocuser 160 may find the petal closest to the coordinates of the new sound activity in order to identify the petal region. For example, the petal autofocuser 160 may find the initial coordinates LO of the petal closest to the new sound activity. i To identify the lobe region, for example, by finding the lobe index i such that the coordinates of the new sound activity are equal to the initial coordinates LO of the lobe i Minimize the distance between: The lobe and its associated lobe region that includes the new sound activity may be determined as the lobe and lobe region identified at step 806 .
[0105] After the lobe region has been identified at step 806, lobe autofocuser 160 may determine whether the coordinates of the new sound activity are outside the apparent radius of the lobe at step 808. If lobe autofocuser 160 determines at step 808 that the coordinates of the new sound activity are outside the apparent radius of the lobe, process 800 may end at step 820 and the position of the lobe of array microphone 700 is not updated. In other words, if the new sound activity is outside the apparent radius of the lobe, the new sound activity may be ignored and may be considered to be outside the coverage of the lobe. As an example, Figure 7 Point A in is within the petal region 5 associated with petal 5, but outside the apparent radius of petal 5. Figure 9 and 10 Describes the details of determining whether the coordinates of new sound activity are outside the appearance radius of the lobe.
[0106] However, if at step 808, the lobe autofocuser 160 determines that the coordinates of the new sound activity are not outside the apparent radius of the lobe (i.e., inside it), then the process 800 may continue to step 810. In this case, as described below, the lobe may be moved toward the new sound activity based on evaluating the coordinates of the new sound activity relative to other parameters (such as the movement radius and the boundary pad). At step 810, the lobe autofocuser 160 may determine whether the coordinates of the new sound activity are outside the movement radius of the lobe. If at step 810, the lobe autofocuser 160 determines that the coordinates of the new sound activity are outside the movement radius of the lobe, then the process 800 may continue to step 816, where the movement of the lobe may be restricted or limited. In particular, at step 816, the new coordinates to which the lobe may be temporarily moved may be set to no greater than the movement radius. As described below, the new coordinates may be temporary because the movement of the lobe may still be evaluated relative to the boundary pad parameters. In an embodiment, the movement of the petal at step 816 may be limited based on a scaling factor α (where 0<α≤1) in order to prevent the petal from moving away from its initial coordinate LO i Too far. As an example, Figure 7 Point C in is outside the movement radius of petal 5, so the farthest distance petal 5 can move is the movement radius. After step 816, process 800 can continue to step 812. Figure 11 and 12 Describe the details of limiting the movement of the flap to within its radius of motion.
[0107] If at step 810, the lobe autofocuser 160 determines that the coordinates of the new sound activity are not outside (i.e., inside) the lobe's movement radius, the process 800 may also continue to step 812. As an example, Figure 7 Point B in is inside the movement radius of petal 5, so petal 5 can be moved to point B. At step 812, the petal autofocuser 160 may determine whether the coordinates of the new sound activity are close to the boundary pad and therefore too close to the adjacent petal. If the petal autofocuser 160 determines at step 812 that the coordinates of the new sound activity are close to the boundary pad, then the process 800 may continue to step 818, where the movement of the petal may be restricted or limited. Specifically, at step 818, the new coordinates to which the petal can be moved may be set to be just outside the boundary pad. In an embodiment, the movement of the petal at step 818 may be limited based on a scaling factor β (where 0<β≤1). As an example, Figure 7 Point D in is outside the boundary pad between adjacent lobe region 8 and lobe region 7. Process 800 may continue to step 814 after step 818. Figures 13 to 15 Describes details about the boundary pad.
[0108] If the lobe autofocuser 160 determines at step 812 that the coordinates of the new sound activity are not close to the boundary pad, then the process 800 may also continue to step 814. At step 812, the lobe autofocuser 160 may transmit the new coordinates of the lobe to the beamformer 170 so that the beamformer 170 may update the position of the existing lobe to the new coordinates. In an embodiment, the new coordinates of the lobe Can be defined as in is the motion vector, and is a constrained motion vector, as described in more detail below. In an embodiment, the lobe autofocuser 160 may store the new coordinates of the lobe in the database 180 .
[0109] Depending on the steps of process 800 described above, when the lobe moves due to detection of new sound activity, the new coordinates of the lobe may be: (1) the coordinates of the new sound activity if the coordinates of the new sound activity are within the apparent radius of the lobe, within the moving radius of the lobe, and not close to the boundary pad of the associated lobe area; (2) a point in the direction of the motion vector toward the new sound activity, and the point is limited to the range of the moving radius if the coordinates of the new sound activity are within the apparent radius of the lobe, outside the moving radius of the lobe, and not close to the boundary pad of the associated lobe area; or (3) just outside the boundary pad if the coordinates of the new sound activity are within the apparent radius of the lobe and close to the boundary pad.
[0110] The process 800 may be continuously performed by the array microphone 700 as the audio activity locator 150 discovers new sound activity and provides the coordinates and confidence scores of the new sound activity to the lobe autofocuser 160. For example, the process 800 may be performed as an audio source (e.g., a human speaker) moves around a conference room so that one or more lobes can focus on the audio source to optimally pick up its voice.
[0111] exist Figure 9 An embodiment of a process 900 for determining whether the coordinates of a new sound activity are outside the apparent radius of a lobe is shown in FIG. For example, process 900 may be used by lobe autofocuser 160 at step 808 of process 800. Specifically, process 900 may begin at step 902, where a motion vector may be Calculated as The motion vector can be the original coordinate LO of the lobe i The center of the connection is connected to the coordinates of the new sound activity For example, if Figure 10 As shown in FIG, the new sound activity S exists in lobe region 3, and the motion vector is shown between the original coordinates LO3 of the lobe 3 and the coordinates of the new sound activity S. The apparent radius of the lobe 3 is also depicted at Figure 10middle.
[0112] At step 902, the motion vector is calculated Thereafter, process 900 may continue to step 904. At step 904, lobe autofocuser 160 may determine whether the magnitude of the motion vector is greater than the apparent radius of the lobe, as in the following equation: If the motion vector If the magnitude of is greater than the apparent radius of the lobe at step 904, then at step 906 the coordinates of the new sound activity may be represented as being outside the apparent radius of the lobe. Figure 10 As shown in FIG, since the new sound activity S is outside the appearance radius of lobe 3, the new sound activity S will be ignored. However, if the motion vector If the magnitude of is less than or equal to the apparent radius of the lobe, then the coordinates of the new sound activity at step 908 may be represented as being inside the apparent radius of the lobe.
[0113] exist Figure 11 An embodiment of a process 1100 for limiting the movement of a petal within its movement radius is shown in FIG. For example, process 1100 may be used by petal autofocus 160 at step 816 of process 800. Specifically, process 1100 may begin at step 1102, where a motion vector Calculated as Similar to the above Figure 9 For example, as shown in step 902 of process 900. Figure 12 As shown in FIG, the new sound activity S exists in lobe region 3, and the motion vector is shown between the original coordinates LO3 of the lobe 3 and the coordinates of the new sound activity S. The apparent radius of the lobe 3 is also depicted at Figure 12 middle.
[0114] At step 1102, the motion vector is calculated Thereafter, process 1100 may continue to step 1104. At step 1104, flap autofocuser 160 may determine the motion vector Is the value of less than or equal to the moving radius of the petal, as in the following formula: If at step 1104 the motion vector If the magnitude of is less than or equal to the movement radius, then at step 1106, the new coordinates of the lobe may be temporarily moved to the coordinates of the new sound activity. Figure 12 As shown in , since the new sound activity S is inside the movement radius of the lobe 3, the lobe will temporarily move to the coordinates of the new sound activity S.
[0115] However, if at step 1104 the motion vector If the magnitude of the motion vector is greater than the movement radius, then at step 1108, the magnitude of the motion vector can be scaled to the maximum value of the movement radius by a scaling factor α while maintaining the same direction, as in the following formula: where the scaling factor α can be defined as:
[0116]
[0117] Figures 13 to 15 It relates to a boundary pad of the lobe region, which is a part of the space adjacent to the boundary or edge of another lobe region close to the lobe region. Specifically, a vector i connecting the original coordinates of two lobes (i.e., LOand LO j ) can be used to indirectly describe the boundary pad near the boundary between two lobes i and j. Therefore, this vector can be described as: The midpoint of this vector can be a point at the boundary between two lobe regions. Specifically, moving along the direction of the vector from the original coordinates LO of lobe i i is the shortest path towards the adjacent lobe j. In addition, moving along the direction of the vector from the original coordinates LO of lobe i i but keeping the amount of movement as half of the magnitude of the vector will be the exact boundary between the two lobe regions. Based on the above, moving along the direction of the vector
[0118] from the original coordinates LO of lobe i i but restricting the amount of movement based on the value A (where 0 < A < 1) (i.e., ) will be within (100 * A)% of the boundary between the lobe regions. For example, if A is 0.8 (i.e., 80%), then the new coordinates of the moving lobe will be within 80% of the boundary between the lobe regions. Therefore, the value A can be used to create a boundary pad between two adjacent lobe regions. Generally, a larger boundary pad can prevent the lobe from moving into another lobe region, while a smaller boundary pad can allow the lobe to move closer to another lobe region.
[0119] In addition, it should be noted that if lobe i moves in the direction towards lobe j due to the detection of new voice activity (e.g., along the direction of the motion vector described above), then there is a movement component in the direction of lobe j (i.e., along the direction of the vector ). To find the movement component in the direction of the vector , the motion vector can be projected onto the unit vector (which has the same direction as the vector with unit magnitude)same direction) to calculate the projection vector As an example, Figure 13 Shows the vector connecting petals 3 and 2 This is also the shortest path from the center of the petal 3 towards the petal area 2 . Figure 13 The projection vector shown in is the motion vector In the unit vector Projection on .
[0120] Figure 14 An embodiment of a process 1400 for creating a boundary pad of a lobe region using vector projection is shown in FIG. For example, process 1400 may be used by lobe autofocuser 160 at step 818 of process 800. Process 1400 may result in limiting motion vectors The magnitude is such that the movement of the lobe in the direction of any other lobe region does not exceed a certain percentage of the size that characterizes the boundary pad.
[0121] Prior to executing process 1400, vectors may be calculated for all pairs of active lobes. and unit vectors As previously described, the vector The original coordinates of petals i and j can be connected. The parameter A can be determined for all active petals i (where 0 i <1), where the parameter characterizes the size of the boundary pad of each lobe region. As previously described, before performing process 1400 (i.e., before step 818 of process 800), lobe regions of new sound activity can be identified (i.e., at step 806) and motion vectors can be calculated (i.e., using process 1100 / step 810).
[0122] At step 1402 of process 1400, projection vectors may be calculated for all lobes not associated with the lobe region identified for new sound activity. Projection vector The size (as mentioned above about Figure 13 The projection vector This magnitude can be calculated as a scalar, such as by the motion vector With unit vector The dot product is calculated so that the projection PM ij =M y Du ij,z +M y Du ij,y +M z Du ij,z .
[0123] When PM ij <0, the motion vector With vector This means that the motion of petal i will be in the opposite direction to the boundary of petal j. In this case, the boundary pad between petals i and j is not a problem because the motion of petal i will be away from the boundary with petal j. However, in PM ij >0, the motion vector With vector This means that the movement of petal i will be in the same direction as the boundary of petal j. In this case, the movement of petal i can be restricted outside the boundary pad so that Among them A i (where 0 i <1) is a parameter characterizing the boundary pad of the lobe area associated with lobe i.
[0124] The scaling factor β can be used to ensure The scaling factor β can be used to scale the motion vector and defined as Therefore, if new sound activity is detected outside the boundary pad of the lobe area, the scaling factor β may be equal to 1, indicating that no motion vector At step 1404, a scaling vector β may be calculated for all lobes that are not associated with a lobe region identified for new sound activity.
[0125] At step 1406, a minimum scaling factor β corresponding to the boundary pad of the nearest lobe region may be determined, as in the following equation: After the minimum scaling factor β has been determined at step 1406, then at step 1408, the minimum scaling factor β may be applied to the motion vector To determine the constrained motion vector
[0126] For example, Figure 15 Shows the new sound activity S present in the lobe area 3 and the motion vector between the initial coordinates LO3 of the lobe 3 and the coordinates of the new sound activity S vector and the projection vector is depicted between petal 3 and each of the other petals not associated with petal region 3 (ie, petals 1, 2, and 4). Specifically, the vector φ can be calculated for all pairs of active petals (ie, petals 1, 2, 3, and 4). and calculate projections PM for all lobes not associated with lobe region 3 (identified for new sound activity S) 31 、PM 32 、PM 34 The magnitude of the projection vector can be used to calculate the scaling factor β, and the minimum scaling factor β can be used to scale the motion vector Therefore, the motion vector It may be restricted to outside the boundary pad of lobe region 3 because the new sound activity S is too close to the boundary between lobe 3 and lobe 2. Based on the restricted motion vector, the coordinates of lobe 3 can be moved to coordinates S outside the boundary pad of lobe region 3. r .
[0127] Figure 15 The projection vector depicted in is negative, and the corresponding scaling factor β4 (for petal 4) is equal to 1. The scaling factor β1 (for petal 1) is also equal to 1, because The scaling factor β2 (for lobe 2) is less than 1 because the new sound activity S is inside the boundary pad between lobe 2 and lobe 3 (i.e., ). Therefore, a minimum scaling factor β2 can be used to ensure that the flap 3 moves to the coordinate S r .
[0128] Figure 16 and 17 Schematic diagram of array microphones 1600 and 1700 that can detect sounds from audio sources of various frequencies. Figure 16 The array microphone 1600 can automatically focus the beamforming lobe in response to the detection of sound activity, while being able to suppress the autofocus of the beamforming lobe when the activity of the remote audio signal from the far end exceeds a predetermined threshold. In an embodiment, the array microphone 1600 may include some or all of the same components as the array microphone 100 described above, such as the microphone 102, the audio activity locator 150, the lobe autofocuser 160, the beamformer 170, and / or the database 180. The array microphone 1600 may also include a sensor 1602, such as a speaker, and an activity detector 1604 in communication with the lobe autofocuser 160. The remote audio signal from the far end may be communicated with the sensor 1602 and the activity detector 1604.
[0129] Figure 17The array microphone 1700 can automatically configure beamforming lobes in response to detection of sound activity, while also being able to suppress automatic configuration of the beamforming lobes when the activity of the remote audio signal from the far end exceeds a predetermined threshold. In an embodiment, the array microphone 1700 may include some or all of the same components as the array microphone 400 described above, such as the microphone 402, the audio activity locator 450, the lobes automatic configurator 460, the beamformer 470, and / or the database 480. The array microphone 1700 may also include a sensor 1702, such as a speaker, and an activity detector 1704 in communication with the lobes automatic configurator 460. The remote audio signal from the far end may be communicated with the sensor 1702 and the activity detector 1704.
[0130] Sensors 1602 and 1702 can be used to play the sound of the remote audio signal in the local environment where the array microphones 1600 and 1700 are located. Activity detectors 1604 and 1704 can detect the amount of activity in the remote audio signal. In some embodiments, the amount of activity can be measured as the energy level of the remote audio signal. In other embodiments, methods in the time domain and / or frequency domain can be used to measure the amount of activity, such as by applying machine learning (e.g., using cepstral coefficients), measuring signal non-stationarity in one or more frequency bands, and / or searching for characteristics of a desired sound or voice.
[0131] In one embodiment, activity detectors 1604 and 1704 may be voice activity detectors (VADs) that can determine the presence of speech in a remote audio signal. For example, speech can be detected by analyzing spectral variations of the remote audio signal, using linear predictive coding, applying machine learning or deep learning techniques, and / or using methods such as ITU G.729 VAD, ETSI standards for VAD calculation included in the GSM specification, or long-term pitch prediction.
[0132] Based on the amount of detected activity, automatic lobe adjustment may be performed or suppressed. As described herein, automatic lobe adjustment may include, for example, automatic focusing of the lobe, automatic focusing of the lobe within a region, and / or automatic configuration of the lobe. Automatic lobe adjustment may be performed when the detected activity of the remote audio signal does not exceed a predetermined threshold. Conversely, automatic lobe adjustment may be suppressed (i.e., not performed) when the detected activity of the remote audio signal exceeds a predetermined threshold. For example, exceeding the predetermined threshold may indicate that the remote audio signal includes speech, voice, or other sounds that are preferably not picked up by the lobe. By suppressing automatic lobe adjustment in this case, the lobe will not be focused or configured to avoid picking up sounds from the remote audio signal.
[0133] In some embodiments, the activity detectors 1604 and 1704 may determine whether the amount of activity detected in the remote audio signal exceeds a predetermined threshold. When the amount of activity detected does not exceed the predetermined threshold, the activity detectors 1604 and 1704 may transmit an enable signal to the flap autofocuser 160 or the flap autoconfigurer 460, respectively, to allow the flap to be adjusted. Additionally or alternatively, when the amount of activity detected in the remote audio signal exceeds the predetermined threshold, the activity detectors 1604 and 1704 may transmit a pause signal to the flap autofocuser 160 or the flap autoconfigurer 460, respectively, to prevent the flap from being adjusted.
[0134] In other embodiments, activity detectors 1604, 1704 may transmit the amount of activity detected in the remote audio signal to flap autofocuser 160 or flap autoconfigurator 460, respectively. Flap autofocuser 160 or flap autoconfigurator 460 may determine whether the amount of activity detected exceeds a predetermined threshold. Based on whether the amount of activity detected exceeds the predetermined threshold, flap autofocuser 160 or flap autoconfigurator 460 may execute or pause adjustments to the flaps.
[0135] The various components included in the array microphones 1600, 1700 may be implemented using software that may be executed by one or more servers or computers, such as a computing device having a processor and memory, a graphics processing unit (GPU), and / or by hardware (e.g., discrete logic circuits, application specific integrated circuits (ASICs), programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.).
[0136] Figure 18 , an embodiment of a process 1800 for automatically adjusting the beamforming lobes of an array microphone based on far-end audio signal suppression is shown. Process 1800 can be performed by array microphones 1600 and 1700 to enable automatic focusing or automatic configuration of the beamforming lobes to be performed or suppressed based on the amount of activity of the far-end audio signal from the far end. One or more processors and / or other processing components (e.g., analog-to-digital converters, encryption chips, etc.) internal or external to array microphones 1600 and 1700 can perform any, some, or all of the steps of process 1800. One or more other types of components (e.g., memory, input and / or output devices, transmitters, receivers, buffers, drivers, discrete components, etc.) can also be used in conjunction with processors and / or other processing components to perform any, some, or all of the steps of process 1800.
[0137] At step 1802, a remote audio signal may be received at the array microphones 1600 and 1700. The remote audio signal may originate from a remote end (e.g., a remote location) and may include sounds from the remote end (e.g., voice, speech, noise, etc.). The remote audio signal may be output on the sensors 1602 and 1702 (e.g., speakers in a local environment) at step 1804. Thus, the sounds from the remote end may be played in the local environment, such as during a conference call, so that local participants can hear the remote participants.
[0138] The remote audio signal may be received by activity detectors 1604 and 1704, which may detect the amount of activity in the remote audio signal at step 1806. The detected amount of activity may correspond to the amount of voice, speech, noise, etc. in the remote audio signal. In embodiments, the amount of activity may be measured as the energy level of the remote audio signal. At step 1808, if the detected amount of activity in the remote audio signal does not exceed a predetermined threshold, process 1800 may continue to step 1810. The fact that the detected amount of activity in the remote audio signal does not exceed the predetermined threshold may indicate that a relatively small amount of voice, speech, noise, etc. is present in the remote audio signal. In embodiments, the detected amount of activity may specifically indicate the amount of speech or speech in the remote audio signal. At step 1810, lobe adjustment may be performed. Step 1810 may include, for example, processes 200 and 300 for automatically focusing a beamforming lobe, process 400 for automatically configuring a beamforming lobe, and / or process 800 for automatically focusing a beamforming lobe within a lobe region, as described herein. Lobe adjustment may be performed in this case because, even though the lobe may be focused or configured, there is a smaller likelihood that such lobe will pick up undesirable sounds from the remote audio signal being output in the local environment. After step 1810, process 1800 may return to step 1802.
[0139] However, if the amount of activity of the remote audio signal detected at step 1808 exceeds a predetermined threshold, process 1800 may continue to step 1812. At step 1812, no lobe adjustment is performed, that is, lobe adjustment may be suppressed. The amount of activity of the remote audio signal detected exceeding the predetermined threshold may indicate that a relatively high amount of voice, speech, noise, etc. is present in the remote audio signal. In this case, suppressing lobe adjustment from occurring may help ensure that the lobe is not focused or configured to pick up sound from the remote audio signal output from the local environment. In some embodiments, process 1800 may return to step 1802 after step 1812. In other embodiments, process 1800 may wait for a specific duration at step 1812 before returning to step 1802. Waiting for the specific duration may allow reverberation in the local environment (e.g., caused by the sound of the remote audio signal being played) to dissipate.
[0140] Process 1800 can be continuously performed by array microphones 1600 and 1700 when a far-end audio signal is received from the far end. For example, the far-end audio signal may include a low amount of activity (e.g., no speech or voice) that does not exceed a predetermined threshold. In this case, lobe adjustment can be performed. As another example, the far-end audio signal may include a high amount of activity (e.g., speech or voice) that exceeds a predetermined threshold. In this case, lobe adjustment may be suppressed. Therefore, whether lobe adjustment is performed or suppressed can change as the amount of activity in the far-end audio signal changes. Process 1800 can result in better sound pickup in the local environment by reducing the possibility of undesirable sound pickup from the far end.
[0141] Any process descriptions or blocks in the figures should be understood to represent modules, segments, or portions of code, which include one or more executable instructions for implementing specific logical functions or steps in the process, and alternative implementations are included within the scope of embodiments of the present invention, in which functions may be performed out of the order from that shown or discussed, including substantially concurrently or in reverse order depending on the functions involved, as will be understood by one skilled in the art.
[0142] The present invention is intended to explain how to make and use various embodiments in accordance with the present technology, not to limit the true, intended, and fair scope and spirit thereof. The foregoing description is not intended to be exhaustive or limited to any precise form disclosed. Modifications or variations are possible in light of the above teachings. The embodiments are selected and described to provide the best illustration of the principles of the described technology and its practical application, and to enable those skilled in the art to use the technology in various embodiments and with various modifications suitable for the specific use contemplated. All such modifications and variations are within the scope of the embodiments as determined by the appended claims and all equivalents thereof, when interpreted in accordance with the breadth to which they are fairly, legally, and equitably entitled, as the appended claims may be amended during the pendency of this patent application.
Claims
1. A method for configuring a flap, comprising: determining whether an inactive middle lobe among a plurality of lobes of an array microphone in an environment is available for deployment; When it is determined that the inactive middle flap is available, positioning the inactive middle flap based on the position data of the sound activity; as well as When it is determined that the inactive valve is unavailable: selecting one of the plurality of deployed petals to move; as well as The selected deployed flap is repositioned based on the position data of the acoustic activity. 2 . The method of claim 1 , wherein the location data for the sound activity comprises coordinates of the sound activity in the environment.
3. The method of claim 1, wherein selecting the one of the plurality of deployed petals comprises selecting the one of the plurality of deployed petals based on a timestamp associated with the plurality of deployed petals.
4. The method of claim 3, wherein the timestamps comprise a first timestamp associated with the location data at which the sound activity was received, and a second timestamp associated with the selected deployed flap.
5. The method of claim 1, wherein selecting the one of the plurality of deployed petals comprises selecting the one of the plurality of deployed petals based on a metric associated with the plurality of deployed petals.
6. The method according to claim 5: wherein one of the metrics comprises a confidence score for the selected deployed flap; and Wherein the confidence score represents one or more of the certainty of the position of the selected deployed flap or the quality of the sound of the selected deployed flap.
7. The method according to claim 1, further comprising: determining, based on the location data of the acoustic activity, whether an existing petal of the plurality of deployed petals is proximate to the acoustic activity; and When it is determined that the existing flap is not in the vicinity of the acoustic activity, the steps of determining whether the inactive flap is available for deployment, positioning the inactive flap, selecting the one of the multiple deployed flaps to move, and repositioning the selected deployed flap are performed.
8. A method according to claim 1, wherein the inactive petals include petals among the multiple petals that are not positioned at specific coordinates in the environment, petals among the multiple petals that have not yet been deployed, or one or more petals among the multiple petals that are inactive based on measurements.
9. A method according to claim 2, wherein the selection of the one of the multiple deployed lobes to move is based on one or more of the following: (1) the difference between the azimuth angle of the coordinate of the sound activity and the azimuth angle of the selected deployed lobe relative to an azimuth threshold, or (2) the difference between the elevation angle of the coordinate of the sound activity and the elevation angle of the selected deployed lobe relative to an elevation threshold.
10. The method of claim 9, wherein selecting the one of the plurality of deployed lobes to move is based on a distance of the coordinates of the sound activity from the array microphone. The method of claim 10 , further comprising setting the azimuth threshold based on the distance of the coordinates of the voice activity from the array microphone.
12. The method of claim 9, wherein selecting the one of the plurality of deployed lobes to move comprises selecting the selected deployed lobe when (1) the absolute value of the difference between the azimuth angle of the coordinates of the sound activity and the azimuth angle of the selected deployed lobe is not greater than the azimuth angle threshold; and (2) the absolute value of the difference between the elevation angle of the coordinates of the sound activity and the elevation angle of the selected deployed lobe is greater than the elevation angle threshold.
13. The method of claim 1, further comprising storing the location data of the acoustic activity in a database as a new location of the selected deployed flap.
14. The method of claim 1, further comprising: Receive remote audio signals from the far end; detecting an amount of activity in the remote audio signal; as well as When the amount of activity of the remote audio signal exceeds a predetermined threshold, the steps of determining whether the inactive flap is available, positioning the inactive flap, selecting the one of the plurality of deployed flaps, and repositioning the selected deployed flap are suppressed.
15. An array microphone system comprising: a plurality of microphone elements, each of the plurality of microphone elements being configured to detect sound and output an audio signal; a beamformer in communication with the plurality of microphone elements, the beamformer configured to generate one or more beamformed signals based on the audio signals of the plurality of microphone elements, wherein the one or more beamformed signals correspond to one or more lobes, each lobe being positioned at a location in an environment; an audio activity locator in communication with the plurality of microphone elements, the audio activity locator being configured to determine coordinates of new sound activity in the environment; and a lobe autoconfigurator in communication with the audio activity locator and the beamformer, the lobe autoconfigurator being configured to: receiving the coordinates of the new sound activity; determining whether the coordinates of the new sound activity are near an existing lobe, wherein the existing lobe comprises one of the one or more lobes; When it is determined that the coordinates of the new sound activity are not near the existing lobe: Determine whether the inactive valve is usable; When it is determined that the inactive middle valve is available, selecting the inactive middle valve; When it is determined that the inactive valve is unavailable, selecting one of the one or more valves; as well as The coordinates of the new sound activity are transmitted to the beamformer to cause the beamformer to update the position of the selected lobe to the coordinates of the new sound activity.
16. The system of claim 15, wherein the inactive lobes comprise one or more of lobes of the beamformer that are not positioned at specific coordinates in the environment, lobes of the beamformer that have not yet been deployed, or lobes of the beamformer that are inactive based on a metric.
17. A system according to claim 15, wherein the lobe automatic configurator is constructed to determine whether the coordinates of the new sound activity are near the existing lobe based on one or more of the following: (1) the difference between the azimuth of the coordinates of the new sound activity and the azimuth of the position of the existing lobe relative to an azimuth threshold, or (2) the difference between the elevation of the coordinates of the new sound activity and the elevation of the position of the existing lobe relative to an elevation threshold.
18. The system of claim 17, wherein the lobe autoconfigurator is constructed to determine whether the coordinates of the new sound activity are near the existing lobe based on their distance from the system.
19. The system of claim 18, wherein the lobe autoconfigurator is further configured to set the azimuth threshold based on the distance of the coordinates of the new sound activity from the system.
20. A system according to claim 17, wherein the lobe automatic configurator is constructed to: determine that the coordinates of the new sound activity are near the existing lobe when (1) the absolute value of the difference between the azimuth angle of the coordinates of the new sound activity and the azimuth angle of the position of the existing lobe is not greater than the azimuth angle threshold; and (2) the absolute value of the difference between the elevation angle of the coordinates of the new sound activity and the elevation angle of the position of the existing lobe is greater than the elevation angle threshold.
21. The system of claim 15, further comprising a database in communication with the flap autoconfigurator, wherein the flap autoconfigurator is further constructed to store a first timestamp associated with the coordinates of receiving the new sound activity in the database.
22. A system according to claim 21, wherein the flap autoconfigurator is further constructed to update the second timestamp associated with the existing flap in the database to the first timestamp when the coordinates of the new sound activity are determined to be near the existing flap.
23. A system according to claim 21, wherein the flap autoconfigurator is further constructed to update the third timestamp associated with the selected flap in the database to the first timestamp when it is determined that the coordinates of the new sound activity are not near the existing flap.
24. A system according to claim 15, wherein the flap autoconfigurator is further constructed to select the one or more flaps based on a timestamp associated with the one or more flaps when it is determined that the coordinates of the new sound activity are not near the existing flap and when it is determined that the inactive flap is not available.
25. The system of claim 15, wherein the lobe autoconfigurator is further constructed to assign a metric associated with the selected lobe when it is determined that the coordinates of the new sound activity are not near the existing lobe.
26. A system according to claim 15, wherein the flap autoconfigurator is further constructed to select the one or more flaps based on a metric associated with the one or more flaps when it is determined that the coordinates of the new sound activity are not near the existing flap and when it is determined that the inactive flap is not available.
27. The system according to claim 25: wherein the metric comprises a confidence score for the selected flap; and Wherein the confidence score represents one or more of the certainty of the coordinates of the selected flap or the quality of the selected flap.
28. The system of claim 15, further comprising a database in communication with the flap autoconfigurator, wherein the flap autoconfigurator is further constructed to store the coordinates of the new sound activity as the new position of the selected flap when it is determined that the coordinates of the new sound activity are not near the existing flap.
29. The system according to claim 15: It further comprises an activity detector in communication with the distal end and the flap autoconfigurator, the activity detector being configured to: receiving a remote audio signal from the remote end; detecting an amount of activity in the remote audio signal; and transmitting the detected amount of activity to the flap autoconfigurator; and wherein the flap autoconfigurator is further configured to: When the amount of activity of the remote audio signal exceeds a predetermined threshold, suppressing the lobe autoconfigurator from performing the following steps: determining whether the coordinates of the new sound activity are near the existing lobe; determining whether the inactive lobe is available, selecting the inactive lobe, selecting one of the one or more lobes, and The coordinates of the new voice activity are transmitted to the beamformer.
30. The system according to claim 15: It further comprises an activity detector in communication with the distal end and the flap autoconfigurator, the activity detector being configured to: receiving a remote audio signal from the remote end; detecting an amount of activity in the remote audio signal; and When the amount of activity of the remote audio signal exceeds a predetermined threshold, transmitting a signal to the lobe autoconfigurator to cause the lobe autoconfigurator to cease performing the following steps: determining whether the coordinates of the new sound activity are in the vicinity of the existing lobe; A determination is made as to whether the inactive lobe is available, the inactive lobe is selected, one of the one or more lobes is selected, and the coordinates of the new sound activity are transmitted to the beamformer.
Citation Information
Patent Citations
Array microphone system and method of assembling the same
US9565493B2
Signal-enhancing beamforming in augmented reality environment
CN104106267A
Microphone array with automated adaptive beam tracking
US10210882B1