Conference system and method for speaker tracking and camera positioning

The method uses multiple microphones to determine and transmit precise speaker locations to cameras, adjusting both camera and microphone positions, addressing the challenge of flexible speaker identification and positioning in conference systems.

JP2025520891APending Publication Date: 2025-07-03SHURE ACQUISITION HLDG INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024577281
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-30
Filing Date
2023-06-30
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In conference environments with multiple microphones and cameras, accurately identifying and positioning cameras towards speakers is challenging due to unknown relative positions and the need for manual, labor-intensive configuration, which becomes inflexible when seating arrangements change.

Method used

A method using multiple microphones to determine speaker locations in different coordinate systems, combining these to estimate a precise speaker location, and transmitting this to cameras to adjust their positioning, along with adjusting microphone lobes based on camera-determined speaker locations.

Benefits of technology

Automatically and accurately positions cameras and microphones to capture speakers, improving video and audio coverage without manual setup, adapting to seating changes, and enhancing beamforming accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520891000001_ABST
    Figure 2025520891000001_ABST
Patent Text Reader

Abstract

A conference system and method configured to generate speaker coordinates for orienting a camera towards a speaker location within an environment are disclosed along with speaker tracking using a plurality of microphones and a plurality of cameras. One method uses a first microphone array (104) to determine a first speaker location (P1) in a first coordinate system with respect to the first microphone array based on acoustics associated with a speaker (102), uses a second microphone array (104) to determine a second speaker location (P2) in a second coordinate system with respect to the second microphone array based on acoustics associated with the speaker (102), determines an estimated speaker location (P3) in a third coordinate system with respect to a camera (106) based on the first speaker location and the second speaker location, and transmits the estimated speaker location in the third coordinate system to the camera to orient the camera towards the estimated speaker location.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference This patent application claims priority to U.S. Provisional Patent Application No. 63 / 367,438, filed on June 30, 2022, the entire content of which is incorporated herein by reference.

[0002] The present disclosure generally relates to speaker tracking and camera positioning in a conferencing environment, and more particularly to a conferencing system and method for positioning a camera towards a speaker location determined using one or more microphones and / or one or more cameras.

Background Art

[0003] Conference environments such as conference rooms, executive offices, video conferencing situations, etc. typically involve microphones (including microphone arrays) for capturing audio from various acoustic sources within the environment (also known as the "near end"), and speakers for presenting acoustics from remote locations (also known as the "far end"). For example, a person in a conference room may be conducting a conference call with a person at a remote location. Typically, speech and audio from the conference room can be captured by a microphone and transmitted to the remote location, while speech and audio from the remote location can be received and reproduced on a speaker within the conference room. Multiple microphones may be used to optimally capture speech and audio within the conference room.

[0004] Such a meeting environment may also include one or more image capture devices such as cameras, which can be used to capture and provide images and videos of people and objects in the environment that are transmitted for viewing at a remote location. However, for example, if the camera is configured to show the entire room, or if the camera is fixed to a particular preconfigured part of the room and the speaker moves in and out of that part during the meeting or event, it may be difficult for viewers at a remote location to see a particular speaker. The speaker may include, for example, a human in the environment who is speaking or making other sounds.

[0005] In addition, in an environment where it is desirable to have multiple cameras and / or multiple microphones (or microphone arrays) for sufficient video and acoustic coverage, it may be difficult to accurately identify the specific speaker in the environment and / or to identify which camera and / or microphone should be directed towards the speaker. Moreover, in some environments with multiple cameras and / or multiple microphones, the relative positions of the cameras and microphones may not be known or predefined. In such an environment, it may be difficult to accurately correlate the camera angles with the speaker positions. A professional installer or integrator can manually configure the camera zones or presets based on the location information from the microphone array, but this is often a time-consuming, labor-intensive, and inflexible process. For example, if the seating arrangement in the room is changed after the initial setup of the meeting system, the preconfigured camera zones may not adequately cover the participants, and such zones may be difficult to modify after setup and / or may only be modifiable by a professional installer or integrator. SUMMARY OF THE INVENTION

[0006] The techniques of the present disclosure are designed to perform, among other things: (1) determining coordinates for positioning a camera towards a speaker based on a speaker location identified by using two or more microphones (or microphone arrays); (2) adjusting a lobe location of a microphone or other acoustic pickup coverage area based on a speaker location identified by using a camera; and (3) selecting a camera for positioning towards a speaker from a plurality of cameras based on a location of a microphone lobe or other acoustic beam that is directed towards the speaker, wherein the selected lobe location is based on a speaker location identified by using two or more microphones.

[0007] In one embodiment, a method implemented by one or more processors that communicate with each of a first microphone, a second microphone, and a camera includes: using a first microphone array to determine a first speaker location in a first coordinate system with respect to the first microphone array based on acoustics associated with a speaker; using a second microphone array to determine a second speaker location in a second coordinate system with respect to the second microphone array based on acoustics associated with the speaker; using at least one processor to determine an estimated speaker location in a third coordinate system with respect to the camera based on the first speaker location and the second speaker location; and transmitting the estimated speaker location in the third coordinate system to the camera to direct an image capturing component of the camera towards the estimated speaker location.

[0008] In another embodiment, the system includes a first microphone array configured to determine a first speaker location in a first coordinate system with respect to the first microphone array based on acoustics associated with a speaker, a second microphone array configured to determine a second speaker location in a second coordinate system with respect to the second microphone array based on acoustics associated with the speaker, a camera comprising an image capture component, and one or more processors communicatively coupled to each of the first microphone array, the second microphone array, and the camera. The one or more processors are configured to determine an estimated speaker location in a third coordinate system with respect to the camera based on the first speaker location and the second speaker location, and to transmit the estimated speaker location in the third coordinate system to the camera. The camera is configured to direct the image capture component towards the estimated speaker location received from the one or more processors.

[0009] In a further embodiment, a non-transitory computer-readable storage medium, when executed by one or more processors communicatively coupled to each of the first microphone array, the second microphone array, and the camera, causes the one or more processors to use the first microphone array to determine a first speaker location in a first coordinate system with respect to the first microphone array based on acoustics associated with a speaker, use the second microphone array to determine a second speaker location in a second coordinate system with respect to the second microphone array based on acoustics associated with the speaker, determine an estimated speaker location in a third coordinate system with respect to the camera based on the first speaker location and the second speaker location, and transmit the estimated speaker location in the third coordinate system to the camera to cause the camera to direct the image capture component of the camera towards the estimated speaker location.

[0010] In another embodiment, a method implemented by one or more processors that communicate with each of a first microphone, a second microphone, and a camera uses a microphone array to determine a first speaker location of the microphone array in a first coordinate system for the microphone array based on acoustics associated with a speaker, uses at least one processor to transform the first speaker location from the first coordinate system to a second coordinate system for the camera, transmits the first speaker location in the second coordinate system to the camera to direct an image capture component of the camera toward the first speaker location, receives from the camera a second speaker location in the second coordinate system identified using a talker detection component of the camera, and uses the microphone array to adjust a lobe location of the microphone array based on the second speaker location received from the camera.

[0011] According to some aspects, adjusting the lobe location includes adjusting a distance coordinate of the lobe location in the second coordinate system based on a distance coordinate of the second speaker location in the second coordinate system and transforming the adjusted lobe location from the second coordinate system to the first coordinate system.

[0012] According to some aspects, adjusting the lobe location includes transforming the second speaker location from the second coordinate system to the first coordinate system and using at least one processor to adjust a distance coordinate of the lobe location in the first coordinate system based on a distance coordinate of the second speaker location in the first coordinate system.

[0013] According to some aspects, determining the first speaker location includes determining the location of the voice generated near the microphone array using an audio localization algorithm executed by an acoustic activity localizer.

[0014] In a further embodiment, a method implemented by one or more processors communicating with each of a plurality of microphone arrays and a plurality of cameras includes using the plurality of microphone arrays to determine a speaker location in a first coordinate system for a first microphone array of the plurality of microphone arrays based on sound associated with the speaker, selecting a lobe location of a selected microphone array of the plurality of microphone arrays in the first coordinate system based on the speaker location in the first coordinate system, selecting a first camera of the plurality of cameras based on the lobe location, converting the lobe location to a second coordinate system for the first camera, and transmitting the lobe location in the second coordinate system to the first camera to direct an image capture component of the first camera towards the lobe location.

[0015] According to some aspects, selecting the lobe location includes determining the distance between the speaker location in the first coordinate system and each of the plurality of microphone arrays, and identifying the selected microphone array of the plurality of microphone arrays as the one closest to the speaker location in the first coordinate system.

[0016] According to some aspects, selecting the first camera includes converting the lobe location from the first coordinate system to a common coordinate system, identifying a first region of a plurality of regions in the common coordinate system as including the lobe location in the common coordinate system, where each region is assigned to one or more of the plurality of cameras, and identifying the first camera as the one assigned to the first region.

[0017] In some embodiments, determining the first speaker location involves using an acoustic localization algorithm executed by an audio activity localizer to determine the location of the voice generated near the plurality of microphone arrays.

[0018] These and other embodiments, as well as various alternatives and aspects, will become apparent and be more fully understood from the following detailed description and the accompanying drawings which set forth exemplary embodiments showing various ways in which the principles of the invention may be utilized. [Brief Description of the Drawings]

[0019]

Figure 1

Figure 1

Figure 2

Figure 2

Figure 3

Figure 3

Figure 4

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Best Mode for Carrying Out the Invention

[0020] The systems and methods described herein can improve the configuration and usage of a conferencing system by using acoustic localization information collected by multiple microphones (or microphone arrays) to position a camera towards an active speaker or other acoustic source within an environment. For example, each microphone can use an acoustic localization algorithm to detect the location of a speaker within the environment and provide the detected speaker location, or corresponding acoustic localization coordinates, to a common aggregator. Typically, the acoustic localization information obtained by the microphones is relatively accurate with respect to azimuth and elevation coordinates, but not as accurate with respect to the radial coordinate, or the distance between the acoustic source and the microphone array. In an embodiment, the aggregator can improve the accuracy of the radial or distance information by aggregating or combining the time-synchronized acoustic localization coordinates obtained by multiple microphones for the same acoustic source (or acoustic event) to determine an estimated speaker location. The estimated speaker location can be provided to a camera to position the camera's image capture component towards the speaker. Prior to the above transmission, the aggregator can transform the coordinates of the estimated speaker location into a coordinate system relative to the camera or a previously determined common coordinate system (e.g., a coordinate system relative to the room), such that the camera receives the estimated speaker location in a format that is understandable and useful to the camera. The camera can utilize the received speaker location to adjust the images and videos captured by the camera in terms of movement, zoom, pan, framing, or other modalities. In this way, the systems and methods described herein can be used by a conferencing system, for example, to enable the camera to more accurately capture images and / or videos of the active speaker.

[0021] The systems and methods described herein can also be used to obtain a speaker location that can be used to steer an acoustic beam or lobe of a microphone (or microphone array) towards an active speaker in the environment, by using the speaker detection component of a camera to improve the configuration and usage of a conferencing system. For example, a microphone array can use an acoustic localization algorithm to detect a first location of an active speaker in the environment and direct the lobe of the microphone array towards the perceived acoustic source or the direction of the first speaker location. As mentioned above, the radius or distance coordinates in acoustic localization can be less accurate than other coordinates (e.g., azimuth and elevation angles), and as a result, the speaker can be located somewhere along a line formed towards the acoustic source perceived based on the azimuth and elevation angle coordinates from the microphone array. The systems and methods described herein can be used by a conferencing system to improve the distance coordinates of acoustic localization by directing the image capture component of a camera towards the first speaker location and identifying a human face, e.g., the face of the speaker, along the line where the speaker can be located, using the speaker detection component or other suitable image processing algorithms. In this way, the camera can determine that the speaker is actually at a second location near or generally in the vicinity of the first speaker location. The second, more accurate location can be provided to the microphone array, and the coordinates of that location can be converted to the coordinate system of the microphone array or other coordinate systems recognized by the microphone array. Based on the second location, the microphone array can adjust the location of the lobe that is directed towards the first speaker location or, in other ways, steer the lobe towards the second location. Thus, the systems and methods described herein can be used by a conferencing system, for example, to enable improving the beamforming accuracy of the array for the microphone array to capture an active speaker.

[0022] In addition, the systems and methods described herein can improve the configuration and usage of a conferencing system by determining which cameras and which microphones are most suitable for capturing the respective video and audio of a given speaker in an environment with multiple cameras and multiple microphones. Typically, when there are multiple microphones, each having multiple acoustic beams or lobe locations, and multiple cameras, it can be difficult to identify the individual speakers within the environment and which camera, if any, should be pointed at a speaker. The systems and methods described herein can be used by a conferencing system to determine or identify the location of an active speaker based on acoustic source localization information obtained by two or more microphones. Based on the speaker location, the conferencing system can select the lobe and corresponding microphone that are most suitable for capturing the sound generated at the identified speaker location. The conferencing system can then select a camera that can optimally video capture the speaker, particularly the speaker's face, at the selected lobe location. Also, the speaker location can be converted to the camera's coordinate system or other common coordinate system and sent to the camera to orient the camera's image capture component towards the speaker location. In this way, the systems and methods described herein can automatically, i.e., without the need for manual installation or setup by one or more users, identify which microphones and / or lobes and cameras should be focused on each individual speaker within the environment for use by a conferencing system.

[0023] As used herein, the terms " lobe " and " microphone lobe " refer to an acoustic beam generated by a given microphone array (or array microphone) to pick up an acoustic signal at a selected location, such as a location where the lobe is directed. The techniques disclosed herein are described with reference to microphone lobes generated by an array microphone, but the same or similar techniques may be utilized by other forms or types of microphone coverage (e.g., cardioid pattern, etc.) and / or by microphones that are not array microphones (e.g., handheld microphones, boundary microphones, lavalier microphones, etc.). Thus, the term " lobe " is intended to cover any type of acoustic beam or coverage.

[0024] Figures 1 and 2 illustrate an exemplary environment 10 in which one or more of the systems and methods disclosed herein may be used. As shown, environment 10 includes a conference system 100 that can be utilized to determine the location of a speaker 102 within environment 10 for the purposes of beamforming and / or camera positioning, according to an embodiment. Figures 1 and 2 illustrate one possible environment, but it should be understood that the systems and methods disclosed herein may be utilized in any suitable environment, including, but not limited to, conference rooms, offices, control rooms, theaters, arenas, concert venues, etc.

[0025] As shown in the illustration, the conference system 100 includes a plurality of microphones 104, at least one camera 106, and an aggregator 108. The system 100 may also include various components not shown in FIGS. 1 and 2, such as, for example, one or more speakers, desktop microphones, display screens, and / or computing devices. In an embodiment, one or more of the components within the system 100 may include one or more digital signal processors or other processing components, controllers, wireless receivers, wireless transceivers, and the like. Additionally, the environment 10 may include one or more other people in addition to the speaker 102, and / or other objects not shown (e.g., musical instruments, telephones, tablets, computers, HVAC equipment, etc.). It should be understood that the components shown in FIGS. 1 and 2 are merely exemplary, and that any number, type, and arrangement of various components within the environment 10 are contemplated and possible.

[0026] Microphone 104 may be any other type of microphone, including a microphone array (also referred to as an "array microphone"), or a non-array microphone such as a directional microphone (e.g., lavalier, boundary, handheld, etc.). The type of transducer (e.g., microphone and / or speaker), and their placement within a particular environment, may depend on the location of the sound source, the location of the listener, physical space requirements, aesthetics, room layout, stage layout, and / or other considerations. For example, one or more microphones may be placed on a table, desk, or other surface near the sound source, or attached to the sound source, e.g., a performer. The microphone may also be mounted overhead or on a wall to capture sound from a larger area, e.g., the entire room. The microphone 104 shown in FIGS. 1 and 2 may be placed within any suitable location, including on a wall, ceiling, table, and / or any other surface within the environment 10. Similarly, speakers may be placed on a wall, ceiling, or table surface to emit sound, such as voice from a remote end of a meeting, pre-recorded audio, streaming audio, etc., to a listener within the environment 10. The microphones and speakers can conform to various sizes, form factors, mounting options, and wiring options to suit the needs of a particular environment. In the illustrated embodiment, the microphone 104 may be positioned at multiple different locations within the environment 10 to accurately capture sound throughout the environment 10.

[0027] In the case where environment 10 is a conference room, environment 10 can be used for meetings, conference calls, or other events where local participants within the room communicate with each other and / or with remote participants. In such a case, microphone 104 can detect and capture voice from an acoustic source within environment 10. The acoustic source may be a local participant, such as human speaker 102 shown in FIG. 1, and the voice may be speech spoken by a local participant, or music or other sounds generated by the same. In a general situation, local participants may be sitting on chairs at a table, but other configurations and locations of the acoustic source are contemplated and possible.

[0028] Camera 106 can capture still images and / or video of environment 10 in which conference system 100 is located. In some embodiments, camera 106 may be a stand-alone camera, while in other embodiments, camera 106 may be a component of an electronic device, such as a smartphone, tablet, etc. In some cases, camera 106 may be included within the same electronic device as one or more of aggregator 108 and microphone 104. Camera 106 may be a pan-tilt-zoom (PTZ) camera that can be physically moved and zoomed to capture desired images and videos, or a virtual PTZ camera that can digitally trim and zoom images and videos to one or more desired portions. System 100 may also include a display, such as a television or computer monitor, for showing other images and / or videos, such as remote participants of a meeting or other image or video content. In some embodiments, the display may include one or more microphones, cameras, and / or speakers, in addition to or including, for example, microphone 104 and / or camera 106.

[0029] Referring additionally to FIG. 3, an exemplary microphone array 200 is shown, which may be any one of the microphones 104 and may be used with the conference system 100 shown in FIGS. 1 and 2 according to an embodiment. The microphone array 200 includes a plurality of microphone elements 202a, b,... zz (or microphone transducers) and can form one or more pickup patterns having lobes, and as a result, can detect and capture sound from an acoustic source, such as the voice of a human speaker 102 or other object or speaker in the environment 10. In some embodiments, the microphone elements 202a, b,... zz may be MEMS (Micro-Electro-Mechanical System) microphones having an omnidirectional pickup pattern. In other embodiments, the microphone elements 202a, b,... zz may have other pickup patterns and / or may be electret condenser microphones, dynamic microphones, ribbon microphones, piezoelectric microphones, and / or other types of microphones. In various embodiments, the microphone elements 202a, b,... zz may be arranged in one or multiple dimensions.

[0030] In some embodiments, each of the microphone elements 202a, b,... zz can detect sound and convert the detected sound into an analog acoustic signal. In such cases, other components within the microphone array 200, such as an analog-to-digital converter, a processor, and / or other components (not shown), can process the analog acoustic signal and ultimately generate one or more digital acoustic output signals. The digital acoustic output signals may conform to appropriate standards and / or transmission protocols for transmitting sound. In other embodiments, each of the microphone elements 202a, b,... zz within the microphone array 200 can detect sound and convert the detected sound into a digital acoustic signal.

[0031] As shown in FIG. 3, one or more digital acoustic output signals 202a, b,... z can be generated to correspond to each of one or more pickup patterns. The pickup pattern can consist of, or include, one or more lobes, such as main lobes, side lobes, and back lobes, and / or one or more nulls. The pickup pattern that can be formed by the microphone array 200 can depend on the type of beamformer used with the microphone elements, such as the beamformer 206. For example, a delay and sum beamformer can form a frequency-dependent pickup pattern based on its filter structure and the layout geometry of the microphone elements. As another example, a difference beamformer can form a cardioid, sub-cardioid, super-cardioid, hyper-cardioid, or bidirectional pickup pattern. Other suitable types of beamformers can include, for example, a minimum variance distortionless response (MVDR) beamformer. In an embodiment, the beamformer 206 can communicate with the microphone elements 202a, b,... zz, either wired or wirelessly. In some embodiments, the beamformer 206 can be a stand-alone device that communicates with the microphone array 200.

[0032] Microphone array 200 can also determine the location of a speaker or other object within environment 10 relative to array 200, or more specifically, the coordinate system of array 200, based on the speech or acoustic activity detected by microphone elements 202a, b, ..., zz. For example, microphone array 200 may include an acoustic activity localizer 208 that communicates with microphone elements 202a, b, ..., zz either wired or wirelessly. Acoustic activity localizer 208 detects or identifies the position or location of acoustic activity detected within an environment, such as environment 10 of FIG. 1, based on the acoustic signals received from microphone elements 202a, b, ..., zz, or, in other ways, can localize the acoustics detected by microphone array 200. In an embodiment, acoustic activity localizer 208 determines the direction of arrival of the detected acoustic activity and uses a steered response power phase transform (SRP-PHAT) algorithm, a generalized cross-correlation phase transform (GCC-PHAT) algorithm, a time of arrival (TOA)-based algorithm, a time difference of arrival (TDOA)-based algorithm, a multiple signal classification (MUSIC) algorithm, an artificial intelligence-based algorithm, a machine learning-based algorithm, or another suitable acoustic or sound source localization algorithm to generate acoustic localization or other data representing the location or position of the detected speech relative to microphone array 200. The detected acoustic activity may include an acoustic source such as a human speaker, e.g., speaker 102 of FIG. 1, or an acoustic trigger from or near a camera, e.g., camera 106 of FIG. 1. As will be appreciated, the location obtained by the sound source localization algorithm may represent the perceived location of the acoustic activity or other estimate obtained based on the acoustic signals received from microphone elements 202a, b, ..., zz, which may or may not coincide with the actual or true location of the acoustic activity.

[0033] The acoustic activity localizer 208 may be configured to indicate the location of the detected acoustic activity as a set of three-dimensional coordinates in a coordinate system relative to the location of the microphone array 200 or in a coordinate system where the microphone array 200 is the origin of the coordinate system. The coordinates may be Cartesian coordinates (i.e., x, y, z) or spherical coordinates (i.e., azimuth φ (phi or “az”), elevation θ (theta or “elev”), radial distance / magnitude (R)). It should be noted that, if desired, Cartesian coordinates can be easily converted to spherical coordinates and vice versa. Spherical coordinates may be used in various embodiments to determine additional information regarding the conference system 100 of FIGS. 1 and 2, such as, for example, the distance between the acoustic source 102 and a given microphone 104, the distance between two microphones 104, the distance between a given microphone 104 and the camera 106, and / or the relative location of the camera 106 and / or microphone 104 within the environment 10.

[0034] In the illustrated embodiment, the acoustic activity localizer 208 is included within the microphone array 200. In other embodiments, the acoustic activity localizer 208 may be included within another component of the conference system 100 or may be a stand-alone component. In various embodiments, the detected speaker location, or more specifically, the localization coordinates representing each location, may be transmitted to one or more other components of the conference system, such as the aggregator 108 and / or the camera 106 of FIG. 1.

[0035] In various embodiments, the location data generated by the acoustic activity localizer 208 also includes a timestamp or other timing information indicating the time at which the coordinates were generated, the order in which the coordinates were generated, and / or any other information that helps to identify coordinates generated simultaneously or nearly simultaneously for the same acoustic source. In some embodiments, the microphones 104 within the conference system 100 have synchronized clocks (e.g., using the Network Time Protocol or the like). In other embodiments, the timing or simultaneous output of the coordinates may be determined using other techniques, such as setting up a set of time-synchronized data channels for transmitting the localization coordinates from the microphones 104 to the aggregator 108.

[0036] As shown in FIG. 3, the microphone array 200 may also include a lobe selector 210 that communicates with the acoustic activity localizer 208, either wired or wirelessly. The microphone array 200 may be capable of forming one or more pickup patterns having lobes that can be steered to sense sound within a particular location in the environment. The acoustic activity localizer 208 can send location data including the detected speaker's speaker location, or acoustic localization coordinates obtained by the localizer 208, to the lobe selector 210 to select a lobe that can optimally capture the speaker location. For example, the lobe selector 210 can determine which of the plurality of lobes of the microphone array 200 is most suitable for picking up or capturing sound at the coordinates determined by the acoustic activity localizer 208, based on distance (e.g., the distance between the microphone array 200 and the speaker location), lobe availability, the coverage area assigned to each lobe, and / or any other appropriate factors. The location of the selected lobe can be provided to the beamformer 206 to deploy an acoustic pickup beam towards the indicated lobe location. The lobe location may be a set of coordinates within a coordinate system relative to the microphone array 200, or any other appropriate format. In various embodiments, location data including the location of the selected lobe may also be transmitted to one or more other components of the conference system, such as the aggregator 108 and / or the camera 106 of FIGS. 1 and 2.

[0037] Referring back to FIGS. 1 and 2, aggregator 108 can be configured to receive time-synchronized speaker location and / or lobe location from one or more of microphones 104 and, based thereon, provide coordinates or other location information to camera 106 to position camera 106 towards speaker 102. In some embodiments, the location data (or localization coordinates) received at aggregator 108 may be relative to the coordinate system of the microphone 104 that generated the data. In such cases, if the relative positions and orientations of microphones 104 and / or cameras 106 within environment 10 are known, a coordinate system transformation can be derived based thereon to transform the location data generated by microphone 104 into, for example, the coordinate system of a receiving component such as camera 106 or another common coordinate system of environment 10.

[0038] In various embodiments, aggregator 108 may include a conversion unit (not shown) configured to convert the location data from its original coordinate system (e.g., relative to microphone 104) to another coordinate system that can be readily used by camera 106 before transmitting the location data to camera 106. For example, the conversion unit may be configured to convert localization coordinates within a first coordinate system relative to a first microphone array of a plurality of microphones 104 (e.g., the first microphone array 104 is the origin of the first coordinate system) to localization coordinates within a second coordinate system relative to camera 106 (e.g., camera 106 is the origin of the second coordinate system). In other embodiments, the conversion unit may be included in each of microphones 104 such that the localization coordinates generated by each microphone 104 can be converted to the coordinate system of the intended recipient (e.g., camera 106) before being transmitted to aggregator 108.

[0039] In some cases, the conversion unit may be configured to convert the location data received at the aggregator 108 into a common coordinate system associated with the environment 10, such as a coordinate system for the room in which the conferencing system 100 is located. As a result, the location data can be readily used by any component of the conferencing system 100. In such an embodiment, each component of the conferencing system 100 (e.g., the microphone 104 and the camera 106) may also include a conversion unit for converting any received location data (or coordinates) into the coordinate system of that component for easier processing and usefulness.

[0040] In some cases, the conversion unit included in the aggregator 108 may convert the location data received from the camera 106 into another coordinate system of the environment 10 before transmitting the received data to another component of the system 100. For example, the speaker location within the coordinate system for the camera 106 may be converted into the speaker location within the coordinate system for one of the microphones 104 and / or the common coordinate system of the environment 10.

[0041] In various cases, the aggregator 108 can combine time-synchronized triangulations from two different microphones 104 to obtain a more accurate estimate of the speaker location, as described herein. In such cases, if the relative positions and orientations of the microphones 104 within the environment 10 are known, a coordinate system transformation can be derived based thereon to convert the location data generated by a first microphone array of the plurality of microphones 104 into the coordinate system of a second microphone array of the plurality of microphones 104 or another common coordinate system of the environment 10. For example, the aggregator 108 may convert the location data within the coordinate system (e.g., x, y, z) of the first microphone 104 into the location data within the coordinate system (e.g., x', y', z') of the second microphone 104 before combining the location coordinates.

[0042] In some embodiments, the conversion unit may be configured to convert the lobe location received from a given microphone 104 into another coordinate system of the environment 10 before transmitting the lobe location to the camera 106 or other components of the system 100, regardless of whether the conversion unit is located within the aggregator 108 or another component of the system 100. For example, the lobe location within the coordinate system for the first microphone 104 may be converted into the lobe location within the coordinate system for the camera 106.

[0043] Referring now to FIGS. 1 and 2, during operation, the first microphone of the microphones 104 can send to the aggregator 108 a first speaker location (e.g., x1, y1, z1) that represents the localization of the speech emitted by a given sound source (e.g., speaker 102) at a first point in time, as perceived or detected by the first microphone 104. Similarly, the second microphone of the microphones 104 can send to the aggregator 108 a second speaker location (e.g., x2, y2, z2) that represents the localization of the same speech corresponding to the same acoustic event, i.e., the same speech at approximately the same point in time, as perceived by the second microphone 104. As will be appreciated, in some cases, for example, as shown in FIGS. 1 and 2, both microphones 104 can localize the speaker 102 to the same location or position within the environment 10, while in other cases, the microphones 104 can localize the speaker 102 to different locations within the environment 10. In the latter case, the aggregator 108 can combine the time-synchronized localizations from the multiple microphones 104 to obtain (or triangulate) an estimated speaker location with a higher accuracy than the individual speaker locations. The aggregator 108 can calculate the estimated speaker location using one or more different techniques, depending on how close the two localizations are to each other, how much overlap there is between them, and / or other relevant considerations. Thus, the aggregator 108 can provide a more precise location of the speaker 102 to the camera 106, minimizing or avoiding incorrect camera positioning.

[0044] More specifically, in FIG. 1, the localization of the voice emitted by the speaker 102 and detected by the first microphone 104 is represented by a first line L1 drawn from the first microphone 104 to the first speaker location (e.g., point P1 in FIG. 1). Similarly, the localization of the same voice as detected by the second microphone 104 is represented by a second line L2 drawn from the second microphone 104 to a second speaker location (e.g., point P2 in FIG. 1) different from the first speaker location. For example, the first speaker location (az1, elev1, R1) and the second speaker location (az2, elev2, R2) may have different radius or distance coordinates ( "R"), but may have the same or similar azimuth ( "az") and elevation ( "elev") coordinates. Thus, the two lines L1 and L2 may terminate at different positions within the environment 10, but may not be the true or actual location of the speaker 102 in either case. Instead, the true location of the speaker 102 (e.g., point P3 in FIG. 1) may be located at different radius coordinates along each of the lines L1 and L2.

[0045] The straight lines shown in FIGS. 1 and 2 are intended to represent the detected speaker locations in three-dimensional space and it should be understood that they can be constructed based on the coordinates received at the aggregator 108. For example, each acoustic localization n can be represented using a straight line starting from the microphone 104 that generated the localization, i.e., the coordinates (az, elev, R), and extending into three-dimensional space at an angle specified by the azimuth and elevation coordinates of the localization, and the line terminates at the point indicated by the radius coordinate of the same localization. The aggregator 108 can calculate each line L(n) using the following formula: L(n)=P(n)+t(n)*V(n), where P(n) is a point on the line L(n) and V(n) is the direction vector of the line L(n). Before such a calculation, the aggregator 108 may first transform the received coordinates, if necessary, into a common coordinate system (e.g., the coordinate system of one of the microphones 104 or the coordinate system of the environment 10, etc.) and / or into spherical coordinates.

[0046] Speaker locations with imprecise radial coordinates can still be used to relatively successfully steer the microphone lobe towards the approximate location of the active sound source, but a higher level of accuracy is required for camera positioning. For example, in FIG. 1, orienting camera 106 towards point P1 or point P2 is not appropriate for capturing the face of speaker 102 who may be located at point P3. As another example, in some cases, camera 106 can be configured to follow the active speaker such that the active speaker is always within the frame or field of view angle of camera 106 when the active speaker moves around the room, including when the active speaker sits or stands, or vice versa.

[0047] In an embodiment, aggregator 108 is configured to improve the acoustic localization accuracy of conference system 100 by determining an estimated location of speaker 102 within environment 10 based on two or more time-synchronized localizations (or location coordinates) generated by two or more different microphones 104 for the same acoustic activity or event. Using various linear algebra techniques, as shown in FIGS. 1 and 2 and described herein, an estimated speaker location can be calculated or determined. Other techniques for estimating the true speaker location may also be used in addition to, or instead of, aggregating the acoustic localizations obtained by microphones 104. For example, in some cases, aggregator 108 uses voice activity detection techniques to prevent ambient noise within environment 10 from being localized, and / or to improve the quality of the localization data used for camera positioning by transmitting only coordinates that are within a predetermined vicinity of the current lobe position and / or within a predefined acoustic coverage area of environment 10. As another example, aggregator 108 may be configured to improve localization quality by filtering out (or not transmitting) continuous coordinates of nearby positions, or other subtle changes due to autofocus activity of a particular lobe, in order to minimize jitter of camera 106 when an active speaker moves their head or mouth. In such cases, a significant movement or substantial change in the detected speaker location may be required before camera 106 moves to a new location, and a constant value may be provided to camera 106 to prevent or minimize jitter until that occurs. As another example, aggregator 108 may be configured to improve localization quality by transmitting coordinates for multiple localizations at once at a predefined rate, and may also be structured according to lobe proximity, coverage area membership, priority (e.g., table head, keynote speaker, company CEO, etc.), and / or other relevant factors.

[0048] Referring initially to FIG. 1, in some embodiments, aggregator 108 can determine an estimated speaker location by identifying a common point based on a first speaker location and a second speaker location, such as, for example, the intersection of a first line L1 and a second line L2. For example, aggregator 108 can formulate an equation of a line (e.g., P1 + t1*V1 = P2 + t2*V2) and solve for t1 and t2 with t1 = t2 to determine the intersection point P3 of lines L1 and L2. As will be appreciated, other techniques for combining or aggregating the first speaker location and the second speaker location to identify a common point may also be used.

[0049] In other embodiments, for example, as shown in FIG. 2, aggregator 108 can determine an estimated speaker location by identifying the nearest or optimal point based on the detected speaker locations, or the point with the minimum distance between regions bounded by the azimuth and elevation coordinates of the speaker locations detected by microphone 104. This technique can be particularly useful in situations where the lines representing the detected speaker locations are inclined and thus do not intersect due to inaccuracies in the azimuth and / or elevation angles detected by microphone 104 and / or other inconsistencies. For example, in FIG. 2, a third line L3 constructed based on a third speaker location (az3, elev3, R3) detected by a first microphone 104 does not intersect a fourth line L4 constructed based on a fourth speaker location (az4, elev4, R4) detected by a second microphone 104. Moreover, the third line L3 terminates at point P4, while the fourth line L4 terminates at point P5, neither of which corresponds to the true speaker location (e.g., point P3).

[0050] In such cases, the aggregator 108 can determine an estimated speaker location based on the detected speaker location by locating the point closest to both lines L3 and L4, or in other ways, by finding the point of minimum distance (or error) that takes into account the inaccuracy of the radial coordinate R. For example, the aggregator 108 can find the point on line L3 closest to line L4 (e.g., point P4 in FIG. 2) and the point on line L4 closest to line L3 (e.g., point P5), and draw a new line segment L5 connecting the two closest points between line L3 and line L4 to resolve the inaccurate radial coordinate. The line segment L5 can be perpendicular to both lines L3 and L4. The aggregator 108 can calculate the line segment L5 using the formula L5 = P4 + t4*V4 + t6*V6, where V6 = V2*V1, which is a vector perpendicular to both P4 and P5. Since it is known that the line segment L5 contains a point such that P4 + t4*V4 + t6*V6 = P5 + t5*V5, the aggregator 108 can solve for t4, t5, and t6 with t4 = t5 = t6 and determine the unknown values. Once the line segment L5 is drawn, the aggregator 108 identifies the midpoint of the line segment L5 (e.g., point Pm in FIG. 2) as the point on the line segment L5 that is most likely to be the estimated speaker location or the location of the acoustic source. In various embodiments, the acoustic localization received at the aggregator 108 may be transformed, if necessary, into a common coordinate system (e.g., the coordinate system of the first microphone 104 or the environment 10, etc.) and / or from Cartesian coordinates (x, y, z) to spherical coordinates (az, elev, R) before calculating the estimated speaker location.

[0051] In some cases, where the coordinates of the acoustic localization including the azimuth angle and / or the elevation angle are not very precise or contain slight inaccuracies, the acoustic localization obtained by a given microphone 104 (or microphone array) may represent a three-dimensional area or region within the environment 10 rather than a straight line. For example, the acoustic localization may refer to a wider area or region, such as an elliptical, cylindrical, or conical region that includes (or is centered on) the corresponding straight line L(n). In various embodiments, the techniques described herein may also be used to calculate an estimated speaker location based on such acoustic localization (or "localized region"). For example, the aggregator 108 may be configured to identify the intersections and deviations of the localized regions, which are three-dimensional areas bounded by the coordinates of the corresponding time-synchronized acoustic localizations. In some embodiments, the aggregator 108 can provide the intersection or overlapping region as the estimated speaker location. In other embodiments, the aggregator 108 can further identify points within the overlapping region using one or more of the techniques described herein and provide the identified points as the estimated speaker location. For example, the estimated speaker location may be the center point of the overlapping region, or the point within the overlapping region that is closest to both localized regions (e.g., the closest point).

[0052] FIG. 4 shows a system 300 (e.g., a conference system) that can be used as any of the conference systems 100 shown in the environment 10 of FIGS. 1 and 2, or any of the other conference systems described herein (e.g., the conference systems 700 of FIGS. 8-10) according to an embodiment. The system 300 can include a plurality of microphones 302a, ..., z (e.g., the microphone 104 of FIG. 1) that can detect and capture audio from acoustic sources (e.g., the speaker 102 of FIG. 1) within the environment. The microphones 302a, ..., z can also detect or determine the location of the detected acoustic sources within the environment (e.g., the environment 10 of FIG. 1). The system 300 can also include an aggregator unit 304 (e.g., the aggregator 108 of FIG. 1) that can receive the locations detected from the microphones 302a, ..., z and determine an estimated speaker location based on the received locations. The system 300 can capture an image and / or video of the environment 10, including the active speaker (e.g., the speaker 102 of FIG. 1) within the environment, and can further include one or more cameras 306a, ..., z (e.g., the camera 106 of FIG. 1) that can be controlled by the camera controller 308 of the system 300. The aggregator unit 304 can provide the estimated speaker location to the camera controller 308 to position at least one of the cameras 306a, ..., z towards the estimated speaker location. The camera controller 308 can provide, for example, appropriate signals to the cameras 306a, ..., z to move and / or zoom towards the active speaker. In some embodiments, the camera controller 308 and / or the cameras 306a, ..., z may be configured to move with the active speaker when moving around the environment (or room).

[0053] In some embodiments, one of microphones 302a, ..., z may act as aggregator unit 304. In other embodiments, aggregator unit 304 and camera controller 308 may be included within the same device (e.g., a computing device). In still other embodiments, camera controller 308 and one of cameras 306a, ..., z may be integrated together. The components of system 300 may communicate with each other and / or with other components of system 100, either wired and / or wirelessly.

[0054] Each of microphones 302a, ..., z can detect sound in environment 10 and, for example, determine the location of the sound relative to itself or within a coordinate system having a given microphone 302a, ..., z as its origin. The sound can be emitted by a speaker (e.g., speaker 102 of FIG. 1), an object (e.g., one or more of cameras 306a, ..., z), or any other acoustic source within environment 10. Each microphone 302a, ..., z can send the location of the detected sound (or "detected speaker location") within its respective coordinate system to aggregator unit 304. In some instances, one or more of microphones 302a, ..., z can send the location of one or more lobes deployed by one or more of microphones 302a, ..., z in their respective coordinate systems to aggregator unit 304. The lobes may be deployed, for example, based on the detected speaker location.

[0055] Accordingly, the aggregator unit 304 can receive the detected speaker location and / or the lobe locations of the microphones 302a, ..., z from each of the microphones 302a, ..., z. Each location received by the aggregator unit 304 may be in the respective coordinate system of the microphone 302a, ..., z that provided the location. Accordingly, the aggregator unit 304 may transform the received locations into a common coordinate system that is readily usable by the aggregator unit 304 to perform one or more calculations and / or is readily usable by one or more of the cameras 306a, ..., z for camera positioning. The common coordinate system may be the coordinate system for a given camera 306a, ..., z, the coordinate system of the room or environment in which the cameras 306a, ..., z are located, or the coordinate system of another component of the system 300, such as, for example, the aggregator unit 304 or one of the microphones 302a, ..., z. In other embodiments, the transformation of the locations to a common coordinate system may be performed by another component included in or communicating with the system 300, such as, for example, a computing device (not shown), a remote computing device (e.g., a cloud-based device), and / or any other suitable device.

[0056] For example, as described herein, aggregator unit 304 may calculate or determine an estimated speaker location based on the time-synchronized detected speaker locations received from two or more microphones 302a, ..., z. In such an instance, aggregator unit 304 may transform the detected speaker locations received from the individual microphones 302a, ..., z into a first common coordinate system (e.g., the coordinate system of the first microphone 302a) associated with one of the microphones 302a, ..., z, and calculate the estimated speaker location in the first common coordinate system. Before transmitting the estimated speaker location to camera controller 308, aggregator unit 304 may transform the estimated speaker location from the first common coordinate system into a second common coordinate system (e.g., the coordinate system of camera 306a) associated with one of cameras 306a, ..., z, such that the estimated speaker location is readily usable by that camera 306a, ..., z. In some instances, aggregator unit 304 may also transmit to camera controller 308 the locations of the microphones 302a, ..., z that detected the voice or otherwise generated the detected voice location. In such an instance, aggregator unit 304 may also transform the location of each microphone 302a, ..., z into a common coordinate system that is readily usable by cameras 306a, ..., z.

[0057] In an embodiment, the aggregator unit 304 and the camera controller 308 can communicate via a suitable application programming interface (API) that enables the camera controller 308 to query the aggregator unit 304 about the locations of specific microphones 302a, ..., z, enables the aggregator unit 304 to send a signal to the camera controller 308, and / or enables the camera controller 308 to send a signal to the aggregator unit 304. For example, in some cases, the aggregator unit 304 can respond to a query from the camera controller 308 via the API by sending the converted location to the camera controller 308. Similarly, each microphone 302a, ..., z can be configured to communicate with the aggregator unit 304 using a suitable API that enables the microphones 302a, ..., z to receive a query from the aggregator unit 304 and send the localization coordinates to the aggregator unit 304.

[0058] The camera controller 308 can receive from the aggregator unit 304 the locations of the microphones 302a, ..., z, the lobe locations, and / or the estimated speaker location. Based on the received location, the camera controller 308 can select, for example, which of the cameras 306a, ..., z to utilize to capture an image and / or video of the specific location where the active speaker is located. The camera controller 308 can provide a suitable signal to the selected camera 306a, ..., z to move and / or zoom the camera 306a, ..., z. For example, the camera controller 308 can utilize the locations of the microphones 302a, ..., z, the received lobe locations, and / or the estimated speaker location within the coordinate system of the cameras 306a, ..., z to generate optimized camera parameters that enable more accurate zooming, panning, and / or framing of the speaker 102.

[0059] The components shown in FIGS. 1-4 are merely exemplary, and it should be understood that various components of any number, type, and arrangement of System 100, Array 200, and / or System 300 are contemplated and possible. For example, in FIG. 1, there may be a plurality of cameras 106 and / or, as shown in FIG. 3, a camera controller coupled between camera 106 and aggregator 108.

[0060] FIG. 5 shows an exemplary method or process 400 for providing an estimated speaker location to a camera based on speaker coordinates obtained by a plurality of microphones, according to an embodiment. Cameras (e.g., cameras 106 of FIGS. 1 and 2 and / or cameras 306a,...z of FIG. 3) and microphones (e.g., microphones 104 of FIGS. 1 and 2 and / or microphones 302a,...z of FIG. 3) may form part of a conference system (e.g., systems 100 of FIGS. 1 and 2 and / or system 300 of FIG. 3) located within an environment (e.g., environment 10 of FIGS. 1 and 2). The environment may be a conference room, event space, or other area including one or more speakers (e.g., speaker 102 of FIGS. 1 and 2) or other acoustic sources. The cameras may be configured to detect and capture images and / or videos of the speakers. The microphones may be configured to detect and capture the sound emitted by the speakers and determine the location of the detected sound. Method 400 may be implemented by one or more processors of the conference system, such as a processor of a computing device included within the system and communicatively coupled to the cameras and microphones. In some instances, method 400 may be implemented by an aggregator (e.g., aggregator 108 of FIG. 1 or aggregator unit 304 of FIG. 4) and / or a camera controller (e.g., camera controller 308 of FIG. 4) included within the conference system.

[0061] As shown in FIG. 5, method 400 can include, at step 402, using a first microphone (or microphone array) to determine a first speaker location within a first coordinate system for the first microphone array based on the acoustics associated with a speaker. For example, the first coordinate system may be a coordinate system in which the first microphone array is at the origin, or any other suitable coordinate system. In an embodiment, determining the first speaker location includes using the first microphone array to detect the voice generated near the first microphone array, such as the acoustics associated with the speaker, and using an acoustic localization algorithm executed by an acoustic activity localizer of the first microphone array (e.g., acoustic activity localizer 208 in FIG. 2) to determine the location of the detected voice.

[0062] Step 404 includes using a second microphone (or microphone array) to determine a second speaker location within a second coordinate system for the second microphone array based on the same acoustics associated with the same speaker (e.g., speaker 102 in FIG. 1). For example, the second coordinate system may be a coordinate system in which the second microphone array is at the origin, or any other suitable coordinate system. In an embodiment, determining the second speaker location includes using the second microphone array to detect the voice generated near the second microphone array, such as the acoustics associated with the speaker, and using an acoustic localization algorithm executed by an acoustic activity localizer of the second microphone array (e.g., acoustic activity localizer 208 in FIG. 2) to determine the location of the detected voice.

[0063] Step 406 includes using at least one processor (e.g., the processor of aggregator 108 in FIG. 1) to determine an estimated speaker location in a third coordinate system for a camera (e.g., camera 106 in FIG. 1) based on a first speaker location and a second speaker location. In some embodiments, determining the estimated speaker location includes using at least one processor to identify a common point based on the first speaker location and the second speaker location, as described herein with respect to FIG. 1. In other embodiments, determining the estimated speaker location includes using at least one processor to identify a closest point based on the first speaker location and the second speaker location, as described herein with respect to FIG. 2.

[0064] In embodiments, method 400 further includes converting the detected speaker locations to a common coordinate system before calculating the estimated speaker location. The common coordinate system may be with respect to one of a camera, a microphone array, the room or environment in which the conferencing system is located, or another component of the system. For example, in some embodiments, method 400 includes using at least one processor to convert the first speaker location from a first coordinate system to a second coordinate system, using at least one processor to determine an estimated speaker location in the second coordinate system based on the first speaker location and the second speaker location in the second coordinate system, and using at least one processor to convert the estimated speaker location from the second coordinate system to a third coordinate system. In other embodiments, method 400 includes using at least one processor to convert the first speaker location from the first coordinate system to the third coordinate system, using at least one processor to convert the second speaker location from a second coordinate system to the third coordinate system, and using at least one processor to determine an estimated speaker location in the third coordinate system based on the first speaker location and the second speaker location in the third coordinate system.

[0065] Step 408 includes transmitting, from at least one processor to the camera, an estimated speaker location in a third coordinate system in order to direct the camera's image capture components towards the estimated speaker location in the third coordinate system. The camera can direct the image capture components towards the estimated speaker location in the third coordinate system by adjusting one or more of the camera's angle, tilt, zoom, and framing, or any other relevant parameter of the camera.

[0066] FIG. 6 shows an exemplary environment 50 in which one or more of the systems and methods disclosed herein may be used. As shown, environment 50 includes a conference system 500 that can be used to detect the location of a speaker 502 within environment 50, direct the lobe of microphone 504 towards the detected speaker location, and refine the lobe location based on a second speaker location determined by camera 506, according to an embodiment. FIG. 6 shows one possible environment, but it should be understood that the systems and methods disclosed herein may be used in any suitable environment, including but not limited to conference rooms, offices, boardrooms, theaters, arenas, concert venues, etc. System 500 may include various components not shown in FIG. 6, such as, for example, one or more speakers, desktop microphones, display screens, and / or computing devices. In an embodiment, one or more of the components within system 500 may include a digital signal processor, a wireless receiver, a wireless transceiver, etc. Additionally, environment 50 may include one or more other people in addition to speaker 502, and / or other objects not shown (e.g., musical instruments, telephones, tablets, computers, HVAC equipment, etc.). It should be understood that the components shown in FIG. 6 are merely exemplary, and that any number, type, and arrangement of various components within environment 50 are contemplated and possible.

[0067] According to an embodiment, the environment 50 may be substantially the same as the environment 10 in FIG. 1, the conference system 500 may be substantially the same as the conference system 100 in FIG. 1, and / or may be implemented using the conference system 300 in FIG. 4. For example, as shown in the figure, the conference system 500 may include a microphone 504 that may be substantially the same as the microphone 104 in FIG. 1, the microphones 302a, ..., z in FIG. 4, and / or the microphone array 200 in FIG. 3, and at least one camera 506 that may be substantially the same as the camera 106 in FIG. 1 and / or the cameras 306a, ..., z in FIG. 4. Therefore, the camera 506 and the microphone 504 are not described in more detail for the sake of brevity. The components of the conference system 500 may communicate with each other and / or with remote devices (e.g., cloud computing servers, etc.) either wired or wirelessly.

[0068] In some embodiments, the system 500 further includes a camera controller (e.g., the camera controller 308 in FIG. 4) that receives location information from the microphone 504 and positions the camera 506 towards the received location. In other embodiments, the camera 506 may include a camera controller. The camera controller can direct or position the camera towards a specific location (e.g., the speaker location) by adjusting one or more of the camera's angle, tilt, zoom, and framing, or any other appropriate settings.

[0069] In some embodiments, the system 500 further includes an aggregator (e.g., the aggregator unit 304 in FIG. 4) that converts the location information received from the microphone 504 into a common coordinate system before providing the location information to the camera 506 for camera positioning. In other embodiments, the aggregator may be included within the microphone 504 or the camera 506. The common coordinate system may be a coordinate system centered on the microphone 504, the camera 506, and / or the environment 50, or any other of the aforementioned coordinate systems.

[0070] The microphone 504 can detect voice or acoustic activities within the environment 50 and determine the location of the voice or the acoustic source (e.g., the speaker 502) that emits the voice within the coordinate system of the microphone 504 or with respect to the microphone 504. For example, the microphone 504 is an acoustic activity localizer (e.g., the acoustic activity localizer 208 in FIG. 3) configured to determine a set of coordinates (or "localization coordinates") representing the location of the acoustic activity (or "detected speaker location") within the environment 50 based on the acoustic signals received from the microphone elements (e.g., the microphone elements 202a, b,..., zz in FIG. 3) within the microphone 504, or can include other acoustic source localization algorithms.

[0071] The microphone 504 can also deploy a lobe for capturing the sound emitted by the active speaker 502 towards the detected speaker location. For example, the microphone 504 can include a lobe selector (e.g., the lobe selector 210 in FIG. 3) configured to determine which of the plurality of lobes of the microphone 504 is most suitable for picking up or capturing sound at the localization coordinates determined by the acoustic activity localizer. Thus, the acoustic activity localizer can send the detected speaker location or the corresponding acoustic localization coordinates to the lobe selector for optimal lobe selection based thereon. The microphone 504 can also include a beamformer (e.g., the beamformer 206 in FIG. 3) for deploying an acoustic pickup beam towards the location of the selected lobe, which may be provided as a set of coordinates within the coordinate system of the microphone 504.

[0072] As described herein, the location obtained by the acoustic source localization algorithm may represent the perceived location of the acoustic activity and, in fact, may be an estimated value of the acoustic source location that may not coincide with the actual or true location of the acoustic activity. For example, when the localization coordinates are provided as spherical coordinates (az, elev, R), the radius component R may not be very accurate, and as a result, the localization point (e.g., point Pa in FIG. 6) may not coincide with the true speaker location (e.g., point Pb in FIG. 6). As a result, the lobe deployed towards the speaker location detected by the microphone 504 may only partially capture, or may not capture at all, the speech emitted by the speaker 502, as shown, for example, in FIG. 6.

[0073] In an embodiment, the conference system 500 may be configured to use the camera 506 to identify the true location of the active speaker or, in other ways, to improve at least the radial coordinates of the detected speaker location obtained by the microphone 504. For example, the microphone 504 can transmit the detected speaker location and / or the selected lobe location to the camera 506 either directly or via an aggregator and / or camera controller, as described herein. The location information (e.g., a set of coordinates representing the lobe location or the detected speaker location) may be with respect to the coordinate system of the microphone 504 and, thus, may be converted to location information with respect to the coordinate system of the camera 506. The conversion may be performed by the microphone 504, the camera 506, or any other suitable component of the system 500. In some embodiments, the location information may be converted, for example, by the microphone 504 or other components of the system 500 before being transmitted to the camera 506.

[0074] Upon receiving location information, camera 506 can direct the image capture component of camera 506 towards the received location (e.g., the detected speaker location or lobe location). The image capture component can be configured to capture still images, videos, and / or video. In an embodiment, camera 506 can include a speaker detection component that uses a face detection algorithm, a human head detection algorithm, or any other suitable image processing algorithm to identify the face, head, or any other distinguishable part of speaker 502, or otherwise determine the camera angle that best corresponds to a human face or the front of a person within the vicinity of the received location information. For example, while camera 506 is moving the image capture component along a straight line extending from a received location (e.g., point Pa) that has the same azimuth and elevation angles as the coordinates of the received location, the speaker detection component can track the motion vector and scan the face or head of speaker 502. In other cases, camera 506 can scan the vicinity area or region of environment 50 in a grid pattern until the face or head of speaker 502 is identified. Other known techniques for obtaining a more precise location of speaker 502 may also be used.

[0075] As shown in FIG. 6, the camera 506 can identify the face of the speaker 502 at a second location (e.g., point Pb) that is distant from the reception location. In an embodiment where the azimuth and elevation components are kept the same, the coordinates of the second location (or “second speaker location”) can differ from the coordinates of the initial location only with respect to the radius or distance component. In other embodiments, the second speaker location can differ from the reception location (e.g., the first speaker location or lobe location) with respect to the azimuth, elevation, and / or distance components. In any case, the lobe location of the microphone 504 can be adjusted with respect to the new lobe location based on the second speaker location, and thus the microphone lobe can be made capable of sufficiently capturing the sound generated by the speaker 502. In particular, the camera 506 can provide the second speaker location (or the coordinates representing it) to the microphone 504. The microphone 504 can adjust the lobe location determined by the microphone 504 by changing one or more coordinates of the lobe location based on the received coordinates. For example, the microphone 504 can change or adjust the distance coordinate of the lobe location based on the distance coordinate of the second speaker location. If the lobe is already deployed towards the active speaker 502, the beamformer of the microphone 504 can steer the lobe towards the adjusted lobe location. If the lobe is not deployed, the beamformer can deploy or direct the lobe towards the adjusted lobe location.

[0076] In some cases, the second speaker location may be within the coordinate system of camera 506. In such cases, camera 506 can convert the second speaker location to the coordinate system of microphone 504 before transmitting the location information to microphone 504. In other cases, the second speaker location may be transmitted to microphone 504 as coordinates within the coordinate system of camera 506. In such cases, microphone 504 can convert the received coordinates to the coordinate system of microphone 504 before adjusting the lobe location. In some embodiments, the conversion step may be performed by one or more other components of system 500, such as an aggregator, for example.

[0077] It should be noted that the techniques described herein for refining the lobe location of a microphone using speaker coordinates obtained by a camera can also be used in a conference system that includes a plurality of microphone arrays, although not shown in FIG. 6. In such cases, the speaker coordinates determined using the face detection algorithm of camera 506 can be used to steer a selected lobe of one of the microphones (or microphone arrays). In some cases, the lobe can be selected based on which lobe and / or microphone array is most suitable for picking up sound from the speaker location detected by camera 506. In other cases, the lobe may be pre-selected based on the localization coordinates obtained by the microphone array based on the detected acoustic activity.

[0078] FIG. 7 shows an exemplary method or process 600 for refining the lobe location of a microphone (or microphone array) based on the speaker location determined by a camera according to an embodiment. The camera (e.g., camera 506 of FIG. 6) and the microphone (e.g., microphone 504 of FIG. 6) may form part of a conferencing system (e.g., system 500 of FIG. 6) located within an environment (e.g., environment 50 of FIG. 6). The environment may be a conference room, an event space, or other area that includes one or more speakers (e.g., speaker 502 of FIG. 6) or other acoustic sources. The camera can be configured to detect and capture images and / or video of the speaker. The microphone can be configured to detect and capture the sound emitted by the speaker and determine the location of the detected sound. Method 600 may be implemented by one or more processors of the conferencing system, such as a processor of a computing device included within the system and communicatively coupled to the camera and the microphone array. In some instances, method 600 may be implemented by an aggregator (e.g., aggregator unit 304 of FIG. 4) and / or a camera controller (e.g., camera controller 308 of FIG. 4) included within the conferencing system.

[0079] As shown in FIG. 7, method 600 can include, at step 602, using a microphone (or microphone array) to determine a first speaker location of the microphone within a first coordinate system with respect to a first microphone based on acoustics associated with a speaker. For example, the first coordinate system may be a coordinate system in which the first microphone is at the origin, or any other suitable coordinate system. In an embodiment, determining the first speaker location includes using the first microphone to detect speech generated near the first microphone, such as acoustics associated with the speaker, and using an acoustic localization algorithm executed by an acoustic activity localizer of the first microphone array (e.g., acoustic activity localizer 208 of FIG. 3) to determine the location of the detected speech.

[0080] Step 604 includes using at least one processor to transform the first speaker location from the first coordinate system to a second coordinate system with respect to a camera. For example, the second coordinate system may be a coordinate system in which the camera is at the origin, or any other suitable coordinate system.

[0081] Step 606 includes transmitting, from at least one processor to the camera, the first speaker location within the second coordinate system to direct the camera's image capture component towards the first speaker location. As shown in FIG. 6, if the first speaker location is not the actual location of the speaker, the camera may be facing away from the active speaker and / or may not properly capture the speaker. In such cases, the camera can use the camera's speaker detection component to identify a second speaker location within the second coordinate system of the camera that is near the first speaker location and includes or is directed towards the location of the speaker's face. For example, the speaker detection component can use a face detection algorithm or the like, as described herein, to identify an image similar to a human face or otherwise detect the face of the active speaker.

[0082] Step 608 includes receiving, from the camera, a second speaker location within a second coordinate system identified using the camera's speaker detection component. Step 610 includes using a microphone to adjust the lobe location of the microphone based on the second speaker location received from the camera. For example, the distance coordinates of the initial lobe location of the microphone may be adjusted or refined based on the distance coordinates of the second speaker location determined by the camera. In some cases, the microphone may adjust the lobe location by deploying a new lobe towards the second speaker location. In other cases, the microphone may adjust the lobe location by steering an existing lobe towards the second speaker location.

[0083] In some embodiments, the second speaker location may be converted to the first coordinate system of the microphone before adjusting the lobe location of the microphone. For example, in such cases, step 610 may include using at least one processor to convert the second speaker location from the second coordinate system to the first coordinate system and using at least one processor to adjust the distance coordinates of the lobe location in the first coordinate system based on the distance coordinates of the second speaker location in the first coordinate system.

[0084] In other embodiments, the location of the microphone lobe may be converted to the first coordinate system after adjusting the lobe location in the second coordinate system based on the second speaker location. For example, in such cases, step 610 may include using at least one processor to adjust the distance coordinates of the lobe location in the second coordinate system based on the distance coordinates of the second speaker location in the second coordinate system and using at least one processor to convert the adjusted lobe location from the second coordinate system to the first coordinate system.

[0085] Figures 8-10 illustrate an exemplary environment 70 in which one or more of the systems and methods disclosed herein may be used. As shown, environment 70 may be utilized to improve camera positioning for one or more speakers 702 by determining which of a plurality of cameras 706 is most suitable for capturing an image and / or video of each speaker 702 using a lobe location selected based on speaker coordinates obtained by one or more microphones 704, according to an embodiment. Figures 8-10 illustrate one possible environment, but it should be understood that the systems and methods disclosed herein may be utilized in any suitable environment, including but not limited to offices, boardrooms, theaters, arenas, concert venues, etc. System 700 may include various components not shown in Figures 8-10, such as, for example, one or more speakers, desktop microphones, display screens, and / or computing devices. In an embodiment, one or more of the components within system 700 may include a digital signal processor, wireless receiver, wireless transceiver, etc. Additionally, environment 70 may include one or more other people in addition to speaker 702 and / or other objects not shown (e.g., musical instruments, telephones, tablets, computers, HVAC equipment, etc.). It should be understood that the components shown in Figures 8-10 are merely exemplary and that any number, type, and arrangement of various components within environment 70 are contemplated and possible.

[0086] According to an embodiment, the environment 70 may be substantially the same as the environment 10 in FIG. 1, the conference system 700 may be substantially the same as the conference system 100 in FIG. 1, and / or may be implemented using the conference system 300 in FIG. 4. For example, as shown in the figure, the conference system 700 may include a plurality of microphones 704, each of which may be substantially the same as the microphone array 200 in FIG. 3, the microphone 104 in FIG. 1, and / or the microphones 302a, ..., z in FIG. 4. The conference system 700 may also include a plurality of cameras 706, each of which may be substantially the same as the camera 106 in FIG. 1 and / or the cameras 306a, ..., z in FIG. 4. In addition, the system 700 includes an aggregator 708 that may be substantially the same as the aggregator unit 304 in FIG. 4 and / or the aggregator 108 in FIG. 1. Therefore, the microphones 704, cameras 706, and aggregator 708 will not be described in more detail for the sake of brevity.

[0087] In some embodiments, the system 700 further includes a camera controller (e.g., the camera controller 308 in FIG. 4) for receiving location information from the microphones 704 and / or the aggregator 708 and positioning the cameras 706 towards the received locations. The camera controller may be a stand-alone device, may be integrated with the aggregator 708, or may be included in one of the cameras 706. The camera controller can direct or position the camera towards a specific location (e.g., the speaker location) by adjusting one or more of the camera's angle, tilt, zoom, and framing, or any other appropriate setting.

[0088] The components of the conference system 700 may communicate with each other and / or with remote devices (e.g., cloud computing servers, etc.) either wired or wirelessly. In some embodiments, one or more components of the system 700 may be integrated together or into a single device. For example, aggregator 708 may be included in one of microphones 704 or one of cameras 706.

[0089] Each microphone 704 can detect voice or acoustic activity in environment 70 and determine the location of the voice or the acoustic source (e.g., speaker 702) that emitted the voice or sound within the coordinate system of microphone 704 or with respect to microphone 704. For example, each microphone 704 can be based on the acoustic signals received from microphone elements within microphone 704 (e.g., microphone elements 202a, b,..., zz in FIG. 3) to determine a set of coordinates (or "localization coordinates") representing the location of the acoustic activity (or "detected speaker location") in environment 70, or an acoustic activity localizer (e.g., acoustic activity localizer 208 in FIG. 3) or other acoustic source localization algorithms for determining the location of the acoustic activity in environment 70.

[0090] Each microphone 704 (or microphone array) can also deploy a lobe for capturing the sound emitted by the active speaker 702 towards the detected speaker location. For example, each microphone 704 can include a lobe selector (e.g., lobe selector 210 in FIG. 3) configured to determine which of the plurality of lobes of the microphone 704 is most suitable for picking up or capturing sound at the localization coordinates determined by the acoustic activity localizer. The acoustic activity localizer can send the detected speaker location or the corresponding acoustic localization coordinates to the lobe selector for selecting the optimal lobe based thereon. The microphone 704 can also include a beamformer (e.g., beamformer 206 in FIG. 3) for deploying an acoustic pickup beam (or lobe) towards the selected lobe location, which may be provided as a set of coordinates within the coordinate system of the microphone 704.

[0091] As shown in the figure, the environment 70 includes a plurality of microphones 704 located in different areas of the room or space. Each microphone 704 can be configured to deploy a plurality of lobes to various locations within the environment 70 using, for example, auto lobe tracking techniques, static lobe techniques, and / or other suitable techniques. The use of the plurality of microphones 704 and lobes can improve the detection and capture of speech from acoustic sources within the environment 70 and provide a more accurate estimate of the location of the active speaker within the environment 70, as described herein. The environment 70 also includes a plurality of cameras 706 located in different areas of the room or space. The use of the plurality of cameras 706 can enable the capture of more various types of images and / or videos of the environment 70. For example, the camera 706 located at the front of the environment 70 can be utilized to capture a wider view of the room, while the camera 706 located on the wall of the room can be utilized to capture a close-up of the speaker 702 within the environment 70.

[0092] The presence of multiple microphones and multiple cameras within a given environment can also complicate camera selection and / or microphone or lobe selection for an active speaker. For example, when multiple microphone arrays are present, there may be two or more possible acoustic beam or lobe locations to cover the active speaker, which can make it difficult to identify a particular speaker and / or lobe location for the purpose of positioning a camera. As another example, when multiple cameras are present, it can be difficult to determine whether any of the cameras should be oriented or directed towards each speaker due to overlapping coverage areas and / or opposing camera angles.

[0093] In an embodiment, the conference system 700 of FIGS. 8-10 can be configured to optimize the selection of appropriate microphones from a plurality of microphones 704 and / or the selection of appropriate cameras from a plurality of cameras 706 for an active sound source (e.g., speaker 702 in FIG. 8) within environment 70 by at least partially coordinating microphone 704 with camera 706. As shown in FIG. 8, environment 70 can be divided into a plurality of regions 72 and 74 such that each region 72, 74 contains only one of cameras 706. In various embodiments, regions 72 and 74 can be assigned to respective cameras 706 depending on, for example, the angle or orientation of camera 706 as to which camera 706 has the best view or coverage of the region, or in other ways, is most suitable for capturing a speaker located within that region of environment 70. Regions 72 and 74 can be preconfigured by a technician or installer of system 700 based on various aspects of environment 70, including, for example, the size and shape of the room or space, the number of cameras 706 present, the location or placement of cameras 706, and the presence and / or location of other known objects (e.g., tables, chairs, desks or podiums, display screens and / or whiteboards, microphones 704, speakers, etc.). FIGS. 8-10 show environment 70 as including only two regions, but it should be understood that any number of regions may be created within a given environment depending on, for example, the physical characteristics of the environment, the total number of cameras present, and / or other relevant considerations.

[0094] In some embodiments, the installer may manually determine the location of the microphone 704 relative to the environment 70, the respective placement and orientation of the camera 706 relative to the microphone 704, and / or the location of other objects within the environment 70. In other embodiments, the installer may use a graphical tool or other software of the graphical tool or system 700 configured to automatically determine the location of each microphone 704, the location and orientation information of each camera 706, and / or the location or placement information of other objects within the environment 70. For example, the graphical tool may scan the environment 70 for one or more objects using acoustic excitation, acoustic localization, and / or triangulation techniques to identify the location of the camera 706, microphone 704, and / or speaker (not shown) relative to a common coordinate system (e.g., the coordinate system of one of the components of the environment 70 or system 700).

[0095] In an embodiment, the known locations of the microphones 704 and cameras 706, and the parameters of regions 72 and 74 including camera assignments, can be provided to the aggregator 708 and stored in a memory communicatively coupled to the aggregator 708. In some cases, the aggregator 708 may receive the parameters of regions 72 and 74 from a separate computing device internal or external to the system 700. In some cases, the aggregator 708 may also receive the known microphone and camera locations from a separate computing device, which can be used, for example, to operate a graphical tool to automatically determine array and camera locations. In other cases, the aggregator 708 may receive the microphone location from each of the microphones 704 and the camera location and orientation from each of the cameras 706.

[0096] During operation, aggregator 708 can receive location information such as localization coordinates generated when acoustic activity associated with speaker 702 is detected from a plurality of microphones 704. For example, aggregator 708 can combine time-synchronized acoustic localization coordinates (or detected speaker locations) received from two or more of microphones 704 for the same acoustic activity to determine or estimate the location of the active speaker within environment 70 with higher accuracy than the individual coordinates. The localization coordinates can be combined, for example, by determining the common point or intersection of the vectors formed by the localization coordinates, or the closest point between the detected speaker locations as described herein, or by using any other suitable technique.

[0097] Before calculating the estimated speaker location, aggregator 708 can first transform the received coordinates (or detected speaker locations) into a common coordinate system, such that the estimated speaker location is provided within the common coordinate system. The common coordinate system can be one of microphones 704, one of cameras 706, and / or a coordinate system centered on environment 70, or any other pre-determined coordinate system. In some cases, aggregator 708 can transform the estimated speaker location from the common coordinate system into a second coordinate system, such as the coordinate system of camera 706 that will receive the transformed information, such that the estimated speaker location is readily usable by that camera 706. In other cases, a given camera 706 and / or camera controller can be configured to transform the location information received from aggregator 708 into the second coordinate system of the receiving camera 706.

[0098] Aggregator 708 can identify or select the lobes of a plurality of microphones 704 that optimally capture the active speaker and / or the estimated speaker location using the estimated speaker location. For example, aggregator 708 may select a lobe having a lobe location that corresponds to or overlaps with the estimated speaker location. If multiple lobe locations correspond to the speaker location, aggregator 708 can determine the distance between the speaker location within a common coordinate system (e.g., the coordinate system of one of microphones 704 or the environment 70) and each of microphones 704, and determine or identify the microphone 704 having the lobe with the lobe location closest to the identified speaker location and / or closest to the speaker location within the common coordinate system. In some instances, aggregator 708 can determine the direction the speaker 702 is facing (e.g., using the face detection technology of camera 706) and identify the microphone 704 having an available lobe closest to the face of speaker 702 and / or that can be oriented towards the face of the speaker. Aggregator 708 can be configured to calculate the distance between the identified speaker location and each of microphones 704 using the voice localization and triangulation techniques described herein or any other known techniques.

[0099] In other embodiments, the aggregator 708 can be configured to identify a unique speaker location within the environment 70 using network-enabled auto-mixer techniques or the like and select an optimal lobe across multiple microphones 704 to capture the identified speaker location. In such cases, the microphones 704 can be connected together to form a network and receive a common gating control signal from the aggregator 708 that indicates which microphone lobes are gated on and / or which lobes are gated off across the network. As an example, a network auto-mixer can determine which lobe detected the strongest voice signal (or other acoustic signal) for a given acoustic event and generate a common gating control signal by selecting that lobe and the corresponding microphone 704 for the common gating control signal. If the microphone 704 is a microphone array, the microphone 704 can be configured to generate a beamformed acoustic signal based on the acoustics detected by the microphone elements within the microphone 704 and the common gating control signal. The aggregator 708 can then generate a final mixed acoustic signal of the detected acoustic event by aggregating the beamformed acoustic signals received from the microphone 704. Thus, the final mixed acoustic signal can reflect a desired acoustic mix where the acoustics from a particular channel of the microphone 704 are emphasized while the acoustics from other channels are attenuated or suppressed.

[0100] The common gating control signal can also be used by aggregator 708 to determine the estimated location and / or orientation of speaker 702. For example, aggregator 708 can use the common gating control signal and / or other gating determination information to determine which lobe of microphone 704 is available (e.g., gated on). Moreover, given that the common gating control signal is derived based on the strongest vocal signal detected by the network automixer, aggregator 708 can use the arrangement or location of the gated-on lobe to determine the estimated speaker orientation. For example, aggregator 708 can use the network-compatible automixer signal to determine the direction that speaker 702 is facing, and use the gating determination to identify the active lobe and / or microphone 704 that is directed towards the speaker's face and the estimated speaker location, or otherwise can capture sound better in the direction that speaker 702 is facing. As an example, using the gating determination from the network-compatible automixer, it can be determined that although speaker 702 is physically closer to the first microphone 704, the second microphone 704 is better positioned to pick up the vocalization of speaker 702 because speaker 702 is facing the second microphone 704.

[0101] In some cases, the network auto mixer can receive location information directly from one or more microphones. For example, in an embodiment where the microphone 704 includes one or more directional microphones or other non-array microphones (such as lavalier microphones, handheld microphones, boundary microphones, etc.), the location of such microphones, and / or the acoustic pickup pattern or other type of microphone coverage arrangement and / or orientation used by such microphones (collectively referred to herein as "location information") may be known in advance or pre-determined. In such cases, the aggregator 708 can receive the location information of one or more directional microphones 704 as pre-stored data or transmitted at runtime and received from the network auto mixer, and the aggregator 708 can use the location information of the base to determine (or triangulate) the estimated location of the speaker 702 and select an appropriate camera 706 based thereon. Thus, in some cases, the aggregator 708 can perform camera selection based not only on acoustic source localization but also on easily available or known location information obtained from a plurality of microphones 704.

[0102] As shown in FIG. 8, when aggregator 708 selects lobe 705 and / or microphone 704 to optimally capture the acoustics emitted by speaker 702, one or more of the techniques described herein can be used to cause the selected lobe 705 to be deployed towards speaker 702 by a corresponding microphone 704 (e.g., second microphone 704). Additionally, the location of the selected lobe 705 (or "lobe location") can be used by aggregator 708 to determine or identify which of a plurality of cameras 706 is most suitable for capturing an image and / or video of speaker 702 and / or the selected lobe location. In some cases, the gating determination of a network-enabled automixer can also derive which camera 706 should be used. For example, since the network-enabled automixer has a global view of environment 100, it can use the gating determination described above to identify the camera 706 facing speaker 702's face, thus avoiding the selection of a camera facing behind the speaker's head.

[0103] In some cases, aggregator 708 can select the appropriate camera 706 by identifying regions 72, 74 of environment 70 that include or correspond to the selected lobe location and selecting the camera 706 assigned to or configured to capture the identified regions 72, 74. For example, in FIG. 8, since the selected lobe 705 is located within the first region 72 and the first camera 706 is assigned to the first region 72, aggregator 708 can select the first camera 706 to capture speaker 702. Aggregator 708 can transform the selected lobe location from a first coordinate system of a corresponding microphone 704 (e.g., second microphone 704 in FIG. 8) to a second coordinate system associated with environment 70 to identify which of regions 72 and 74 of environment 70 encompasses the selected lobe location.

[0104] Aggregator 708 can position or direct the image capture component of camera 706 towards speaker 702, or more specifically, towards a selected lobe location, by sending or transmitting the coordinates of the selected lobe location to the selected camera 706. In some embodiments, when an appropriate camera 706 is selected, aggregator 708 can convert the lobe location from a second coordinate system of environment 70 to a third coordinate system associated with the selected camera 706, such that the lobe location (or set of coordinates) is received and readily usable by that camera 706. In other embodiments, aggregator 708 can send the lobe location within the first or second coordinate system to the selected camera 706, and the camera 706 and / or camera controller can be configured to convert the received lobe location to a third coordinate system.

[0105] Although FIG. 8 shows only one speaker 702, it should be understood that similar techniques can be applied to cases where multiple speakers 702 are present within environment 70. In particular, aggregator 708 can use one or more of the techniques described herein to determine the best lobe location for each individual speaker 702. For example, as shown in FIG. 9, aggregator 708 can determine that the first lobe 705 of the first microphone 704 is optimal for capturing the audio generated by the first speaker 702, while the second lobe 705 of the second microphone 704 is optimal for capturing the audio generated by the second speaker 702.

[0106] FIG. 9 also shows that two speakers 702 are located at different locations within the same region 72 of the environment 70 and thus cannot be captured by the same camera 706. In such cases, the aggregator 708 can further analyze the camera 706, the microphone 704, and the relative locations of each of the regions 72 and 74 to identify the camera 706 that is most suitable for capturing each of the speakers 702 or has the best angle for that purpose. For example, in FIG. 9, the aggregator 708 can determine that the first camera 706 is closest to the second speaker 702 and has the clearest view of the second speaker 702, and thus, even if both speakers 702 are located within the region 72 assigned to the first camera 706, the second camera 706 can remain capturing the first speaker 702.

[0107] In various embodiments, the microphone 704 and / or the camera 706 can track one or more active speakers 702 in real time as the speaker moves around the environment 70 while speaking or otherwise emitting sound. For example, each of the cameras 706 can be configured to scan its assigned regions 72, 74 for an image similar to a human face and can stay on the identified face as the corresponding speaker 702 moves around (e.g., sits, stands, gestures, changes position, etc.). As another example, the microphone 704 can be configured to continuously or periodically determine the speaker location of newly detected sound and provide the newly identified speaker location to the aggregator 708. The aggregator 708 can combine the new speaker locations from each of the microphones 704 corresponding to the same acoustic activity and, based thereon, determine a more precise estimate of the current location of the speaker.

[0108] In some cases, the speaker 702 may be in an environment 70 that moves around so dramatically or moves to such a new location that the camera 706 and / or microphone 704 that is directed towards the speaker 702 can no longer capture the speaker 702 and / or is no longer the most suitable for doing so. For example, FIG. 10 shows an exemplary scenario in which a second speaker 702 moves from a first location within a first region 72 (e.g., as shown in FIG. 9) to a second location within a second region 74 where the first camera 706 no longer has a clear line of sight to the second speaker 702 (e.g., because the first speaker 702 is in the way).

[0109] In such cases, the aggregator 708 can change the assignment of cameras and / or microphones within the environment 70 based on or in response to the new speaker location. For example, the aggregator 708 can identify or estimate the new location of the second speaker 702 using acoustic localization and / or triangulation techniques as described herein based on new acoustic signals detected by one or more microphones 704. The aggregator 708 can determine that the new speaker location is closer to the first microphone 704 than to the second microphone 704 at this point and that the acoustics emitted at the new speaker location are better captured by the second lobe 705 of the first microphone 704. Accordingly, the first microphone 704 can be directed, for example, by the aggregator 708 to deploy and / or steer the second lobe 705 towards the new speaker location of the second speaker 702. Moreover, the lobe of the second microphone 704 that was originally directed towards the second speaker 702 (e.g., as shown in FIG. 9) can be gated off or deactivated.

[0110] Similarly, aggregator 708 can determine that the new speaker location is closer to the second camera 706 than the first camera 706, and thus can orient the second camera 706 outward from the first speaker 702 (e.g., as shown in FIG. 8) and toward the new speaker location as shown in FIG. 10. Aggregator 708 can also determine that the first camera 706 is optimal at this point for capturing the first speaker 702, and thus can move or steer the first camera 706 toward the first speaker 702, or an estimated speaker location corresponding to the first speaker 702, as shown in FIG. 10.

[0111] In some embodiments, the conferencing system 700 shown in FIGS. 8-10 can be configured to allow a speaker 702 to capture an image of the speaker using one or more personal cameras (not shown) instead of the camera 706 disposed within the environment 70. For example, one or more of the speakers 702 may choose to participate in a conference call or other event using a camera installed on a laptop, tablet, or other personal device that can be brought into the environment 70 by the speaker 702. In such cases, the speaker 702 can connect their personal device to the aggregator 708 or other controller within the environment 70 and use their device to contribute individual videos to the call. In some cases, the personal camera can be used in conjunction with one or more of the cameras 706. For example, a first camera 706 can be used to capture a first speaker 702 located within a first region 72 of the environment 70, while a personal camera can be used to individually capture a second speaker 702 and other speakers (not shown) located within a second region 74 of the environment 70. In any case, the aggregator 708 can be configured to stitch together the received videos or otherwise combine the images captured in other manners into a single video output. For example, the individual videos of each speaker may be presented within separate tiles, as separate stripes (e.g., vertical or horizontal), or within any other compartment of the combined video.

[0112] In embodiments using a personal camera, one or more of the microphones 704 can still provide audio localization information for estimating a speaker location using the techniques described herein and can be used to capture acoustic signals generated within the environment 70 (e.g., for transmission to remote participants in a conference call or other event). The estimated speaker location may be used for purposes other than steering the appropriate (e.g., nearest) camera 706 towards the corresponding speaker 702. For example, when a given speaker 702 is using their own camera, the aggregator 708 can determine the estimated location of that speaker 702 based on, e.g., audio localization provided by one or more of the microphones 704 and can be configured to assign the estimated speaker location to the personal camera or device of the given speaker. In this way, the output of each personal camera can be associated with a particular location within the environment 70. Additionally, the aggregator 708 can be configured to pair each captured acoustic signal with the appropriate personal camera, i.e., the camera that is pointed towards the speaker 702 (or a speaker 702 located at or near the estimated speaker location) that emitted the acoustic signal, using the estimated speaker location. This prevents the aggregator 708 from enabling one of the cameras 706 installed within the environment 70 based on the acoustics detected at the location of the personal camera user.

[0113] In various instances, the aggregator 708 can be configured to identify the speaker 702 (or their device) located within the environment 70 based on an identifier or other identifying information associated with the speaker 702 (or their device). For example, the aggregator 708 may receive an identifier from the speaker's personal device or camera when the speaker 702 enters the environment 70, or it may be provided in advance as part of an ongoing conference call or other event. The identifier may be a device identifier, or other information uniquely associated with the speaker's personal camera (or device), a user identifier, or other information uniquely associated with a given speaker, or any other type of identifier. In some instances, the aggregator 708 can be further configured to use the identifier to assign the estimated speaker location as corresponding to a particular speaker, personal camera, and / or personal device. For example, the aggregator 708 can assign an appropriate identifier (i.e., an identifier corresponding to a particular speaker) to the estimated speaker location.

[0114] FIG. 11 shows an exemplary method or process 800 for determining camera positioning within an environment based on lobe locations selected according to speaker coordinates obtained by one or more microphones (or microphone arrays) according to an embodiment. A camera (e.g., camera 706 of FIG. 8) and one or more microphones (e.g., microphone 704 of FIG. 8) may form part of a conferencing system (e.g., system 700 of FIG. 8) located within an environment (e.g., environment 70 of FIG. 8). The environment may be a conference room, an event space, or other area that includes one or more speakers (e.g., speaker 702 of FIG. 8) or other acoustic sources. The camera may be configured to detect and capture images and / or video of the speaker. The microphone may be configured to detect and capture sound emitted by the speaker and determine the location of the detected sound. Method 800 may be performed by one or more processors of the conferencing system, such as a processor of a computing device included within the system and communicatively coupled to the camera and the microphone. In some instances, method 800 may be performed by an aggregator (e.g., aggregator unit 304 of FIG. 4) and / or a camera controller (e.g., camera controller 308 of FIG. 4) included within the conferencing system.

[0115] As shown in FIG. 11, method 800 can include, at step 802, using a plurality of microphones (or a microphone array) and at least one processor to determine a speaker location in a first coordinate system for a first microphone of the plurality of microphones based on acoustics associated with a speaker. For example, the first coordinate system can be a coordinate system in which the first microphone is at the origin, or any other suitable coordinate system. In an embodiment, determining the speaker location can include detecting, using the first microphone, speech generated near the first microphone, such as acoustics associated with the speaker, and determining the location of the detected speech using an acoustic localization algorithm executed by an acoustic activity localizer of the first microphone (e.g., acoustic activity localizer 208 of FIG. 3).

[0116] Step 804 includes using at least one processor to select a lobe location for a selected microphone of the plurality of microphones in the first coordinate system based on the speaker location in the first coordinate system. In some embodiments, selecting the lobe location can include determining the distance between the speaker location in the first coordinate system and each of the plurality of microphones, and identifying the selected microphone of the plurality of microphones as the one closest to the speaker location in the first coordinate system.

[0117] Step 806 involves using at least one processor to select a first camera among a plurality of cameras based on the lobe location. In some embodiments, selecting the first camera may involve using at least one processor to convert the lobe location from a first coordinate system to a common coordinate system, and identifying a first region among a plurality of regions within the common coordinate system as including the lobe location within the common coordinate system, where each region is assigned to one or more of the plurality of cameras, and identifying the first camera as being assigned to the first region. As an example, the common coordinate system may be the coordinate system of the environment (e.g., environment 70 of FIG. 8), or any other mutually agreed-upon coordinate system.

[0118] Step 808 involves using at least one processor to convert the lobe location to a second coordinate system relative to the first camera so that the lobe location can be readily used by the first camera. For example, the second coordinate system may be a coordinate system with the first camera at the origin, or any other suitable coordinate system. In some cases, the lobe location may be converted from the first coordinate system of the first microphone array to the second coordinate system of the first camera. In other cases, the lobe location may be converted from a third coordinate system of the environment to the second coordinate system of the first camera.

[0119] Step 810 involves transmitting the lobe location within the second coordinate system from at least one processor to the first camera in order to direct the image capture component of the first camera towards the lobe location within the second coordinate system. The first camera can direct the image capture component towards the lobe location within the second coordinate system by adjusting one or more of the camera's angle, tilt, zoom, and framing, or any other relevant parameters of the camera.

[0120] As described herein, in some embodiments, the positions of the cameras and / or microphones included in the conference system described herein relative to each other and / or relative to a given environment are known in advance, for example, by another component of the system or by an external device communicating with the system. However, in other embodiments, the relative positions of the cameras and / or microphones are not initially known. In such cases, the conference system can use acoustic localization algorithms, triangulation techniques, and / or other tools to automatically determine the locations of the cameras, microphones, and / or other acoustic devices within the environment relative to each other.

[0121] In some cases, the location and / or orientation information of a given microphone (e.g., microphone 302a in FIG. 4) can be further improved or refined based on the location and / or orientation information obtained by a camera (e.g., camera 306a in FIG. 4) within the same environment, thereby, for example, improving the coordinate information provided to the camera for positioning towards the active speaker.

[0122] More specifically, the aggregator (e.g., aggregator unit 304 in FIG. 4) and / or camera controller (e.g., camera controller 308 in FIG. 4) of the conference system (e.g., conference system 300 in FIG. 4) can use the pre-determined location of the array within the first coordinate system relative to the camera to direct the camera towards the microphone and configure that first location to be identified as the center or "0" point of the second coordinate system relative to the microphone. Next, the aggregator and / or camera controller can direct towards a second location within the first coordinate system that is known or intended to be the center (0,0,0) of the second coordinate system of the microphone.

[0123] Based on the first location and the second location within the first coordinate system, the discrepancy between the actual center of the second coordinate system of the microphone and the intended center of the coordinate system of the microphone can be calculated. That discrepancy can be sent to the microphone to correct for the deviation of the center of the second coordinate system. The discrepancy can also be used to improve the coordinate transfer or transformation from the second coordinate system to the first coordinate system and vice versa. Additionally, the discrepancy identified by the camera can be used to correct or refine the information indicating the orientation of the camera with respect to the microphone. In some embodiments, the aggregator and / or the camera controller, or other components of the conferencing system, may be configured to automatically calculate the discrepancy between the intended center and the true center of the second coordinate system. In other embodiments, the installer or other user may manually calculate the difference between the two locations and accordingly correct the center coordinates of the second coordinate system, or in other ways, for example, enter this value into the microphone array using an appropriate user interface.

[0124] FIG. 12 shows another exemplary environment 90 in which one or more of the systems and methods disclosed herein may be used. As shown, environment 90 includes a conferencing system 900 that can be utilized to improve the camera positioning for speakers 902 located in different but proximate regions 92 and 94 of environment 90, based on speaker coordinates obtained by microphones 904 located in different regions 92 and 94 of environment 90. In an embodiment, conferencing system 900 includes a first microphone and a second microphone 904, a first camera and a second camera 906, and a first aggregator and a second aggregator 908. The components of conferencing system 900 may be substantially similar to the components of conferencing system 700 shown in FIGS. 8-10. For example, microphone 904 may be similar to microphone 704, camera 906 may be similar to camera 706, and aggregator 908 may be similar to aggregator 708. Accordingly, the individual components of conferencing system 900 are not described in more detail for the sake of brevity.

[0125] In the illustrated embodiment, the environment 90 includes a first region 92 and a second region 94 adjacent to the first region 92. Regions 92 and 94 may be adjacent areas as shown, or there may be a gap or space between them (not shown). In some cases, regions 92 and 94 may be physically separated workspaces or areas (e.g., using one or more partitions, walls, etc.). In other cases, regions 92 and 94 may be "virtual" workspaces or other areas of a shared space that have a boundary of a specified known to aggregator 908 and speaker 902, but may not have a physical wall or other structure between the areas. As shown, the first region 92 includes or encompasses a first speaker 902, a first microphone 904, a first camera 906, and a first aggregator 908. Similarly, the second region 94 includes or encompasses a second speaker 902, a second microphone 904, a second camera 906, and a second aggregator 908. According to an embodiment, regions 92 and 94 can be configured to allow the speakers 902 to individually work or otherwise operate within their respective regions 92 and 94 even though they are part of a shared space. For example, the virtual workspace can allow the speakers 902 to individually participate simultaneously in different video conference calls or other audiovisual events without disturbing each other.

[0126] In an embodiment, to improve the camera positioning or speaker tracking of camera 906, aggregator 908 can be configured to calculate an estimated speaker location of a given speaker 902 using speaker coordinates obtained by various microphones 904 in environment 90, including one or more microphones 904 located in regions 92, 94 different from the given speaker 902. For example, a first microphone 904 can send a first set of speaker coordinates (e.g., x1, y1, z1) of a first estimated location p1 of a first speaker 902 to a first aggregator 908 using the positioning techniques described herein. For the same event, a second microphone 904 can also send a second set of speaker coordinates (e.g., x2, y2, z2) of a second estimated location p2 of the same first speaker 902 to the same first aggregator 908 using the positioning techniques. Using one or more of the techniques described herein, the first aggregator 908 can combine the two sets of coordinates to determine a more accurate estimated speaker location of the first speaker 902. Thus, the conference system 900 can be configured to obtain (or triangulate) the estimated speaker location with higher accuracy by using microphones 904 located outside a given region 92, 94 of environment 90 to increase the number of time-synchronized positionings available for estimating the speaker location.

[0127] Thus, the techniques described herein can help reduce manual measurements typically performed by installers or integrators during the configuration of a conference system, such as measurements of the distance and location between a camera and a microphone. Thus, the amount of time and effort by installers, integrators, and users can be reduced, resulting in increased satisfaction with the installation and use of the conference system.

[0128] Any component of the microphone array 200 and / or the conference systems 100, 300, 500, 700, and 900 may be implemented in hardware (e.g., discrete logic circuits, application specific integrated circuits (ASICs), programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), digital signal processors (DSPs), microprocessors, etc.), using software executable by one or more computers such as computing devices having a processor and memory (e.g., personal computers (PCs), laptops, tablets, mobile devices, smart devices, thin clients, etc.), or through a combination of both hardware and software. For example, some or all components of the microphone array 200 and / or the systems 100, 300, 500, 700, and 900 may be implemented using discrete circuit devices and / or using one or more processors (e.g., an acoustic processor and / or a digital signal processor) that execute program code stored in a memory (not shown), the program code being configured to execute one or more processes or operations described herein, such as the methods shown in FIGS. 5, 7, and 11. Thus, in embodiments, the microphone array 200 and / or any of the systems 100, 300, 500, 700, and 900 may include one or more processors, memory devices, computing devices, and / or other hardware components not shown in the drawings.

[0129] All or portions of the processes described herein, including method 400 of FIG. 5, method 600 of FIG. 7, and method 800 of FIG. 11, may be implemented by one or more processing devices or processors (e.g., analog-to-digital converters, cryptographic chips, etc.) within or external to the corresponding conferencing system (e.g., system 300 of FIG. 4). Additionally, one or more other types of components (e.g., memory, input and / or output devices, transmitters, receivers, buffers, drivers, discrete components, logic circuits, etc.) may also be used, along with the processor and / or other processing components, to perform any, some, or all of the steps of methods 400, 600, and / or 800. As an example, in some embodiments, each of the methods described herein may be performed by a processor executing software stored in a memory. The software may include, for example, program code or computer program modules that include software instructions executable by the processor. In some embodiments, the program code may be a computer program stored on a non-transitory computer-readable medium that is executable by a processor of the associated device.

[0130] The terms "non-transitory computer-readable medium" and "computer-readable medium" include a single medium or multiple media, such as a centralized or distributed database, and / or associated caches and servers that store one or more sets of instructions. Further, the terms "non-transitory computer-readable medium" and "computer-readable medium" include any tangible medium that can store, encode, or carry a set of instructions for execution by a processor or cause a system to perform any one or more of the methods or operations disclosed herein. As used herein, the term "computer-readable medium" is explicitly defined to exclude propagating signals and includes any type of computer-readable storage device and / or storage disk.

[0131] Any process description or block in the drawings should be understood as representing a module, segment, or portion of code that includes one or more executable instructions for performing a particular logical function or step within the process. As would be understood by one of ordinary skill in the art, depending on the functions involved, alternative embodiments may be within the scope of the embodiments of the present invention, where the functions may be performed out of the order illustrated or discussed, including substantially simultaneously or in reverse order.

[0132] It should be understood that the examples disclosed herein may refer to computing devices and / or systems having components that may or may not be physically proximate to each other. Particular embodiments may take the form of cloud-based systems or devices, and the term "computing device" should be understood to include distributed systems and devices (such as cloud-based ones) configured to perform one or more of the functions described herein, as well as those including software, firmware, and other components. Further, as mentioned above, one or more features of a computing device may be physically remote (e.g., a stand-alone microphone) and may be communicatively coupled to the computing device.

[0133] Note that in this specification and the drawings, similar or substantially similar elements may sometimes be labeled with the same reference numerals. However, sometimes these elements may be labeled with different numerals, for example, when such labeling facilitates a clearer description. Additionally, the drawings described in this specification are not necessarily to scale and in some cases, the proportions may be exaggerated to more clearly depict certain features. Such labeling and drawing practices do not necessarily imply a fundamental substantive purpose. As described above, this specification is intended to be taken as a whole and interpreted in accordance with the principles of the invention as taught and understood by those skilled in the art.

[0134] In this disclosure, the use of disjunctive language is intended to include conjunctive language. The use of the definite or indefinite article is not intended to indicate quantity. In particular, references to "the" object or "a" or "an" object are also intended to indicate one of a possible plurality of such objects.

[0135] This disclosure describes, illustrates, and exemplifies one or more specific embodiments of the present invention in accordance with its principles. The disclosure is not intended to limit the true, intended scope and spirit thereof, but rather is intended to explain how to form and use various embodiments in accordance with the technology. That is, the foregoing description is not intended to be exclusive or to limit to the precise forms disclosed herein. Rather, the principles of the invention are intended to be explained and taught such that those skilled in the art can understand these principles and, by that understanding, apply them not only to the embodiments described herein but also to other embodiments that may be contemplated in accordance with these principles. The embodiments provided herein are selected to best explain the principles of the described technology and its practical application and to enable those skilled in the art to make and use various modifications to the technology as adapted to various embodiments and to the particular uses contemplated. All such modifications and variations are within the scope of the embodiments as determined by the appended claims and all equivalents thereof, as may be amended during the pendency of this patent application when construed fairly, legally, and equitably in accordance with the scope of rights accorded.

Claims

1. A method implemented by one or more processors communicating with each of a first microphone, a second microphone, and a camera, the method comprising: using a first microphone array to determine a first speaker location in a first coordinate system relative to the first microphone array based on acoustics associated with a speaker; using a second microphone array to determine a second speaker location in a second coordinate system relative to the second microphone array based on the acoustics associated with the speaker; determining an estimated speaker location in a third coordinate system relative to the camera based on the first speaker location and the second speaker location; transmitting the estimated speaker location in the third coordinate system to the camera to direct an image capture component of the camera towards the estimated speaker location. A method as described above.

2. Converting the first speaker location from the first coordinate system to the second coordinate system; determining an estimated speaker location in the second coordinate system based on the first speaker location and the second speaker location in the second coordinate system; converting the estimated speaker location from the second coordinate system to the third coordinate system. The method according to claim 1, further comprising the above steps.

3. Converting the first speaker location from the first coordinate system to the third coordinate system; Converting the second speaker location from the second coordinate system to the third coordinate system; determining the estimated speaker location in the third coordinate system based on the first speaker location and the second speaker location in the third coordinate system. The method according to claim 1, further comprising the above steps.

4. Determining the estimated speaker location includes: identifying a common point based on the first speaker location and the second speaker location. The method according to claim 1.

5. Determining the estimated speaker location includes: identifying a closest point based on the first speaker location and the second speaker location. The method according to claim 1.

6. Determining the first speaker location includes: The method according to claim 1, comprising determining the location of the voice generated near the first microphone array using an acoustic localization algorithm executed by an acoustic activity localizer.

7. A first microphone array configured to determine a first speaker location in a first coordinate system with respect to the first microphone array based on acoustics associated with a speaker. A second microphone array configured to determine a second speaker location in a second coordinate system with respect to the second microphone array based on the acoustics associated with the speaker. A camera comprising an image capture component. One or more processors communicatively coupled to each of the first microphone array, the second microphone array, and the camera. Comprising The one or more processors Determining an estimated speaker location in a third coordinate system with respect to the camera based on the first speaker location and the second speaker location. Transmitting the estimated speaker location in the third coordinate system to the camera. Are configured to perform The camera is configured to direct the image capture component towards the estimated speaker location received from the one or more processors. A system.

8. The one or more processors Converting the first speaker location from the first coordinate system to the second coordinate system. Determining an estimated speaker location in the second coordinate system based on the first speaker location and the second speaker location in the second coordinate system. Converting the estimated speaker location from the second coordinate system to the third coordinate system. The system according to claim 7, further configured to perform.

9. The one or more processors Converting the first speaker location from the first coordinate system to the third coordinate system. Converting the second speaker location from the second coordinate system to the third coordinate system. Determining the estimated speaker location in the third coordinate system based on the first speaker location and the second speaker location in the third coordinate system The system according to claim 7, further configured to perform the above.

10. The system according to claim 7, wherein determining the estimated speaker location includes identifying a common point based on the first speaker location and the second speaker location.

11. The system according to claim 7, wherein determining the estimated speaker location includes identifying a closest point based on the first speaker location and the second speaker location.

12. The system according to claim 7, further comprising an acoustic activity localizer, wherein determining the first speaker location includes determining the location of the voice generated near the first microphone array using an acoustic localization algorithm executed by the acoustic activity localizer.

13. The system according to claim 7, wherein the camera is configured to direct the image capture component towards the estimated speaker location by adjusting one or more of the angle, tilt, zoom, and framing of the camera.

14. A non-transitory computer-readable storage medium including instructions, which, when executed by one or more processors communicating with each of a first microphone array, a second microphone array, and a camera, cause the one or more processors to Use the first microphone array to determine a first speaker location in a first coordinate system with respect to the first microphone array based on the sound associated with the speaker; Use the second microphone array to determine a second speaker location in a second coordinate system with respect to the second microphone array based on the sound associated with the speaker; Determine an estimated speaker location in a third coordinate system with respect to the camera based on the first speaker location and the second speaker location; Transmit the estimated speaker location in the third coordinate system to the camera to direct the image capture component of the camera towards the estimated speaker location A non-transitory computer-readable storage medium for causing the following to be performed.

15. Causing the one or more processors to convert the first speaker location from the first coordinate system to the second coordinate system; determine an estimated speaker location in the second coordinate system based on the first speaker location and the second speaker location in the second coordinate system; convert the estimated speaker location from the second coordinate system to the third coordinate system The non-transitory computer-readable storage medium according to claim 14, further comprising instructions for causing the above to be performed.

16. Causing the one or more processors to convert the first speaker location from the first coordinate system to the third coordinate system; convert the second speaker location from the second coordinate system to the third coordinate system; determine the estimated speaker location in the third coordinate system based on the first speaker location and the second speaker location in the third coordinate system The non-transitory computer-readable storage medium according to claim 14, further comprising instructions for causing the above to be performed.

17. Determining the estimated speaker location includes identifying a common point based on the first speaker location and the second speaker location. The non-transitory computer-readable storage medium according to claim 14.

18. Determining the estimated speaker location includes identifying a closest point based on the first speaker location and the second speaker location. The non-transitory computer-readable storage medium according to claim 14.

19. Determining the first speaker location includes determining the location of the voice generated near the first microphone array using an acoustic localization algorithm executed by an acoustic activity locator. The non-transitory computer-readable storage medium according to claim 14.

20. The camera directs the image capture component towards the estimated speaker location by adjusting one or more of the angle, tilt, zoom, and framing of the camera. The non-transitory computer-readable storage medium according to claim 14.