Video fence system and method
By adjusting camera parameters based on the audio coverage area defined by the microphone and the speaker's location information, the image focusing problem in the audiovisual system is solved, enabling clear video capture of the speaker's surroundings and effective exclusion of external images, thus improving the professionalism and clarity of video calls.
Patent Information
- Application Number
- CN202480048329.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-07
- Filing Date
- 2024-07-05
- Publication Date
- 2026-02-17
AI Technical Summary
Existing audiovisual systems struggle to effectively capture and focus images within the audio coverage area around the active speaker, while excluding images of non-participants and external environments, leading to unnecessary interference and distraction during video calls.
The visual boundaries of the video are defined by using the microphone's audio coverage area. Based on the speaker location information and audio coverage area information provided by the microphone, the camera parameters are adjusted so that the video focuses on the active speaker and the surrounding audio coverage area, while excluding unwanted images.
It enables image capture to be focused on the desired audio source in an audiovisual system and effectively removes other images located outside the audio coverage area, improving the clarity and professionalism of video calls and reducing interference from non-participants and external environments.
Smart Images

Figure CN121548988A_ABST
Abstract
Description
[0001] Cross-references
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 512,389, filed July 7, 2023, the contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] This disclosure generally relates to focusing a camera on an active speaker, and more specifically, to systems and methods for defining visual boundaries around an active speaker based on speaker location information and corresponding audio coverage area information provided by one or more microphones. Background Technology
[0004] Various audiovisual environments (such as conference rooms, boardrooms, classrooms, video conferencing venues, performance venues, etc.) typically involve the use of microphones (including microphone arrays) for capturing sound from one or more audio sources in the environment (e.g., human speakers) and one or more image capture devices (e.g., cameras) for capturing images and / or video of one or more audio sources or other people and / or objects in the environment. The captured audio and video can be propagated to a local audience in the environment via speakers (for sound enhancement) and displays (for visual enhancement), and / or transmitted to remote locations for remote audiences to listen to and watch (e.g., via television broadcasting, webcasting, etc.). For example, the transmitted audio and video can be used by people in a conference room to conduct telephone conferences with others at remote locations.
[0005] One or more microphones can be used to optimally capture speech and sounds produced by people in the environment. Some existing audio systems ensure optimal audio coverage for a given environment by defining an "audio coverage area," which represents an area in the environment designated for capturing audio signals such as speech produced by human speakers. For example, an audio coverage area defines the space in which beamforming audio pickup lobes can be deployed by microphones. A given environment or room may include one or more audio coverage areas, depending on the size, shape, and type of the environment. For example, the audio coverage area of a typical conference room may include the seating area around the conference table, while a typical classroom may include one coverage area around the blackboard and / or podium at the front of the room and another coverage area around the tables and chairs or other audience areas facing the front of the room. Some audio systems have fixed audio coverage areas, while others are configured to dynamically create audio coverage areas for a given environment.
[0006] Some existing camera systems are configured to point the camera in the direction of a moving speaker, such as a human in the environment who is speaking, singing, or otherwise making sounds, allowing local or remote viewers to see who is speaking. Some cameras use motion sensors and / or facial recognition software to guess which person is speaking for camera tracking purposes. Some camera systems use multiple cameras to optimally capture people located in different parts of the environment or otherwise capture video of the entire environment. Summary of the Invention
[0007] The technology disclosed herein provides systems and methods designed for: (1) using the audio coverage area of a microphone to define one or more visual boundaries of a video captured by a camera; (2) adjusting one or more parameters of the camera based on speaker location information and audio coverage area information provided by the microphone, such that the captured video is focused on the active speaker and the surrounding audio coverage area; and (3) excluding unwanted images from outside the one or more visual boundaries from the captured video.
[0008] An exemplary embodiment includes a method executed by one or more processors communicating with each of at least one microphone and at least one camera, the method comprising: receiving boundary information from at least one microphone defining one or more boundaries of an audio pickup area; receiving sound location information from at least one microphone indicating the detected sound location of an audio source located within the audio pickup area; identifying a first boundary of the one or more boundaries as being near the detected sound location based on the sound location information and the boundary information; calculating a first distance between the detected sound location and the first boundary; determining a depth parameter of at least one camera based on the first distance; and providing the depth parameter and the sound location information to at least one camera.
[0009] Another exemplary embodiment includes a system comprising: at least one microphone configured to provide: boundary information defining one or more boundaries of an audio pickup area, and sound location information indicating the detected sound location of an audio source located within the audio pickup area; at least one camera configured to capture an image of the audio pickup area; and one or more processors communicatively coupled to each of the at least one microphone and the at least one camera, the one or more processors being configured to: receive the boundary information and the sound location information from the at least one microphone; identify a first boundary among the one or more boundaries located near the detected sound location based on the sound location information and the boundary information; calculate a first distance between the detected sound location and the first boundary; determine a depth parameter of the at least one camera based on the first distance; and provide the depth parameter and the sound location information to the at least one camera.
[0010] Another exemplary embodiment includes a method executed by one or more processors communicating with a first camera, a second camera, and at least one microphone, the method comprising: receiving boundary information from at least one microphone defining one or more first boundaries of a first audio pickup area and one or more second boundaries of a second audio pickup area; receiving sound location information from at least one microphone, the sound location information indicating: a first detected sound location of a first audio source located within the first audio pickup area, and a second detected sound location of a second audio source located within the second audio pickup area; identifying the first camera as being near the first audio pickup area and the second camera as being near the second audio pickup area based on the boundary information; and configuring the first camera to capture the first audio pickup area. The system uses an image or video, and configures a second camera to capture an image or video of a second audio pickup area; identifies a first boundary of one or more first boundaries as being near a first detected sound location based on sound location information and boundary information, and identifies a second boundary of one or more second boundaries as being near a second detected sound location; calculates a first distance between the first detected sound location and the first boundary, and a second distance between the second detected sound location and the second boundary; determines a first depth parameter of the first camera based on the first distance; determines a second depth parameter of the second camera based on the second distance; provides the first detected sound location and the first depth parameter to the first camera; and provides the second detected sound location and the second depth parameter to the second camera.
[0011] Another exemplary embodiment includes a system comprising: a first camera; a second camera; at least one microphone configured to provide: boundary information: a first audio pickup area defined by one or more first boundaries, and a second audio pickup area defined by one or more second boundaries; and sound location information indicating: a first detected sound location of a first audio source located within the first audio pickup area, and a second detected sound location of a second audio source located within the second audio pickup area; and one or more processors communicatively coupled to each of the first camera, the second camera, and the at least one microphone, the one or more processors being configured to: receive the boundary information and the sound location information from the at least one microphone; and identify the first camera as being near the first audio pickup area and the second camera as being in the first audio pickup area based on the boundary information. The system is configured to capture images or videos of the second audio pickup area; a first camera is configured to capture images or videos of the second audio pickup area; based on sound location information and boundary information, a first boundary of one or more first boundaries is identified as being located near a first detected sound location, and a second boundary of one or more second boundaries is identified as being located near a second detected sound location; a first distance between the first detected sound location and the first boundary, and a second distance between the second detected sound location and the second boundary are calculated; a first depth parameter of the first camera is determined based on the first distance; a second depth parameter of the second camera is determined based on the second distance; the first detected sound location and the first depth parameter are provided to the first camera; and the second detected sound location and the second depth parameter are provided to the second camera.
[0012] These and other embodiments, as well as various arrangements and aspects, will become apparent and more fully understood from the following detailed description and accompanying drawings, which illustrate exemplary embodiments that may demonstrate various ways in which the principles of the invention can be applied. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of an exemplary environment including an audiovisual system according to one or more embodiments, the audiovisual system being configured to capture and focus images on one or more audio sources within an audio coverage area and exclude unwanted images outside the audio coverage area.
[0014] Figure 2 This is a schematic diagram of another exemplary environment including an audiovisual system according to one or more embodiments, the audiovisual system being configured to focus a first camera on a first audio source in a first audio coverage area and a second camera on a second audio source in a second audio coverage area.
[0015] Figure 3 This is a schematic diagram of another exemplary environment including another audiovisual system according to one or more embodiments, which can be used to focus image capture on a selected audio source within an audio coverage area and exclude unwanted images outside the audio coverage area.
[0016] Figure 4 This is a schematic diagram of another exemplary environment including another audiovisual system according to one or more embodiments, which can be used to individually focus image capture on each of a plurality of audio sources within an audio coverage area.
[0017] Figure 5 This is a block diagram of an exemplary audiovisual system according to one or more embodiments.
[0018] Figure 6 This is a flowchart illustrating exemplary operations for using an audiovisual system to focus image capture on a selected audio source, according to one or more embodiments.
[0019] Figure 7 is a schematic diagram illustrating an exemplary image field of an existing camera. Detailed Implementation
[0020] Generally, an audio system can use an audio coverage area to focus one or more beamformed audio pickup lobes onto sound generated by an audio source located within a predefined area or receiving area of a given environment (e.g., a room), and the audio signal captured by the audio pickup lobes can be provided to the appropriate channels of an automatic mixer to generate the desired audio mix. In some cases, an audio system can form an "audio fence" around one or more audio coverage areas to prevent or block unwanted sounds generated by audio sources located outside the audio coverage area from entering the desired audio output (e.g., mixing of audio signals captured within the audio coverage area). For example, an audio fence can be formed by muting any lobes deployed toward audio sources located outside the audio coverage area, such that audio signals captured outside the audio coverage area are not included in the desired audio output.
[0021] In hybrid work environments, meeting spaces and other workspaces shared by multiple individuals have become increasingly common and popular. Some shared workspaces have public audio systems and can therefore benefit from the use of technologies such as audio fencing to help prevent video calls, conference calls, etc., from disturbing others working in neighboring workspaces or attending individual meetings. In workspaces that also utilize public camera systems, it can also be unintended or disruptive if non-participants or others in neighboring workspaces are accidentally caught in a video call, etc., due to proximity.
[0022] The systems and methods described herein improve the configuration and use of audiovisual systems (such as conferencing systems, stage performance systems, gaming systems, etc.) by defining a “video fence” around one or more audio sources using an audio coverage area to focus image capture on the desired audio source and exclude or remove other images located outside the corresponding audio coverage area. In embodiments, the video fence can be configured based on the boundary of the audio coverage area and further based on the distance between the desired audio source and the boundary (or boundary line) of the audio coverage area behind the audio source. For example, the boundary of the audio coverage area can be used to define the image field of a camera for an audiovisual system used to capture video of the audio source. The distance from the audio source to the boundary line can be used to determine a depth-of-field parameter used to adjust the camera's focus area so that the desired audio source is clearly visible or in focus, and areas outside the boundary line or outside the video fence are blurred and out of focus, or otherwise excluded from the video. In some embodiments, instead of blurring the external image or other than blurring the external image, additional image enhancement can be applied to the corresponding portion of the image field, such as, for example, overlaying a selected image or video on top of the external image, so that areas outside the video fence are completely excluded from the video.
[0023] As used herein, the terms "lobe" and "microphone lobe" refer to an audio beam generated by a given microphone array (or array microphone) to pick up audio signals at a selected location, such as the location the lobe points to. While the techniques disclosed herein are described with reference to microphone lobes generated by array microphones, the same or similar techniques can be used with other forms or types of microphone coverage (e.g., cardioid patterns, etc.) and / or with microphones that are not array microphones (e.g., handheld microphones, boundary microphones, lavalier microphones, etc.). Therefore, the term "lobe" is intended to encompass any type of audio beam or coverage.
[0024] Referring first to Figure 7, an exemplary image field of a conventional camera 10 is shown. Figure 7 is provided to aid in the illustration and explanation of various terms generally used with cameras. As shown, the image field (or field of view) extends a selected distance or length L in front of the camera 10. For example, the image field may extend from the lens of the camera 10 to a far end located above an object 12 (e.g., a person) to be captured. As shown, not all areas of the image field are in focus, or sharp and clearly visible. Instead, the image field includes areas known as “depth of field,” which indicate a specific depth or range of the image field with an acceptable level of sharpness, or otherwise appearing to be in focus to the human eye. Depth of field includes a focal plane that may be centered on the focal point or subject to be captured in the image, such as, for example, object 12 in Figure 7. Other areas of the image field within the depth of field may also appear to be in focus, such as, for example, in-focus areas 1 and 2 located on either side of the focal plane in Figure 7. Conversely, areas outside the depth of field but still within the image field of the camera 10 may appear out of focus or blurred, such as, for example, out-of-focus areas 1 and 2 in Figure 7. For example, the out-of-focus area may be too close to or too far from the camera 10.
[0025] Therefore, depth of field can create a "focus area" (or apparent focus) around object 12. Generally, a shallow depth of field creates a narrower focus area, while a deeper depth of field creates a wider focus area. As shown in Figure 7, the focus area extends a first distance or length L1 in front of or toward camera 10 and a second length L2 behind or away from camera 10. Furthermore, the first length L1 includes a focus area 2 located between the focal plane and camera 10, and the second length L2 includes a focus area 1 located between the focal plane and the far end or boundary of the image field. It should be understood that in some cases, the first length L1 and the second length L2 may not be equal distances, and the distance on either side of the focal plane may depend on the aperture size of camera 10 and / or other camera settings.
[0026] Now for reference Figure 1An exemplary audiovisual environment 100 according to an embodiment is shown, in which one or more of the systems and methods disclosed herein may be utilized. As shown, environment 100 includes a microphone 102, a camera 104, and one or more audio sources 106 located within an audio coverage area 108 of the microphone 102. Environment 100 may be a conference room, boardroom, classroom, or other meeting room; a theater, stadium, auditorium, or other performance or event venue; or any other space. One or more audio sources 106 may be human speakers or talkers participating in a teleconference, television broadcast, webcast, class, seminar, performance, sporting event, or any other activity, and may be located at different locations around environment 100. For example, one or more audio sources 106 may be local participants in a teleconference sitting in corresponding chairs 110 arranged around a table 112 (such as...). Figure 1 (as shown), or local audience members seated in chairs arranged in front of the podium or other presentation space (e.g., as shown). Figure 2 (As shown). Although Figure 1 This paper presents one potential environment, but it should be understood that the systems and methods disclosed herein can be utilized in any applicable environment.
[0027] Microphone 102 can be configured to detect sounds from audio source 106, such as human voices or speech and / or music spoken by audio source 106, clapping, or other sounds generated by that audio source, and convert the detected sounds into one or more audio signals. Although Figure 1Only one microphone 102 is shown, but microphone 102 may include one or more of an array microphone, a non-array microphone (e.g., a directional microphone, such as a lavalier, boundary, etc.), or any other type of audio input device capable of capturing speech and other sounds. As an example, microphone 102 may include, but is not limited to, SHURE MXA310, MX690, MXA910, MXA920, MXW1 / 2 / 8, ULX-D, etc. Microphone 102 may be placed in any suitable location, including walls, ceilings, tables, podiums, and / or any other surface in environment 100, and may conform to various sizes, form factors, mounting options, and wiring options to suit the needs of a particular environment. For example, one or more microphones may be placed on a table, podium, or other surface near the audio source in a classroom or conference room environment, or may be attached to an audio source, such as a performer or speaker, in an auditorium, stadium, or concert hall environment. In some cases, one or more microphones may also be mounted overhead or on a wall to capture sound from a larger area, such as an entire room or hall. The exact type, number, and placement of microphones in a given environment can depend on the audio source, the location of the audience, physical space requirements, aesthetics, room layout, stage layout, and / or other considerations. In the illustrated embodiment, microphone 102 may be positioned at a selected location within environment 100 to adequately capture sound throughout environment 100.
[0028] Camera 104 can be configured to capture still images or pictures, moving images, videos, or other images of the environment 100 visible within the image field of camera 104. Various parameters or settings can be used to control and / or configure one or more aspects of camera 104. For example, image field parameters define the image field (or visible frame) of camera 104. Image field parameters can be adjusted such that the image field includes a selected area of environment 100, such as, for example, an area including one or more audio sources 106 located around table 112, or more generally, an audio coverage area 108. In embodiments, image field parameters may include distance values for configuring the length or other dimension of the image field (e.g., length L in FIG. 7), which determines how far the visible frame extends in front of camera 104. As another example, depth-of-field parameters define the depth of field of camera 104. Depth-of-field parameters can be configured or adjusted such that only a selected portion of the image field is part of the depth of field, or appears to be in focus. In embodiments, depth-of-field parameters may include distance values for configuring the length or other dimension of the depth of field, which determines how much of the audio coverage area 108 is in focus (or within the focus area).
[0029] In some embodiments, camera 104 may be a standalone camera, while in other embodiments, camera 104 may be a component of an electronic device (e.g., a smartphone, tablet, etc.). In some cases, camera 104 may be included in the same electronic device as one or more other components of environment 100 (such as, for example, microphone 102). Camera 104 may be a pan-tilt-zoom (PTZ) camera capable of physically moving and zooming to capture desired images and video, or it may be a virtual PTZ camera capable of digitally cropping and scaling images and video to one or more desired portions. Environment 100 may also include a display (such as a television or computer monitor) for displaying images and / or video or other image or video content, for example, associated with remote participants in a telephone conference. In some embodiments, for example, in addition to or including microphone 102 and / or camera 104, the display may also include one or more microphones, cameras, and / or speakers.
[0030] Audio coverage area 108 (also referred to herein as "audio pickup area") represents the area where microphone 102 accepts audio pickup. Specifically, audio coverage area 108 defines an area or space within which microphone 102 may deploy or focus beamforming audio lobes (not shown) for capturing or detecting desired audio signals, such as sound generated by one or more audio sources 106 located within audio coverage area 108. In embodiments, microphone 102 may be part of an audio system (or audiovisual system) configured to define audio coverage area 108 based on, for example, predetermined audio coverage information, the known or calculated location of microphone 102, the known or expected location of one or more audio sources 106, and / or the real-time location of audio sources 106.
[0031] For example, such as Figure 1 As shown, environment 100 may also include one or more other audio sources 114 located outside the audio coverage area 108. The one or more other audio sources 114 (also referred to herein as “out-of-area audio sources”) may be human speakers or talkers located in nearby workspaces or other areas of environment 100, sufficiently close to the audio coverage area 108 to be within the image field of camera 104, or otherwise included in the visible frame during image capture of the one or more audio sources 106. For example, as Figure 1 As shown, another audio source 114 could be a human speaker sitting at another table 113 near table 112 in environment 100, but not participating in a telephone conference or other meeting taking place within the audio coverage area 108. In some cases, another audio source 114 could be located within another audio coverage area (not shown) of microphone 102 (see example...). Figure 2In some cases, environment 100 may include another audio source 114 configured to capture environment 100 (see, for example...). Figure 2 The second camera (not shown) in the second area of the second region.
[0032] like Figure 1 As shown, environment 100 may further include control module 116 for enabling one or more aspects of a conference or event, such as teleconferencing, webinars, television broadcasting, or otherwise conducting a meeting or event, and / or implementing one or more of the techniques described herein. Control module 116 may be implemented in hardware, software, or a combination thereof. In some embodiments, control module 116 may be a standalone device, such as a controller, control device, computing device, or other electronic device, or may be included in such a device. In other embodiments, all or part of control module 116 may be included in microphone 102 and / or camera 104. In one exemplary embodiment, control module 116 may be a general-purpose computing device including a processor and storage devices. In another exemplary embodiment, control module 116 may be part of a cloud-based system or otherwise reside on an external network.
[0033] It should be understood that Figure 1 The components shown are merely exemplary, and any number, type, and placement of various components in environment 100 are contemplated and possible, including, for example, different arrangements of audio source 106, audio source 106 moving around the room, different arrangements of audio coverage area 108, different locations of microphone 102 and / or camera 104, different numbers of audio source 106, microphone 102, camera 104, and / or audio coverage area 108, etc.
[0034] In various embodiments, the control module 116, microphone 102, and camera 104 can form an audiovisual system (such as, for example, Figure 5 The audiovisual system 500 shown, or part of such an audiovisual system, is configured to define a “video fence” around one or more audio sources 106, such that other audio sources 114 and any other person or object located outside the audio coverage area 108 are not visible in the images and / or videos captured by camera 104. The video fence can be implemented by camera 104 using one or more parameters provided by control module 116, such as, for example, image field parameters for defining the image field or field of view of camera 104 and / or depth field parameters for adjusting the depth of field or focus area of camera 104. One or more parameters can be configured by control module 116 based on information received from microphone 102, such as, for example, information defining the location of the active speaker 106 and / or one or more boundaries of the audio coverage area 108.
[0035] More specifically, according to an embodiment, microphone 102 can be configured to provide boundary information that defines one or more boundaries or boundary lines 118 of audio coverage area 108. Boundary lines 118 may depict the outer limit of audio coverage area 108, or the point where coverage area 108 ends. The number of boundary lines 118 used to create a given audio coverage area can vary depending on the general shape of the area. For example, Figure 1 The audio coverage area 108 is configured to have a generally rectangular shape and therefore four boundary lines 118. In some embodiments, the boundary information may also indicate the general shape of the audio coverage area 108 (e.g., rectangle, square, triangle, octagon, polygon, circle, ellipse, etc.) such that, for example, the expected number of boundary lines 118 and / or other information (e.g., the angle at which the boundary lines 118 intersect, etc.) can be predetermined. For example, if the audio coverage area has a circular or elliptical shape, the boundary information will define only one (continuous) boundary line. Although the illustrated embodiment depicts boundary “lines,” in other embodiments, the audio coverage area 108 may be defined by other types of boundaries 118 (such as, for example, one or more points, shapes, or other indicators).
[0036] In some embodiments, microphone 102 may provide control module 116 with additional information about audio coverage area 108, such as identification information for identifying each coverage area associated with microphone 102 (e.g., area 1, area 2, etc.), location information for indicating the relative location of audio coverage area 108 or any other coverage area within environment 100, activity information for indicating which coverage area is currently active, or any other relevant information. In some embodiments, control module 116 may store boundary information of each coverage area previously identified by or associated with microphone 102 in memory, and upon receiving activity information from microphone 102, control module 116 may retrieve the corresponding boundary information from memory.
[0037] Boundary information can be used to define each boundary or boundary line 118 of the audio coverage area 108 using one or more coordinates (e.g., a set of endpoint coordinates), vectors, or any other suitable format. In some embodiments, each boundary line 118 can be defined by coordinates (e.g., Cartesian or rectangular coordinates, spherical coordinates, etc.) representing one or more points along the line 118 (e.g., the start point, end point, and / or center point of the line 118). For example, in Figure 1In this context, the first boundary line 118 (or "first boundary line") can be defined by a first set of coordinates (a1, b1, c1) representing the center point p of the first boundary line 118. The boundary information may be previously known and stored in the memory of the microphone 102 and / or may be determined at least in part by the microphone 102, for example, using a processor of the microphone 102 based on other known information about the audio coverage area 108 (e.g., location, size, center point, number of boundary lines, shape of the coverage area, etc.). In some embodiments, the boundary line coordinates received at the control module 116 may be relative to the coordinate system of the microphone 102. In other embodiments, the boundary line coordinates may be relative to the coordinate system of the environment 100 and can be translated or transformed to the coordinate system of the microphone 102 by the control module 116 or the microphone 102, or vice versa.
[0038] Microphone 102 can also be configured to provide sound location information indicating the detected sound location of an active audio source 106 (or “active speaker”) located within the audio coverage area 108. The detected sound location, or the location where microphone 102 detects audio or sound generated by the active speaker 106, can be relative to microphone 102 and can be provided as a set of coordinates. For example, microphone 102 can be configured to generate the localization of the detected sound and determine coordinates (or “localization coordinates”) representing the position of the detected sound relative to microphone 102. Various methods for generating sound localization are known in the art, including, for example, generalized cross-correlation (“GCC”). Localization coordinates can be Cartesian or rectangular coordinates representing a three-dimensional location point, or x, y, and z values. For example, using localization techniques, microphone 102 can identify the location of the active speaker 106 as a detected sound location s with coordinates (x1, y1, z1). In some embodiments, the localization coordinates can be converted to polar or spherical coordinates, i.e., azimuth (phi), elevation (theta), and radius (r), as known in the art, for example using transformation formulas. Spherical coordinates can be used in various embodiments to determine additional information about the audio system, such as, for example, the distance between the active speaker 106 and the microphone 102. In some embodiments, the location coordinates of the detected sound location can be relative to the coordinate system of the microphone 102 and can be transformed or translated to the coordinate system of the environment 100, or vice versa.
[0039] In some embodiments, in addition to or instead of the audio source location coordinates, the control module 116 may receive other types of information for identifying the speaker's location. For example, the environment 100 may further include one or more other sensors (i.e., in addition to the microphone 102) configured to detect or determine the current location of a human speaker or other audio source within the audio coverage area. Such additional sensors may include thermal sensors, time-of-flight (“ToF”) sensors, optical sensors, and / or any other suitable sensors or devices.
[0040] In an embodiment, the control module 116 can be configured to use boundary information to determine or adjust the image field parameters of the camera 104 such that the audio coverage area 108 and the audio source 106 located therein fall within the image field or visible frame of the camera 104. For example, the control module 116 can determine or calculate one or more distance values of the image field parameters (e.g., length L in FIG. 7) based on the location of one or more boundaries 18 relative to the camera 104, or otherwise configure the image field parameters such that at least the boundary 118 of the audio coverage area 108 is included within the visible frame of the camera 104. The control module 116 can provide the image field parameters to the camera 104, and the camera 104 can adjust its image field accordingly.
[0041] In some embodiments, the position of camera 104 relative to microphone 102 and / or environment 100 may be previously known and stored in the memory of camera 104. In such cases, camera 104 may be configured to provide camera location information to control module 116, and control module 116 may be configured to use both camera location information and boundary information to optimize or determine image field parameters. For example, control module 116 may first use the camera location information to determine the location of camera 104 relative to audio coverage area 108, or more specifically, relative to each of one or more boundaries 118. This may include, for example, determining the distance from camera 104 to each boundary line 118, determining the orientation of the audio coverage area 108 relative to the lens or field of view of camera 104, and / or determining which boundary line 118 is closest to or adjacent to camera 104 and which boundary line 118 is opposite or on the opposite side of camera 104. Control module 116 may use the relative location of camera 104 to determine or adjust the image field parameters of camera 104 such that the entire audio coverage area 108 is visible within the image field.
[0042] In some cases, although the image field is configured based on one or more boundaries 118 of the audio coverage area 108, the camera 104 can still capture images located outside the audio coverage area 108, such as areas near or adjacent to the boundary line 118, or areas otherwise visible at a distance and / or far to the target area. For example, in Figure 1 In this context, camera 104 can capture another audio source 114 and / or one or more other tables 113 located behind the speaker 106 of the event and outside the audio coverage area 108. In various use cases, it may not be desirable to include people and / or objects not intended to be part of a teleconference or other audiovisual activity. For example, others may disagree with having their images captured, or may be perceived as a nuisance or distraction by participants in the event.
[0043] In embodiments, control module 116 may be configured to adjust or optimize the image field, or more specifically, adjust or optimize the focus area within the image field, such that any images of unwanted people, objects, and / or areas (e.g., areas outside audio coverage area 108) have limited visibility or no visibility, or are otherwise excluded from the captured images and / or video. To achieve this, control module 116 may first determine the position of the active speaker 106 relative to microphone 102 and / or the audio coverage area 108 of microphone 102. According to various embodiments, control module 116 may be configured to use the location coordinates of detected sound locations s (or "speaker locations") to determine the relative location of the active speaker 106 within audio coverage area 108 or the location of the speaker 106 relative to the boundary of coverage area 108. For example, based on the location coordinates and boundary information of the audio coverage area 108, the control module 116 can determine which boundary line 118 is closest to or closest to the active speaker 106 by calculating the distance between the detected sound location s and a known point on each boundary line 118 and comparing the calculated distances to identify the minimum distance. In the example shown, the control module 116 can determine that the detected sound location s is closest to the first boundary or boundary line 118a directly behind the active speaker 106 based on the distance between the first set of coordinates (a1, b1, c1) of the first boundary line 118a and the location coordinates (x1, y1, z1) (or "speaker coordinates").
[0044] The control module 116 can further determine the relative location of the active speaker 106 by calculating the amount of audio coverage area 108 remaining between the active speaker 106 and the first or nearest boundary line 118a, or otherwise extending beyond the detected sound location s. In some embodiments, the control module 116 can quantify this amount by calculating the proximity or first distance d between the detected sound location s and the nearest boundary line 118a. For example, the control module 116 can determine that a second point p2 of the first boundary line 118a, represented by a second set of coordinates (a2, b2, c2), is closest to the detected sound location s. Figure 1 As shown, the second point p2 can be used to calculate the first distance d (e.g., by calculating the distance between the speaker coordinates (x1, y1, z1) and the second set of coordinates (a2, b2, c2)). As another example, the control module 116 can find the distance between the detected sound location s and the first boundary line 118a, i.e., the first distance d, by calculating the second distance between the microphone 102 and the center point p of the first boundary line 118a (e.g., using the first set of coordinates (a1, b1, c1)), the third distance between the microphone 102 and the detected sound location s (e.g., using the speaker coordinates), and quantifying the remaining amount of coverage area 108 by subtracting the third distance from the second distance. In either case, the control module 116 can use the first distance d to determine the amount of audio coverage area 108 remaining behind or far from the active speaker 106.
[0045] Once the position of the active speaker 106 relative to the audio coverage area 108 is determined, the control module 116 can use the relative position information to optimize the image field of the camera 104, such that only areas of the environment 100 falling within the audio coverage area 108 appear sharp and in focus, while any areas outside the audio coverage area 108 are out of focus, blurred, or otherwise limited in visibility. According to various embodiments, the control module 116 can achieve this by using a first distance d between the detected sound location s and a first (or nearest) boundary line 118a to determine or adjust the depth-of-field parameters of the camera 104. For example, the control module 116 can be configured to select or calculate distance values for the depth-of-field parameters based on the first distance d, or otherwise configure the depth of field of the camera 104 to extend further backward from the detected sound location s by no more than the first distance d. In this way, the control module 116 can use a first distance d to adjust the focus area of the camera 104 to include a first region 120 of the image field between the active speaker 106 and the first boundary line 118 (e.g., focus area 1 in FIG. 7), and exclude a second region 122 of the image field that extends beyond the first boundary line 118a (e.g., out-of-focus area 1 in FIG. 7), or a distance s from the detected sound location that exceeds the first distance d. In some embodiments, the depth-of-field parameter can be configured such that a second length of the depth of field (e.g., L2) or a distance value of the portion of the focus area extending behind the active speaker 106 is substantially equal to or less than the first distance d. Other techniques for configuring the depth-of-field parameter based on the relative position of the active speaker 106 within the audio coverage area 108 are also contemplated and can be used in place of or in addition to the techniques described above.
[0046] In embodiments where the camera location is known or determined, the control module 116 can be configured to use camera location information and / or relative camera location information to optimize the selection of boundary lines for configuring depth-of-field parameters of camera 104. For example, using the location of camera 104 relative to audio coverage area 108, the control module 116 can determine which boundary line 118 is located at or near the far end of the camera's image field (or relative to camera 104), and is therefore most likely to be positioned behind the active speaker 106. For example, in Figure 1 In this context, control module 116 can use camera location information to identify the first boundary line 118a as being located near the far end of the image field of camera 104 (or on the opposite side of the camera's field of view). Control module 116 can then use the identified boundary line 118a to configure depth parameters. For example, control module 116 can calculate a distance value for the depth parameters based on the distance between camera 104 and the identified boundary line 118a (e.g., using a second set of coordinates (a2, b2, c2)), or more specifically, calculate a second length behind the active speaker 106.
[0047] In some embodiments, in addition to configuring the depth of field of camera 104 such that only areas of the image field coinciding with the audio coverage area 108 are in focus, control module 116 may also be configured to apply image enhancement to the image depicting the second region (e.g., out-of-focus region 1 in FIG. 7) or the portion of the image field extending beyond the first boundary line 118a and outside the audio coverage area 108. For example, video captured by camera 104 may include an image of an active speaker 106 sitting in the foreground of the video in a first region 120 and an image of the second region 122 in the background of the video. Control module 116 can alter or enhance such video by adding image enhancement to cover, exclude, or block the second region 122 from the captured video. Image enhancement may be applied only to the portion of the captured image depicting the second region 122 and may leave the rest of the image unaffected. In some embodiments, image enhancement may be a selected image displayed above or over the portion of the captured image showing the second region 122, such that the image of the second region 122 is no longer visible in the video output by camera 104. In other embodiments, image enhancement may be a blurring effect or other type of visual effect applied to the image of the second region 122 to further reduce the visibility of the second region 122 within the video output by the camera 104. In some cases, the blurring effect may be an additional blurring or masking layer applied to the captured image after adjusting the focus area to exclude the second region 122. That is, the blurring effect may blur or adjust the captured image in a different manner or to a different degree than the out-of-focus blur that occurs due to placing the region outside the depth of field. Other types of image enhancement are contemplated and may be included in addition to or instead of the examples provided herein.
[0048] Figures 2 to 4 Additional audiovisual environments 200, 300, and 400 according to various embodiments are shown, in which the systems and methods disclosed herein can be utilized. Several aspects of these environments or use cases may be similar to... Figure 1 Those aspects of the audiovisual environment 100 shown are described herein. For example, each of environments 200, 300, and 400 includes those related to... Figure 1 At least one microphone that is substantially similar to microphone 102, and Figure 1 At least one camera that is substantially similar to camera 104 and with Figure 1 The control module 116 is substantially similar to at least one control module. Therefore, for the sake of brevity, aspects of environments 200, 300, and 400, which are common to environment 100, will not be described in detail in the following paragraphs.
[0049] Now for reference Figure 2The document illustrates an exemplary use case where multiple video fences are provided to allow the use of multiple cameras to capture images and / or video of an audio source located within a separate audio coverage area of a microphone. In such cases, the techniques described herein can be used to create a first video fence around a first audio coverage area covering a first audio source, and a second video fence around a second audio coverage area covering a second audio source. The first video fence can be used by a first camera to capture an image of the first audio source and exclude or blur images of areas outside the first audio coverage area. Similarly, the second video fence can be used by a second camera to capture an image of a second audio source and exclude or blur images of areas outside the second audio coverage area.
[0050] More specifically, the audiovisual environment 200 includes a microphone 202, a first camera 204, a second camera 205, a first audio source 206 located within a first audio coverage area 208 of the microphone 202, and a second audio source 207 located within a second audio coverage area 209 of the microphone 202. Environment 200 can be a classroom, lecture hall, auditorium, courtroom, church, or other place of worship, or any other activity space having a first designated area for a presenter or performer (e.g., the first audio source 206) and a second designated area for one or more audience members (e.g., the second audio source 207 and / or one or more other audio sources 214), as illustrated. The first area (or “presenter space”) may include a podium, desk, stage, etc. The second area (or “audience space”) may include one or more tables and / or chairs or other types of seating. As an example, environment 200 can be used to capture audio and / or video of lectures, meetings, performances, or other events.
[0051] According to an embodiment, microphone 202 can be configured to assign each of audio coverage areas 208 and 209 to a corresponding area of environment 200, such that audio or sound generated in each area of environment 200 can be captured as a separate audio signal. For example, first audio coverage area 208 can be used to capture sound generated in the presenter space, and second audio coverage area 209 can be used to capture sound generated in the audience space (or vice versa). In most cases, presenter 206 can be the primary audio source in environment 200, and therefore, first audio coverage area 208 can be active or "on" for most activities. In such cases, second audio coverage area 209 can be inactive or "off" or otherwise used to prevent audio generated in the audience space from being included in the presenter's audio. In some cases, one or more of audio sources 207 and 214 in the audience space (or audience members) can also speak or otherwise generate audio. For example, at some times, second audio source 207 can actively speak simultaneously with or in place of first audio source 206. In such cases, the second audio coverage area 209 can be turned on or otherwise activated so that audio generated by the second audio source 207 (and / or others) can also be captured.
[0052] According to various embodiments, environment 200 also includes a control module 216 communicatively coupled to each of microphone 202, first camera 204, and second camera 205. Like control module 116, control module 216 can be configured to implement one or more aspects of a meeting or activity occurring within environment 200 and / or implement one or more of the techniques described herein.
[0053] In an embodiment, the control module 216, microphone 202, first camera 204, and second camera 205 can form an audiovisual system (such as, for example, Figure 5The audiovisual system 500 shown, or part of an audiovisual system, is configured to define a first video fence around presenter 206 such that second audio source 207 and / or other audio sources 214, as well as any other person or object located outside the first audio coverage area 208, are not visible in the images and / or videos captured by first camera 204. In some embodiments, control module 216 is further configured to define a second video fence around one or more of audience members 207 and 214 such that presenter 206, as well as any other person or object located outside the second audio coverage area 209, are not visible in the images and / or videos captured by second camera 205. The first and second video fences may be implemented by first camera 204 and second camera 205 respectively using one or more parameters provided by control module 216, such as, for example, a first field parameter for defining a first image field of first camera 204, a second field parameter for defining a second image field of second camera 205, a first depth-of-field parameter for adjusting a first depth of field of first camera 204, and / or a second depth-of-field parameter for adjusting second camera 205. One or more parameters can be configured by the control module 216 based on information received from the microphone 202, such as information defining, for example, the location of the first audio source 206, the location of the second audio source 207, one or more boundaries of the first audio coverage area 208 and / or one or more boundaries of the second audio coverage area 209.
[0054] More specifically, like microphone 102, microphone 202 can be configured to provide boundary information defining one or more first boundaries or boundary lines 218 of a first audio coverage area 208 and one or more second boundaries or boundary lines 219 of a second audio coverage area 209. The boundary information may include one or more sets of coordinates for defining boundaries 218 and 219. For example, a first boundary or boundary line 218a of one or more boundaries 218 may be defined by a first set of coordinates (a1, b1, c1) representing a first center point p1 of the first boundary line 218a. As another example, a second boundary or boundary line 219a of one or more second boundaries 219 may be defined by a second set of coordinates (a2, b2, c2) representing a second center point p2 of the second boundary line 219a. In the illustrated embodiment, the first boundary line 218a is located behind and closest to the first audio source 206, while the second boundary line 219a is located behind and closest to the second audio source 207.
[0055] Similar to microphone 102, microphone 202 can also provide sound location information to control module 216. For example, the sound location information may include a first set of coordinates (x1, y1, z1) representing a first detected sound location s1 of the first audio source 206 and a second set of coordinates (x2, y2, z2) representing a second detected sound location s2 of the second audio source 207. In some embodiments, control module 216 may be configured to combine the sound location information with other types of sensor information (such as, for example, thermal, ToF, optical, etc.) to more accurately identify the speaker's location.
[0056] Using the techniques described herein, the first boundary line 218a and the first detected sound position s1 can be used to define or adjust the first depth-of-field parameters of the first camera 204, such that the camera's focus area includes only the first audio coverage area 208, or does not extend beyond the first boundary line 218a. For example, the control module 216 can determine that a third point p3 of the first boundary line 218a, represented by a third set of coordinates (a3, b3, c3), is closest to the first detected sound position s1, and can use this third point p3 to calculate a first distance d1 between the first audio source 206 and the first boundary line 218a.
[0057] Similarly, using the techniques described herein, the second boundary line 219a and the second detected sound position s2 can be used to define or adjust the second depth-of-field parameters of the second camera 205 so that its focus area includes only the second audio coverage area 209, or does not extend beyond the second boundary line 219a. For example, the control module 216 can determine that the fourth point p4 of the second boundary line 219a, represented by the fourth set of coordinates (a4, b4, c4), is closest to the second detected sound position s2, and can use the fourth point p4 to calculate the second distance d2 between the second audio source 207 and the second boundary line 219a.
[0058] In some embodiments, a similar technique can be used to set a front limit for the depth of field of each camera 204, 205, such that the area in front of the first audio coverage area 208 is not included in the focus area of the first camera 204, and vice versa. This ensures, for example, that the image captured by the first camera 204 is focused on the presenter 206 and does not include the backs of audience members 207 and 214 or other unwanted areas of the second audio coverage area 209, and that the image captured by the second camera 205 is focused on the audience and does not include the back of the presenter 206. The control module 216 can be configured to implement the front limit by adjusting the depth of field parameters based on the "front" boundary of each audio coverage area. More specifically, the control module 216 can be configured to determine a first depth of field parameter of the first camera 204 based on the front boundary of the first audio coverage area 208 or the boundary line 218 (or "front boundary line") located in front of the first audio source 206. For example, microphone 202 can be configured to calculate the distance from the first audio source 206 to the front boundary line of the first audio coverage area 208, and control module 216 can be configured to use this distance to adjust the front length (e.g., L1 in FIG. 7) of the first depth-of-field parameter of the first camera 204. Similarly, the second depth-of-field parameter of the second camera 205 can be configured based on the front boundary of the second audio coverage area 209 or the boundary line 219 located in front of the audience members 207 and 214.
[0059] Therefore, the second audiovisual environment 200 can be used to focus the first camera 204 on the first audio source 206 and exclude any area outside the first audio coverage area 208 from the focus area of the first camera 204, and similarly to focus the second camera 205 on the second audio source 207 and / or other audio sources 214 and exclude any area outside the second audio coverage area 209 from the focus area of the second camera 205.
[0060] Although Figure 2 Specific use cases are illustrated, but it should be understood that the techniques described herein can be used in any environment with two or more audio coverage areas and two or more cameras. For example, in [the context of audio coverage and cameras]. Figure 1 In a meeting environment similar to that shown, a second audio coverage area can be added around another audio source 114 located outside the audio coverage area 108, and a second camera can be added to the environment 100 to capture images and / or videos of the other audio source 114 and / or the second audio coverage area.
[0061] Now for reference Figure 3 This illustrates another exemplary use case with multiple cameras used to capture different views or angles of a given room or other activity space. More specifically, Figure 3An exemplary audiovisual environment 300 is shown, which includes a microphone 302, a first camera 304, a second camera 305, and one or more audio sources 306 located in an audio coverage area 308 of the microphone 302. Environment 300 can be a theater, auditorium, place of worship, lecture hall, or any other event space having designated areas for performers or presenters (e.g., one or more audio sources 306) and one or more other areas (e.g., other audio sources 314) for audience members watching the event. Designated areas may include, for example, a stage or other performance space, and environment 300 can be used to capture audio and / or video of lectures, performances (e.g., plays, comedies, musicals, etc.), or other performances.
[0062] According to various embodiments, the audio coverage area 308 can be configured to include only a designated area (or stage area) and exclude the audience space, as shown. Similarly, the first camera 304 and the second camera 305 can be configured to capture images and / or video only of the designated stage area and / or the audio source 306 located thereon. For example, as... Figure 3 As shown, the first camera 304 can be located near the back of the stage area and can point towards the front of the stage area or behind the performer 306, while the second camera 305 can be located near the front of the stage area and can point towards the back of the stage area or in front of the performer 306. Therefore, both cameras 304 and 305 can point (but from different angles) at the audio source 306.
[0063] According to an embodiment, environment 300 also includes a control module 316 communicatively coupled to each of microphone 302, first camera 304, and second camera 305. Like control module 116, control module 316 can be configured to implement one or more aspects of a performance or activity occurring in environment 300 and / or implement one or more of the techniques described herein.
[0064] In an embodiment, the control module 316, microphone 302, first camera 304, and second camera 305 can form an audiovisual system (such as, for example, Figure 5The audiovisual system 500 shown, or part of an audiovisual system, is configured to define a video fence around the performer 306 and / or a designated stage area, such that audience members 314 or any other person or object 308 located outside the audio coverage area are not visible in the images and / or videos captured by the first camera 304. The video fence may be implemented by the first camera 304 using one or more parameters provided by the control module 316, such as, for example, image field parameters for defining the image field of the first camera 304 and / or depth field parameters for adjusting the depth of field of the first camera 304. One or more parameters may be configured by the control module 316 based on information received from the microphone 302, such as, for example, information defining the location of the audio source 306 and / or one or more boundaries of the audio coverage area 308.
[0065] More specifically, like microphone 102, microphone 302 can be configured to provide boundary information for a plurality of boundary lines 318 defining an audio coverage area 308. The boundary information may include one or more sets of coordinates for defining the boundary lines 318. For example, a first boundary line 318a of the plurality of boundary lines 318 may be defined by a first set of coordinates (a1, b1, c1) representing the center point p1 of the first boundary line 318a. In the illustrated embodiment, the first boundary line 318a (or “front boundary line”) is located in front of the audio source 306 and is used to exclude audience members 311 from the image captured by the first camera 304. Like microphone 102, microphone 302 may also provide sound localization information to control module 316. For example, the sound localization information may include a first set of coordinates (x1, y1, z1) representing the detected sound location s of the audio source 306.
[0066] Using the techniques described herein, the front boundary line 318a and the detected sound position s can be used to define or adjust the depth-of-field parameters of the first camera 304 such that the camera's focus area includes only the audio coverage area 308, or does not extend beyond the front boundary line 318a. For example, the control module 316 can determine that a second point p2 on the front boundary line 318a, represented by a second set of coordinates (a2, b2, c2), is closest to the detected sound position s, and can use this second point p2 to calculate the distance d between the audio source 306 and the front boundary line 318a. Therefore, the audiovisual environment 300 can be used to focus the first camera 304 on the audio source 306 and exclude any area outside the audio coverage area 308 from the focus area of the first camera 304, including audience members 314 located in front of the performer 306.
[0067] Figure 4An exemplary use case is illustrated where multiple video fences are provided to the same designated area, allowing a given camera to alternate its focus between different audio sources or areas within the designated area. For example, a first video fence may cover a first audio source, and a second video fence may cover a second audio source located at a short distance from the first audio source. The camera may switch or alternate between the two video fences, for example, based on which audio source is actively generating sound.
[0068] More specifically, Figure 4 An exemplary audiovisual environment 400 is illustrated, which includes a first microphone 402, a second microphone 403, a camera 404, a first audio source 406 located in a first audio coverage area 408 of the first microphone 402, and a second audio source 407 located in a second audio coverage area 409 of the second microphone 403. Environment 400 can be a theater, auditorium, place of worship, lecture hall, or any other event space having designated areas for performers or presenters (e.g., audio sources 406 and 407) and one or more other areas (e.g., other audio sources 414) for audience members watching the event. Designated areas may include, for example, a stage or other performance space, and environment 400 can be used to capture audio and / or video of lectures, performances (e.g., plays, comedies, musicals, etc.), or other performances.
[0069] According to various embodiments, the first audio coverage area 408 and the second audio coverage area 409 can be configured to include or cover different portions of a designated area (or stage area). For example, the first audio coverage area 408 may include a first portion of the stage area, and the second audio coverage area 408 may include a second portion of the stage area adjacent to the first portion, as shown in the figures. However, both areas 408 and 409 can be configured to exclude audience space, also as shown in the figures. Similarly, the camera 404 can be configured to capture images and / or video only of the designated stage area. For example, as... Figure 4 As shown, camera 404 can be located in front of or near the stage area and can be pointed at the stage area, or away from the audience space.
[0070] According to an embodiment, environment 400 also includes a control module 416 communicatively coupled to each of the first microphone 402, the second microphone 403, and the camera 404. Like control module 116, control module 416 can be configured to implement one or more aspects of a performance or activity occurring in environment 400 and / or implement one or more of the techniques described herein.
[0071] In an embodiment, the control module 416, the first microphone 402, the second microphone 403, and the camera 404 can form an audiovisual system (such as, for example, Figure 5The audiovisual system 500 shown, or part of an audiovisual system, is configured to define a separate video fence around each of the audio sources 406 and 407 within a designated stage area, such that the second audio source 407 is clearly visible in the image captured by the first video fence, and vice versa. The video fence can be implemented by the camera 404 using one or more parameters provided by the control module 416, such as, for example, image field parameters for defining the image field of the camera 404 and / or depth field parameters for adjusting the depth of field of the camera 404. One or more parameters can be configured by the control module 416 based on information received from each of the first microphone 402 and the second microphone 403, such as, for example, information defining the location of each of the first audio sources 406 and the second audio sources 407 and / or one or more boundaries of each of the first audio coverage area 408 and the second audio coverage area 409.
[0072] More specifically, the first microphone 402 may be configured to provide boundary information for a first plurality of boundary lines 418 defining a first audio coverage area 408, and the second microphone 403 may be configured to provide boundary information for a second plurality of boundary lines 419 defining a second audio coverage area 409. The boundary information may include one or more sets of coordinates for defining the boundary lines 418. For example, a first boundary line 418a of the first plurality of boundary lines 418 may be defined by a first set of coordinates (a1, b1, c1) representing the center point p1 of the first boundary line 418a. In the illustrated embodiment, the first boundary line 418a (or “left boundary line”) is located near the left side of the first audio source 406, or towards the second audio source 407, and is therefore used to exclude the second audio source 407 from the image captured by the camera 404 using a first video fence. As another example, a second boundary line 419a of the second plurality of boundary lines 419 may be defined by a second set of coordinates (a2, b2, c2) representing the second center point p2 of the second boundary line 419a. In the illustrated embodiment, the second boundary line 419a (or “right boundary line”) is located near the right side of the second audio source 407, or toward the first audio source 406, and is therefore used to exclude the first audio source 406 from the image captured by the camera 404 using the second video fence.
[0073] Similar to microphone 102, each of microphones 402 and 403 can also provide sound location information to control module 416. For example, the first microphone 402 can provide sound location information including a first set of coordinates (x1, y1, z1) representing a first detected sound position s1 of the first audio source 406. Similarly, the second microphone 403 can provide sound location information including a second set of coordinates (x2, y2, z2) representing a second detected sound position s2 of the second audio source 407.
[0074] Using the techniques described herein, control module 416 can implement a first video fence by defining or adjusting the depth-of-field parameters of camera 404 using a first boundary line 418a and a first detected sound location s1, such that the camera's focus area includes only the first audio coverage area 408 and therefore does not extend beyond the first boundary line 418a. For example, control module 416 can determine that the first boundary line 418a is closest to a second video fence and can use a first point p1 on the first boundary line 418a to calculate a first distance d1 between the first audio source 406 and the first boundary line 418a.
[0075] Similarly, using the techniques described herein, control module 416 can define or adjust the depth-of-field parameters of camera 404 using the second boundary line 419a and the second detected sound position s2 to achieve a second video fence, such that the camera's focus area includes only the second audio coverage area 409 and therefore does not extend beyond the second boundary line 419a. For example, control module 416 can determine that the second boundary line 419a is closest to the first video fence and can use a second point p2 on the second boundary line 419a to calculate a second distance d2 between the second audio source 407 and the second boundary line 419a.
[0076] Therefore, the audiovisual environment 400 can be used to create two separate video fences using the same camera 404, configuring each video fence to focus on a selected audio source 406 / 407 and exclude any other area from the focus area of the camera 404, except for the other audio source 407 / 406 and the corresponding audio coverage area 408 / 409. In some cases, the environment 400 may include multiple cameras, and each video fence may be assigned to a separate camera. In such cases, the video output may include an image captured using the first video fence, displayed adjacent to an image captured using the second video fence, for example, as side-by-side video strips or video blocks, or the two videos may otherwise be stitched together to be displayed as one.
[0077] Figure 5 An exemplary audiovisual system 500 is shown, which can be used as an audiovisual system according to embodiments of any of the environments 100, 200, 300, and 400 described herein. As shown, system 500 includes components that can be similar to or used as... Figure 1 The microphone 102 or any other microphone described herein, at least one microphone 502a..., 502n, and microphones that can be similar to or used as microphones. Figure 1 The system 500 includes at least one camera 504a…, 504n, of camera 104 or any other camera described herein. The system 500 also includes cameras that can be similar to or used as… Figure 1The controller 506 is the control module 116 or any other control module described herein. In particular, like the control module 116, the controller 506 may be implemented in hardware, software, or a combination thereof.
[0078] As shown in the figure, at least one microphone 502a..., 502n can be configured to provide information to the controller 506, such as, for example, audio information (e.g., audio signals captured by the microphone), boundary information of one or more audio coverage areas, and / or sound location information of one or more audio sources. In some embodiments, the controller 506 may also receive other types of sensor information (e.g., thermal, ToF, optical, etc.) from the microphone and / or one or more other sensors to determine the location of a human speaker or other audio source. The controller 506 can be configured to generate one or more parameters, images or image data, and / or control signals based on the received information and provide them to at least one camera 504a..., 504n, such as, for example, image field parameters, depth parameters, and / or image enhancement data. According to various embodiments, components of the audiovisual system 500 can use wired or wireless connections to transmit information to or receive information from the controller 506.
[0079] It should be understood that Figure 5 The components shown are merely exemplary, and any number, type, and placement of the various components of the audiovisual system 500 are contemplated and possible. For example, in Figure 5 In this configuration, multiple controllers may be coupled between cameras 504a..., 504n and microphones 502a..., 502n. As shown, in some embodiments, controller 506 may be a standalone device (such as a control device, computing device, or other electronic device) or may be included in such a device. In other embodiments, all or part of controller 506 may be included in one or more of microphones 502a..., 502n and / or one or more of cameras 504a..., 504n.
[0080] Now for reference Figure 6 An exemplary method or process 600 according to an embodiment is illustrated, which includes operations for focusing an image capture onto a selected audio source using an audiovisual system. Process 600 can be implemented using at least one processor communicating with at least one microphone and at least one camera, or otherwise using an audiovisual system. For ease of explanation, reference will be made below. Figure 5The process 600 is described using an audiovisual system 500. This process includes at least one microphone 502a…502n, at least one camera 504a…504n, and / or a controller 506; however, it should be understood that the process 600 may also be implemented using other audiovisual systems, processors, or devices. In embodiments, one or more processors and / or other processing components within the audiovisual system 500 may perform any, some, or all of the steps of the process 600. For example, the process 600 may be implemented entirely or at least partially by the controller 506. In some embodiments, the process 600 may be implemented by a computing device included in the audiovisual system, or more specifically by a processor of the computing device executing software stored in memory. In some cases, the computing device may further implement the operation of the process 600 by interacting with or engaging one or more other devices, internal or external to the audiovisual system 500 and communicatively coupled to the computing device. One or more other types of components (e.g., memory, input and / or output devices, transmitters, receivers, buffers, drivers, discrete components, etc.) may also be used in conjunction with the processor and / or other processing components to perform any, some, or all of the steps in process 600.
[0081] like Figure 6 As shown, process 600 may begin at step 602, wherein multiple boundary lines defining an audio pickup area (such as, for example,) are received from at least one microphone. Figure 1 The boundary information of the boundary line 118 of the audio coverage area 108 in the image. In some embodiments, the process 600 further includes determining the image field parameters of at least one camera based on the boundary information, and providing the image field parameters to at least one camera. The image field parameters can be configured to define the image field of at least one camera such that the image field includes the audio pickup area.
[0082] In some embodiments, process 600 further includes receiving camera location information indicating the position of at least one camera. In such cases, determining image field parameters may include further determining image field parameters based on the camera location information, and identifying the first boundary line may include further identifying the first boundary line based on the camera location information.
[0083] In some embodiments, process 600 further includes causing at least one camera to apply image enhancement to a portion of the image field extending beyond the first boundary line into the audio pickup area. Image enhancement may be a selected image displayed on a portion of the image field, a blurring effect applied to a portion of the image field, or any other visual effect that covers or blurs the portion of the image field extending beyond the audio pickup area.
[0084] At step 604, process 600 includes receiving from at least one microphone an audio source indicating that it is located within the audio pickup area (e.g., Figure 1 The detected sound location (e.g., audio source 106) in the audio source 106 Figure 1 The sound location information of the location s in the middle. At step 606, process 600 includes, based on the sound location information and boundary information, identifying the first boundary line (e.g., position s) among a plurality of boundary lines. Figure 1 Line 118a) is identified as being near the detected sound location. At step 608, process 600 includes calculating a first distance (e.g., between the detected sound location and the first boundary line) between the detected sound location and the first boundary line. Figure 1 The distance d in the middle.
[0085] At step 610, process 600 includes determining depth parameters of at least one camera based on a first distance. According to an embodiment, the depth parameters adjust the focus area of at least one camera such that the focus area includes the audio source and a first region between the audio source and a first boundary line (e.g., Figure 1 Region 120 in the middle), and a second region excluding the audio pickup region (e.g., Figure 1 (Region 122 in the image). At step 612, process 600 includes providing depth parameters and sound location information to at least one camera. Once step 612 is completed and / or video output has been generated accordingly by at least one camera, process 600 may terminate.
[0086] In other embodiments, process 600 may be adapted to accommodate multiple cameras and / or multiple microphones according to one or more use cases or environments described herein. For example, in some cases, the process or method may be coupled with a first camera (e.g., Figure 2 Camera 204), second camera (e.g., Figure 2 (camera 205) and at least one microphone (e.g., Figure 2 The process can be performed by one or more processors communicating with microphone 202 in the microphone. Such a process may include, for example, at step 602, receiving from at least one microphone one or more first boundaries or boundary lines defining a first audio pickup area (e.g., ...). Figure 2 One or more first boundaries 218 of the audio coverage area 208 and one or more second boundaries or boundary lines of the second audio pickup area (e.g., Figure 2 The process may further include, for example, receiving sound location information from at least one microphone at step 604, indicating a first audio source located within the first audio pickup area (e.g., [missing information]). Figure 2 The first detected sound location (e.g., audio source 206) in the audio source 206 Figure 2 The position s1 in the middle), and the second audio source located in the second audio pickup area (e.g., Figure 2 The second detected sound location (e.g., audio source 207) in the audio source 207 Figure 2 The position s2 in the middle). Additionally, the process may also include: identifying a first camera as being near a first audio pickup area and a second camera as being near a second audio pickup area based on boundary information, and configuring the first camera to capture an image of the first audio pickup area and the second camera to capture an image of the second audio pickup area. The process may further include: for example, at step 606, identifying a first boundary or boundary line (e.g., in one or more first boundaries) based on sound location information and boundary information. Figure 2 The first boundary 218a) is identified as being located near the first detected sound location, and the second boundary or boundary line in one or more second boundaries (e.g., Figure 2 Line 219a) is identified as being near the second detected sound location. Furthermore, the process may include, for example, at step 608, calculating a first distance between the first detected sound location and the first boundary line (e.g., Figure 2 The distance d1 in the middle), and the second distance between the second detected sound location and the second boundary line (e.g., Figure 2 The process may further include, for example, at step 610, determining a first depth parameter of the first camera based on a first distance, and determining a second depth parameter of the second camera based on a second distance. Additionally, the process may include, for example, at step 612, providing a first detected sound location and a first depth parameter to the first camera, and providing a second detected sound location and a second depth parameter to the second camera. The process may end there and / or once each camera in the system has generated video output accordingly.
[0087] In other embodiments, a single high-fidelity camera can be used instead of two separate cameras to achieve similar results. For example, process 600 can be performed using a single camera and a controller configured to segment the images captured by the single camera into multiple videos or images. The multiple videos or images may each correspond to multiple audio pickup areas, and the controller may be configured to adjust the focus parameters of each segment based on sound location information and boundary information using the techniques described herein.
[0088] Therefore, the techniques described herein can be used to define video fences that focus a camera on a given audio source and exclude a selected area from the captured image. Video fences can be created based on “speaker tracking information” or output from a microphone indicating the detected location of sound generated by an audio source (e.g., an active speaker) and boundary information of the audio coverage area (or audio pickup area) used by the microphone to capture the detected sound. For example, in some cases, boundary information and speaker location information can be used to determine how much audio coverage area remains behind the active speaker and adjust the camera’s depth-of-field parameters accordingly, such that the camera’s focus area includes only the audio source and the surrounding audio coverage area. In this way, video fences can be used to prevent areas outside the audio coverage area from being included in the image and / or video captured by the camera, thereby excluding any unwanted or unnecessary people, objects, or scenery from the output video. By configuring video fences using speaker tracking information, the techniques described herein can be used to provide a flexible and configurable intelligent activity space for both audio and video settings, regardless of the diversity of configurations.
[0089] Return to reference Figure 5 In various embodiments, the audiovisual system 500 may further include Figure 5 Various components not shown in the diagram (such as, for example, one or more speakers, displays, computing devices, and / or cameras). Furthermore, one or more components in system 500 may include one or more digital signal processors or other processing components, controllers, wireless receivers, wireless transceivers, etc., although not shown or mentioned above.
[0090] One or more components of system 500 may communicate with one or more other components of system 500 via wired or wireless communication. For example, at least one microphone 502a, ... 502n and at least one camera 504a, ... 504n may be connected or coupled to controller 506 via a wired connection (e.g., Ethernet cable, USB cable, etc.) or a wireless network connection (e.g., WiFi, Bluetooth, Near Field Communication (“NFC”), RFID, infrared, etc.). In some cases, at least one microphone 502a, ... 502n may include a network audio device coupled to controller 506 via a network cable (e.g., Ethernet) and configured to process digital audio signals. In other cases, at least one microphone 502a, ... 502n may include an analog audio device or another type of digital audio device and may be connected to controller 506 using a Universal Serial Bus (USB) cable or other suitable connection mechanism. In some embodiments, one or more components of system 500 may communicate with one or more other components of system 500 via a suitable application programming interface (API).
[0091] In some embodiments, one or more components of the audiovisual system 500 may be combined into or reside in a single unit or device. For example, all components of the audiovisual system 500 may be included in the same device, such as at least one of microphones 502a, ... 502n, at least one of cameras 504a, ... 504n, or a computing device that includes all of these components. As another example, the controller 506 may be included in or combined with any of the microphones 502a, ... 502n or any of the cameras 504a, ... 504n. In some embodiments, the system 500 may take the form of a cloud-based system or other distributed system, such that the components of the system 500 may be physically close to each other or may not be physically close to each other.
[0092] The components of system 500 may be implemented in hardware (e.g., discrete logic circuits, application-specific integrated circuits (ASICs), programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), microprocessors, etc.), using software that can be executed by one or more servers or computers or other computing devices having processors and memory (e.g., personal computers (PCs), laptops, tablets, mobile devices, smart devices, thin clients, etc.), or by a combination of hardware and software. For example, some or all of the components of at least one microphone 502a, ... 502n, at least one camera 504a, ... 504n, and / or controller 506 may be implemented using discrete circuit system devices and / or using one or more processors (e.g., audio processors and / or digital signal processors) that execute program code stored in memory (not shown) configured to perform one or more processes or operations described herein, such as, for example Figure 6 The method or process 600 shown is described. Therefore, in an embodiment, one or more components of the audiovisual system 500 may include one or more processors, storage devices, computing devices, and / or other hardware components not shown in the figures.
[0093] The process described in this article (including) Figure 6 Method 600) can be all or part of the following: Figure 5The method is executed by one or more processing devices or processors (e.g., analog-to-digital converters, encryption chips, etc.) within or outside the audiovisual system 500. Additionally, one or more other types of components (e.g., memory, input and / or output devices, transmitters, receivers, buffers, drivers, discrete components, logic circuits, etc.) may also be used in conjunction with a processor and / or other processing components to perform any, some, or all of the steps of method 600. For example, in some embodiments, each method described herein may be executed by a processor that runs software stored in memory. The software may include, for example, program code or computer program modules containing software instructions executable by a processor. In some embodiments, the program code may be a computer program stored on a non-transitory computer-readable medium that may be executed by a processor of the associated device.
[0094] Any processor described herein may include general-purpose processors (e.g., microprocessors) and / or special-purpose processors (e.g., audio processors, digital signal processors, etc.). In some examples, the processor described herein may be any suitable processing device or group of processing devices, such as, but not limited to, microprocessors, microcontroller-based platforms, integrated circuits, one or more field-programmable gate arrays (FPGAs) and / or one or more application-specific integrated circuits (ASICs).
[0095] Any memory or storage device described herein can be volatile memory (e.g., RAM, including non-volatile RAM, magnetic RAM, ferroelectric RAM, etc.), non-volatile memory (e.g., disk storage, flash memory, EPROM, EEPROM, memristor-based non-volatile solid-state memory, etc.), non-replaceable memory (e.g., EPROM), read-only memory, and / or high-capacity storage devices (e.g., hard disk drives, solid-state drives, etc.). In some examples, the memory described herein includes multiple types of memory, particularly volatile and non-volatile memory.
[0096] Furthermore, any memory described herein can be a computer-readable medium on which one or more sets of instructions can be embedded. During execution of the instructions, the instructions can be wholly or at least partially deployed in any one or more of the memory, the computer-readable medium, and / or in one or more processors. In some embodiments, the memory described herein may include one or more data storage devices configured to provide persistent storage for data that needs to be stored and accessed by an end user. In this case, the data storage device may store the data in flash memory or other storage devices. In some embodiments, the data storage device may be implemented using, for example, an SQLite database, UnQLite, BerkeleyDB, BangDB, etc.
[0097] Any computing device described herein can be any general-purpose computing device including at least one processor and one storage device. In some embodiments, the computing device may be a stand-alone computing device included in the audiovisual system 500, or it may reside in another component of the system 500, such as, for example, any of microphones 502a, ... 502n, any of cameras 504a, ... 504n, and / or controller 506. In such embodiments, the computing device may be physically located and / or dedicated to a given environment or room, such as, for example, the same environment in which microphones 502a, ... 502n and cameras 504a, ... 504n are located. In other embodiments, the computing device may not be physically located near microphones 502a, ... 502n and cameras 504a, ... 504n, but may reside in an external network (such as a cloud computing network) or may otherwise be distributed in a cloud-based environment. Furthermore, in some embodiments, the computing device may be implemented using firmware or entirely based on software as part of a network that can be accessed or otherwise communicated with by another device, including other computing devices such as, for example, desktops, laptops, mobile devices, tablets, smart devices, etc. Therefore, the term "computing device" should be understood to include distributed systems and devices (such as, for example, cloud-based ones), as well as software, firmware, and other components configured to perform one or more of the functions described herein. Further, one or more functions of the computing device may be physically remote and may be communicatively coupled to the computing device.
[0098] In some embodiments, any computing device described herein may include one or more components configured to support teleconferencing, meetings, classrooms, or other events, and / or process associated audio signals to improve the audio quality of the event. For example, in various embodiments, any computing device described herein may include a digital signal processor (“DSP”) configured to process audio signals received from various microphones or other audio sources using, for example, automatic mixing, matrix mixing, delay, compressor, parametric equalizer (“PEQ”) functions, acoustic echo cancellation, etc. In other embodiments, the DSP may be a standalone device operatively coupled to or connected to a computing device using a wired or wireless connection. An exemplary embodiment of a DSP (when implemented in hardware) is SHURE’s P300 IntelliMix audio conferencing processor, whose user manual is incorporated herein by reference in its entirety. As further explained in the P300 manual, this audio conferencing processor incorporates algorithms optimized for audio / video conferencing applications and is designed to deliver a high-quality audio experience, including eight-channel acoustic echo cancellation, noise reduction, and automatic gain control. Another typical embodiment of a DSP (when implemented in software) is SHURE's IntelliMix Room, whose user guide is incorporated herein by reference in its entirety. As further explained in the IntelliMix Room user guide, this DSP software is configured to optimize the performance of networked microphones in conjunction with audio and video conferencing software and is designed to run on the same computer as the conferencing software. In other embodiments, it will be understood that other types of audio processors, digital signal processors, and / or DSP software components may be used to perform one or more of the audio processing techniques described herein.
[0099] In addition, any computing device described herein may also include various other software modules or applications (not shown) configured to facilitate and / or control meeting activities, such as internal or proprietary meeting software and / or third-party meeting software (e.g., Microsoft Skype, Microsoft Teams, Bluejeans, Cisco WebEx, GoToMeeting, Zoom, Join.me, etc.). Such software applications may be stored in the memory of the computing device and / or may be stored on a remote server (e.g., locally or as part of a cloud computing network) and accessed by the computing device via a network connection. Some software applications may be configured as distributed cloud software, where one or more portions of the application are deployed in the computing device, while one or more other portions are deployed in a cloud computing network. One or more software applications may be deployed in an external network, such as a cloud computing network. In some embodiments, one or more software applications may be accessed via a web portal architecture or otherwise provided as Software as a Service (SaaS).
[0100] Typically, computer program products according to embodiments described herein include computer-usable storage media (e.g., standard random access memory (RAM), optical disc, universal serial bus (USB) drive, etc.) in which computer-readable program code is embedded, wherein the computer-readable program code is adapted to be executed by a processor (e.g., to work with an operating system) to implement the methods described herein. In this regard, the program code can be implemented in any desired language and can be implemented as machine code, assembly code, bytecode, interpreted source code, etc. (e.g., via C, C++, Java, ActionScript, Python, Objective-C, JavaScript, CSS, XML, and / or others). In some embodiments, the program code may be a computer program stored on a non-transitory computer-readable medium that can be executed by a processor of an associated device.
[0101] The terms "non-transitory computer-readable medium" and "computer-readable medium" include single or multiple media, such as centralized or distributed databases, and / or associated caches and servers storing one or more sets of instructions. Further, the terms "non-transitory computer-readable medium" and "computer-readable medium" include any tangible medium capable of storing, encoding, or carrying a set of instructions for execution by a processor, or causing a system to perform one or more methods or operations disclosed herein. As used herein, the term "computer-readable medium" is explicitly defined to include any type of computer-readable storage device and / or storage disk, but excludes propagating signals.
[0102] Figures (such as, for example) Figure 6 Any process description or block in the document should be understood to represent a module, segment, or code portion comprising one or more executable instructions for implementing a specific logical function or step in the process, and alternative implementations are also included within the scope of the embodiments described herein, wherein functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order, depending on the functions involved, as understood by those skilled in the art.
[0103] It should be noted that in the specification and drawings, identical or substantially similar elements may be labeled with the same reference numerals. However, sometimes these elements may be labeled with different numerals, for example, where such labeling facilitates a clearer description. Furthermore, system components may be arranged in various ways known in the art. Additionally, the drawings described herein are not necessarily drawn to scale, and in some cases, the scale may be exaggerated to more clearly describe certain features and / or related elements may be omitted to emphasize and clearly illustrate the novel features described herein. Such labeling and drawing practices do not necessarily imply any underlying substantial purpose. The above description should be considered as a whole and interpreted in accordance with the principles taught herein and understandable to those skilled in the art.
[0104] In this disclosure, the use of disjunctive conjunctions should be understood to include conjunctive conjunctions. The use of definite or indefinite articles is not intended to indicate cardinality. Specifically, references to "the" object or "one" object are also intended to indicate one of a plurality of such objects.
[0105] This disclosure describes, illustrates, and exemplifies one or more specific embodiments of the invention based on its principles. This disclosure is intended to explain how various embodiments can be designed and used according to the technology, and not to limit its true, intended, and fair scope and spirit. That is, the foregoing description is not intended to be exhaustive or limited to the precise forms disclosed herein, but rather to explain and teach the principles of the invention in a way that enables those skilled in the art to understand these principles and apply them to practice not only the embodiments described herein, but also other embodiments that may conceive based on these principles. The embodiments provided herein were chosen and described to best illustrate the principles of the technology and its practical application, and to enable those skilled in the art to use the technology in various embodiments and make various modifications according to the intended particular use. All such modifications and variations are within the scope of the embodiments defined in the appended claims, may be modified during the pending period of this patent application, and are equivalent to their equivalents when interpreted to the fullest extent they are fairly, legally, and justly enjoyed.
Claims
1. A method performed by one or more processors in communication with each of at least one microphone and at least one camera, the method comprising: receiving, from the at least one microphone, boundary information defining one or more boundaries of an audio pickup area; receiving, from the at least one microphone, sound location information indicative of a detected sound location of an audio source located within the audio pickup area; identifying, based on the sound location information and the boundary information, a first boundary of the one or more boundaries as being located in proximity to the detected sound location; calculating a first distance between the detected sound location and the first boundary; determining, based on the first distance, a depth of field parameter of the at least one camera; and providing the depth of field parameter and the sound location information to the at least one camera.
2. The method of claim 1, wherein the depth of field parameter adjusts a focus zone of the at least one camera such that the focus zone includes the audio source and a first region between the audio source and the first boundary, and excludes a second region outside of the audio pickup area.
3. The method of claim 1, further comprising: determining, based on the boundary information, an image field parameter of the at least one camera; and providing the image field parameter to the at least one camera, wherein the image field parameter is configured to define an image field of the at least one camera such that the image field includes the audio pickup area.
4. The method of claim 3, further comprising: causing the at least one camera to apply an image enhancement to a portion of the image field that extends beyond the first boundary line to outside of the audio pickup area.
5. The method of claim 4, wherein the image enhancement is a selected image displayed on the portion of the image field.
6. The method of claim 4, wherein the image enhancement is a blur effect applied to the portion of the image field.
7. The method of claim 3, further comprising: receiving, from the at least one camera, camera location information indicative of a location of the at least one camera, wherein determining the image field parameter includes determining the image field parameter further based on the camera location information, and identifying the first boundary includes identifying the first boundary further based on the camera location information.
8. A system comprising: at least one microphone configured to provide boundary information defining one or more boundaries of an audio pickup area, and sound location information indicative of a detected sound location of an audio source located within the audio pickup area; at least one camera configured to capture images or video of the audio pickup area; and one or more processors communicatively coupled to each of the at least one microphone and the at least one camera, the one or more processors configured to: receive the boundary information and the sound location information from the at least one microphone; identify, based on the sound location information and the boundary information, a first boundary of the one or more boundaries as being located in proximity to the detected sound location; computing a first distance between the detected sound location and the first boundary; determining a depth of field parameter for the at least one camera based on the first distance; and providing the depth of field parameter and the sound location information to the at least one camera.
9. The system of claim 8, wherein the at least one camera is further configured to adjust a focus zone of the at least one camera based on the depth of field parameter such that the focus zone includes the audio source and a first region between the audio source and the first boundary, and excludes a second region outside of the audio pickup region.
10. The system of claim 8, wherein the at least one camera is configured to capture the image or video of the audio pickup region based on an image field parameter configured to define an image field of the at least one camera to include the audio pickup region, and wherein the one or more processors are further configured to determine the image field parameter based on the boundary information, and provide the image field parameter to the at least one camera.
11. The system of claim 10, wherein the at least one camera is configured to apply an image enhancement to a portion of the image field that extends beyond the first boundary to outside of the audio pickup region.
12. The system of claim 11, wherein the image enhancement is a selected image displayed on the portion of the image field.
13. The system of claim 11, wherein the image enhancement is a blur effect applied to the portion of the image field.
14. The system of claim 10, wherein the at least one camera is further configured to provide camera location information indicating a location of the at least one camera to the one or more processors, and the one or more processors are configured to determine the image field parameter further based on the camera location information, and identify the first boundary further based on the camera location information.
15. The system of claim 8, further comprising a second camera, wherein the boundary information further defines a second audio pickup region of one or more second boundaries, the sound location information further indicates a second detected sound location of a second audio source located within the second audio pickup region, and the one or more processors are further configured to: identify the second camera as being in proximity to the second audio pickup region based on the boundary information; configure the second camera to capture an image or video of the second audio pickup region; identify a second boundary of the one or more second boundaries as being in proximity to the second detected sound location based on the sound location information and the boundary information; compute a second distance between the second detected sound location and the second boundary; determine a second depth of field parameter for the second camera based on the second distance; and provide the second detected sound location and the second depth of field parameter to the second camera.
16. A method performed by one or more processors in communication with a first camera, a second camera, and at least one microphone, the method comprising: receiving, from the at least one microphone, boundary information defining one or more first boundaries of a first audio pickup area and one or more second boundaries of a second audio pickup area; receiving, from the at least one microphone, sound location information indicating a first detected sound location of a first audio source located within the first audio pickup area and a second detected sound location of a second audio source located within the second audio pickup area; identifying, based on the boundary information, the first camera as being in proximity to the first audio pickup area and the second camera as being in proximity to the second audio pickup area; configuring the first camera to capture images or video of the first audio pickup area and the second camera to capture images or video of the second audio pickup area; identifying, based on the sound location information and the boundary information, a first boundary of the one or more first boundaries as being in proximity to the first detected sound location and a second boundary of the one or more second boundaries as being in proximity to the second detected sound location; calculating a first distance between the first detected sound location and the first boundary and a second distance between the second detected sound location and the second boundary; determining a first depth of field parameter for the first camera based on the first distance; determining a second depth of field parameter for the second camera based on the second distance; providing the first detected sound location and the first depth of field parameter to the first camera; and providing the second detected sound location and the second depth of field parameter to the second camera.
17. The method of claim 16, wherein the first depth of field parameter adjusts a first focus zone of the first camera such that the first focus zone includes the first audio source and a first region between the first audio source and the first boundary and excludes a second region outside of the first audio pickup area, and wherein the second depth of field parameter adjusts a second focus zone of the second camera such that the second focus zone includes the second audio source and a third region between the second audio source and the second boundary and excludes a fourth region outside of the second audio pickup area.
18. The method of claim 16, further comprising: determining a first image field parameter for the first camera based on the boundary information; providing the first image field parameter to the first camera; determining a second image field parameter for the second camera based on the boundary information; and providing the second image field parameter to the second camera, wherein the first image field parameter is configured to define a first image field of the first camera such that the first image field includes the first audio pickup area, and wherein the second image field parameter is configured to define a second image field of the second camera such that the second image field includes the second audio pickup area. 19. The method of claim 18, further comprising: causing the first camera to apply first image enhancements to a first portion of the first image field that extends beyond the first boundary line to outside the first audio pickup area, and causing the second camera to apply second image enhancements to a second portion of the second image field that extends beyond the second boundary line to outside the second audio pickup area.
20. The method of claim 18, further comprising: receiving, from the first camera, first camera location information indicative of a first position of the first camera; and receiving, from the second camera, second camera location information indicative of a second position of the second camera, wherein determining the first image field parameters comprises determining the first image field parameters further based on the first camera location information, and identifying the first boundary comprises identifying the first boundary further based on the first camera location information, and wherein determining the second image field parameters comprises determining the second image field parameters further based on the second camera location information, and identifying the second boundary comprises identifying the second boundary further based on the second camera location information.