Video view based on tracked characteristics of conference participants
Through smart cameras or multi-camera systems combined with artificial intelligence technology, automatically analyzing video streams and assigning roles, solving the problem that traditional video conferencing systems are difficult to show non-speakers and provide interactive contexts, achieving a richer and more interactive video conferencing experience.
Patent Information
- Application Number
- CN202380080784.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-22
- Filing Date
- 2023-09-08
- Publication Date
- 2025-07-01
AI Technical Summary
Traditional video conferencing systems have difficulty effectively identifying and presenting nonspeakers in meetings, making it difficult for remote participants to participate in the conversation, and the system lacks the ability to provide contextual information about interactions between conference participants and nonspeakers’ participation.
Through smart cameras or multi-camera systems, artificial intelligence technology is used to automatically analyze video streams in the conference environment, identify and assign roles of conference participants, and dynamically adjust the framing and layout of video streams based on roles and interactions to provide a richer user experience.
Achieve a more informative and interactive video conferencing experience to remote meeting participants, enhance contextual understanding of the conference environment and interactions between participants, and improve the sense of participation of remote participants.
Smart Images

Figure CN120239966A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to single-camera and multi-camera video conferencing systems. Background Art
[0002] In traditional video conferencing, the experience of participants may be static. Cameras used in conference rooms may not take into account social cues (e.g., reactions, body language, and other non-verbal communication), speaker awareness, or direction of attention in the meeting situation. For meeting participants located in the corners of the meeting environment or far from the speaker (e.g., remote participants), the video conferencing experience may lack engagement, making it difficult to participate in the conversation. Single-camera systems can display the meeting environment from a limited number of angles, which may lack the ability to feature non-speaking meeting participants. Additionally, in large video conference rooms, it may be difficult to frame some or all of the meeting participants and maintain the display or representation of meeting participants located farther away from the camera. Meeting participants watching a streamed video conference may not be able to see the facial expressions of the meeting participants in the meeting environment and may therefore not be able to interact proactively with the meeting participants present in the meeting environment.
[0003] In traditional video conferencing systems (even multi-camera systems), the user experience may be limited to the display of the meeting participants determined to be speaking. Such systems may lack the ability to change the shot of the detected speaker (e.g., by selecting a different camera output source to feature the detected speaker in the video stream of the shot, by selectively including other meeting participants in the shot, etc.). Such systems may also lack the ability to feature shots of non-speaking meeting participants who are actively listening to or reacting to the speaker (either together with or isolated from the shot featuring the speaker). Thus, the user experience provided by traditional video conferencing systems may lack a certain degree of depth and interactivity due to the display of the representation of the speaking meeting participants without conveying information (information associated with, for example, reactions, interactions, spatial relationships, etc. between the speaker and other meeting participants).
[0004] There is a need for a single-camera and / or multi-camera system that can increase the user experience and interactivity by identifying meeting participants and roles associated with the represented meeting participants, selectively framing the conversation and interactions between the speaker and the meeting participants based on the determined roles of the meeting participants, and / or adaptively framing and laying out the meeting participants based on their roles, to create a more robust user experience that conveys a greater degree of contextual information about the physical meeting environment and the meeting participants. Summary of the Invention
[0005] The disclosed embodiments can address one or more of these challenges. The disclosed cameras and camera systems can include an intelligent camera or a multi-camera system that understands the dynamics of meeting participants (e.g., using artificial intelligence (AI), such as a trained network), determines or infers the roles associated with the meeting participants represented in one or more camera outputs, and provides a video output stream for display to remote meeting participants, where the framed representation of the meeting participants is selected, or based on relevant context information about the roles and interactions of the meeting participants in the physical meeting environment, such as the determined roles of the meeting participants and other factors. In this way, a video stream focused on the most relevant meeting participants (e.g., the speaker or speakers) can be provided to remote meeting participants, while also providing information about other meeting participants in a balanced manner to provide further context information about the interactions between meeting participants and the participation of non-speaking or non-presenting meeting participants. Thus, the disclosed embodiments can provide a more informative and engaging video conferencing experience.
[0006] In some embodiments, by dividing the meeting room into multiple zones and identifying the zone in which the speaking participant is located, the disclosed systems and methods can alternate between showing a speaker shot and a listening shot to give a closer view of the speaker, create a better flow in the session, and provide spatial context for remote participants. This can also provide a more dynamic viewing experience for remote participants, similar to how meeting participants would naturally look around the meeting environment and interact with other meeting participants.
[0007] According to a first aspect of the present disclosure, there is provided a video conferencing system, comprising: at least one camera, each camera being configured to generate an overview video output stream representing a meeting environment; and at least one video processing unit, configured to: automatically analyze at least one overview video output stream to identify at least a first meeting participant and a second meeting participant represented in the at least one overview video output stream; assign a first role name to the first meeting participant and a second role name to the second meeting participant based on one or more determined characteristics associated with at least one of the first meeting participant or the second meeting participant; determine a relative framing priority associated with the first meeting participant and the second meeting participant based on the priorities associated with the first role name and the second role name; generate one or more main video streams featuring the first meeting participant and / or the second meeting participant as an output (e.g., for display) based on the relative framing priority.
[0008] Generating one or more primary video streams characterized by a first conference participant and / or a second conference participant as output based on relative framing priorities may include: selecting a framing representation of the first conference participant or the second conference participant based on the relative framing priorities; and generating a primary video stream characterized by the selected framing representation of the first conference participant or the second conference participant as output (e.g., for display).
[0009] The video processing unit may be configured to select a framing representation of a conference participant associated with a higher relative framing priority.
[0010] The framing priority associated with the first conference participant may be higher than the framing priority associated with the second conference participant; and the selected framing representation of the first conference participant or the second conference participant may be the framing representation of the first conference participant.
[0011] The selected framing representation may be the framing representation of the second conference participant, and at least one video processing unit may be configured to: determine the likelihood that the first conference participant will start speaking based on a first role name assigned to the first conference participant; in response to the likelihood that the first conference participant will start speaking meeting a likelihood condition, perform one or more actions to prepare the video processing unit to output a primary video stream characterized by the framing representation of the first conference participant; determine whether the first conference participant has started speaking based on an analysis of at least one overview video output stream; and in response to determining that the first conference participant has started speaking, output a primary video stream including the framing representation of the first conference participant for display.
[0012] Performing one or more actions to prepare the video processing unit to output a primary video stream characterized by the first conference participant may include one or more of the following: selecting a framing representation of the first conference participant for the primary video stream; and generating a primary video stream characterized by the framing representation of the first conference participant.
[0013] The determination of the likelihood that the first conference participant will start speaking may also be based on one or more determined quantities characterizing the session dynamics between the first conference participant and one or more other conference participants characterized in one or more video output streams.
[0014] One or more determined quantities may include one or more determined vectors that represent one or more session exchanges between the first conference participant and one or more other conference participants characterized in one or more video output streams.
[0015] Each vector can represent a speaking episode during which a first meeting participant or another meeting participant among one or more other meeting participants is determined to have spoken.
[0016] The magnitude of each vector can represent the duration of the speaking episode. The direction of each vector can represent the meeting participant who will speak next.
[0017] One or more determined quantities can include at least one of the following: the identity of one or more other meeting participants; one or more durations characterizing one or more episodes during which the first meeting participant or one or more other meeting participants are determined to have spoken; one or more frequencies characterizing the frequency of the episodes during which the first meeting participant or one or more other meeting participants are determined to have spoken; one or more time intervals characterizing one or more gaps between the episodes during which the first meeting participant or one or more other meeting participants are determined to have spoken.
[0018] Determining the likelihood that the first meeting participant will start speaking can also be based on one or more determined attributes associated with the first meeting participant.
[0019] One or more determined attributes can include at least one of the following: the time since the meeting participant last spoke, the frequency at which the meeting participant speaks, the direction of head turning to face the speaker, gestures, facial expressions, changes in facial expressions, nods, gaze direction, non-verbal sounds, and inhalations.
[0020] Generating one or more primary video streams characterized by the first meeting participant and / or the second meeting participant as output based on relative framing priorities can include: generating a first primary video stream and a second primary video stream as output (e.g., for display), the first primary video stream including a framed representation of the first meeting participant during a first time period, and the second primary video stream including a framed representation of the second meeting participant during a second time period; wherein the relative durations of the first time period and the second time period are determined based on the relative framing priorities associated with the first meeting participant and the second meeting participant.
[0021] The framing priority associated with the first meeting participant can be higher than the framing priority associated with the second meeting participant; and the duration of the first time period can be longer than the duration of the second time period.
[0022] The first time period and / or the second time period can be discontinuous.
[0023] One or more detected characteristics may include one or more of the following: whether a participant is speaking, the length of time a participant is speaking, the percentage of the duration of a participant's speech, whether a participant is standing, the viewing direction of a participant, a change in the viewing direction of a participant, a gesture performed by a participant, or a reaction performed by a participant.
[0024] At least one video processing unit may also be configured to assign observed properties to each of a first conference participant and a second conference participant based on an analysis of at least one overview video output stream.
[0025] The observed properties may include demeanor.
[0026] Demeanor may include at least one of serious, humorous, questioning, listening, or disengaged.
[0027] A relative framing priority associated with the first conference participant and the second conference participant may be determined based on a priority associated with the assigned observed property.
[0028] The priority associated with at least one of the observed properties may be user-assignable.
[0029] At least one video processing unit may also be configured to assign observed actions to each of a first conference participant and a second conference participant based on an analysis of at least one overview video output stream.
[0030] The observed actions may include at least one of standing, sitting, raising a hand, interacting with a whiteboard, applauding, speaking, sleeping, writing, or laughing.
[0031] A relative framing priority associated with the first conference participant and the second conference participant may be determined based on a priority associated with the assigned observed action.
[0032] The priority associated with at least one of the observed actions may be user-assignable.
[0033] At least one video processing unit may be configured to: receive a user-assigned priority associated with a first conference participant; and determine a relative framing priority based on the user-assigned priority associated with the first conference participant.
[0034] At least one video processing unit may be configured to: receive a user-assigned role name associated with a first meeting participant and a user-assigned priority associated with the user-assigned role name; assign the user-assigned role name to the first meeting participant; and determine a framing priority for the first meeting participant based on the user-assigned priority associated with the user-assigned role name.
[0035] At least one camera may include a plurality of cameras.
[0036] At least one video processing unit may also be configured to automatically analyze an overview video output stream from each of the plurality of cameras and, for example, based on at least one identifier, track and correlate representations of meeting participants across the overview video output streams.
[0037] At least one video processing unit may be included in one of the plurality of cameras.
[0038] At least one video processing unit may include one or more microprocessors on-board the camera and / or one or more microprocessors located remotely with respect to the camera.
[0039] The roles assigned to each of the first meeting participant and the second meeting participant based on the analysis of at least one overview video output stream may be updated periodically.
[0040] The role name may include one or more of the following: presenter, contributor, audience, and / or user-assigned.
[0041] The role name may be selected from: presenter, contributor, audience, and / or user-assigned.
[0042] At least one video processing unit may be configured to implement at least one trained network that is configured to output, based on an input including one or more captured video frames from at least one overview video output stream, the role name of the identified meeting participant.
[0043] At least one video processing unit may be configured to cause at least one display (e.g., a user display) to show one or more primary video streams.
[0044] According to a second aspect of the present disclosure, there is provided a video processing method performed by at least one video processing unit, the method comprising: automatically analyzing at least one overview video output stream representing a meeting environment generated by at least one camera to identify at least a first meeting participant and a second meeting participant represented in the at least one overview video output stream; assigning a first role name to the first meeting participant and a second role name to the second meeting participant based on one or more detected characteristics associated with at least one of the first meeting participant or the second meeting participant; determining a relative framing priority associated with the first meeting participant and the second meeting participant based on the priorities associated with the first role name and the second role name; and generating, based on the relative framing priority, one or more main video streams featuring the first meeting participant and / or the second meeting participant as output.
[0045] According to a third aspect of the present disclosure, there is provided a video conferencing system comprising: at least one camera, each camera being configured to generate an overview video output stream representing a meeting environment; and at least one video processing unit configured to: automatically analyze the at least one overview video output stream to identify at least a first meeting participant represented in the at least one overview video output stream; assign a first role name to the first meeting participant based on one or more detected characteristics associated with the first meeting participant; determine the likelihood that the first meeting participant will start speaking based on the first role name assigned to the first meeting participant; in response to the likelihood that the first meeting participant will start speaking meeting a likelihood condition, perform one or more actions to prepare the video processing unit to output a main video stream featuring a view of the first meeting participant for display; determine whether the first meeting participant has started speaking based on the analysis of the at least one overview video output stream; and in response to determining that the first meeting participant has started speaking, output a main video stream featuring a view of the first meeting participant for display.
[0046] According to a fourth aspect of the present disclosure, there is provided a video processing method performed by at least one video processing unit, the method comprising: automatically analyzing at least one overview video output stream representing a meeting environment generated by at least one camera to identify at least a first meeting participant represented in the at least one overview video output stream; assigning a first role name to the first meeting participant based on one or more detected characteristics associated with the first meeting participant; determining a likelihood that the first meeting participant will start speaking based on the first role name assigned to the first meeting participant; in response to the likelihood that the first meeting participant will start speaking satisfying a likelihood condition, performing one or more actions to cause the video processing unit to prepare to output a main video stream characterized by a view of the first meeting participant for display; determining, based on the analysis of the at least one overview video output stream, whether the first meeting participant has started speaking; and in response to determining that the first meeting participant has started speaking, outputting a main video stream characterized by a view of the first meeting participant for display.
[0047] According to a fifth aspect of the present disclosure, there is provided a video conferencing system comprising: at least one camera, each camera being configured to generate an overview video output stream representing a meeting environment; and at least one video processing unit configured to: automatically analyze the at least one overview video output stream to identify a plurality of meeting participants characterized in the at least one overview video output stream; identify a plurality of non-speaking meeting participants among the identified meeting participants; assign a corresponding role name to each of the identified non-speaking meeting participants based on one or more detected characteristics associated with each of the non-speaking meeting participants; and generate, as an output for display, a main video stream characterized by one or more of the non-speaking meeting participants, wherein the selection of the one or more non-speaking meeting participants to be included in the main video stream is based on the roles assigned to each of the non-speaking meeting participants.
[0048] The at least one video processing unit may be configured to: identify speaking meeting participants among the identified meeting participants; and the main video stream may be characterized by the speaking meeting participants.
[0049] According to a sixth aspect of the present disclosure, there is provided a processing method performed by at least one video processing unit, the method comprising: automatically analyzing at least one overview video output stream representing a meeting environment generated by at least one camera to identify a plurality of meeting participants characterized in the at least one overview video output stream; identifying a plurality of non-speaking meeting participants among the identified meeting participants; assigning a corresponding role name to each non-speaking meeting participant among the identified non-speaking meeting participants based on one or more detected characteristics associated with each non-speaking meeting participant; and generating a main video stream characterized by one or more non-speaking meeting participants among the non-speaking meeting participants as an output for display, wherein the selection of one or more non-speaking meeting participants to be included in the main video stream is based on the roles assigned to each non-speaking meeting participant among the non-speaking meeting participants.
[0050] According to a seventh aspect of the present disclosure, there is provided at least one video processing unit configured to perform any one of the methods described herein.
[0051] According to an eighth aspect of the present disclosure, there is provided one or more computer-readable storage media storing instructions that, when executed by at least one video processing unit, cause the at least one video processing unit to perform any one of the methods described herein.
[0052] Unless otherwise indicated, the features of the different disclosed embodiments may be combined interchangeably. For example, unless otherwise indicated, features or aspects described in the context of one embodiment may be used with or introduced into different disclosed embodiments.
[0053] The methods disclosed herein, such as those performed by a video processing unit, may be described as computer-implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The drawings incorporated in and constituting a part of this specification illustrate the disclosed embodiments and, together with the specification, are used to explain the disclosed embodiments. The details shown are by way of example and for purposes of illustrative discussion of the embodiments of the present disclosure. The description in conjunction with the drawings makes it apparent to those skilled in the art how the embodiments of the present disclosure may be practiced.
[0055] Figure 1 is a graphical representation of an example of a multi-camera system consistent with some embodiments of the present disclosure.
[0056] Figure 2A-2F is an example of a meeting environment consistent with some embodiments of the present disclosure.
[0057] Figure 3AIs a graphical representation of a multi-camera system in a meeting environment consistent with some embodiments of the present disclosure.
[0058] Figure 3B And Figure 3C Is an example of an output stream including a representation of a common meeting participant consistent with some embodiments of the present disclosure.
[0059] Figure 4 Is a flowchart of an example method for analyzing multiple video output streams and generating a main video stream consistent with some embodiments of the present disclosure.
[0060] Figure 5 Is a depiction of a facial visibility score consistent with some embodiments of the present disclosure.
[0061] Figure 6 Is an illustration of an overview stream, an output stream, and a viewfinder representation consistent with some embodiments of the present disclosure.
[0062] Fig. 7A And Figure 7B Is an example of a main video stream consistent with some embodiments of the present disclosure.
[0063] Figure 8 Is a flowchart of another example method for analyzing multiple video output streams and generating a main video stream consistent with some embodiments of the present disclosure.
[0064] Fig. 9 Is an example of viewfinder composition in an identified video output stream consistent with some embodiments of the present disclosure.
[0065] Fig.10 Is an illustration of a camera including a video processing unit consistent with some embodiments of the present disclosure.
[0066] Fig.11 Is a graphical representation of an area in a meeting environment consistent with some embodiments of the present disclosure.
[0067] Fig.12 Is an illustration of an overview stream and a main stream consistent with some embodiments of the present disclosure.
[0068] Fig.13 Illustrates an example of a golden rule consistent with some embodiments of the present disclosure.
[0069] Fig.14 Is a graphical representation of the components and connections of an example multi-camera system consistent with some embodiments of the present disclosure.
[0070] Fig.15 Illustrates examples of different lens types consistent with some embodiments of the present disclosure.
[0071] Figure 16A-16BAn example of a shot that includes other meeting participants and auxiliary items is illustrated, which is consistent with some embodiments of the present disclosure.
[0072] Fig.17 It is a graphical flow chart that illustrates a discovery process consistent with some embodiments of the present disclosure.
[0073] Fig.18 It is a graphical representation of an example DirectorWits architecture consistent with some embodiments of the present disclosure.
[0074] Fig.19 It is a flow chart that shows the main concepts and relationships in an example of DirectorWits consistent with some embodiments of the present disclosure.
[0075] Fig. 20 It is a graphical representation of an image processing technique consistent with some embodiments of the present disclosure.
[0076] Fig.21 A flow chart is provided that describes an example image processing pipeline consistent with some embodiments of the present disclosure.
[0077] Fig. 22 An example of a tile layout consistent with some embodiments of the present disclosure is provided.
[0078] Figure 23A-23B An example of a composite layout / group shot consistent with some embodiments of the present disclosure is illustrated.
[0079] Fig.24A An example of a gallery view type consistent with some embodiments of the present disclosure is illustrated.
[0080] Fig. 24B Examples of different overview view layouts consistent with some embodiments of the present disclosure are illustrated.
[0081] Fig.24C Examples of different split view layouts consistent with some embodiments of the present disclosure are illustrated.
[0082] Fig.25 An example of reframing consistent with some embodiments of the present disclosure is illustrated.
[0083] Figure 26A-26B Examples of overview layout configurations consistent with some embodiments of the present disclosure are illustrated.
[0084] Fig.26C Examples of layout configurations with objects of interest consistent with some embodiments of the present disclosure are illustrated.
[0085] Figure 26D-Figure 26F Examples of split view layout configurations consistent with some embodiments of the present disclosure are illustrated.
[0086] Figure 27A-27B is a flowchart representation of an example layout engine process consistent with some embodiments of the present disclosure.
[0087] Fig.28 is a graphical representation of an example of a multi-camera system implementing a layout engine consistent with some embodiments of the present disclosure.
[0088] Fig.29 is a graphical representation of an example layout engine process consistent with some embodiments of the present disclosure.
[0089] Fig.30 Illustrates examples of two types of modes within a gallery view consistent with some embodiments of the present disclosure.
[0090] Fig.31 Illustrates an example of a gallery view within an example video conferencing platform consistent with some embodiments of the present disclosure.
[0091] Figure 32A-Figure 32D Illustrates an example of a person viewfinder corresponding to a participant in a meeting environment consistent with some embodiments of the present disclosure.
[0092] Fig.33 Illustrates an example of a transition between frames or sub-frames consistent with some embodiments of the present disclosure.
[0093] Fig.34 Illustrates another example of a transition between frames or sub-frames consistent with some embodiments of the present disclosure.
[0094] Figure 35A-Figure 35B Illustrates an example of an output tile view consistent with some embodiments of the present disclosure.
[0095] Fig.36 Illustrates an example meeting environment with meeting participants consistent with some embodiments of the present disclosure.
[0096] Figure 37A-Figure 37Y Illustrates various example layouts based on Fig.36 the example meeting environment shown.
[0097] Figure 38A-Figure 38D Illustrates various example layouts involving clusters and group shots consistent with some embodiments of the present disclosure.
[0098] Fig.39 Illustrates various examples of a floating tile layout consistent with some embodiments of the present disclosure.
[0099] Fig.40 Illustrates various examples of an adjusted grid layout consistent with some embodiments of the present disclosure.
[0100] Fig.41 Illustrate various example geometric layouts consistent with some embodiments of the present disclosure.
[0101] Fig.42 Illustrate various example pie chart layouts consistent with some embodiments of the present disclosure.
[0102] Fig.43 Illustrate various example soft rock tile layouts consistent with some embodiments of the present disclosure.
[0103] Fig.44 Illustrate various example organic layouts consistent with some embodiments of the present disclosure.
[0104] Fig.45 Illustrate an example determination of a layout to be displayed consistent with some embodiments of the present disclosure.
[0105] Fig.46 Illustrate an example determination of a layout to be displayed consistent with some embodiments of the present disclosure.
[0106] Fig.47 Is a graphical representation of an example of a multi-camera system implementing a role assignment feature consistent with some embodiments of the present disclosure.
[0107] Fig.48 Provide an exemplary table illustrating possible meeting scenarios and example role and priority assignments consistent with some embodiments of the present disclosure.
[0108] Fig.49 Provide an exemplary table showing possible meeting scenarios and example role and priority assignments of user assignments consistent with some embodiments of the present disclosure.
[0109] Fig.50A 、 Fig.50B and Fig.50C Illustrate using vectors to predict which meeting participant will speak next. Detailed Description
[0110] Embodiments of the present disclosure may include features and techniques for presenting the video of participants on a display. Conventional video conferencing systems and associated software may have the ability to present the video of participants on a display. In some cases, static images of one or more conference participants may be displayed on the display. In other cases, one or more conference participants may be selectively featured on the display (e.g., based on detected audio from one or more microphones). However, with such systems, remote users may have difficulty seeing certain conference participants adequately or interacting with certain conference participants. For example, in a video shot showing a conference room table and all the participants sitting around the table, a remote user may have difficulty seeing and interacting with a conference participant sitting at the other end of the table. Additionally, it may be difficult to see and interact with a conference participant presented in a profile or from behind in the video feed. Further, even in systems that are able to highlight one or more conference participants, a remote user may have difficulty or be unable to determine how the featured speaker is received by others in the room, especially in cases where the others are not presented with the featured speaker. At least for these reasons, a remote user may feel isolated, detached, and / or not integrated during a video conferencing event.
[0111] The disclosed systems and methods may provide single-camera or multi-camera video conferencing systems that naturally and dynamically follow conversations and meeting interactions that occur between participants sitting in a conference room and more broadly between conference participants distributed across multiple environments. In some embodiments, the disclosed systems and methods may detect what is happening within an environment (e.g., a conference room, a classroom, a virtual distributed meeting environment, etc.) and adapt the video feed view based on an analysis of detected events, interactions, movements, audio, etc. In some cases, the adapted video feed (e.g., including the framing and display of selected conference participants in one or more output video streams) may be based on the determined or inferred roles of the detected conference participants. The adapted video feed may also be based on the determined characteristics of the detected conference participants (e.g., level of participation, etc.).
[0112] In this way, remote users can feel and be more involved and included in the group of meeting participants, including when the group is together in a common physical environment. For example, the disclosed video conferencing camera system can assist the end user in determining who is speaking, who the discussion is directed at, how non-speaking participants are receiving the discussion, etc. As a result of being provided with this contextual information about the meeting environment and meeting participants, remote users participating in the meeting may virtually feel closer to being in the same room as the group of meeting participants. Such remote users can more easily follow the flow of the conversation, understand more of the discussion (even information conveyed subtly through body language, facial expressions, gestures, etc. - common features of multi-person discussions that may be missed or not discernible using traditional video conferencing platforms), more easily identify speakers and listeners, and gain more context from the meeting session. Such features can help remote participants play a more active role in conversations with other meeting participants.
[0113] In addition, the disclosed embodiments can provide systems and methods for event detection and analysis with respect to the bodies of meeting participants (e.g., in the head region), the direction of received audio (e.g., in combination with video analysis / detection), the movement patterns / history of meeting participants, speaker tracking, the predicted roles of meeting participants, etc. Such event detection and analysis can be used to determine which subjects to feature in a synthetic video (or a series of video stream outputs), the positions of meeting participants to display on a display, the relative positioning between meeting participants on the display, what type of highlighting technique to use with respect to a selected meeting participant, which audio feed to select, how long to display certain video streams, how to transition between different video frames, etc.
[0114] Traditional video conferencing platforms may be associated with additional challenges. For example, when a call enters a meeting, remote video conferencing participants may have difficulty feeling integrated into the physical meeting environment. At the same time, for meeting participants located in the physical meeting environment, it may be difficult to divide their attention between the screen and the meeting environment.
[0115] Depending on the activity and the number of attendees, meetings may require different spaces. They may also require different types and levels of concentration, attention, and presence. Examples of meeting participants can include a primary speaker, contributors, or an audience listening in. For example, in an educational setting, meeting participants can include a teacher or lecturer and students, where the teacher or lecturer is typically the primary speaker and the students are typically contributors or members of the listening audience. In either case, it can be beneficial for all participants to feel connected and have the opportunity to contribute to the meeting. Additionally, a hybrid office situation may continue to be prevalent, and participants may attend meetings where some participants are located in a physical meeting environment while other participants join the video conference from elsewhere. The disclosed embodiments can provide an experience in which all meeting participants, including those physically present and remote participants, can contribute and participate at the same level in the meeting. Further, the disclosed embodiments can provide context information to remote meeting participants regarding the physical meeting environment and the meeting participants present in the meeting environment.
[0116] The present disclosure provides a video conferencing system and a camera system for video conferencing. Thus, where a camera system is mentioned herein, it should be understood that this can alternatively be referred to as a video conferencing system, a video conferencing camera system, or a camera system for video conferencing. As used herein, the term "video conferencing system" refers to a system (such as a video conferencing camera) that can be used for video conferencing and can alternatively be referred to as a system for video conferencing. A video conferencing system does not need to be able to provide video conferencing capabilities on its own and can interface with other devices or systems, such as a laptop, a PC, or other network-enabled devices, to provide video conferencing capabilities.
[0117] A video conferencing system / camera system according to the present disclosure can include at least one camera and a video processor for processing the video output generated by the at least one camera. The video processor can include one or more video processing units.
[0118] According to an embodiment, a video conferencing camera may include at least one video processing unit. The at least one video processing unit may be configured to process a video output generated by the video conferencing camera. As used herein, a video processing unit may include any electronic circuitry designed to read, manipulate, and / or alter computer-readable memory to create, generate, or process video images and video frames intended for output (e.g., in a video output or video feed) to a display device. The video processing unit may include one or more microprocessors or other logic-based devices configured to receive digital signals representative of the acquired images. The disclosed video processing unit may include an application specific integrated circuit (ASIC), a microprocessor unit, or any other suitable structure for analyzing the acquired images, selectively framing a subject based on the analysis of the acquired images, generating an output video stream, and the like.
[0119] In some cases, the at least one video processing unit may be located within a single camera. In other words, the video conferencing camera may include the video processing unit. In other embodiments, the at least one video processing unit may be located remotely from the camera or may be distributed among multiple cameras and / or devices. For example, the at least one video processing unit may include more than one video processing unit distributed among a group of electronic devices including one or more cameras (e.g., a multi-camera system), a personal computer, a mobile device (e.g., a tablet, a phone, etc.), and / or one or more cloud-based servers. Accordingly, a video conferencing system, such as a video conferencing camera system, is disclosed herein that includes at least one camera and at least one video processing unit as described herein. The at least one video processing unit may or may not be implemented as part of the at least one camera. The at least one video processing unit may be configured to receive a video output generated by one or more video conferencing cameras. The at least one video processing unit may generate a video output (e.g., one or more video output streams) for display. The at least one video processing unit may decode digital signals to display video and / or may store image data in a memory device. In some embodiments, the video processing unit may include a graphics processing unit. It should be understood that where a video processing unit is referred to in the singular herein, more than one video processing unit is also contemplated. The various video processing steps described herein may be performed by the at least one video processing unit, and the at least one video processing unit may thus be configured to perform a method as described herein, such as a video processing method, or any step of such a method. Where a determination of a parameter, value, or quantity is disclosed herein with respect to this method, it should be understood that the at least one video processing unit may perform the determination and may thus be configured to perform the determination.
[0120] Single-camera and multi-camera systems are described herein. Although some features may be described for a single camera and other features may be described for a multi-camera system, it should be understood that any and all of the features, embodiments, and elements herein may relate to both single-camera and multi-camera systems, or may be implemented in both single-camera and multi-camera systems. For example, some features, embodiments, and elements may be described as relating to a single-camera system. It should be understood that those features, embodiments, and elements may relate to and / or be implemented in a multi-camera system. Additionally, other features, embodiments, and elements may be described as relating to a multi-camera system. It should also be understood that these features, embodiments, and elements may relate to and / or be implemented in a single-camera system.
[0121] Embodiments of the present disclosure include multi-camera systems. As used herein, a multi-camera system may include two or more cameras employed in an environment (such as a conference environment) and may simultaneously record or broadcast one or more representations of the environment. The disclosed cameras may include any device that includes one or more photosensitive sensors configured to capture a stream of image frames. Examples of cameras may include, but are not limited to L1 or S1 cameras, IQ cameras, digital cameras, smartphone cameras, compact cameras, digital single-lens reflex (DSLR) video cameras, mirrorless cameras, action (adventure) cameras, 360-degree cameras, medium format cameras, webcams, or any other device for recording visual images and generating a corresponding video signal.
[0122] Referring to Figure 1, provides a graphical representation of an example of a multi-camera system 100 consistent with the present invention. The multi-camera system 100 may include a main camera 110, one or more peripheral cameras 120, one or more sensors 130, and a host computer 140. In some embodiments, the main camera 110 and the one or more peripheral cameras 120 may be of the same camera type, such as but not limited to the examples of cameras discussed above. Additionally, in some embodiments, the main camera 110 and the one or more peripheral cameras 120 may be interchangeable such that the main camera 110 and the one or more peripheral cameras 120 may be located together in a conferencing environment and any one of the cameras may be selected to be used as the main camera. Such a selection may be based on various factors such as but not limited to the position of the speaker, the layout of the conferencing environment, the position of auxiliary items (e.g., whiteboard, presentation screen, television), etc. In some cases, the main camera and the peripheral cameras may operate in a master-slave arrangement. For example, the main camera may include most or all of the components for video processing associated with the multiple outputs of the various cameras included in the multi-camera system. In other cases, the system may include a more distributed arrangement where the video processing components (and tasks) are more evenly distributed across the various cameras of the multi-camera system.
[0123] As Figure 1 shown, the main camera 110 and the one or more peripheral cameras 120 may each include an image sensor 111, 121. Additionally, the main camera 110 and the one or more peripheral cameras 120 may include a Directional Audio (DOA / Audio) unit 112, 122. The DOA / Audio units 112, 122 may detect and / or record audio signals and determine the direction from which one or more audio signals originate. In some embodiments, the DOA / Audio units 112, 122 may determine or be used to determine the direction of a speaker in a conferencing environment. For example, the DOA / Audio units 112, 122 may include a microphone array that may detect audio signals from different positions relative to the main camera 110 and / or the one or more peripheral cameras 120. The DOA / Audio units 112, 122 may use the audio signals from the different microphones and determine the angle and / or position from which the audio signal (e.g., speech) originates. Additionally or alternatively, in some embodiments, the DOA / Audio units 112, 122 may distinguish between a situation in a conferencing environment where a conference participant is speaking and other situations where there is silence in the conferencing environment. In some embodiments, determining the direction from which one or more audio signals originate and / or distinguishing between different situations in a conferencing environment may be determined by a unit other than the DOA / Audio units 112, 122 (such as one or more sensors 130).
[0124] The main camera 110 and one or more peripheral cameras 120 may include vision processing units 113, 123. The vision processing units 113, 123 may include one or more hardware-accelerated programmable convolutional neural networks with pre-trained weights, which may detect different attributes from video and / or audio. For example, in some embodiments, the vision processing units 113, 123 may use a vision pipeline model to determine the positions of meeting participants in a meeting environment based on the representations of the meeting participants in an overview stream. As used herein, the overview stream may include a video recording of the meeting environment at the standard zoom and perspective of the camera used to capture the recording or at the maximum zoom-out perspective of the camera. In other words, the overview shot or stream may include the maximum field of view of the camera. Alternatively, the overview shot may be a scaled or cropped portion of the full video output of the camera, but may still capture an overview shot of the meeting environment. Generally, the overview shot or overview video stream may capture an overview of the meeting environment and may be framed, for example, to feature the representations of all or substantially all of the meeting participants within the camera's field of view, or present in the meeting environment and detected or recognized by the system (e.g., by one or more video processing units based on an analysis of the camera output). The main stream or focused stream may include a focused, enhanced, or magnified recording of the meeting environment. In some embodiments, the main stream or focused stream may be a sub-stream of the overview stream. As used herein, a sub-stream may refer to a video recording that captures a portion or sub-frame of the overview stream. Additionally, in some embodiments, the vision processing units 113, 123 may be trained to be unbiased towards various parameters, including but not limited to gender, age, race, scene, light, and size, thus allowing for a robust meeting or video conferencing experience.
[0125] As Figure 1As shown, the main camera 110 and one or more peripheral cameras 120 may include virtual director units 114, 124. In some embodiments, the virtual director units 114, 124 may control the main video stream that can be consumed by the connected host computer 140. In some embodiments, the host computer 140 may include one or more of a television, a laptop computer, a mobile device, or a projector or any other computing system. The virtual director units 114, 124 may include software components that may use inputs from the vision processing units 113, 123 and determine the video output stream and from which camera (e.g., among the main camera 110 and one or more peripheral cameras 120) to stream to the host computer 140. The virtual director units 114, 124 may create an automated experience that may be similar to a television talk show production or an interactive video experience. In some embodiments, the virtual director units 114, 124 may frame the representation of each meeting participant in a meeting environment. For example, the virtual director units 114, 124 may determine that a camera (e.g., a camera of the main camera 110 and / or one or more peripheral cameras 120) may provide an ideal picture or shot of the meeting participants in the meeting environment. The ideal picture or shot may be determined by various factors including, but not limited to, the angle of each camera relative to the meeting participants, the position of the meeting participants, the level of participation of the meeting participants, or other attributes associated with the meeting participants. More non-limiting examples of attributes associated with a meeting participant that may be used to determine the ideal picture or shot of the meeting participant may include: whether the meeting participant is speaking, the duration for which the meeting participant has been speaking, the direction of the meeting participant's gaze, the percentage of the meeting participant that is visible in the picture, the reaction and body language of the meeting participant, or other meeting participants visible in the picture.
[0126] The multi-camera system 100 may include one or more sensors 130. The sensors 130 may include one or more intelligent sensors. As used herein, an intelligent sensor may include a device that receives input from a physical environment and, upon detecting a particular input, uses built-in or associated computing resources to perform a predefined function and processes the data before sending the data to another unit. In some embodiments, one or more sensors 130 may send data to the main camera 110 and / or one or more peripheral cameras 120, or to the at least one video processing unit. Non-limiting examples of sensors may include level sensors, current sensors, humidity sensors, pressure sensors, temperature sensors, proximity sensors, thermal sensors, flow sensors, fluid velocity sensors, and infrared sensors. Additionally, non-limiting examples of intelligent sensors may include touchpads, microphones, smart phones, GPS trackers, echolocation sensors, thermometers, humidity sensors, and biometric sensors. Further, in some embodiments, one or more sensors 130 may be placed throughout the conferencing environment. Additionally or alternatively, the sensors in one or more sensors 130 may be of the same type or different types of sensors. In other cases, the sensors 130 may generate an (or more) original signal output and send it to one or more processing units, which may be located on the main camera 110 or distributed between two or more cameras included in the multi-camera system. The processing unit may receive the (or more) original signal output, process the received signal, and use the processed signal to provide various features of the multi-camera system (such features are discussed in more detail below).
[0127] As Figure 1 shown, one or more sensors 130 may include an application programming interface (API) 132. Additionally, also as Figure 1 shown, the main camera 110 and one or more peripheral cameras 120 may include APIs 116, 126. As used herein, an API may refer to a set of defined rules that may enable different applications, computer programs, or units to communicate with each other. For example, the API 132 of one or more sensors 130, the API 116 of the main camera 110, and the API 126 of one or more peripheral cameras 120 may be interconnected, as Figure 1As shown, and allows one or more sensors 130, the main camera 110, and one or more peripheral cameras 120 to communicate with each other. It is contemplated that the APIs 116, 126, 132 can be connected in any suitable manner, such as but not limited to, via Ethernet, local area network (LAN), wired or wireless network. It is also contemplated that each sensor in one or more sensors 130 and each camera in one or more peripheral cameras 120 can include an API. In some embodiments, the host computer 140 can be connected to the main camera 110 via the API 116, and the API 116 can allow communication between the host computer 140 and the main camera 110.
[0128] The main camera 110 and one or more peripheral cameras 120 can include stream selectors 115, 125. The stream selectors 115, 125 can receive the overview stream and the focused stream of the main camera 110 and / or one or more peripheral cameras 120, and provide the updated focused stream (e.g., based on the overview stream or the focused stream) to the host computer 140. The selection of the stream to be displayed to the host computer 140 can be performed by the virtual director units 114, 124. In some embodiments, the selection of the stream to be displayed to the host computer 140 can be performed by the host computer 140. In other embodiments, the selection of the stream to be displayed to the host computer 140 can be determined by user input received via the host computer 140, where the user can be a conference participant.
[0129] In some embodiments, an autonomous video conferencing (AVC) system is provided. The AVC system can include any combination of any one or all of the features described above with respect to the multi-camera system 100. Additionally, in some embodiments, one or more peripheral cameras and intelligent sensors of the AVC system can be placed in a separate video conferencing space (or meeting environment) as a secondary space for video conferencing (or meeting). These peripheral cameras and intelligent sensors can be networked with the main camera and adapted to provide image and non-image inputs from the secondary space to the main camera. In some embodiments, the AVC system can be adapted to generate an automated television studio production for a combined video conferencing space based on inputs from cameras and intelligent sensors in both spaces.
[0130] In some embodiments, the AVC system can include intelligent cameras adapted with different angular fields of view. For example, in a small video conferencing (or meeting) space with fewer intelligent cameras, the intelligent cameras can have a wide field of view (e.g., approximately 150 degrees). As another example, in a large video conferencing (or meeting) space with more intelligent cameras, the intelligent cameras can have a narrow field of view (e.g., approximately 90 degrees). In some embodiments, the AVC system can be equipped with intelligent cameras having various angular fields of view, thus allowing optimal coverage of the video conferencing space.
[0131] In addition, in some embodiments, at least one image sensor of the AVC system can be adapted to zoom up to 10X to achieve a close-up image of an object at the far end of the video conferencing space. Additionally or alternatively, in some embodiments, at least one smart camera in the AVC system can be adapted to capture content on or around an object that can be a non-person item within the video conferencing space (or meeting environment). Non-limiting examples of non-person items include whiteboards, television (TV) displays, posters, or presentation stands. The camera adapted to capture content on or around an object can be smaller and placed differently from other smart cameras in the AVC system, and can be mounted, for example, on the ceiling to provide effective coverage of the target content.
[0132] At least one audio device (e.g., a DOA audio device) in the smart camera of the AVC system can include a microphone array that is adapted to output an audio signal representing sounds originating from different positions and / or directions around the smart camera. Signals from different microphones can allow the smart camera to determine the direction of arrival (DOA) of the audio signal associated with it, and to discern, for example, whether there is silence at a particular location or direction. Such information can be available for the vision pipeline and virtual director included in the AVC system. In some embodiments, a computing device with high computing power can be connected to the AVC system via an Ethernet switch. The computing device can be adapted to provide additional computing power to the AVC system. In some embodiments, the computing device can include one or more high-performance CPUs and GPUs, and can run portions of the vision pipeline for the main camera and any designated peripheral cameras.
[0133] Figure 2A-2F A graphical representation including various examples of meeting environments consistent with some embodiments of the present disclosure. Figure 2A An example of a meeting room 200a is depicted. The meeting room 200a can include a table 210, three meeting participants 212, seven cameras 214, and a display unit 216. Figure 2B An example of a meeting room 200b is depicted. The meeting room 200b can include two tables 220, two meeting participants 222, and four cameras 224. Figure 2C An example of a video conferencing space 200c is depicted. The video conferencing space 200c can include a table 230, nine meeting participants 232, nine cameras 234, and two display units 226. Figure 2D An example of a boardroom 200d is depicted. The boardroom 200d can include a table 240, eighteen meeting participants 242, ten cameras 244, and two display units 246. Figure 2EDepicts an example of a classroom 200e. The classroom 200e can include a plurality of conference participants 252, seven cameras 254, and a display unit 256. Figure 2F Depicts an example of a lecture hall 200f. The lecture hall 200f can include a plurality of conference participants 262, nine cameras 264, and a display unit 266. Although specific numbers are used to refer to, for example, the number of tables, conference participants, cameras, and display units, it is contemplated that the meeting environment can contain any suitable number of tables, furniture (sofas, benches, meeting pods, etc.), conference participants, cameras, and display units. It is also contemplated that the tables, conference participants, cameras, and display units can be organized in any location within the meeting environment and are not limited to the depictions herein. For example, in some embodiments, the cameras can be placed "in a line" or in the same horizontal and / or vertical plane relative to each other. Additionally, it is contemplated that the meeting environment can include any other components not discussed above, such as but not limited to whiteboards, presentation screens, shelves, and chairs. For purposes of description, a name is given to each meeting environment (e.g., conference room, meeting room, video conferencing space, boardroom, classroom, lecture hall), and Figure 2A-2F each meeting environment shown in is not limited to the name associated with it herein.
[0134] In some embodiments, by placing a plurality of wide - field single - lens cameras, a multi - system camera can create a variety of, flexible, and interesting experiences, where these wide - field single - lens cameras cooperate to frame conference participants in the meeting environment as the conference participants engage and participate in the session from different camera angles and zoom levels. This can give remote participants (e.g., participants farther from the camera, participants participating remotely or via video conferencing) a natural sense of what is happening in the meeting environment.
[0135] The disclosed embodiments may include a multi-camera system including a plurality of cameras. Each camera may be configured to generate a video output stream representing a meeting environment. Each video output stream may be characterized by one or more meeting participants present in the meeting environment. In this context, "characterized" means that the video output stream includes or is characterized by a representation of one or more meeting participants. For example, a first representation of a meeting participant may be included in a first video output stream from a first camera included in the plurality of cameras, and a second representation of the meeting participant may be included in a second video output stream from a second camera included in the plurality of cameras. As used herein, a meeting environment may refer to any space in which there is a gathering of people interacting with each other. Non-limiting examples of a meeting environment may include a meeting room, a classroom, a lecture hall, a video conferencing space, or an office space. As used herein, a representation of a meeting participant may refer to an image, video, or other visual rendering of the meeting participant that may be captured, recorded, and / or displayed to, for example, a display unit. A video output stream or video stream may refer to a media component (which may include visual and / or audio rendering) that may be delivered via a wired or wireless connection to, for example, a display unit and played back in real time. Non-limiting examples of a display unit may include a computer, a tablet, a television, a mobile device, a projector, a projection screen, or any other device that may display or show an image, video, or other rendering of the meeting environment.
[0136] Referring Figure 3A , a graphical representation of a multi-camera system in a meeting environment 300 consistent with some embodiments of the present disclosure is provided. Cameras 310a - 310c (e.g., a plurality of cameras) may record the meeting environment 300. The meeting environment 300 may include a table 340 and meeting participants (such as meeting participants 330, 332). As Figure 3A shown, cameras 310a - 310c may capture portions of the meeting environment 300 in their respective flow directions (such as flow directions 320a - 320c).
[0137] Referring Figure 3B-Figure 3C , output streams 370a, 370b may include representations 380, 390 of a common meeting participant 360. For example, representation 380 may be included in output stream 370a from camera 310b in flow direction 320b. Similarly, representation 390 may be included in output stream 370b from camera 310c in flow direction 320c. As Figure 3A shown, cameras 310b and 310c may be located at different positions and include different flow directions 320b, 320c such that output streams 370a, 370b from each of cameras 310b, 310c may include different representations 380, 390 of a common or the same meeting participant 360.
[0138] In some embodiments, it is contemplated that the output stream may display representations of more than one conference participant, and the representations may include representations of common or identical conference participants. It is also contemplated that in some embodiments, the output stream may include representations of different conference participants. For example, in some embodiments, the output video stream generated by cameras 310a - 310c may include an overview stream that includes a field of view that is wider or larger than that of the examples of Figure 3B and Figure 3C . In some cases, for example, the output streams provided by cameras 310b and 310c may include a representation of conference participant 360 and a representation of one or more (or all) of the other conference participants included in the conference environment (e.g., any or all of the participants located around table 340). From the overview video stream, a focused main video stream can be selected and generated based on shot selection criteria (such as those shown in Figure 3B and Figure 3C ), as discussed further below.
[0139] In some embodiments, the multi - camera system may include one or more video processing units. In some embodiments, the video processing unit may include at least one microprocessor deployed in a housing associated with one of the plurality of cameras. For example, the video processing unit may include vision processing units 113, 123; virtual director units 114, 124; or both vision processing units 113, 123 and virtual director units 114, 124. The video processing unit may also implement the processing pipelines (such as the vision pipeline) described herein. As shown by the examples in Figure 1 , video processing may occur on the camera and be displayed by the video processing unit located on the camera of the multi - camera system. Additionally, in some embodiments, the (one or more) video processing units may be located remotely with respect to the plurality of cameras. For example, referring to Figure 1 , the (one or more) video processing units may be located in host computer 140. As another example, the (one or more) video processing units may be located on a remote server (such as a server in the cloud). In some embodiments, the video processing unit may include multiple logical devices distributed across two or more of the plurality of cameras. Additionally, the (one or more) video processing units may be configured to execute method 400 of analyzing multiple video output streams and generating a main video stream, as shown in Figure 4 .
[0140] Referring to Figure 4, as shown in step 410, method 400 may include automatically analyzing multiple video streams to determine whether multiple representations correspond to the same conference participant, each representation being included in a video stream among the multiple video streams. For example, the video processing unit may be configured to automatically analyze a first video output stream and a second video output stream to determine whether a first representation of a conference participant and a second representation of the conference participant correspond to a common conference participant (i.e., the same person). A common conference participant may refer to a single specific conference participant represented in the outputs of two or more output video streams. As an example, and with reference to Figure 3A-3C , the video processing unit (e.g., the vision processing unit 113 and / or the virtual director unit 114, the video processing unit located in the host computer 140, etc.) may automatically analyze the video output streams 370a, 370b to determine whether the representations 380, 390 correspond to the common conference participant 360. Similar analysis may be performed with respect to multiple output video streams, each output video stream including representations of multiple individuals. Using various identification techniques, the video processing unit may determine which individual representations across multiple camera outputs correspond to Participant A, which correspond to Participant B, which correspond to Participant C, and so on. This identification of conference participants across the video outputs of multiple camera systems may provide the system with the ability to select one representation of a particular conference participant rather than another representation of the same conference participant to feature in the output of the multi-camera system. Only as an example, in some cases and based on various criteria related to the individuals, the interactions between conference participants, and / or the conditions or characteristics of the conference environment, the system may select the representation 380 of the conference participant 360 rather than the representation 390 of the conference participant 360 to feature in the output of the multi-camera system.
[0141] The analysis for determining whether two or more conference participant representations correspond to a common conference participant may be based on at least one identity indicator. The identity indicator may include any technique or may be based on any technique suitable for associating the identities of individuals represented in the video output streams. In some embodiments, the at least one identity indicator may include an embedding determined for each of the first representation and the second representation. As used herein, an embedding may include a digital representation of a video stream (e.g., one or more frames associated with the output streams 370a, 370b), a section or segment of a video stream (e.g., a sub-section associated with one or more captured frames included in the video stream), an image, a region of a captured image frame including a representation of a particular individual, etc. In some cases, the embedding may be expressed as an N-dimensional vector (e.g., a feature vector). For example, the embedding may include at least one feature vector representation. In Figure 3B and Figure 3CIn the example of, at least one identity indicator may include a first feature vector embedding determined with respect to a first representation of the conference participant (e.g., representation 380) and a second feature vector determined with respect to a second representation of the conference participant (e.g., representation 390).
[0142] At least one feature vector representation may include a series of numbers generated based on characteristics specific to or representing the subject being represented. Factors that may contribute to the generated series of numbers may include eye color, hair color, clothing color, body contour, skin color, eye shape, facial shape, presence / color / type of facial hair, etc. Notably, the generation of the feature vector is repeatable. That is, repeatedly exposing the feature vector generator to the same image or image segment will result in the repeated generation of the same feature vector.
[0143] Such an embedding can also be used as a basis for identification. For example, in the case where feature vectors are determined for each of individuals A, B, and C represented in a first image frame derived from the output of a first camera, those feature vectors can be used to determine whether any of individuals A, B, or C is represented in a second image frame derived from the output of a second camera. That is, feature vectors can be generated for each of individuals X, Y, and Z represented in the second image frame. In the vector space, the distances between various feature vectors can be determined as a basis for comparing the feature vectors. Thus, although the feature vector determined for individual A may not be exactly the same as any of the feature vectors generated for individuals X, Y, or Z, the A feature vector can closely match one of the X, Y, or Z feature vectors. If, for example, the distance between the feature vectors for individual A is within a predetermined distance threshold of the feature vectors generated for individual Z, it can be determined that individual A in the first frame corresponds to individual Z in the second frame. Similar comparisons can be performed with respect to other conference participants and for multiple frames from the outputs of multiple different cameras. Based on this analysis, the system can: determine and track which individuals are represented in the outputs of which cameras; and also identify various individuals across the available camera outputs. This identification, correlation, and tracking can allow the system to compare the available shots of a particular individual and select a particular shot of the individual rather than another shot of the individual based on various criteria for output as part of the output of the camera system.
[0144] Other types of identifiers or recognition techniques can also be used to associate representations of an individual across multiple camera outputs. Such alternative techniques can be used alone or in combination with the feature vector embedding method or any other recognition technique described herein. In some cases, at least one identity indicator can include one or more of a body contour or silhouette shape associated with the individual, at least one body dimension, and / or at least one color indicator. Such techniques can be helpful in situations where the representation of the face includes an invisible or only partially visible face relative to one or more image frames. As used herein, the body contour can refer to the shape of the body of a meeting participant. The silhouette shape can refer to the shape of the body, face, etc. (or any subpart of the face or body) of the meeting participant represented in the image frame. The body dimensions can include, but are not limited to, the height, width, or depth of any feature associated with the body of the meeting participant. The color indicator can be associated with the color and / or shade of the representation of the skin, hair, eyes, clothing, jewelry, or any other part of the body of the meeting participant. It is contemplated that at least one identity indicator can include any unique feature of the meeting participant, such as unique facial features and / or body features.
[0145] The identifier / recognition technique can be based on a corresponding analysis of a series of captured images and image streams. For example, in some embodiments, at least one identity indicator can include tracked lip movement. For example, as Figure 3B and Figure 3C shown, both representation 380 and representation 390 can show the mouth or lips of the common meeting participant 360 closing and / or not moving. The video processing unit can determine that representation 380, 390 corresponds to the common meeting participant 360, in whole or in part, by tracking the lip movement of the common meeting participant 360 across a corresponding series of captured images (e.g., the image stream associated with the output of camera 310b and another image stream associated with the output of camera 310c). The video processing unit can determine that the lips of the common meeting participant 360 are closed and / or not moving in both representations 380, 390 and identify the common meeting participant 360 as the common meeting participant rather than two different meeting participants. Additionally, the video processing unit can track the movement of the detected lips or mouth across different image streams. Temporal correspondence between the movement of the lips and / or mouth represented across two or more image streams can indicate the representation of the common meeting participant across two or more image streams. As another example, if the first representation of a meeting participant shows lip movement and the second representation of the meeting participant shows no lip movement (or is in a different position from the lips of the first representation), the video processing unit can determine that the first representation and the second representation correspond to two different meeting participants.
[0146] Techniques other than image analysis can also be used to identify common conference participants across multiple camera outputs. For example, in some embodiments, an audio track can be associated with each camera video output stream. The audio track associated with each camera video output stream can be recorded by an audio recording device or audio source (such as a microphone or microphone array) associated with the camera from which the corresponding video output stream is output. The audio track can relate to a stream of recorded sounds or audio signals. By way of example, and with reference to Figure 3A-3C , a first audio track can be associated with output stream 370a, and a second audio track can be associated with output stream 370b. The video processing unit can be configured to determine whether representations 380, 390 correspond to common conference participant 360 based on an analysis of the first audio track and the second audio track. Such an analysis can be based on a time synchronization analysis of the audio signals, and can also include a time synchronization analysis with tracked lip / mouth movement obtainable from image analysis techniques. Additionally or alternatively, a single audio track can be associated with the conference environment 300, and the video processing unit can be configured to determine whether representations 380, 390 correspond to common conference participant 360 based on an analysis of the single audio track in combination with tracked lip / mouth movement.
[0147] It is contemplated that at least one identity indicator can include any combination of the non-limiting examples of the identity indicators discussed previously. For example, at least one identity indicator can include at least one embedding determined for each of a first representation (e.g., representation 380) and a second representation (e.g., representation 390), and at least one color indicator (e.g., hair color, eye color, skin color) associated with each of the first representation and the second representation.
[0148] It is also contemplated that the video processing unit can determine that a first conference participant representation and a second conference participant representation do not correspond to a common conference participant. For example, using any one or combination of the techniques described above, the video processing unit can determine that a first conference participant representation and a second conference participant representation correspond to different conference participants.
[0149] Using information including which conference participants are represented in which camera outputs and which representations in those camera outputs correspond to which participants, the video processing unit can select a particular camera output for generating a feature shot of a particular conference participant (e.g., a preferred or best shot of a particular conference participant from the available representations from multiple cameras). The shot selection can depend on various shot selection criteria. For example, as Figure 4As shown in step 420, method 400 may include evaluating multiple representations of co-conference participants relative to one or more predetermined criteria that may be used in shot selection. The predetermined criteria may include, but are not limited to, the viewing direction (e.g., head pose) of the co-conference participants determined for each of the first video output stream and the second video output stream, and / or a facial visibility score associated with the co-conference participants (e.g., associated with a facial visibility level) determined for each of the first video output stream and the second video output stream, and / or a determination of whether the conference participant is speaking. It is contemplated that any one or all of these criteria may be analyzed relative to any number of video output streams including representations of co-conference participants. For example, shot selection may be based on any one of these criteria individually or in any combination.
[0150] A co-conference participant may be detected as speaking based on an audio track that includes the speech of the co-conference participant (or speech originating from the direction of the co-conference participant) and / or tracked lip movements identified according to an analysis of the video output from a camera of the video conferencing system. As used herein, head pose may relate to the degree to which the head of the conference participant is angled or turned, and / or the position of the head of the conference participant relative to other anatomical body parts of the conference participant (e.g., hands, arms, shoulders). Facial visibility level may relate to the percentage or proportion of the face of the conference participant that is visible in a particular output stream (e.g., facial visibility score).
[0151] As an example, and with reference to Figure 3A-3C , the positions of cameras 310b, 310c and the position of co-conference participant 360 may be used to determine the viewing direction of co-conference participant 360. The relative distances and angles between cameras 310b, 310c and their stream directions 320b, 320c may be used to calculate the angles at which the viewing direction or profile of the face of co-conference participant 360 should be represented in each output stream 370a, 370b. This calculation may be considered a predetermined criterion and may be used to evaluate representations 380, 390 of co-conference participant 360. It is contemplated that the viewing direction of the co-conference participant and / or the angle associated with the viewing direction may be used for any suitable number of representations of co-conference participant 360 included in any number of video output streams from any number of cameras in a multi-camera system. In Figure 3B and Figure 3C 's example, conference participant 360 has a viewing direction of approximately 30 degrees relative to the normal of the capture camera in Figure 3B (meaning that representation 380 of participant 360 looks to the left of the capture camera from the body's reference frame). In contrast, Figure 3CThe participant 360 therein has a viewing direction of approximately 0 degrees relative to the normal of the capture camera (meaning that the representation 390 of the participant 360 directly views the capture camera). In some cases, a viewing direction of 0 degrees may be preferred over other viewing directions. In other cases, a head pose that provides an indirect gaze (such as Figure 3B a 30-degree head pose) may be preferred and can be used as a basis for the shot of the participant 360.
[0152] As another example, a facial visibility score can be used to evaluate the representation of co-meeting participants, which in turn can be used as a basis for shot selection with respect to a particular meeting participant. Figure 5 Multiple examples of individuals with different head poses are provided, resulting in various degrees of facial visibility. A facial visibility score consistent with some embodiments of the present disclosure can be assigned to each of the different captured image frames. As Figure 5 shown, the facial visibility score can be determined based on the percentage or proportion of the face of the subject or meeting participant that can be seen in the output stream or frame. In some embodiments, the representations of the meeting participants can be shown in different output streams 500a - 500e. In other embodiments, the representations of the meeting participants can be the representations of the meeting participants in different frames displayed within a single output stream or at different times. Embodiments of the present disclosure can provide the facial visibility score as a percentage of the face of the subject or meeting participant visible in the frame or output stream. As Figure 5 shown, in output stream 500a, 93% of the face of the meeting participant can be visible. In output stream 500b, 75% of the face of the meeting participant can be visible. In output stream 500c, 43% of the face of the meeting participant can be visible. In output stream 500d, 25% of the face of the meeting participant can be visible. In output stream 500e, 1% of the face of the meeting participant can be visible. In some embodiments, the facial visibility score can be a score between 0 and 1, or any other indicator that can be used to convey the amount of the face of the meeting participant represented in a particular image frame or stream of image frames (e.g., an average facial visibility score over a series of images captured from the output of a particular camera).
[0153] The evaluation of the shot selection criteria described above can enable the video processing unit to select a camera output from which to generate a desired shot of a particular meeting participant. Returning to Figure 4, in step 430, method 400 may include selecting an output stream from among multiple output streams to serve as the source of the framed representation of a common conference participant in the multiple camera outputs. As used herein, the term "framed representation" of an individual may be understood to mean a video representation of the individual that is framed (e.g., included or characterized) within a video stream. In other words, the video stream is framed or cropped such that it is characterized by or includes the representation of the individual. Thus, the framed representation of an individual is a representation associated with or defined by a video framing or cropping scheme that includes the representation. The selected output stream may then be used as the basis for outputting a primary video stream that is characterized by the framed representation of the common conference participant (e.g., the desired shot of the common conference participant). The framed shot of the primary video stream may include only a sub-region of the overview video captured as the camera output. However, in other cases, the framed shot of the primary video stream may include the entire overview video. Generally, the framed representation of an individual may be derived from the camera output, such as the overview video output stream of a camera from a video conferencing or camera system.
[0154] In one example, the video processing unit may be configured to select the first video output stream or the second video output stream (e.g., from the first camera and the second camera respectively) as the source of the framed representation of a common conference participant. The framed representation may include a "close-up" shot of the common conference participant and may be output as the primary video stream. For example, referring to Figure 3A-3C , the video processing unit may select the output associated with camera 310c and stream direction 320c as the source of the framed representation (e.g., desired shot) of common conference participant 360. The output of camera 310c may be selected instead of the output of camera 310b, for example, based on any combination of the shot selection criteria described above. Using the selected camera output, the video processing unit may then proceed to generate the framed representation of the common conference participant as the primary video stream. In Figure 3C 's example, the framed representation (desired shot) may include a sub-region of the output of camera 310c that is primarily characterized by the head / face of conference participant 360. In other cases, conference participant 360 may be displayed in the primary video stream in combination with the representations of one or more other conference participants and / or one or more objects (e.g., whiteboard, microphone, display screen, etc.).
[0155] Figure 6 A graphical representation of the relationship between the overview video stream and the various sub-picture representations that may be output as one or more primary or focused video streams is provided. For example, Figure 6Represents an overview stream 610. Based on the overview video, various sub-picture videos can be generated, featuring one or more of the meeting participants. The first sub-picture representation includes two meeting participants and can be output by the multi-camera system as the main video stream 620. Alternatively, the second sub-picture representation includes only one meeting participant (i.e., participant 360) and can be output by the multi-camera system as the main video stream 630. As Figure 6 shown, the common meeting participant 360 can be represented in both the main video output stream 620 and the main video output stream 630. Whether the video processing unit generates the main video output stream 620 or the main video output stream 630 as the output of the camera system can depend on the lens selection criteria described above, the proximity of the meeting participant 360 to other meeting participants, the detected interaction between the meeting participant 360 and other participants, and other factors.
[0156] In some embodiments, and with reference to Figure 3B and Figure 3C as an example, the output stream 610 can correspond to the overview output stream obtained from camera 310c. The video processing unit can select the output stream 610 from camera 310c based on lens selection criteria such as face visibility score, head pose, whether the meeting participant 360 is detected as speaking, etc. For example, based on determining that the meeting participant is facing camera 310c (e.g., based on determining that the meeting participant 360 has a higher face visibility score relative to camera 310c than other cameras such as camera 310a or 310b), the output stream 610 can be selected instead of other camera outputs (such as the output from camera 310b or 310a). As a result, in this particular example, a sub-picture representation as shown in Figure 3C can be generated as the main output video 630.
[0157] In some embodiments, a camera among the multiple cameras can be designated as the preferred camera for a specific meeting participant. For example, the first or second camera associated with the selected first or second video output stream can be designated as the preferred camera associated with the common meeting participant. With reference to Figure 3A-3C , the camera 310c associated with the output stream 370b can be designated as the preferred camera associated with the common meeting participant 360. Thus, in some embodiments, when the common meeting participant 360 is determined to be speaking, actively listening, or moving, the output stream 370b can be used as the source of the main video stream. In some embodiments, the common meeting participant can be centered in the output associated with the preferred camera. In such a case, the preferred camera can be referred to as the "center" camera associated with a specific meeting participant.
[0158] As Figure 4As shown in step 440 of, method 400 may include generating a primary video stream as an output. For example, a video processing unit may be configured to generate a primary video stream as an output of a multi-camera system. In some embodiments, the generated primary video stream may include a framed representation of co-conference participants. Referring to Figure 6 , the primary video stream may include a framed representation 630 of co-conference participant 360, and the primary video stream may be sent to a display unit (e.g., host computer 140; display units 216, 236, 246, 256, 266) or displayed on the display unit.
[0159] In some embodiments, it may be determined that a co-conference participant is speaking, listening, or reacting. Such characteristics of a conference participant may be used to determine whether and when to feature the conference participant in the primary video output generated by the multi-camera system. It may be determined that a co-conference participant is speaking based on, for example, an audio track and / or tracked lip movement. In some embodiments, it may be determined that a co-conference participant is listening based on, for example, a head pose (e.g., tilted head), a viewing direction (e.g., viewing a speaking conference participant), a facial visibility score (e.g., a percentage associated with the direction of viewing a speaking conference participant), and / or based on determining that the conference participant is not speaking. Additionally, in some embodiments, it may be determined that a co-conference participant is reacting based on a detected facial expression associated with an emotion (such as, but not limited to, anger, disgust, fear, happiness, neutral, sadness, or surprise). A trained machine learning system (such as a neural network) implemented by the video processing unit may be used to identify the emotion or facial expression of a conference participant. As used herein, a neural network may refer to a series of algorithms that mimic the operation of an animal brain to identify relationships between large amounts of data. As an example, a neural network may be trained by providing a data set including multiple video recordings or captured image frames, where the data set includes images representing emotions of interest. For a particular image, the network may be penalized (e.g., as indicated by a predetermined annotation) for generating an output inconsistent with the emotion represented by the particular image. Additionally, each time the network generates an output that correctly identifies the emotion represented in an annotated image, the network may be rewarded. In this way, the network may "learn" by iteratively adjusting the weights associated with one or more models including the network. The performance of the trained model may increase with the number of training examples (especially difficult case examples) provided to the network during training. This method is generally referred to as "supervised learning".
[0160] As described above, multiple conference participants can be tracked and correlated across outputs generated by two or more cameras included in the described multi-camera system. In some embodiments, conference participants can be tracked and identified in each of a first, second, and third video output stream respectively received from a first, second, and third camera among a plurality of cameras included in the multi-camera system. In such an example, the video processing unit can be configured to analyze the third video output stream received from the third camera and, based on an evaluation of at least one identity indicator (as described above), determine whether a representation of a conference participant included in the third video stream corresponds to a common conference participant represented in the outputs of the first and second cameras. For example, referring to Figure 3A-3C , a third representation (not shown) of common conference participant 360 can be included in the third output stream from camera 310a. The video processing unit can analyze the third output stream and determine that the third representation corresponds to common conference participant 360. In this way, the described system can correlate and track a single conference participant across three or more camera outputs.
[0161] Using similar identification techniques, the described system can track multiple different conference participants across multiple camera outputs. For example, the described system can receive an output from a first camera and an output from a second camera, where both outputs include representations of a first conference participant and a second conference participant. Using the disclosed identification techniques, the video processing unit can correlate the first representation and the second representation with the first conference participant and the second conference participant. In the example, the first camera output and the second camera output can also include representations of one or more other conference participants (e.g., a third representation of a conference participant included in the first video output stream from the first camera and a fourth representation of a conference participant included in the second video output stream from the second camera). The video processing unit can also be configured to analyze the first video output stream and the second video output stream based on at least one identity indicator to determine whether the third representation of the conference participant and the fourth representation of the conference participant correspond to another common conference participant (e.g., a common conference participant different from both the first conference participant and the second conference participant).
[0162] Based on determining that the first camera output and the second camera output each include a representation of three common meeting participants (e.g., meaning that a representation of each of the three common meeting participants appears in both the output from the first camera and the output from the second camera), the video processing unit may select the first camera or the second camera as the source of the primary video stream characterized by any one of the first common meeting participant, the second common meeting participant, or the third common meeting participant. In other words, the video processing unit may be configured to evaluate a third representation and a fourth representation of another common meeting participant (e.g., a representation of the third common meeting participant included in the outputs of the first and second camera outputs) with respect to one or more predetermined shot selection criteria. Based on the shot selection evaluation, the video processing unit may select the first video output stream or the second output stream as the source of the framed representation of another common meeting participant (e.g., the third common meeting participant) for output as an alternative to the primary video stream. The video processing unit may be configured to generate an alternative primary video stream including the framed representation of the other / third common meeting participant as the output of the multi-camera system. The alternative primary video stream may be a video stream that is displayed in addition to or in place of the first primary video stream. As an example, referring to Figure 3A-3C , the output stream 370a may be selected as the source of the alternative primary video stream based on an evaluation of the representation (not shown) of the second meeting participant 332 displayed in the output streams 370a, 370b.
[0163] In Figure 3B and Figure 3C 's example, only a single meeting participant is shown in the primary video output. However, in some cases, multiple meeting participants may be shown together in a single primary video output. For example, under some conditions, the first common meeting participant and the second common meeting participant may be shown together in the primary video stream. Such conditions may include whether the first meeting participant and the second meeting participant are determined to be both speaking, actively participating in a back-and-forth conversation, looking at each other, etc. In other cases, whether to include both the first meeting participant and the second meeting participant together in the same primary video output stream may depend on other criteria, such as the physical distance between the two meeting participants, the number of intervening meeting participants located between the first meeting participant and the second meeting participant, etc. For example, if the number of intervening meeting participants between the first common meeting participant and the second common meeting participant is four or fewer, and / or if the distance between the first common meeting participant and the second common meeting participant is less than two meters, then the first meeting participant and the second meeting participant may be shown together in the same primary video output stream. For example, in cases where more than four meeting participants separate the first and second meeting participants and / or where the first meeting participant and the second meeting participant are separated by more than 2 meters, the first meeting participant and the second meeting participant may be characterized separately in the respective primary video output streams.
[0164] Fig. 7A and Figure 7B are examples that can generate the primary video stream that is the output of the described multi-camera system. As Fig. 7A shown, the first co-meeting participant 710, the second co-meeting participant 720, and the third co-meeting participant 730 can all be shown together in the primary video output stream 700. In this example, the first co-meeting participant 710 and the third co-meeting participant 730 can be shown together because the number of interleaved meeting participants between them is four or less (or less than another predetermined interleaved participant threshold) and / or because the distance between them is less than 2 meters (or less than another predetermined separation threshold distance). In Figure 7B the example of, the second co-meeting participant 720 and the third co-meeting participant 730 are shown together in the primary video stream 700, but do not include the first co-meeting participant 710. Such a framing determination can be based on determining that the first co-meeting participant 710 and the third co-meeting participant 730 are separated by more than the threshold distance d. It should be noted that other shot selection criteria can also be relied upon to feature participants 720 and 730 together, excluding participant 710. For example, participants 720 and 730 can be determined to be talking to each other, looking at each other, or otherwise interacting with each other, while participant 710 is determined not to be talking or otherwise interacting with participants 720 and 730.
[0165] The following sections describe examples of various shot selection scenarios and corresponding primary video outputs. In one example, the video processing unit can be configured to determine whether the first co-meeting participant or the second co-meeting participant (e.g., the co-meeting participant corresponding to the meeting participant featured in two or more camera output video streams) is talking. Additionally, the video processing unit can be configured to: based on determining whether the first co-meeting participant or the second co-meeting participant is talking, if the first co-meeting participant is determined to be talking, generate a first primary video stream as the output of the multi-camera system, and if the second co-meeting participant is determined to be talking, generate a second (alternate) primary video stream as the output of the multi-camera system. The first primary video stream can feature the first co-meeting participant, and the second / alternate primary video stream can feature the second co-meeting participant.
[0166] Additionally or alternatively, the video processing unit may also be configured to determine, based on the speaking participant, to generate the primary video stream as the output of the multi-camera system if the first co-conference participant is determined to not be speaking, and to generate the second / alternate primary video stream as the output of the multi-camera system if the second co-conference participant is determined to not be speaking. These options can be useful, for example, for providing a listener shot featuring the first co-conference participant or the second co-conference participant. Such video output can enable the display of conference participants who may be actively listening to the speaking conference participant or otherwise reacting to the speaking conference participant.
[0167] In some embodiments, determining whether the first co-conference participant or the second co-conference participant is speaking may be based on directional audio information received at the video processing unit from one or more directional microphones. As discussed with respect to Figure 1 the DOA / audio units 112, 122 may include a directional microphone array that can detect audio signals originating from different positions or directions relative to the cameras (e.g., the primary camera 100 and / or one or more peripheral cameras 120). The DOA / audio units 112, 122 may use the audio signals from different microphones and determine the angle and / or position from which the sound corresponding to the audio signal (e.g., speech) originates. The (one or more) video processing units may receive the directional audio information, e.g., a directional audio track or signal corresponding to different directions relative to the cameras, and may determine whether the featured conference participant is speaking based at least in part on the received directional audio information. For example, the video processing unit may determine that the direction or position associated with the directional audio signal or track representing speech corresponds to or is related to the determined direction or position of the featured conference participant, as determined from an image analysis of the video output from the cameras. In such a case, the video processing unit may determine that the featured conference participant is the speaker. More generally, the video processing unit may be configured to correlate the direction or position associated with one or more directional audio signals with the determined direction or position of one or more conference participants featured in the video output stream received from the cameras, and identify one or more conference participants as speakers based on that correlation. The video processing unit may also be configured to track the lip movements of the featured conference participant and identify the speaking conference participant based on the tracked lip movements and the directional audio information. For example, a conference participant may be identified as a speaker if the conference participant is identified as being associated with the tracked lip movements and the directional audio signal.
[0168] Additionally or alternatively, in some embodiments, determining whether the first co-meeting participant or the second co-meeting participant is speaking may be based on the output of a trained network, such as a neural network, that is configured to detect speech activity based on inputs including one or more captured images and / or one or more audio signals, such as directional audio signals. Speech activity may include the start position, decibel level, and / or pitch of an audio signal corresponding to speech or an audio track. In some embodiments, the speech processing unit may be configured to associate the speech or audio track with a particular meeting participant. The association may be based on, for example, the original position or orientation or the pitch of the audio signal. Additionally, in some embodiments, determining whether the first co-meeting participant or the second co-meeting participant is speaking may be based on lip movement detection across multiple captured image frames.
[0169] In some embodiments, the output of the multi-camera system may also include an overview video stream that includes representations of the first co-meeting participant and one or more other meeting participants. For example, referring to Figure 6 , the output of the multi-camera system may include an overview stream 610 that may include co-meeting participant 360 and one or more other meeting participants. As another example, referring to Figure 7A-7B , in an overview shot, meeting participant 720 may be shown with meeting participants 710 and 730, as Fig. 7A shown. Alternatively, meeting participant 720 may be shown with meeting participant 730 but not shown with the first meeting participant 710 in another type of overview shot. In some cases, the overview shot output from the camera system may be commensurate with the output of any one of the multiple cameras included in the camera system.
[0170] The outputs of a multi-camera system can be shown together on a single display. For example, in some embodiments, an overview video stream can be output from the camera system along with one or more primary video streams, and either the overview video stream or one of the one or more primary video streams can be shown in corresponding tiles on the display. In other embodiments, the outputs of the multi-camera system can include multiple primary video streams for inclusion in corresponding tiles to be shown on the display. In one example, the outputs of the multi-camera system can include an overview video stream captured by a first camera, a first primary video stream captured from a second camera and featuring a conference participant determined to be speaking, and a second primary video stream captured from a third camera and featuring a conference participant determined to be listening or otherwise reacting to the speaking conference participant. The overview video stream, the first primary video stream, and the second primary video stream can all be featured simultaneously in different tiles shown on the display. Any number of additional tiles can be featured with additional primary video streams. Additionally, the tile layout and / or the timing associated with the displayed tiles can be varied such that in some cases, the first primary video output stream is shown on the display together with the second primary video output stream, and in other cases, the first primary video output stream can be shown alternately on the display relative to the second primary video output stream.
[0171] As used herein, a tile can refer to a portion of a display unit (e.g., a square, rectangular, or other shaped area) in which a video output stream can be displayed. Multiple video streams can be displayed on a display unit (e.g., a tablet, mobile device, television), where each video stream is included in a corresponding tile, and the collection of tiles can form a gallery view. As used herein, a gallery view can refer to the simultaneous display of multiple video streams and / or system outputs on a display unit. As an example, a first primary video stream and a second primary video stream can be shown in corresponding tiles on the display. In another example, a first primary video stream, a second primary video stream, and an overview stream can be shown in corresponding tiles on the display. In some embodiments, a primary video stream can be generated for each conference participant in a conference environment, and the generated primary video streams can be shown in corresponding tiles on the display (e.g., each conference participant is shown in a tile on the display).
[0172] As noted, the timing for displaying various video outputs can vary. In some examples, the output of a multi-camera system can alternate between a first primary video stream and a second primary video stream. Alternatively, the output of a multi-camera system can include multiple video outputs (e.g., one or more overview video outputs and one or more primary video outputs), and the particular output selected for display on a display can vary (e.g., alternate with each other). It is contemplated that the output of the system (or video output selection) can alternate between a first primary video stream and a second primary video stream, between a first primary video stream and any one of a second, third, or other primary video streams, between a first primary video stream and an overview video stream, or between any other combination of video outputs.
[0173] The alternation between a primary video stream and an alternate primary video stream can include any suitable type of transition. In some examples, the transition between video stream outputs can include a hard cut transition or a smooth transition. A hard cut transition can include an immediate (or near immediate) transition from a first video stream to a second video stream. A hard cut transition can involve little or no smoothing between video signals. A smooth transition can include processing such as blending, fading, etc. from one video stream to another. In some cases, a smooth transition can involve a non-linear transition (e.g., where the associated picture change first occurs slowly and accelerates from the picture associated with the first video stream to the picture associated with the second video stream). In some embodiments, the alternation between a primary video stream and an alternate primary video stream can be based on whether a common meeting participant or another common meeting participant is determined to be speaking (e.g., speaker view) or listening (e.g., listener view). That is, the transition from one video stream to another can be based on determining that a first meeting participant has started speaking (which can be used to initiate a transition to the video output characteristics of the first meeting participant) and / or determining that a second meeting participant has stopped speaking.
[0174] The disclosed embodiments can also utilize "over-the-shoulder" shots. Such a shot can be displayed in a primary video stream that includes a representation of the face of a first meeting participant and at least a portion of the back of the head of a second meeting participant. Such a shot can be combined with other multi-participant shots such as Figure 7Bis contrasted with a shot over the shoulder), where the primary video stream can include representations of both the face of a first meeting participant and at least the face of a second meeting participant. Although a shot over the shoulder may not show the face of the second participant (or at least not a significant part of the face), this type of shot can convey a great deal of information to the viewer. For example, a shot over the shoulder can represent the interaction between two participants engaged in a direct conversation (and can precisely convey to the viewer who is participating in the conversation), can indicate the reaction of the listener (the second participant) to the speaker (the first participant), and / or can convey other information associated with the body language of the participants.
[0175] In other examples, the primary video stream can include a group shot in which multiple meeting participants are characterized together in a single frame. In some cases, two or more of the participants can face each other to form an above-the-shoulder arrangement, while one or more additional participants can face in a direction common to the other meeting participants. In some examples, the primary video stream can include a representation of the face of each of three different co-meeting participants (e.g., Fig. 7A in the primary video stream 700).
[0176] Figure 8 is a flowchart of another example method 800 for analyzing multiple video output streams and generating a primary video stream that is consistent with some embodiments of the present disclosure. In some embodiments, the video processing unit can be configured to execute method 800.
[0177] As Figure 8 shown in step 810 of
[0178] As Figure 8 shown in step 820 of
[0179] As Figure 8 shown in step 830 of, method 800 may include generating a primary video output stream based on the identified video output stream. For example, a video processing unit may be configured to generate a primary video stream based on the identified video output stream. In some embodiments, the primary video stream may include a view composition that includes a representation of at least the face of a first subject and at least a portion of the shoulders and back of a second subject. As used herein, a view composition may include a sub-picture of a video output stream, which may include more than one subject, meeting participant, or object. In some embodiments, at least a portion of the back of the head of the second subject may also be visible in the identified video output stream.
[0180] Fig. 9 is an example of a view composition 910 in an identified video output stream 900 that is consistent with some embodiments of the present disclosure. Composition 910 represents a shot on the shoulder as described above. For example, view composition 910 may include a representation of the face of a first subject 920 and a portion of the shoulders and back of a second subject 930. Additionally, as Fig. 9 shown, view composition 910 may optionally include a representation of a portion of the back of the head of the second subject 930. In some embodiments, view composition 910 may be included in the primary video stream.
[0181] In some embodiments, a view composition may be determined based on one or more of a head box, head pose, or shoulder position. As used herein, a head box may refer to a view or box that identifies the head of a subject. Referring to Fig. 9 , head boxes 940, 950 of subjects 920, 930 are shown. Additionally or alternatively, as Fig. 9 shown in, the shoulder position 960 of the second subject 930 may be used to determine the view composition 910 associated with a shot on the shoulder.
[0182] As noted in the above section, the identification of the video output stream to be featured on a display may be determined based on an evaluation of multiple output streams from multiple cameras. At least a portion of the multiple output streams (e.g., two or more of the output streams) may include a representation of a first meeting participant. The identified video stream may be selected based on one or more predetermined criteria (e.g., a shot selection criterion). The predetermined criteria may include, but are not limited to, the viewing direction of the first meeting participant as represented in the multiple output streams and / or a facial visibility score associated with the first meeting participant as represented in the multiple output streams. Referring to Fig. 9, it can be determined that the first subject 920 is watching the second subject 930. This can indicate a conversation or session occurring between the first subject 920 and the second subject 930. Additionally, the representation of the first subject 920 in the identified output stream 900 can include a 93% face visibility score, and the representation of the second subject 930 in the identified output stream 900 can include a 1% face visibility score. This can indicate that the first subject 920 and the second subject 930 are facing each other, and the output stream 900 can be used to capture the conversation or session between the first subject 920 and the second subject 930.
[0183] Fig.10 is a graphical representation of the camera 1000 including a video processing unit 1010 (e.g., virtual director unit). As Fig.10 shown, the video processing unit 1010 can process video data from the sensor 1020. Additionally, the video processing unit 1010 can split the video stream or video data into two streams. These streams can include an overview stream 1030 and an enhanced and scaled video stream (not shown). Using dedicated hardware and software, the camera 1000 can detect the positions of meeting participants using a wide-angle lens (not shown) and / or a high-resolution sensor such as the sensor 1020. Additionally, in some embodiments, the camera 1000 can determine the person speaking, detect facial expressions, and determine where the attention is focused based on the (one or more) head directions of the meeting participants. This information can be sent to the virtual director 1040, and the virtual director 1040 can determine the appropriate video setting selections for the (one or more) video streams.
[0184] Fig.11 is a graphical representation of the areas 1110, 1120, 1130 in the meeting environment 1100. As Fig.11 shown, in some embodiments, the camera 1140 can be located at the top or bottom of the monitor or screen 1150. In some embodiments, the camera 1140 can be located at the short end of a table (e.g., table 1160). The meeting environment 1100 can be divided into areas 1110, 1120, 1130 by the camera 1140. In some embodiments, the camera 1140 can use the areas 1110, 1120, 1130 to determine the positions of meeting participants.
[0185] In some embodiments, a video camera or a camera with a lens having a wide enough field of view to capture the entire space of a meeting environment can be provided. The camera can be equipped with machine learning techniques (e.g., learning models / algorithms, trained networks, etc.). The machine learning techniques can enable the camera to determine where a person is located within the camera's field of view, who is speaking, who is listening, and in what direction the heads of the people in the camera view are pointing. The camera can use algorithms suitable for including a flexible image pipeline to capture relevant views from the room.
[0186] For example, the output of the algorithm can identify a portion of the camera's field of view recommended to be shown in a video client. In response, the camera can change the video stream or content provided via a host stream. The desired view of the host video stream can be managed by a virtual director or any other software component.
[0187] Such operations can be used in single-camera or multi-camera systems. In a multi-camera system, for example, the cameras can communicate with each other via a defined application programming interface (API) provided on an internal network bus. In some embodiments of the system, the communication can include information about the state of the cameras, what the cameras are currently detecting, audio / video streams, potential framing from the virtual director, and camera settings.
[0188] In some embodiments, such as when the camera is placed in a smaller meeting environment, each camera can have a field of view of approximately 150 degrees. In other embodiments, such as when the camera is placed in a larger meeting environment, each camera can have a field of view of approximately 90 degrees. It is contemplated that a camera with a field of view of approximately 150 degrees can be used in a larger meeting environment, and a camera with a field of view of approximately 90 degrees can be used in a smaller meeting environment. It is also contemplated that any combination of cameras with different fields of view can be used in any meeting environment. Additionally, cameras with any field of view can be used and are not limited to the examples provided herein.
[0189] In some embodiments, each camera can include an optical lens with an appropriate field of view and a high-resolution image sensor, allowing the camera to zoom in without loss of perceived resolution. The camera can have the ability to process video data from its sensor and, in some embodiments, split the video data into two or more streams. These streams can include one stream that is scaled down (e.g., an overview stream) and at least one other stream that provides an enhanced and zoomed video stream (e.g., a main stream).
[0190] Fig.12 is a graphical representation of the overview stream 1210 and the main stream 1220. As Fig.12As shown, each intelligent camera in the system can internally include two video streams: a high-resolution stream (e.g., the main stream 1220), where video settings can be applied to scale and change other video stream attributes; and a secondary stream (e.g., the overview stream 1210), which captures the entire scene and is consumed by the vision pipeline.
[0191] Fig.13 An example of a rule-based method for lens determination is shown. As Fig.13 shown, the eyes 1310, 1320 of one or more persons can be aligned in the top third of the images 1300a, 1300b, 1300c.
[0192] Embodiments of the present disclosure can include a multi-camera system, and the multi-camera system can include any suitable number of cameras. In some embodiments, a single camera can be provided. The single camera can be configured to generate an overview video stream representing a region of the environment. Based on the analysis of the overview stream, one or more main video streams can be generated relative to the overview stream. For example, individual participants, objects, etc. can be detected in the overview stream, and based on those detections, one or more main video streams can be generated, each main video stream featuring at least one of the detected participants, objects, etc. The main video streams can each represent a subset of the overview video stream. Additionally or alternatively, the main video streams can have different camera characteristics relative to the overview video stream. For example, each main video stream can have a pan value, a tilt value, and / or a zoom value that is different from the pan, tilt, and / or zoom values associated with the overview video stream. The final video display layout can include any combination of the main video stream(s), optionally together with the overview video stream, with each main video stream featuring a separate tile in the layout. The main / overview video streams shown in each tile, the number of tiles in the layout, the size of the tiles in the layout, and the orientation of the tiles in the layout can be controlled based on the analysis of the overview video stream (e.g., using any of the techniques described above).
[0193] The disclosed embodiments can include multiple cameras. Each of the cameras can generate a corresponding overview video. Similar to the techniques described above for a single-camera system, in some embodiments, the main video streams can be generated as subsets of the overview video stream supplied by any one of the multiple cameras. One or more of these main video streams can optionally be displayed on a display together with the overview video stream of one or more of the multiple cameras. For example, the main video stream(s) and optionally the overview video stream(s) can be shown in the corresponding tiles of a video tile layout shown on the display. The main / overview video streams shown in each tile, the number of tiles in the layout, the size of the tiles in the layout, the orientation of the tiles in the layout, etc. can be controlled based on the analysis of the overview video stream (e.g., using any of the techniques described above).
[0194] In addition, since each camera can be associated with a different field of view, viewing angle, etc., a multi-camera system can provide more options for a main video stream characterized by a particular subject. For example, in a single-camera embodiment, in some cases, a subject can be represented in an overview video stream based on a side (e.g., side profile). In turn, the main video stream derived from the overview video stream and characterized by the subject (in the single-camera case) can also represent the subject based on a side or front profile. However, in a multi-camera embodiment, there may be a possibility that an object is captured by more than one camera and is thus represented in the overview streams of more than one camera. Based on an analysis of one or more of the overview video streams, the system can determine that the subject should be featured in the main video stream displayed in a tile of the video tile layout. In the multi-camera case, there can be multiple options for generating a main video stream characterized by a particular subject, rather than generating the main video stream based on a single overview video stream (as in the single-camera case). In some cases, the main video stream can be derived from an overview video stream in which the subject is represented as facing the camera. In other cases, the main video stream can be derived from an overview video stream in which the subject is represented as not facing the camera but turned to one side or the other. In other embodiments, the main video stream can be derived from an overview video stream in which the subject is represented as facing away from the camera. Non-limiting examples of criteria for selecting an originating overview video stream can include whether the subject is presenting to the audience, whether the subject is interacting with an active meeting participant, and whether the main video stream should exclude or include other participants.
[0195] It is worth noting that in a multi-camera system, there can also be the ability to actively control one or more of the cameras to obtain source video streams that are designed to provide a desired main video stream. For example, based on the overview video streams of one or more of the cameras, a particular subject can be identified to be featured in at least one main video stream. However, rather than deriving the main video stream representing the subject from the overview video stream, one or more of the multiple cameras included in the multi-camera system can be actively controlled to capture a desired shot of the subject. This can include using a camera facing in front of the subject to zoom in on the subject, or panning or tilting the camera towards the subject. In this way, the cameras included in the multi-camera system can operate in an overview video mode, in a main video mode, or operate in an overview video mode during some time periods and in a main video mode during other time periods.
[0196] In addition, the disclosed systems and methods can provide several types of video outputs. In some embodiments, the system can provide a multi-stream tile video layout for display on a monitor. Additionally or alternatively, the system can provide multiple video streams (e.g., one or more overview video streams, one or more primary video streams, layout streams, etc.) as outputs. In such cases, another system (e.g., a server, a web-based system, a cloud-based system, MS Teams, Zoom, Google Meet, WebEx, etc.) can receive the video stream outputs from the disclosed embodiments, such as speaker, presenter, overview, person, group, gesture streams, and / or layouts (such as gallery view) and adaptive layout streams, and display some or all of the video streams on a monitor according to, for example, lens selection criteria specific to the system platform. Additionally, in collaboration with other systems (e.g., Microsoft, Google), it can be further ensured that the layouts or streams output to the system will be properly displayed without being cropped and with sufficient display space.
[0197] The virtual director unit can employ machine learning (ML) vision / audio and information about previous events to determine which image or part of an image (from one or more cameras) should be placed in the composite video stream—whether it includes a composite video (tile layout) or a multi-stream video output. In some embodiments, the virtual director unit determines the layout format based on earlier detected events. Some potential benefits of the system can include using the display real estate in the video stream to better show the participants and bring the remote participants closer to the meeting (e.g., blurring the boundary between the meeting room and the remote participants).
[0198] The disclosed embodiments can operate with respect to a variety of different environments and settings. Such environments and settings can include, for example, classrooms, boardrooms, meeting rooms, home offices, or any other environment where fixed or mobile cameras can be used to capture images of individuals or objects.
[0199] Referring Fig.10 , camera 1000 can be equipped with a directional microphone array 1050, which can capture audio signals from different positions and / or directions relative to the camera. By using the signals from different microphones within the microphone array 1050, camera 1000 can determine the angle or direction from which the voice is coming (e.g., audio direction / direction of arrival or DOA). Additionally, in some embodiments, based on the input from the microphone array 1050, camera 100 can be equipped with a system for differentiating between the case where someone is speaking and other cases where there is silence. This information can be transmitted to the convolutional neural network (CNN) pipeline 1010 and the virtual director 1040.
[0200] Camera 1000 may include one or more hardware-accelerated programmable convolutional neural networks, such as CNN pipeline 1010. In some embodiments, CNN pipeline 1010 may implement one or more machine learning models that may use an overview stream as input to allow the associated hardware to provide information about people (such as meeting participants) within the view of the camera. CNN pipeline 1010 may view overview stream 1030 and detect where people are detected within the view of camera 1000. It may also provide information about the people within the view of camera 1000, such as, but not limited to, whether they are speaking, facial expressions, how many people are visible, and head pose. It may also track each person over time to determine where each person was previously within the view and whether they are in motion. Advantages of using CNN pipeline 1010 and machine learning to detect people in overview stream 1030 may include: Machine learning models running on the CNN can be trained to not be biased by parameters such as, but not limited to, gender, age, and race. In some embodiments, CNN pipeline 1010 may also be able to understand people and partial views of people viewed from different angles (e.g., from behind). This may create a robust video conferencing experience. The CNN hardware may run these detections multiple times within a second, allowing camera 1000 to react to changes within the view of camera 1000 at the appropriate times.
[0201] In some embodiments of the system, audio signals from microphone array 1050 may be aggregated and processed by CNN pipeline 1010, and CNN pipeline 1010 may distinguish speech during a meeting in a conferencing environment. Based on this distinction, CNN pipeline 1010 may combine speech feature classification with other information such as angle, a person's location within the room, and / or other relevant detections. Speech that does not belong to any meeting participant may be classified as an artificial sound source (such as a speaker).
[0202] Virtual director 1040 may be a software component that obtains input from CNN pipeline 1010 and determines which region of the camera view should be shown to a host (e.g., host computer 140). In some embodiments, virtual director 1040 may create an automated experience similar to a television (TV) talk show production. Virtual director 1040 may use rules similar to those of TV production and apply those rules to an interactive video call and / or select framing options to be applied to the camera stream that will be relayed to the host stream.
[0203] The virtual director 1040 can perform its functions by evaluating possible viewing angles in the room and by monitoring ongoing events and event history associated with the room. For example, for each participant, the virtual director 1040 can evaluate different crops of an image of a person (e.g., a frame) to find a preferred frame for a particular situation. Attributes that can be evaluated can include: whether a person is speaking, the duration that a person has spoken or is speaking, where a person is looking, how many people are visible in the frame, the reaction and body language that a person is showing, and / or what other people are visible in the frame. People visible in the frame can be placed in a viewing composition that is natural with respect to the direction of attention, the golden rules of viewer comfort, and nearby people to avoid unappealing, unnatural, or chaotic frames. Based on this evaluation, a frame can be selected when something changes in the meeting environment or based on previous events in the meeting environment. Examples of such changes include but are not limited to: a second person starts speaking, someone moves, someone stands up, someone changes the direction they are looking (e.g., changes their viewing direction), someone has a reaction, someone shows an object, and / or someone has been speaking for a long time, as well as other detected events. In some embodiments, a person who has been speaking for a long time may indicate a lack of reaction.
[0204] The virtual director 1040 can determine the video settings required to change the main stream of a camera to the selected frame and then apply the video settings to the main stream of the selected camera. In some embodiments, additional input can be provided to the virtual director unit by intelligent sensors, such as the placement and movement of meeting participants in the meeting environment, the number of meeting participants in the meeting environment, the physiological attributes of meeting participants in the meeting environment, and other physical attributes of the meeting environment.
[0205] A frame to which the video settings are applied can be selected from the full wide-angle field-of-view images captured by each camera. This can be based on principles from TV production. In some embodiments, the system can operate with different lens types, such as but not limited to close-up shots (e.g., framing a person's head and shoulders), medium shots (e.g., framing one or more people, showing their upper bodies), and overall shots (e.g., showing one or more people fully or showing a table as a whole). Each frame can be positioned based on photographic principles, where people can be placed according to the golden rules, which include leaving one-third of the space from the center of their head to the top of the image, leaving space horizontally in the direction they are looking, and leaving space around the subject in the frame. It is envisioned that parameters can be adjusted in the system.
[0206] In some embodiments, the virtual director 1040 may start by displaying a fully minimized view to create an understanding of the context, the visual relationship between the meeting environment (e.g., the room) and the meeting participants (e.g., people) in the room. After a predetermined time, the virtual director 1040 may switch to the area with the best view (e.g., areas 1110, 1120, 1130), where a group of meeting participants in the room can be seen. In some embodiments, the area with the best view may be on the left and / or right side of the table (e.g., table 1160). The virtual director 1040 may continue to frame the person who is speaking. If the speaker speaks for longer than a predetermined time, the virtual director 1040 may switch to framing the other people in the room who are listening. If no one is speaking or the virtual director 1040 determines that the voice is from an artificial sound source (e.g., a speaker), the virtual director 1040 may switch to framing everyone in the room.
[0207] Other embodiments of the system may use more information detected by the camera to construct other activities. For example, when a meeting participant raises an object or otherwise interacts with an object, the camera may switch to framing both the person and the object. The object may be, for example, a whiteboard (such as an interactive or smart whiteboard). The interaction between the meeting participant and the object may include the participant touching the object, holding the object, writing or drawing on the object, making a gesture relative to the object (e.g., pointing at the object), looking at the object, or any other form of interaction with the object. In this case, the video processing unit may generate and output a focused video stream to be displayed on a display representing both the meeting participant and the object. The focused video stream may include at least a partial representation of the meeting participant (e.g., a head and upper body shot) or be characterized by at least a partial representation of the meeting participant (e.g., a head and upper body shot) and a complete representation of the object, such as a representation of the entire visible whiteboard. In other words, the focused video stream may be framed to include a cropped or uncropped representation of the meeting participant and an uncropped representation of the object, such that no part of the object that is visible to the camera or within the camera's view is excluded from the focused stream. In some embodiments, the video processing unit may identify a part of the object with which the meeting participant is interacting. For example, the video processing unit may identify the part of the whiteboard on which the meeting participant is drawing or writing. In this case, that part of the object is considered the "object" as described above. Thus, when the video processing unit determines that the meeting participant is only drawing or writing on a part of the whiteboard, the focused video output stream may be characterized by a representation of the meeting participant and the part of the whiteboard with which the meeting participant is interacting, optionally excluding those parts of the whiteboard with which the meeting participant is not interacting. When more than one person is looking in the same direction, the system may use the framing principles described above to change to frame the people who are watching and / or switch to the person they are watching.
[0208] The virtual director 1040 can switch between the images at a rhythm to maintain the engagement of remote participants (e.g., people watching the video stream from the host computer 140), where the active speaker and content will receive more time than the listeners. Live TV production can switch quickly, with the images lasting less than a second, but for an interactive video call, each image may need to last longer than the common practice in TV production to allow remote participants the opportunity to speak.
[0209] Fig.14 is a graphical representation of the components and connections of an example multi-camera system 1400 in a video conferencing environment. As Fig.14 shown, the multi-camera system can include multiple cameras 1410, 1420, 1430 connected to an Ethernet Power over Ethernet (PoE) switch 1440. The PoE switch 1440 can be connected to a PoE 1450, and the PoE 1450 can be connected to a host machine 1460. The host machine 1460 can be connected to a display (such as a TV screen 1470). The connections between the components of the multi-camera system 1400 are not limited to Fig.14 what is shown, and can include, for example, wired connections, wireless connections, or any combination thereof.
[0210] In some embodiments, multiple intelligent cameras can be provided, which can use artificial intelligence (AI) to understand the dynamics of the meeting room participants and can work together to provide an engaging experience to remote participants based on knowledge of how many people are in the room, who is speaking, who is listening, and where the attendees are focusing their attention. This can make it easier to pick up social cues and increase engagement.
[0211] In some embodiments, one or more intelligent sensors can be connected together and sense what is happening in the room. The intelligent sensors (such as wide-field cameras) can see the entire room and have dedicated hardware to run a vision pipeline that can detect what is happening in the room. This information can be provided to a software component that evaluates the information provided from the room and makes decisions about the best camera angles and images to display from the room. The virtual director unit can control the main video stream consumed by the connected host computer, and this information can also be obtained through an application programming interface (API) available for adjusting the system.
[0212] Referring to Fig.14, the cameras 1410, 1420, 1430 in the multi-camera system 1440 can be connected together via Ethernet through a PoE switch 1440, and the cameras 1410, 1420, 1430 can be able to communicate with each other. In some embodiments, the system 1400 can be connected to a host machine 1460 (e.g., a computer, a mobile device), and the host machine 1460 consumes the video stream (e.g., the host stream) created by the camera system 1400. The video stream can be consumed by a video application on the host machine 1460.
[0213] The host machine 1460 can consume multiple streams, and each of the multiple streams can frame a specific conference participant among the conference participants in the conference environment. If a content camera is connected, the content can also be displayed in a dedicated stream when the content is determined to be relevant. This can allow the video client to mix the video streams into its layout, thus allowing (e.g.) one square (or tile) per person, which shows the most interesting shot and camera angle of that person.
[0214] In some embodiments, the main camera 1420 can be directly connected to a network attached to the Internet. The main camera 1420 can detect that it can access the Internet and establish a connection with a cloud service, and the cloud service can relay the video stream to a selected video content analysis (VCA). This can be through a cloud-to-cloud connection, where the camera talks to the cloud, and the cloud can relay the video service to a selected video conferencing provider. In other embodiments, the cloud connection can relay the video stream to a local application running on the connected computer, and the local application can present the relayed video as a camera.
[0215] The cameras can communicate with each other via messages through a defined API that can be sent on the internal network bus. These messages can include information about the status of the camera, whether the camera is connected, the type of software the camera is running, the current health status, etc. It can also convey what it detects in the image, such as the location of a person or an object detected in the image, the location where they are placed in the room, and other information detected by the vision pipeline. Additionally, the messages can convey the video settings (such as image properties, color / brightness, and / or white balance) that have been applied to the main stream. In some embodiments, the messages can transmit virtual director unit parameters, which can allow the system to automatically adjust the virtual director unit experience and / or allow the users of the system to control / personalize the virtual director unit experience based on their preferences.
[0216] In a smaller meeting room, the camera can typically have a field of view of approximately 150 degrees, and in a larger meeting room, a field of view of approximately 90 degrees. Additionally, the camera can include an optical lens with an appropriate field of view and a high-resolution image sensor, allowing the camera to zoom in on the video without loss of perceived resolution. The camera can have the ability to process video data from the sensor, including splitting the video data into two streams: an overview stream and a main stream.
[0217] In some embodiments, one or more of the cameras can have a smaller field of view but the ability to zoom the main stream up to 10 times (10X). These cameras can be able to frame presenters or participants located further away from the camera (e.g., a presenter in a classroom where the camera is placed away from the board / presenter).
[0218] Furthermore, in some embodiments, the system can have a special camera to capture content or a whiteboard / wall. The camera can have adapted optics, image processing, and mounting to capture the content on the whiteboard / wall as optimally as possible. The content / object camera can be smaller and easy to handle, and can be mounted in the ceiling or held by a person to capture the content / objects that the participants in the meeting are working on, holding, or presenting.
[0219] A computing device can be attached to the system. The computing device can be connected to an Ethernet switch and provide additional services to the system. For example, the computing device can have one or more high-performance central processing units (CPUs) that can run parts of the vision pipeline of the cameras connected to the system. This can enable the vision pipeline to run faster and perform additional tasks.
[0220] In some embodiments, smart sensors can be connected to the system (e.g., wirelessly via Wi-Fi or directly to a wired network). The smart sensors can provide additional inputs to the virtual director unit for decision-making. For example, the smart sensors can include a smartphone that provides data inputs from various sensors, such as movement, location, audio signals, etc. Non-limiting examples of additional inputs can include other types of room occupancy information, such as who booked the room for how long, etc. All these inputs can be provided to the virtual director unit via an API.
[0221] In some embodiments, one of the cameras in the system can be selected as the main camera and can be responsible for controlling the host stream. The main camera can consume the main video stream of each camera and relay the selected main stream to the main stream based on inputs from the virtual director unit.
[0222] The camera can be equipped with a microphone array that receives audio from different locations on top of the camera. By using the signals from different microphones, the camera can determine from what angle the voice is coming (e.g., the direction of the audio or DOA). Additionally, in some embodiments, the camera or multi-camera system can be able to distinguish when someone is speaking, the difference between person A / B, and when there is silence based on the input from the microphones. This information can be transmitted to the vision pipeline and the virtual director unit.
[0223] The disclosed embodiments can include a vision pipeline. For example, each camera can include one or more hardware-accelerated programmable convolutional neural networks with pre-trained weights that are capable of detecting different attributes from video and / or audio (e.g., vision pipeline models). In some embodiments, the vision pipeline model can analyze the overview stream and detect the position of people in the view of the camera. The (one or more) vision pipeline models can also provide information about the people in the camera view, such as but not limited to, whether they are speaking, facial expressions, how many people are visible, the position of the people, and head pose. The (one or more) vision pipeline models can also track each person over time to determine where each person was previously in the view, whether they are moving, and in what direction they are facing.
[0224] One advantage of using a convolutional neural network and machine learning to detect people in the overview stream can include a vision pipeline model that can be trained to be unbiased towards parameters such as gender, age, and race, scene, light, and size. This can allow for a robust experience. The vision pipeline hardware can be a CPU or a dedicated chip with a hardware accelerator for the different mathematical operations used in the convolutional neural network architecture in the (one or more) vision pipeline models. This can allow the (one or more) vision pipeline models to run these detections multiple times within a second, thus allowing the camera to react to changes in the camera view at the appropriate time.
[0225] Embodiments of the present disclosure may include visual pipeline model training. As described above, the visual pipeline model may operate by obtaining an overview image and / or an audio signal from a camera and running the overview image and / or the audio signal through a pre-trained convolutional neural network (CNN). The visual pipeline model may be trained by running thousands of images and videos related to the scene and task objectives. During the training of the visual pipeline model, the model may be evaluated with a loss function that measures the extent to which the model can perform the task. Feedback from the loss function may be used to adjust the parameters and weights of the visual pipeline model until it can perform its task satisfactorily. These methods and other methods may include well-known machine learning tools and best practices that may be applied to training convolutional neural networks. The trained visual pipeline model may then be converted from a selected training tool (such as tensorflow) and optimized for a chipset of the visual pipeline hardware (HW) using a chipset manufacturing conversion tool to utilize the hardware (HW) acceleration block. The trained visual pipeline may be bundled with the camera software. In some embodiments, the bundled visual pipeline model may be fine-tuned for a specific use case. For example, if the system is to be used in a classroom, the visual pipeline model may be fine-tuned based on a training set having images, audio, and video representing the classroom scene.
[0226] For some of the visual pipeline models, it may be necessary to adjust the convolutional neural network architecture to best fit the hardware chipset of the visual pipeline. This may be performed by removing or replacing mathematical operations in the visual pipeline convolutional neural network architecture with equivalent mathematical operations supported by the chipset.
[0227] Consistent with embodiments of the present disclosure, a camera system including at least one camera and a plurality of audio sources is provided. As used herein, an audio source may refer to a device or mechanism for generating an audio output signal representing sensed sound. For example, an audio source may capture sound or an audio signal, which may be stored on a recording medium for preservation and reproduction. An audio source may alternatively be referred to as an audio capture device, a sound sensor, or an audio sensor. The audio source may include one or more microphones configured to convert sensed sound into an audio signal output. The audio sources may be distributed in an environment such as a conference environment to capture sound (in the form of an audio signal) from various locations within the environment. For example, the audio sources may be placed at specific locations within the conference environment such that at least one audio source can capture sound originating from any location within the conference environment. As another example, the audio sources may be placed at specific locations within the conference environment corresponding to the locations of the participants. In some embodiments, at least one of the audio sources may continuously capture sound or an audio signal. In other embodiments, at least one of the audio sources may be configured to capture and record sound or an audio signal in response to a user input, a detected motion, or the proximity of a detected sound (e.g., a sound above a decibel threshold).
[0228] A video processing unit (or video director unit) may be configured to analyze video from at least one camera of the camera system and aggregate audio signals from the audio sources based on one or more detected features of conference participants represented in the video. The detected features may be associated with or indicative of speech. The video processing unit may aggregate the audio signals to generate an aggregated audio output corresponding to the conference participants, such as representing the speech of the conference participants. For example, the video processing unit may analyze video from at least one camera of the camera system and, based on detected mouth movements performed by the conference participants, aggregate the audio signals by (e.g., including in the aggregated audio output) characterizing with audio signals determined to be associated with the speech of the conference participants and at least partially filtering (e.g., at least partially excluding from the aggregated audio signals) audio signals not associated with the speech of the conference participants. Thus, the video processing unit may aggregate and generate audio based on audio signals associated with the speech of the conference participants, ignoring audio signals that may be associated with environmental noise (e.g., white noise, background discussion / speech, music, sound of object movement). In some embodiments, the conference participants may be presenters.
[0229] As another example, a video processing unit may detect mouth or lip movements performed by a first meeting participant and a second meeting participant in a video. The video processing unit may identify the first meeting participant as a presenter and the second meeting participant as a non-presenter. In aggregating the audio signals, the video processing unit may accordingly characterize the audio signals associated with the presenter's speech and at least partially filter, ignore, or exclude the audio signals associated with the non-presenter's speech. In some embodiments, determining whether a meeting participant is a presenter or a non-presenter may be based on one or more of the following: user settings, the location of the meeting participant in the meeting environment (or relative to an object or area of interest), and the duration that the meeting participant spends speaking during the entire meeting (in some embodiments, relative to the duration that other meeting participants spend speaking).
[0230] Non-limiting examples of detected features include gestures, head direction, viewing direction, posture (e.g., whether the meeting participant is standing or sitting), or mouth movements consistent with any recognizable characteristic of the subject determined to be associated with the corresponding audio signal or with the encountered audio. For example, in some cases, one or more detected features may include the mouth movements of the subject, particularly in those cases where the mouth movements are synchronized with or otherwise consistent with the received audio signal. The audio signal may be characterized or included in the aggregated audio signal output for the subject to be synchronized with or otherwise consistent with the detected mouth movements of the subject, and those audio signals may be at least partially filtered or excluded. As another example, the system may detect whether the mouth movements of the subject are consistent with a speech audio stream representing speech. Other detected features may also be relied upon to aggregate the audio signals. Such features may include the gestures of the subject, audio signatures known to be associated with certain individuals, and the like.
[0231] In some embodiments, the disclosed system may include a video processing unit implemented to include one or more microprocessors on-board the camera and / or one or more microprocessors located remotely relative to the camera (e.g., in a server system, a cloud-based system, etc.). The video processing unit may be configured to analyze video from at least one camera and aggregate audio signals from multiple audio sources (e.g., microphones included on one or more cameras, microphones distributed around the environment, etc.). Aggregation of the audio signals may include selecting certain audio signals to be characterized and / or filtering one or more audio signals (e.g., background noise, non-presenter speech, side conversation participant speech, speech reproduced via a speaker, etc.). Aggregation of the audio signals may be based on one or more detected features of at least one subject represented in the video captured by the camera.
[0232] In some cases, selecting an audio signal from a corresponding audio source for inclusion in an aggregated audio output for a subject can depend on the proximity of the subject to the audio source. For example, in a conference room setting where each participant can be paired with or associated with the nearest audio source, when a participant's speech signal is received from the audio source nearest to that participant, the participant's speech signal can be featured in the aggregated audio. However, in other cases, an audio signal from an audio source that is farther away from a particular participant can be featured in the aggregated audio. For example, in a case where a first participant is speaking in the direction of a second participant, the speech signal of the first participant can be selected for the aggregated audio, even if it is primarily received from the audio source that is closest to the second participant but not the nearest to the first participant. In some embodiments, a machine learning vision / audio pipeline can detect people, object speech, movement, pose, or canvas enhancement, document detection, and depth.
[0233] In some embodiments, selecting an audio signal from a corresponding audio source for inclusion in an aggregated audio output for a subject can be based on the direction of the audio or the direction of arrival associated with the audio signal. A video processing unit can receive directional audio information (e.g., directional audio signals corresponding to different directions relative to the camera (e.g., from a directional microphone array)), and can determine which audio signals to include or exclude from the aggregated audio output of conference participants based on the determined direction of the conference participants relative to the camera (e.g., as determined from an analysis of the video output from the camera) and the audio / direction of arrival (DOA) associated with each audio signal. For example, the video processing unit can include audio signals having a DOA corresponding to or related to the determined direction or location of the conference participants, and can at least partially exclude audio signals having a DOA not corresponding to or related to the determined direction or location of the conference participants.
[0234] A camera system according to the present invention consists of one or more cameras and one or more microphones, the cameras having an overview video stream that views the entire FOV from the cameras. The microphones can be part of the cameras, but can also be separate.
[0235] In some embodiments, the video processing unit may use audio signals to distinguish the voices of various meeting participants in a meeting environment. As described above, the audio signals from the microphone may be aggregated and processed by the vision pipeline model. The (one or more) vision pipeline models may be able to distinguish the voices during the meeting and whether they are raised or lowered depending on what is happening in the room. In some embodiments, the vision pipeline model may be able to classify the topic of the conversation. Based on this, the vision pipeline model may combine the voice feature classification with other information such as the angle, the position of the person in the room, and other relevant detections. Voices that do not belong to a person may be classified as artificial sound sources such as speakers. This information may be provided to the virtual director unit, and the virtual director unit may use this information to select the best shot from the room.
[0236] Embodiments of the present disclosure may include features and techniques for identifying and including auxiliary items in a stream. In some embodiments, the vision pipeline model may determine objects of interest in the room. This determination may be based on the input of where the participants are looking and the items being held or pointed at. The input may be generated by using a vision pipeline model that may determine poses such as time graphs of pointing, head poses, object classification, and where people are looking. By knowing where the head is from different angles and by having the head position, depth may be available. From there, a two-dimensional (2D) overview space may be created to project and find the space where the angles cross (e.g., corresponding to where the person / participant is looking).
[0237] In some embodiments, the vision pipeline model may be able to determine that someone is painting on an auxiliary item, such as but not limited to a non-digital whiteboard. This determination may be based on the input that someone is standing in front of the non-digital whiteboard. The input may be generated by a vision pipeline model that may determine the movement patterns and gestures of the person in front of the non-digital whiteboard.
[0238] Embodiments of the present disclosure may implement principles from TV production. For example, the frame to which the video settings are applied may be selected from the full-wide-angle field-of-view images captured by each camera. This may be based on the principles from TV production.
[0239] The disclosed embodiments may include features and techniques for providing an adaptive layout and (one or more) gallery views. The disclosed embodiments may include AI-driven features that may create a more engaging and democratic video conferencing experience. During operation, one or more cameras of the system may dynamically adjust the projected view based on what they see and / or hear in the room or other environment.
[0240] In a static environment such as video, it may be difficult to interpret non-verbal communications such as gestures, body language, and facial expressions. Embodiments of the present disclosure can automatically detect and capture these non-verbal details, and while focusing on the person speaking, these systems can also draw attention to reactions and events in the room. This can provide remote participants with information that is naturally ascertained as being an actual meeting participant (e.g., a meeting participant present in a physical meeting environment) but may be more difficult to receive through traditional video-based solutions. The disclosed systems can employ the principles of live TV production (e.g., different types of camera shots, etc.), which can be used to make the video experience more engaging and inclusive for all meeting participants.
[0241] The following sections describe various features, capabilities, and configurations of the disclosed video system, including but not limited to: genius framing, speaker framing, gallery view, adaptive layout engine, framing transitions, and platform configurations. Although these features and capabilities are referred to by these terms for convenience and practicality, it should be understood that the functions and capabilities of these features are important and that these features and capabilities can be described independently of these terms.
[0242] Genius framing can involve a framing method in which detected objects in a video stream can be featured (e.g., by actually or effectively zooming in, panning, tilting, etc. to provide a desired shot of the subject of interest). Genius framing can refer to features that can particularly generate smooth zooms and scene transitions to capture meeting participants or other objects of interest in a room or environment (e.g., a whiteboard, a table, etc.). Machine learning can enable the detection of the number of people present and where the people are located within the room / environment. Genius framing can smoothly zoom in on and frame specific meeting participants. If people leave the room, or more people enter, the camera can zoom in or out to capture the new group of people.
[0243] The disclosed system can also respond to various types of meeting participants. For example, the system can feature active participants in one or more video frames. Active participants or participating listeners can include any individual or speaker who participates in at least one detectable action or displays at least one characteristic indicating participation in the meeting (e.g., detectable through video or audio analysis). Such a characteristic can be referred to as a "participating listener", "participating participant", "active listener", or "active participant" characteristic. Such actions can include, for example, speaking, moving one or more parts of the body (e.g., mouth movement, hand raising, nodding or shaking the head, change in facial expression), exhaling, generating non-verbal audible sounds, and / or moving into or out of the environment. As described above, the system can also feature inactive participants in one or more video frames. Inactive participants can include any individual who is not currently participating in a detectable (or detected) action. Such inactive participants can, for example, be sitting or standing quietly in the environment without participating in detectable (or detected) movement, speaking, or sound generation.
[0244] Speaker framing can involve techniques that can feature a detected speaker in a video stream. For example, when a person is detected speaking, that person can be featured in the video frame output for presentation as a framed speaker shot on a display. In some embodiments, speaker framing can involve using artificial intelligence to detect speakers in the environment (or meeting environment). To provide an overview for remote participants and track reactions in the environment, speaker framing can also use overview, group, and listener shots. The overview shot can show all the participants in the room, while the listener shot can show the non-verbal reactions of at least one participant who is not currently speaking in the room. A video processing unit (e.g., a virtual director unit) can determine the best possible shot or sequence corresponding to the meeting environment and / or meeting activities. The video processing unit can use artificial intelligence to perform the determination.
[0245] In some embodiments, speaker allocator filtering logic can be employed to identify speakers in a meeting. The factors considered can include determining the meeting participant who has spoken the most in past iterations (e.g., since the start of the meeting). The filter can consider iterations (or durations) with silence and iterations with speakers different from the candidate (potential) speaker. The filter can manage periodic meeting situations where different participants can participate for short periods while the current speaker is speaking. In some embodiments, the filter can consider the micro pauses or durations when people stop speaking. People often take short breaks while speaking. The filter can identify these scenarios and determine who the current speaker is based on these scenarios. Speaker filtering can be used in the layout engine (discussed further below), and the algorithm can determine where the speaker is located in the meeting environment.
[0246] The gallery view can relate to the ability of the disclosed systems and methods to generate multiple video streams for display together (e.g., in a tile layout) on a display.
[0247] The adaptive layout engine can relate to a software-based system component that controls aspects of the gallery view tile layout based on, for example, the detection of various conditions and / or events associated with a meeting environment.
[0248] The view transition can relate to the ability of the disclosed systems and methods to employ various view transition techniques from one shot to the next. The platform configuration can relate to the disclosed systems and methods implemented as a single camera system, a multi-camera system, a fully integrated live video solution, a distributed or cloud-based system, and / or a system that collaborates with and / or generates video output for various video conferencing platforms.
[0249] Video shot generation and selection can relate to different shot types that can be used to make the video experience more engaging. Shot selection can ensure that everyone in the room (e.g., a meeting environment) gets screen time, which can make the meeting more inclusive for participants.
[0250] The disclosed embodiments can refer to 5 types of shots: speaker shots, listening shots, group shots, reaction shots, and context shots (e.g., overview shots). Speaker shots can provide a closer view of the speaker, making it easy to follow the conversation. Listening shots can be intended to provide diversity and capture the reactions of non-speaking participants. Even when they are not speaking, they can ensure that everyone is visually present in the conversation. Using context shots (e.g., overview shots), remote participants can get a complete picture of what is happening in the room. The overview shot can convey information about the physical meeting environment, such as the movement and reactions of meeting participants and who is entering / leaving the meeting environment. Group shots can provide a view of a group / cluster of meeting participants who are located close together or very near each other. Reaction shots can provide a closer view of one or more meeting participants showing a reaction (e.g., smiling, nodding, frowning, or other facial expressions or body language) in response to a speaker or an event in the meeting environment. In other words, a reaction shot can be characterized by or can include a view representation of a meeting participant determined to be showing a reaction, which can be verbal or non-verbal. Context shots (e.g., overview shots) can be shown or displayed when a remote participant is speaking or when there is a lot of movement in the room.
[0251] In some embodiments, a camera may use a directional microphone to determine where sound is originating within a room or environment. Machine learning may enable the camera to detect the number of people present and where they are located. Combining these two types of data, the disclosed embodiments can accurately identify who is speaking and who is listening, and can use this information to provide a video experience that represents all participants in a natural way.
[0252] Embodiments of the present disclosure may relate to speaker framing methods. Speaker framing may be implemented as an AI feature that is aware of what is happening in a room and can dynamically adapt the view based on an understanding of what the camera sees and hears. It can provide a good view of the person speaking while giving the context needed to feel comfortable participating and being part of the conversation.
[0253] Gallery view may refer to a method of viewing a video stream that is intended to provide an overview of conference participants in addition to fairness among the participants in a video conference. For example, Fig.35A An example conference environment including two conference participants is shown. As Fig.35B shown, each conference participant may be represented in a subpicture, video picture, or stream within a tile of the layout. In Fig.35B which, Fig.35A shown, the two conference participants are displayed horizontally (in a row) side by side. The subpicture, video picture, or stream may be adjusted such that each conference participant has a similar size, is horizontally aligned, and shows both the head and upper body. This may promote fairness among the conference participants in a video conference. The gallery view may include a video picture divided into two or more video tiles, each of the video tiles associated with a corresponding video stream featuring one or more objects and / or conference participants. The relative size, orientation, and position of the tiles may be used to highlight various objects or conference participants. In some embodiments, a tile may include a video stream of a presenter or other active participant, and the tile may be highlighted relative to one or more other tiles. Additionally, in some embodiments, a tile may feature an inactive participant or may provide an overview of some or all of the conference participants in a particular environment. The tiles may be oriented, sized, and positioned to highlight the video streams of certain participants (active or inactive) and show the relationships between the participants (e.g., speaker-listener relationships, spatial relationships between participants, etc.). For example, Fig.35C illustrates Fig.35A a gallery view or tile layout of conference participants where the conference participants are not aligned. As Fig.35C shown, the conference participant on the left may be shown as larger, positioned higher, or may show more of their upper body (compared to the other conference participants on the right) to indicate that the conference participant on the left is of interest (e.g., the speaker or an active listener).
[0254] Fig. 22 shows an example of a tile layout. As Fig. 22 shown, the example tile layout may include column layouts 2210a - 2210f, row layouts 2220a - 2220c, gallery layouts 2230a - 2230b, postcard layouts 2240a - 2240d, and Lego layouts 2250a - 2250d. It is contemplated that any arrangement of tiles may be formed on a display, and the tile layout is not limited to the examples shown herein. Additionally, for exemplary purposes, the provided example layouts are given names and / or categories and are not limited to the names and / or categories provided herein.
[0255] Figure 36-Figure 44 shows various non - limiting examples of a tile layout. As Fig.36 shown, a meeting environment may include four meeting participants 3610, 3620, 3630, 3640. Additionally, one of the meeting participants may be identified as the speaker 3610. The meeting participants 3610, 3620, 3630, 3640 may be seated at a table, as Fig.36 shown. Although four meeting participants are shown by way of example in Fig.36 , it is contemplated that the meeting environment may include any number of meeting participants who may be seated in any order or combination.
[0256] FIG. 37A to FIG. 37Y shows various tile layouts based on the Fig.36 shown meeting environment. As Fig.37A shown, a 3×3 matrix may be displayed. The video stream corresponding to the speaker 3610 may occupy the top two rows of tiles (e.g., a total of six tiles). The video stream of each remaining meeting participant 3620, 3630, 3640 may occupy a tile in the bottom row of the matrix. Additionally or alternatively, as Fig.37B shown, the video stream corresponding to the meeting participant 3640 may occupy the top two rows of tiles (e.g., a total of six tiles) in the 3x3 matrix. It is contemplated that in this context, the meeting participant 3640 may be an active participant, a listener, or showing a reaction. The video stream of each remaining meeting participant 3610, 3620, 3630 (including the speaker 3610) may occupy a tile in the bottom row of the matrix. Fig.37C shows a 3x3 matrix display where an overview shot of the meeting environment occupies the top two rows (e.g., six tiles) of the matrix. The video streams corresponding to the meeting participants 3620, 3630, 3640 who are not the speaker 3610 may occupy the tiles on the bottom row of the matrix. Fig.37DA 4×3 matrix is shown, where an overview diagram of the meeting environment occupies the first (top) two rows of the matrix (e.g., eight tiles). Video streams corresponding to each meeting participant 3610, 3620, 3630, 3640 can occupy the tiles in the bottom row of the matrix. Figure 37E-37H Various 4×3 matrix displays are shown, where the video stream corresponding to the speaker 3610 occupies the first (top) two rows and additional tiles in the bottom row of the matrix display. The video streams of each remaining meeting participant 3620, 3630, 3640 can occupy the remaining bottom row tiles in the matrix. Fig.37I A 4x3 matrix display is illustrated, where the video stream corresponding to the speaker 3610 occupies the left three columns of the matrix display. The video streams of each remaining meeting participant 3620, 3630, 3640 can occupy the tiles in the remaining right column of the matrix display. Fig.37J A 4×4 matrix display is shown, where an overview diagram of the meeting environment is shown in the left three column tiles. The video streams of each meeting participant 3610, 3620, 3630, 3640 can be shown in the tiles in the remaining right column of the matrix display. Figure 37K A 2×2 matrix display is illustrated, where the video streams corresponding to each meeting participant 3610, 3620, 3630, 3640 are shown in the tiles. It is envisioned that the arrangement of the video streams in the tiles is shown in any display discussed herein and can correspond to the position(s) of the corresponding meeting participant(s) within the meeting environment. Figure 37L A 3×2 matrix display is illustrated, where the overview shot occupies the bottom row of the matrix display, the speaker shot of the speaker / meeting participant 3610 occupies the right two columns of the top row, and the reaction shot of the meeting participant 3640 occupies the left column of the top row. Figure 37M A 1×2 matrix display is shown, where the overview shot occupies the bottom row and a shot on the shoulder consistent with the embodiments discussed herein occupies the top row. In some embodiments, a whiteboard 3650 can be included in the meeting environment, as Fig.37N shown. Fig.37N A 3×2 matrix display is shown, where the overview shot occupies the bottom row of the matrix display, the reaction shot of the meeting participant 3630 occupies the left column of the top row, a group shot including the speaker / meeting participant 3610 and the meeting participant 3640 occupies the center column of the top row, and the video stream of the whiteboard 3650 occupies the right column of the top row. Although the whiteboard is discussed, it is envisioned that any interesting / important object or item can be included in the meeting environment and can be shown in the video streams and / or tiles on the display. Fig.37O An example layout with diagonal (or triangular) tiles is shown, where the diagonal of the tiles can follow one or more table angles. As Fig.37OAs shown, the layout can include three diagonal (or triangular) tiles, each corresponding to a conference participant 3610, 3630, 3640. Figure 37P Illustrates another example layout with diagonal (or triangular) tiles. As Figure 37P shown, the layout can include two diagonal (or triangular) tiles, each corresponding to conference participants 3610, 3640. Each conference participant 3610, 3640 shown can be speaking, and in some embodiments, Figure 37P the layout can show a conversation between conference participants 3610, 3640. Figure 37Q Illustrates a third example layout with four diagonal (or triangular) tiles. As Figure 37Q shown, each diagonal (or triangular) tile can correspond to a representation of conference participants 3610, 3620, 3630, 3640. Figure 37R Illustrates an example layout including a group shot showing conference participants 3610, 3620; an overview shot in a tile at the bottom corner of the layout; and a floating circular tile showing conference participant 3630. The group shot showing conference participants 3610, 3620 can be a speaker shot. Figure 37S Illustrates an example layout including a group shot showing conference participants 3610, 3620; an overview shot in a tile at the bottom center of the layout; and two floating circular tiles showing conference participants 3630, 3640. As discussed above with respect to Figure 37R the group shot showing conference participants 3610, 3620 can be a speaker shot. Figure 37T Illustrates an example layout showing conference participants 3610, 3620 and an overview shot at the bottom center of the display. As discussed above with respect to Figure 37R-Figure 37S the group shot showing conference participants 3610, 3620 can be a speaker shot. Figure 37U Illustrates an exemplary organic tile layout where the tiles are generally amorphous, as shown by the curves. Figure 37U the layout can include a speaker shot of conference participant 3610 (shown in a tile on the left side of the display) and a group shot of conference participants 3630, 3640 (shown in a tile on the right side of the display). Figure 37V Illustrates a layout with circular tiles of different sizes, each corresponding to a representation of conference participants 3610, 3620, 3630, 3640. As discussed herein, the circular tiles can be the same size or different sizes, and the size of the (one or more) tiles can be determined by various factors including, but not limited to, who is speaking, who spoke previously, important objects, reaction shots, group shots, etc. For example, as Figure 37VAs shown, the tile including the speaker shot of speaker 3610 can be larger than the tiles including the shots of other conference participants 3620, 3630, 3640 determined not to be speaking. Figure 37W Illustrates an example of a geometric layout, including tiles of various sizes showing conference participants 3610, 3620, 3630, 3640, 3650. Conference participant 3650 can be a remote (or far - end) conference participant. Figure 37X Illustrates a layout that combines a geometric layout and an organic layout, where speaker 3610 can occupy a geometric - layout tile and other conference participants 3620, 3630, 3640 can occupy organic - layout tiles. Figure 37Y Illustrates an example layout showing speaker 3610 and three floating circular tiles showing conference participants 3620, 3630, 3640. In some embodiments, as Figure 37Y shown, a speaker silhouette can be used to highlight, for example, the leader (the speaker) 3610 of the conference participants. Although Fig.36 and Figure 37A-Figure 37Y illustrate examples with a specific number of conference participants and a specific matrix display, it should be understood that the matrix display and tile layout can incorporate any combination of the number of conference participants, number of rows, number of columns, and the way of determining the number of tiles in the matrix occupied by each video stream. The number of tiles / cells in the matrix occupied by each video stream can be determined based on importance. For example, video streams corresponding to speakers, active participants, listeners, conference participants showing reactions, and / or objects of interest can occupy more tiles in the matrix than other conference participants or objects in the conference environment.
[0257] Figure 38A-Figure 38D Shows various examples of tile layouts involving cluster and group shots. For example, Fig.38A Illustrates a display involving a 4×2 matrix display, where a cluster / group shot of two conference participants occupies the two upper - left cells / tiles of the display. Individual shots can occupy each of the remaining tiles. Fig.38B Illustrates a display involving a 4×2 matrix, where a cluster / group shot of three conference participants occupies the two upper - left cells / tiles of the display, and another cluster / group shot of three conference participants occupies the two upper - right cells / tiles of the display. Individual shots can occupy each of the remaining tiles. Fig.38C Illustrates a display involving a 4×2 matrix, where a cluster / group shot of two conference participants occupies the two upper - left cells / tiles of the display, and another cluster / group shot of two conference participants occupies the two upper - right cells / tiles of the display. Individual shots can occupy each of the remaining tiles. Fig.38DIllustrates a display involving tiles of various sizes and positions. As Fig.38D shown, the display can include various cluster / group shots and individual shots. Additionally, the display can involve full-body shots.
[0258] Figure 39-Figure 44 Illustrates various examples of tile layouts. As Figure 39-Figure 43 shown, tiles representing speakers, active participants, or other meeting participants / objects of interest are shown as tiles with solid borders. Tiles representing other meeting participants are shown as tiles with dashed borders. Fig.39 Illustrates various examples of floating tile layouts, where the blocks shown on the display can have various sizes, be located at various positions, have various orientations, and / or overlap each other. Tile layout 3910 illustrates a floating tile layout with a single speaker, tile layout 3920 illustrates a floating tile layout with a transition to a new speaker, and tile layout 3930 shows a floating tile layout with two speakers. Fig.40 Illustrates various examples of adjusted grid layouts, where the tiles shown on the display can be in a grid shape but have various sizes. Tile layout 4010 illustrates an adjusted grid layout with a single speaker, tile layout 4020 illustrates an adjusted grid layout with a transition to a new speaker, and tile layout 4030 shows an adjusted grid layout with two speakers. Fig.41 Illustrates various examples of geometric layouts, where the tiles shown on the display can have different shapes (e.g., circular, as Fig.41 shown). As shown in geometric layout 4110, the presenter of whiteboard 4112 can be identified as the speaker, and the size of the corresponding tile can be adjusted such that it is larger than the tiles of other meeting participants 4114. In some embodiments, as shown in geometric layout 4120, a meeting participant (e.g., an audience member) can start speaking or provide a comment 4124. The size of the tile corresponding to the meeting participant can be adjusted such that it is larger than the tiles of other meeting participants and the previous speaker 4122. Fig.42 Illustrates various examples of circular tile layouts, where the shown tiles can be in a circular (or circle) shape. Although shown as circular or a circle in Fig.42 , it is contemplated that the tiles can be of any shape, such as but not limited to square, rectangle, triangle, pentagon, hexagon, etc. Additionally, it is contemplated that the tiles within the display can be of different shapes. As shown in circular tile layout 4210, the speaker can be identified by a corresponding tile that is larger than the tiles corresponding to other meeting participants. As shown in circular tile layout 4220, a new speaker can start speaking, and the tile corresponding to the new speaker can increase in size while the tile corresponding to the previous speaker can decrease in size. Fig.43Illustrates various examples of soft rock tile layouts. As used herein, a soft rock tile layout can involve tiles that are typically amorphous and / or have no specified shape. Tile layout 4310 illustrates a soft rock tile layout with a single speaker, tile layout 4320 illustrates a soft rock tile layout with a transition to a new speaker, and tile layout 4330 illustrates a soft rock tile layout with two speakers.
[0259] Fig.44 Illustrates various examples of organic layouts. Tiles in an organic layout can typically be amorphous and, in some embodiments, can be shaped based on the general profile of the meeting participants or the objects represented in their corresponding video streams. As shown in organic layout 4410, some single camera systems can capture the meeting participants and display them in various clusters within the tiles. As shown in organic layout 4420, a multi-camera system can capture the meeting participants and objects of interest (e.g., a whiteboard) and display them in various tiles.
[0260] The gallery view can show or display certain meeting participants in more than one tile to highlight those participants and provide context on how those participants relate to others in the group. For example, the gallery view can include two or more tiles. In some embodiments, in a first tile, at least one active participant can be represented individually and can also be shown in a second tile together with one or more other participants. The terms "first" and "second" do not specify any particular ordering, orientation, etc. of the tiles on the display. Instead, the first tile and the second tile can designate any of the tiles in the gallery view of two or more tiles.
[0261] In some embodiments, the gallery view can be implemented using AI techniques and can provide individual views of everyone in the room / environment. The camera can detect the people in the room and create a split view based on the detection.
[0262] By using machine learning, the number of people in the room and the location of the people in the room can be detected and / or determined. These detections can be used with a rule set / training method regarding how people should be framed to create a split view with a selected framing for the meeting participants.
[0263] In some embodiments, the gallery view may involve an AI-driven framing experience that can ensure meeting fairness in conference environments of various sizes (e.g., small conference spaces, medium conference spaces). The gallery view may involve a split-screen layout that automatically adjusts the zoom level to give each participant equal representation. In some embodiments, the gallery view may remove (unnecessary) blank space from video streams (e.g., overview video stream, primary video stream), thereby allowing for focusing on the speaker or active meeting participants. Additionally or alternatively, the framing may be adjusted when an individual enters or leaves the room to ensure that all meeting participants are visible during the meeting. In some embodiments, the gallery view may use overview and group layouts to show all participants in the conference environment. For example, when people are close together, the gallery view may place them in a common tile as a group. As another example, when there is frequent movement within the conference environment, people enter or leave the conference environment, or there are poor framing conditions, the gallery view may use the overview to frame all participants without disrupting the layout.
[0264] Figure 23A-23B An example of a composite layout / group shot is illustrated. As Fig.23A shown, four meeting participants are represented in the overview stream 2310. A composite layout 2320 can be selected, determined, and output that groups the meeting participants together and presents the group shot as a tile in the gallery view. As Fig. 23B shown, three meeting participants can be represented in the overview stream 2330. A composite layout 2340 can be selected, determined, and output that groups the meeting participants together and presents the (one or more) group shots as tiles in the gallery view.
[0265] At the start of a new stream (or meeting, e.g., video conference), the gallery view feature may show an overview with the full field of view of the camera. This can provide remote (or far-end) participants with the orientation of the scene in the conference environment, conveying information about the physical conference environment, such as who is in the conference environment and where they are located. Showing an overview shot at the start of a new stream can also give the AI model time to accurately detect the participants in the conference environment (or room). After a period of time (e.g., one minute, five minutes, ten minutes, thirty minutes, one hour), if the required conditions are met, the gallery view feature may transition to a split-view layout. In some embodiments, if one or more conditions are not met, the gallery view feature may continue to present the overview and reframe the best possible view. Once all conditions are met, the gallery view feature may ultimately transition to a split-view layout.
[0266] Consistent with the disclosed embodiments, the zoom level of the split view layout can be automatically adjusted to show meeting participants with equal equity. The zoom level can be limited to a maximum zoom (e.g., 500%) to avoid degradation of the video stream or image quality. Additionally, in some embodiments, the gallery view can be designed to frame all meeting participants in such a way that all meeting participants are centered in their respective frames and horizontally aligned relative to each other. Each frame of each meeting participant can show the upper body of each meeting participant to capture body language. Additionally, based on their position in the room, the meeting participants can be shown in order from left to right (or right to left). If a meeting participant moves or switches positions, the gallery view can adapt and adjust the split view layout once the meeting participant stops moving.
[0267] In some embodiments, when a meeting participant moves, the gallery view can reframe. As Fig.25 shown, meeting participant 2500 can move slightly out of the frame, and the gallery view can adjust to reframe meeting participant 2500. Although shown in Fig.25 as a split view, it is contemplated that the gallery view can reframe for an overview view type.
[0268] As Fig.24A shown, the gallery view can include two general view types: overview 2410 and split view layout 2420. The overview 2410 can frame all meeting participants in the best possible way and can include different layouts. Non-limiting examples of overview layouts 2410a, 2410b, 2410c, 2410d are shown in Fig. 24B based on the framing of meeting environments 2430a, 2430b, 2430c, and 2430d. The split view layout 2420 can include different layouts. Fig.24C Non-limiting examples of split view layouts 2420a, 2420b, 2420c, 2420d are shown in Figure 24A-Figure 24C based on the framing of meeting environments 2440a, 2440b, 2440c, 2420d. As shown in split view layout 2420d, a group shot of an odd number of meeting participants can be generated and displayed based on, for example, an odd number of individuals in meeting environment 2440d. The disclosed embodiments can receive visual input from an artificial intelligence model that detects people. Using this input and a set of rules regarding composition and timing, the gallery view can frame the people in the meeting room with the view types as
[0269] Additionally, as Fig.30As shown in the example in [0], the gallery view can include two general modes: the individual view mode 3020 and the overview mode 3030. The meeting environment can include two meeting participants in the 120-degree field-of-view lens 3010. In the individual view mode 3020, the person viewfinder can identify the representations of the meeting participants and display each representation of each meeting participant in a tile. For example, as shown in the individual view mode 3020, each representation of each meeting participant can be shown in a tile, and the tiles can be side by side in a horizontal manner. It is contemplated that the tiles can be any layout or arrangement, including but not limited to vertical columns, horizontal rows, or an N×M matrix (where N can represent the number of columns in the matrix and M can represent the number of rows in the matrix). It is also contemplated that the tiles can be of different sizes, and the meeting participant or object of interest can be shown in a larger tile than other meeting participants or objects. Additionally, as shown in the individual view mode 3020, the picture corresponding to each meeting participant can be adjusted so that the meeting participants are displayed with equal fairness (as discussed herein). In some embodiments, the overview mode 3030 can provide the best possible view of all participants when, for example, the gallery mode (or individual view mode or split view mode) cannot be achieved. As shown in the overview mode 3030, the meeting participants (and, in some embodiments, the objects of interest) can be viewed together and displayed in a tile or on a display.
[0270] In some embodiments, as discussed herein, the display can be included in a video conferencing system (e.g., Microsoft Teams, Zoom, Google Meet, Skype), as Fig.31 shown. Fig.31The display may include: a gallery view tile layout (or composite layout) 3110 of the meeting environment; additional video streams 3120a, 3120b, 3120c, 3120d of remote meeting participants; and a chat window including chat bubbles 3130a, 3130b, 3130c. As shown, in some embodiments, based on the detection of the overlap of the representations of the meeting participants, multiple meeting participants may be displayed in each tile. As used herein, the terms "overlap" or "overlapping" may be understood to mean that the representations of individuals (e.g., meeting participants) spatially overlap in the video stream. It is contemplated that the display included in the video conferencing system may include any combination of the features of the display discussed herein, including but not limited to the layouts discussed herein. For example, in some embodiments, the video conferencing system (or streaming platform) may design the layout according to the various shot types discussed herein (via user input or video conferencing system requirements). As another example, the layouts discussed herein may be directly streamed to a display within the video conferencing system and shown on the display within the video conferencing system.
[0271] In some embodiments, the gallery view may implement a transition between an overview view layout and a split view layout. The transition may occur when certain conditions are met, and the certain conditions may include: the number of people that must be supported for the split view layout; all people in the room must be reliably detected, there are no large movements in the scene, there is sufficient space between people, and a specific period of time has elapsed since the last transition. If the number of people has changed, but the split view still supports that number, the gallery view may transition directly from the split view layout to another split view layout. If the number of people has changed, but the split view does not support that number, the gallery view may transition from the split view layout to an overview layout. If people move such that it is not possible to frame them without someone (or a head) overlapping into an adjacent tile, the gallery view may transition from the split view layout to an overview layout.
[0272] Embodiments of the present disclosure can combine speaker framing and gallery views such that an overview of the meeting environment is provided while the currently selected shot (or the dominant stream that is important, e.g., due to the speaker, listener, or active participant captured thereby) is prioritized in space. Prioritization in space can include, but is not limited to, where the tile showing the prioritized shot is larger (e.g., in area) (e.g., twice the size of the other tiles shown) than the other tiles shown. Thus, participants of interest (e.g., speakers, listeners, active participants) and all participants in the meeting environment can be viewed simultaneously. Additionally, the roles (e.g., speakers, presenters) and relationships (e.g., speaker and listener, presenter and audience) within the meeting environment can be understood and depicted to remote or distant participants.
[0273] Figure 26A-26F Various examples of layout configurations are illustrated. Fig.26A An example video stream of a meeting environment is shown. Fig.26B An example layout including an overview stream and four dominant streams (each dominant stream corresponding to each meeting participant) is illustrated. Fig.26C An example presenter shot with an object of interest (such as a whiteboard) is illustrated. In the detection of an object of interest, embodiments of the present disclosure can run a neural network or engine to simultaneously process the detection of meeting participants and character detection (e.g., writing, drawing, or other characters written on, e.g., a whiteboard). Fig.26D An example video stream of a meeting environment where a meeting participant (e.g., a presenter) is actively speaking is illustrated. Fig.26E An example split view stream is illustrated where the display of the presenter occupies two rows and three columns of tiles. Fig.26F An example split view stream is illustrated where the display of the listener occupies two rows and three columns of tiles.
[0274] As Fig.26E shown in the example, the speaker in the prioritized tile may no longer be shown in the smaller tile. Additionally, some embodiments of the present disclosure can involve switching between a speaker view (e.g., Fig.26E ) and a listener view (e.g., Fig.26F ). In some embodiments, the speaker can occupy the small tile of the listener while the listener is in the prioritized tile, and vice versa.
[0275] The layout configuration can be described by specifying the position and size of each tile. Additionally, two corner points can be specified to further specify any possible layout (e.g., a rectangle) composed of the tiles, and this information can be sent to the layout engine (discussed below) to compose the output stream.
[0276] Some embodiments of the present disclosure may allow a user to manually select meeting participants to give them priority in the layout. The selected meeting participants may be presenters or any important people. In some embodiments, a user interface (e.g., via a camera app) may allow the user to manually select meeting participants, e.g., by selecting a bounding box associated with the meeting participant.
[0277] The camera may run a neural network for detecting people, a Kalman filter, and a matching algorithm for tracking people or meeting participants. An identifier (ID) may be assigned to each meeting participant. Non-limiting examples of IDs include numerical values, text values, symbols, combinations of symbols, vectors, or any other value that can distinguish a person. The ID may be shared between the camera and, e.g., the camera app. The ID may be used to select a person in the camera app and forward information to the camera. The camera may then prioritize the selected person as long as they are in the stream. Once a person leaves the room, the ID may be deleted or otherwise cease to exist, and the features may be reset, e.g., to automatic speaker detection or equal fairness, until a new person is manually selected.
[0278] In some embodiments, a gallery view may receive trajectories representing meeting participants. The gallery view may group these trajectories based on whether they can be framed together and produce a grouped list. The list may contain single-person clusters and / or group clusters. A single-person cluster may be obtained when a person in the meeting environment is at a sufficient distance from the rest of the participants such that the meeting participant can be framed by themselves. A group cluster may include a group of meeting participants that are close enough together such that they can be framed together (or in one shot).
[0279] In some embodiments, meeting participants may be grouped based on head height. By approximating that the width of a person's shoulders is approximately twice their head height, the head height can be used to determine the distance in pixels from the center of the person's head that is required to frame them based on the gallery view principles discussed herein. The determination may be calculated as shown below, where λ represents a factor by which the shoulder width needs to be multiplied to include body language.
[0280]
[0281] Body language can be an important part of communication, and by framing people by focusing on the face and upper body, a space for connection and understanding can be created. As described above, the gallery view can be designed to reduce white space and frame each meeting participant equally. Thus, participants in a meeting can be framed such that their body expressions or body language can be captured by their respective main video streams and / or tiles. Additionally, in some embodiments, meeting participants can occupy similar portions of the screen. In some embodiments, the gallery view can be used to indicate the layout when there is no speaker. In other embodiments, when there are multiple speakers or too many speakers to determine which speaker needs priority, the gallery view can be used to indicate the layout.
[0282] The gallery view can make everyone in the room appear similar in size and keep people's heads aligned on the screen. For example, if someone appears larger or taller in the image, they may seem more important, and this can create an unnecessary sense of power imbalance among the participants, which may not actually exist.
[0283] If a person moves such that they are cropped or no longer visible in their frame, the camera can adjust the framing to capture their new position. Potential benefits can include any of the following: creating a sense of fairness among all meeting participants; ensuring that remote participants get a closer view of everyone in the meeting room; and / or removing white space (walls, ceiling, floor, etc.) in the room from the image.
[0284] The gallery view can also help by: helping remote participants maintain an overall view of everyone in the meeting room; ensuring that everyone gets the same amount of space and time on the screen; and framing meeting participants more closely (and, in some embodiments, without disturbing or overlapping with other meeting participants).
[0285] The technical implementation of the gallery view can include a machine learning (ML) vision pipeline that can detect people (heads and bodies) in an image. By using ML and filtering techniques (e.g., Kalman filter), person trajectories can be created based on these detections. These trajectories can be based not only on the current input detections but also on the input history and contain additional information (e.g., if a person is moving). The trajectories can provide input data to a virtual director unit. The virtual director unit (which can be implemented as a finite state machine) can determine the layout of each tile and the framing commands based on the input data and its own state.
[0286] As described above, the layout engine director can implement a combination of speaker framing and gallery view features. Speaker framing (SF) and gallery view (GV) can run in parallel, as Fig.27AAs shown, and forward its corresponding output to the layout engine in the form of status information. The layout engine can consider the basic status from the gallery view to determine what basic layout to create, and then add (e.g., overlay) the status of the speaker viewfinder on top to create a prioritized gallery view. In some embodiments, the speaker viewfinder can be replaced with an assigned speaker (as described above) and a low-pass filter, as Fig.27B shown. In such embodiments, the system can involve only tracking the speaker and only giving priority to the active speaker in the gallery.
[0287] In some embodiments, as discussed herein, the layout engine director can be implemented in a multi-camera system, such as Fig.28 shown in the example of. Consistent with some embodiments of the present disclosure, the multi-camera system 2800 can include a main camera 2810 (or primary camera), one or more secondary cameras 2820 (or one or more peripheral cameras), and a user computer 2840 (or host computer). In some embodiments, the main camera 2810 and one or more secondary cameras 2820 can be of the same camera type, such as but not limited to the examples of cameras discussed herein. Additionally, in some embodiments, the main camera 2810 and one or more secondary cameras 2820 can be interchangeable, such that the main camera 2810 and one or more secondary cameras 2820 can be located together in a conference environment, and any of the cameras can be selected to act as the main camera. Such a selection can be based on various factors, such as but not limited to the location of the speaker, the layout of the conference environment, the location of auxiliary items or items of interest (e.g., whiteboard, presentation screen, television), etc. In some cases, the main camera and the secondary cameras can operate in a master-slave arrangement. For example, the main camera can include most or all of the components for video processing associated with the multiple outputs of the various cameras included in the multi-camera system. In other cases, the system can include a more distributed arrangement, where the video processing components (and tasks) are more evenly distributed across the various cameras of the multi-camera system.
[0288] As Fig.28 shown, the multi-camera system 2800 can include components similar to those of the multi-camera system 100, including but not limited to: image sensors 2811, 2821; DOA / audio units 2812, 2822; visual processing units 2813, 2823, virtual director unit 2814; layout engine 2815 and APIs 2816, 2826. These components can perform similar functions herein. Additionally or alternatively, as Fig.28As shown, the main camera 2810 may include components different from one or more secondary cameras 2820. For example, the multi-camera system 2800 may include a layout engine 2815, which may incorporate any and all features discussed herein with respect to layout engine director, layout engine, and adaptive layout.
[0289] The speaker or group of speakers may be located on the left, right, or middle of the image relative to other participants in the room. In some embodiments, once the position of the speaker in the room relative to the camera is determined, the positions of each layout can be mapped based on the other participants. As an example, if the current speaker is on the left side of the image, the layout engine algorithm can determine what is on the left side of the speaker. The remaining clusters can be used to map each scene, considering that each cluster can consist of a single person being framed or a group of people being framed.
[0290] In some embodiments, a set of rules may be employed to determine the potential positions of the clusters. For example, if the position of the speaker or group of speakers is on the left, it may have up to three clusters to its left. As another example, if the position of the speaker or group of speakers is on the right, it may have up to three clusters to its right. As yet another example, if the position of the speaker or group of speakers is in the middle, there may be up to two clusters to its left and up to two clusters to its right. The subject or meeting participant may have from zero to three clusters to its left, and each cluster may include any possible combination of single-person clusters and group clusters. Such rules may be determined based on the size of most meeting rooms (e.g., the average size of meeting rooms) or the size of a particular meeting room.
[0291] Embodiments of the present disclosure may include a part of the state logic or algorithm, which may decide whether the current layout (e.g., the layout currently being displayed) should be changed. To this end, it may track the information provided by the layout engine. The state logic may depend on the assumption that all previous components provide consistent information to make a decision. The current layout may be updated if (i) the candidate layout is approved, or (ii) it is not possible to determine what layout should be displayed.
[0292] In some embodiments, in order to approve a candidate layout, the state logic must verify that the layout engine has returned to the same candidate layout for a certain number of consecutive iterations. To this end, the state logic can maintain a count of the number of times the same candidate layout has been consecutively repeated, and the candidate clusters are also the same. As described above, a cluster can consist of an undefined number of people / objects, and there can be a situation where the layout engine can return the same layout but the members of the cluster have changed. This can lead to framing people / objects in an unexpected or incorrect order. Once the layout and the cluster have been repeated a specific number of times, the state logic can update the state and send the necessary information to the image pipeline to update the layout.
[0293] In some embodiments, the state logic may not be able to map a scene to a specific layout. In these cases, new methods can be employed to (i) update to an overview shot (e.g., if no one is speaking), or (ii) frame only the speaker (e.g., if someone is speaking). This new method can be triggered if the layout engine has returned different candidate layouts during multiple iterations. This can occur in scenarios where people / subjects move around the room and switch positions over an extended period of time and there are clusters that have not been mapped to a layout. In some cases, there may be several people / subjects speaking simultaneously over an extended period of time, or there may be a remote speaker. An algorithm consistent with the disclosed embodiments can employ a gallery view layout.
[0294] Fig.29 Processing of a layout engine consistent with some embodiments of the present disclosure is shown. As shown in step 2910, a camera (or cameras) can first detect four people in a meeting environment. The gallery view feature can determine a four-split view basic layout (e.g., based on equal fairness), as shown in step 2920. The layout engine can combine information from the gallery view with speaker information (e.g., via speaker framing or assigned speakers), and modify the basic layout accordingly to give priority to speaker 2940, as shown in step 2930. As Fig.29 shown, the final layout can be a combination of inputs from both the gallery view and speaker framing. The layout engine can be part of an algorithm, and the layout engine can decide which layout is more suitable for a particular situation based on the current speaker and the current participant clusters (e.g., if the speaker is included in the cluster, the cluster can be given priority for display). With this information, the layout engine can provide a candidate arrangement for each iteration. If the information should change to a new state / layout, this information can later be used by the framing logic for the device.
[0295] In some embodiments, the layout engine can be configured to implement a stream with up to 4 different tiles. The layout engine can be configured to implement multiple different layout variations, such as but not limited to 1×2 split, then 2×2, and 4×4, where each tile can be associated with a separate video stream. Any number of tiles can also be combined and used by one stream, so in some embodiments, one stream can occupy 2 columns and 2 rows. To switch from the default overview (all participants in one tile) to a multi-tile layout, the following conditions may need to be met: the correct number of person tracks (e.g., for the corresponding layout); all tracks need to be valid (e.g., a person needs to be detected within a specified time, such as 5 seconds); people are not moving in the image; including a waiting time (e.g., 5 seconds) after the layout switch to prevent the layout from switching before it can switch again to reduce visual noise.
[0296] In addition, if people overlap in the image, this may cause their corresponding tiles to be merged.
[0297] To frame all participants, the virtual director unit can have at least three different types of viewfinders for its use: an overview viewfinder, a group viewfinder, and a person viewfinder.
[0298] The overview viewfinder can frame all participants in the image, and the person viewfinder and the group viewfinder can be attached to or correspond to specific people and groups of people, respectively. For example, as Fig.32A shown, the meeting environment 3210 can include four (4) meeting participants. Each meeting participant can have a corresponding person viewfinder A, B, C, D. In some embodiments, and as Fig.32A shown, the person viewfinders A, B, C, D can overlap. Thus, as Fig.32B shown, the sub-streams 3220a, 3220b, 3220c, 3220d created based on the person viewfinders A, B, C, D (from Fig.32A ) can include cropped bodies, which may distract or be unpleasant to the user or viewer of the framed video stream.
[0299] Therefore, in some embodiments, the disclosed systems and methods can generate framed shots according to a method in which participants (e.g., participants sitting close together in a meeting) are grouped together to produce a shot / viewfinder that focuses on a specific participant of interest while including adjacent participants. This can produce a visually pleasing shot. For example, as Figure 32C-Figure 32DAs shown, sub-streams 3240a and 3240b can be displayed respectively based on person viewfinders 3230a and 3230b. The person viewfinders 3230a and 3230b and the sub-streams 3240a and 3240b can avoid cropping of the subject by grouping participants together or focusing on a specific participant of interest while including adjacent participants.
[0300] A person viewfinder can be bound to the lifetime of its corresponding person track (and the same is true for a group viewfinder with a selected group of persons). The virtual director unit can be responsible for providing each viewfinder it creates with the correct subset of tracks it receives, and for delegating and arranging the viewfinder output (e.g., viewfinder commands) in the correct manner and order (e.g., according to the active layout).
[0301] In some embodiments, the virtual director unit can (i) manage layout selection, and (ii) manage viewfinders that provide separate viewfinder commands for each tile. The virtual director unit can forward layout information (e.g., the number of tiles, tile arrangement, any tiles that should be merged) and viewfinder commands for each tile (e.g., as pan-tilt-zoom values with additional information about when to reframe) to the layout engine. In some embodiments, the gallery view can use hard cut transitions for layout switching and Bezier interpolation for reframing within a tile. Additionally, the virtual catalog unit can continuously evaluate the input from the visual pipeline to indicate the layout and framing to be sent to the layout engine.
[0302] The prioritization criteria / detection can depend on the scene / activity in the room or meeting environment. The virtual director unit can ensure that the speaker is focused in the layout, and can ensure that if the speaker moves or changes position in the room, the layout will adapt accordingly. In some embodiments, the virtual director can ensure that the camera that shows the most of the person from the front is used in its corresponding tile. As the meeting progresses, it may be necessary to change the layout to give one person more space, or to give each person the same size.
[0303] The virtual director unit considers the duration of the conversation between persons and the duration of the last conversation. For example, during a discussion, the virtual director unit can give each person the same amount of space in the layout and ensure that their relative positions are maintained in the layout. As another example, if person A looks left to look at person B, and person B has to look right to look at person A, then person A can be placed to the right of person B in the layout. In some embodiments, the virtual director unit can also use gestures or body postures to control the layout. For example, if a person stands up and starts a presentation, the visual pipeline can detect this, and their entire body is in the view. The virtual director unit can take this into account and instruct the layout engine that this person should occupy the entire column to give them enough space.
[0304] In some embodiments, when the vision pipeline detects a gesture (such as raising a hand), the virtual director can take this into account and adjust the layout accordingly. For example, a person who raises their hand can receive the same tile size as a person who is speaking in the meeting.
[0305] In some embodiments, the virtual director unit can include software components that can obtain input from the vision pipeline components and determine the layout composition and which part of the primary video stream image should be used in each part of the layout. Attributes that can be evaluated can include, but are not limited to: whether this person is speaking; how long they have been speaking; if someone is engaged in a discussion or a short conversation, where the change in who is speaking occurs (e.g., each person speaks for less than one minute at a time); if someone is presenting or leading the meeting (e.g., one person speaks for most of the meeting or for a total of more than one minute); where they are looking; how many people are visible in the frame; what reactions and body language they are showing (e.g., if they look away, or look at a person, if they smile or laugh, if a person shows signs of sleepiness or closes their eyes); what other people are visible in the frame; where the individual is moving and / or where they have been; what activity they are engaged in (e.g., writing on a whiteboard or drawing on a document); position and orientation; timing (e.g., avoiding frequent switching or reframing between layouts).
[0306] Embodiments of the present disclosure can include additional features and techniques, including an adaptive layout engine. The adaptive layout engine can be implemented by one or more microprocessors associated with the disclosed systems and methods (e.g., one or more microprocessors associated with the video processing unit of a camera or a server or a cloud-based system). Among other operational capabilities, the adaptive layout engine can analyze one or more overview video streams (or any other video stream, audio stream, and / or peripheral sensor output) to detect various conditions, events, movements, and / or sounds in the environment. Based on such detections, the adaptive layout engine can determine the gallery view video layout to be shown on the display. Aspects of the gallery view that can be controlled by the adaptive layout engine can include, but are not limited to: the number of tiles to include; tile orientation; the relative size of the included tiles; the relative positioning of the tiles; the video stream selected, generated, and / or assigned for each tile; the transitions between frames associated with one or more tiles; the framing of individuals or objects within the tiles (e.g., genius framing, speaker framing, etc.); the selection of individuals or groups of individuals to feature in the gallery view tiles (based on detected actions, total cumulative screen time, screen time fairness, etc.); the selection of durations to maintain a particular shot; any other aspects and their combinations.
[0307] In some embodiments, the layout engine may operate by receiving instructions from a virtual director unit and compositing a new video stream with a portion of one or more of the primary video streams in each tile according to the layout instructions. The layout engine may also support different transitions, where the layout may change smoothly or change size according to instructions from the virtual director unit.
[0308] Additionally, the disclosed systems and methods may use different types of transitions, such as but not limited to: hard cut, interpolation transition, and / or fade transition. A hard cut transition may involve directly replacing a previous image or layout with a new image or layout from one frame to another. An interpolation transition may involve a transition between a previous viewfinder position and a new viewfinder position in an image (e.g., in the form of a Bezier curve or other non-linear change in camera parameter values). The viewfinder may not change its position directly within one frame transition. Instead, it may follow a calculated trajectory between the start and end viewfinder positions over a period of time (e.g., not exceeding 1 - 2 seconds). A fade transition may involve placing a new image over a previous image and gradually increasing the intensity of the new image while gradually decreasing the intensity of the previous image.
[0309] For transitions on a merged or split grid layout, a hard cut or fade transition may be used because an interpolation transition may add unnecessary visual noise and may not always be able to find corresponding viewfinder positions in the previous (old) and new layouts. For transitions within a cell when a person moves, an interpolation (or smooth) transition similar to the transition performed for genius viewfinder may be used.
[0310] In some embodiments, the layout engine may provide multiple streams in addition to composing videos in a main stream, and each of the streams may correspond to a tile in the layout. These streams may be provided to a host / computer / client so that each video stream can be processed and adapted to the overall layout in the video client.
[0311] The video client may also indicate to the virtual director the preferences / requirements for which layouts should be provided. In some embodiments, the client may support only one output stream. The client may provide this requirement to the virtual director, and the virtual director may instruct the layout engine to provide only layouts with one output stream. In other scenarios, the client may have preferences regarding which type of layout it wants to display. As discussed herein, the disclosed embodiments may provide a technical or process improvement over traditional systems and methods by providing an adaptive layout engine directed by machine language (or artificial intelligence) such that the layout of a video conference can be changed based on the determined optimal layout (given, e.g., speakers, active listeners, reactions, etc.). Traditional systems and methods simply are not equipped to perform such optimal layout determination based on the various factors discussed herein.
[0312] Potential scenes or situations captured by the disclosed systems and methods may include: a meeting in a normal room, someone talking for a long time; a discussion; a brainstorm; a standup; a presentation; security / surveillance; collaborative drawing on a canvas, or multiple canvases can be stitched together; or any of the scenes listed previously, but with multiple cameras in the room, with and without canvases.
[0313] In some embodiments, such as in a large collaborative room, the multi-camera system may include 6 cameras: 3 cameras pointing at 3 whiteboards attached to the wall, and three cameras on the opposite side facing the whiteboards to frame participants using the whiteboards. When the vision pipeline detects people or movement / changes in the whiteboards and the people in front of the whiteboards, the virtual director unit can create the layout accordingly. For example, a professor may use all three whiteboards for a lecture and may move back and forth between them as they appear. The vision pipeline can detect this and which whiteboard there is activity on. The vision pipeline can then instruct the layout engine to frame the area in one cell of the whiteboard where there is currently activity, while keeping the portion where there was previously activity in another cell, while keeping the professor continuously presenting in a third cell. The camera feed and perspective that best shows the portion of the whiteboard and the professor is always used.
[0314] The virtual director unit can serve multiple roles. It can manage layout selections and manage a viewfinder (e.g., a software component) that can provide separate framing commands for each tile. The virtual director unit can forward layout information (e.g., number of tiles, tile arrangement, tiles that should be merged) and framing commands for each tile (e.g., as pan-tilt-zoom values with additional information when to reframe) to the layout engine. As an example, the gallery view can use a hard cut transition for layout switching and a Bezier interpolation transition for reconstruction within a tile.
[0315] In addition, examples of adaptive layout engines are provided herein with respect to video conferencing scenarios. The multi-camera system can be installed in a medium or large conference room (e.g., a conference room suitable for approximately 8 to 12 persons). The multi-camera system (e.g., a system featuring appropriate software A crew system may include three cameras placed at the front of the room, one camera placed below the TV, one camera to the left of the conference room, one camera to the right of the conference room, and optionally one camera attached to a whiteboard on the back wall. Additionally, in some embodiments, the multi-camera systems disclosed herein may include 5, 6, or 9 cameras. The cameras may be numbered and / or placed so that an over the shoulder shot may be captured / generated / displayed as discussed herein. Over the Shoulder Shot
[0316] When the room is not in a meeting, the system can be inactive. Four people may enter the room, two may sit on the right side of the table, and the other two may sit on the left side of the table. People can interact with the video client and start a video conference. As the system starts up, the video client can start consuming the video stream from the system.
[0317] The vision pipeline can detect that there are four people in the room and can detect the distance between each person. Then, the virtual director unit can pick an overview shot from the most central camera to provide an overview of the room to remote participants.
[0318] Fig.45 and Fig.46 illustrates various step - by - step determinations of the (tile) layout of various shots based on the meeting environment. For example, as Fig.45 shown, the meeting environment 4510 can include three cameras, a presenter at a whiteboard, and three meeting participants at a table. Various shots or frames 4520 can be generated based on the meeting environment. Important shots or frames can be given more weight and can be selected to be shown on a display or be given more prominence (e.g., larger size / tile) when shown on a display. Non - limiting examples of shots given more weight can include speaker shots, group shots, overview shots, and presenter shots. Various layouts 4530 can be selected to be shown on a display, and the display can switch between various layouts 4530 depending on, for example, changing speakers, movement, and reactions (e.g., nodding, smiling, clapping). As another example, as Fig.46 shown, the meeting environment 4610 can include three cameras, a presenter at a whiteboard, and three meeting participants at a table. Consistent with embodiments of the present disclosure, various shots or frames 4620 can be generated, including speaker shots, group shots, and overview shots. Additionally, as Fig.46 shown, two shots or frames 4620 can be selected to be shown on a tile layout on a display 4630. It is contemplated that the selected layout can be output or streamed to various video conferencing services (such as Microsoft Teams, Google Meet, and Zoom) for display, as discussed herein.
[0319] The following sections discuss various examples of the adaptive layout / layout engine concept implemented in example video conference scenarios and meeting environments. Although each example is discussed with a specific number of people (meeting participants), it should be understood that the specific number of people (meeting participants) discussed is exemplary, and each example can be extended to include more people or reduced to include fewer people.
[0320] The first example may involve a vision pipeline that detects various speakers and implements various layouts. When the meeting starts, everyone in the room can introduce themselves. When the first participant starts talking in the room, the vision pipeline can detect that the first participant is talking, and the virtual director unit can check how far apart the participants are. As an example, the virtual director unit can determine that each participant is far enough apart such that each person can be given their own tile. The virtual director unit can instruct the layout engine to transition to a 3x3 layout, and the picture coordinates of the talking person can occupy the first two rows in all 3 columns. And each of the non-talking participants can occupy one column in the last row.
[0321] Meanwhile, the vision pipeline can detect the gaze, head, and body positions of each person. The vision pipeline can select a camera for each person where most of the person's face is visible. For a participant on the left side of the table looking at a person talking on the right side of the table, it can be the right camera. The vision pipeline can detect their positions, and the virtual director can find a fitting picture based on their gaze, previous movements, and body size. Then, the virtual director can instruct the layout engine to frame the corresponding streams from different picked cameras. In this case, the streams from the right camera can be used to frame the two people on the left side, while the two people on the right side are framed from the left side. Each framing can represent each person in the same size. This can happen before the virtual director can apply any changes in the next step and continuously between each step. If the virtual director determines that a person is looking in a different direction than in the selected camera picture and enough time has passed since the previous change, then it can change the camera feed in the corresponding person tile (or cell) to the camera feed where most of the person's face can be seen.
[0322] When the next person starts talking, the vision pipeline can detect this, and after a specified duration, it can switch to that person occupying the first two rows in all columns. Then, the previous speaker can transition to a tile at the bottom row.
[0323] When everyone in the room has introduced themselves, the people participating remotely can introduce themselves. The vision pipeline can detect that no one in the room is talking, but the system can be playing audio from the remote end. Then, the virtual director can transition to showing a 2x2 layout where each participant occupies a 1-cell, and where each person occupies the same size in their cell.
[0324] After the introductions, the group can start discussing the topic. The second person on the left side can introduce the topic. As the vision pipeline detects the speaker, the virtual director can instruct the layout engine to go to a 3x3 layout where the speaker can occupy the first two rows in each column, and the other participants are in the bottom row.
[0325] After the topic has been introduced, the first person on the right can speak, and the system can transition again to the person occupying the largest cell.
[0326] After a short time, the second person on the right may say something. The vision pipeline can detect this, and considering the previous actions, the virtual director can transition to a dialogue setup and instruct the layout engine to transition to a 2x1 grid where the person on the right occupies one cell and the person on the left occupies one cell. The virtual director can take their gaze and head positions and can ensure that the sizes of the views are equal. Spacing can be added asymmetrically in front of where the people are looking.
[0327] After a short discussion in the room, one of the remote location participants can speak, and when the vision pipeline detects that the remote participant is speaking, it can maintain the previous layout. However, when each person is now looking at the screen, the vision pipeline can continue to evaluate the gaze of each person. For example, if they are all looking at the screen above the center camera, it will transition to showing the view from the center camera stream in two cells.
[0328] The discussion can return to the room, and a participant may want to demonstrate their idea and walk to the whiteboard at the back of the room. When the vision pipeline detects that the task is moving by seeing the speed of the trajectory associated with the person, the virtual director can follow the person, and the cell can be dedicated to following the person. The virtual director can instruct the layout engine to display a 2x2 grid where each person occupies one cell.
[0329] When the person reaches the whiteboard, they can start writing on the whiteboard. The vision pipeline can detect the activity on the whiteboard, and the virtual director can instruct the layout engine to change to a 3x3 layout where the stream from the canvas / whiteboard camera can occupy the first two rows of the first two columns, and the person writing on the whiteboard can be framed by the camera that best captures their face in the first two rows of the last column. Each of the other participants can occupy one cell in the bottom column using the camera that best sees their face.
[0330] The people on the whiteboard may have been talking for a few minutes, demonstrating their ideas, and the first person on the right may have a comment. Then they raise their hand so as not to interrupt the person at the whiteboard. When the vision pipeline detects that they may have raised their hand, the virtual director can maintain the same view until a specified duration has passed. Once the threshold is reached, the virtual director unit can instruct the layout engine that the person who raised their hand should occupy two cells in the first two rows of the last column, while the person on the whiteboard moves down to the bottom row.
[0331] When the group starts discussing the comment being presented, the person on the whiteboard can stop writing. When the vision pipeline detects that the whiteboard is no longer changing and a specified duration has elapsed, the virtual director can instruct the layout engine to return to a 2x2 layout, where the person on the whiteboard and one person on the left can each occupy a cell in the first column. The two people still sitting on the right can share the second column as they may have moved closer to better see the whiteboard.
[0332] When the discussion ends, the meeting can adjourn, and the vision pipeline can detect people getting up or waving. The vision pipeline unit can provide instructions to return to a center overview view, showing all the people in the room and them waving goodbye.
[0333] A second example can relate to an instance in a meeting room where three people (or meeting participants) are sitting around a table (Persons A, B, and C). The position of each person can be detected by the vision pipeline and sent to the virtual director unit. If no one is speaking, the virtual director unit can determine that a three-column layout where each person is roughly the same size is the best (ideal) layout. The virtual director unit can then instruct the layout engine to create a three-column layout accordingly (such as layout 2210b in Fig. 22 ), and provide the coordinates of which part of the main stream should be in each column. The coordinates of each column can correspond to the position and size of each person. The representation of Person A can be shown in the first column, the representation of Person B can be shown in the second column, and the representation of Person C can be shown in the third column. Then, when Person B starts speaking, the virtual director unit can determine that the representation of Person B should occupy a larger portion of the display. The virtual director unit can then instruct the layout engine to change the layout accordingly. Examples of the changed layout can include, but are not limited to, increasing the size of the column corresponding to Person B (such as layout 2210f in Fig. 22 ); transitioning to a three-column layout with two rows, where the representation of Person B occupies the first two columns, the representation of Person A occupies the top row of the last column, and the representation of Person C occupies the bottom row of the last column (such as layout 2240a in Fig. 22 ); and / or decreasing the size of the columns corresponding to Person A and Person C (such as layout 2210f in Fig. 22 ).
[0334] A third example can relate to a scenario in a small meeting room with three people (or meeting participants) sitting around a table (Person A, Person B, and Person C). Similar to the second example discussed above, the vision pipeline can detect each person, and the virtual director unit can instruct the layout engine to distribute the participants into three equally sized columns (such as Fig. 22in layout 2210b). Two people (Person A and Person B) can move closer to each other (e.g., read a document together). As a result of the movement, the representation of each person (Person A and Person B) in their corresponding columns may appear to be "cut in half" or otherwise obstructed. This may have a negative impact on the experience of remote participants as it may increase visual noise and be confusing. As used herein, a remote participant may refer to a participant located at the remote end of a meeting environment such that they may have difficulty directly viewing the participants in the room or a participant joining a video conference call remotely. Thus, to avoid this, the visual pipeline can monitor the overlapping regions of the detected people relative to each other and, once a defined limit is reached, can decide to merge their corresponding columns. In this example, Person A and Person B can be jointly framed as a group (such as Fig. 22 in layout 2210d), until they move far enough apart from each other to return to the initial layout.
[0335] A fourth example can relate to a scenario where 6 - 8 people are sitting around a table in a meeting room. The dynamic part of the meeting (e.g., speaking, interacting, active participants, reactions) can be centered around two individuals. For example, two people can be the main speakers leading the conversation. It is envisioned that the two individuals can stand up and move around the meeting room. Thus, the initial situation in this scenario can include an overview framing where all participants are framed together as a group. Once the virtual director unit identifies the individuals leading the conversation (e.g., by voice and / or movement frequency), the virtual director unit can instruct the layout engine to frame these individuals in separate tiles, as well as group - frame them with the remaining seated participants. Additionally, the specific room geometry may result in different layout options. For example, in some embodiments, based on the specific geometry of the meeting room, in addition to having their representations shown in their own tiles, the two individuals can still be part of a group framing.
[0336] A fifth example can relate to a scenario where 4 - 6 people (or meeting participants) are sitting around a table in a room. Multiple cameras (in a multi - camera system) can follow the rules mentioned in the previous examples to appropriately detect and frame the participants. Specifically, in this example, Person A can stand up, walk to the whiteboard, and start drawing. A camera with a specific visual pipeline pointing at the whiteboard can detect that someone is drawing on the whiteboard and send this information to the virtual director unit. A new stream with the whiteboard content can be made available to video clients (e.g., Microsoft Teams, Google Meet, Skype, Zoom), as Fig.26C shown in the example in. The video director unit can also combine the streams from multiple cameras in various layouts that show Person A writing on the whiteboard and the content of the whiteboard, not limited to Fig. 22 Those shown in
[0337] Continuing with the fifth example, in some embodiments, the scenario may require multiple cameras pointing at one or more whiteboards to capture everything. The virtual director unit can combine the streams from multiple whiteboard cameras and present them as one whiteboard stream. Using the vision pipeline, the virtual director unit can determine which areas of the one or more whiteboards are being actively drawn on or interacted with and combine the input streams into one whiteboard stream showing the active areas.
[0338] The sixth example can relate to a video client or other consumer of the output from the virtual director unit that has specific preferences for the types of streams they need or expect. For example, the video client may have more than 25 video streams and may want to show small profile pictures of each participant in a meeting. The consumer can request a "profile picture" stream of all participants from the virtual director unit. Using the vision pipeline, the virtual director unit can detect all participants and select the best camera stream for each person (meeting participant) and select the best camera stream for each person. The virtual director unit can also send multiple streams or a combined stream with a "profile picture" view of all participants to the client.
[0339] The seventh example can relate to a meeting environment that includes four people in a room (Person A, Person B, Person C, and Person D). Representations of Person A and Person B can each be shown in separate streams. Representations of Person C and Person D can be shown in a combined stream due to, for example, being in close proximity. Additionally or alternatively, a layout engine can provide a fourth stream that is an overview shot in which representations of Persons A, B, C, and D are all in the overview shot. The video client (or other consumer) can select which streams will be presented to remote participants.
[0340] The eighth example may relate to a multi-camera system installed in a meeting room suitable for eight people. The multi-camera system may include seven cameras: three cameras placed in the front of the room, one camera placed under the television (TV), one camera placed on the left side of the room, one camera placed on the right side of the room, and one camera attached to the whiteboard on the back wall. When the room is not in a meeting, the multi-camera system may remain inactive. However, four people may enter the room, two sitting on the right side of the table and two sitting on the left side of the client. The meeting participants may interact with the video client and start the meeting. Then, the video client may start consuming the video stream from the system, and the system may start. First, the vision pipeline may detect that there are four people in the room and the distances between them. Next, the virtual director unit may pick an overview shot from the most central camera to show an overview of the room to five remote participants. When the meeting starts, everyone in the room may introduce themselves. As the first participant (sitting on the right side of the table) starts speaking, the vision pipeline may detect that a person is speaking, and the virtual director unit may check how far apart the participants are. In this particular example, the virtual director unit may determine that each participant may have their own corresponding tile. Then, the virtual director unit may instruct the layout engine to transition to a 3x3 layout, and the speaker (the meeting participant who is speaking) may occupy the first two rows in all three columns. Each non-speaking participant may occupy one column in the last row.
[0341] Continuing with this example, the vision pipeline may also detect the gaze, head position, and body position of each meeting participant. Additionally, the vision pipeline may select the camera from which most of each person's face is visible. For the participant on the left side of the table who is looking at the person speaking, the right camera may capture a stream in which most of their face is visible. The vision pipeline may detect their position, and the virtual director may find a suitable frame based on their gaze, previous movement, and body size. Then, the virtual director unit may instruct the layout engine to frame the corresponding streams from different cameras. In this example, the streams from the right camera may be used to frame the two people on the left side, and the streams from the left camera may be used to frame the two people on the right side. Each frame may represent each person in the same size. This framing may occur before the virtual director applies any changes. If the virtual director unit determines that a person is looking in a different direction than the selected camera frame and enough time has passed since the previous change, the virtual director unit may change the camera feed in the tile of the corresponding person to the camera feed that shows most of the person's face that is visible.
[0342] Then, the vision pipeline can detect a second person who starts speaking. After a specified duration has elapsed (e.g., 5 seconds, 10 seconds, 30 seconds, 1 minute), the vision pipeline and the virtual director unit can implement a layout change such that the second person who is speaking occupies the top two rows in all columns, and the previous speaker transitions to a tile located in the bottom row. After all the meeting participants have introduced themselves, the people participating remotely can introduce themselves. The vision pipeline can detect that no one in the room is speaking, but the system may be playing audio from a remote location. The virtual director can then transition to showing a 2x2 layout, where each meeting participant occupies a cell. The representation of each meeting participant can occupy the same size in each cell.
[0343] Then, the group can start discussing the topic. The meeting participants can introduce the topic. The vision pipeline can detect that a meeting participant is speaking and can instruct the layout engine to transition to a 3x3 layout, where the representation of the speaker occupies the top two rows in each column, and each representation of the other meeting participants is shown in the bottom tiles. After the topic has been introduced, a second meeting participant can start speaking. The disclosed multi-camera system can transition to showing the representation of the second meeting participant as occupying the largest tile on the display. Then, the vision pipeline can detect a third person as speaking, and the virtual director can (taking into account previous actions) transition to a conversation setup, instructing the layout engine to transition to a 2x1 grid, where the person on the right side of the table occupies one tile, and the person on the left side of the table occupies the other tile. In some embodiments, the representation of the person on the right side of the table can occupy the right tile, and the representation of the person on the left side of the table can occupy the left tile. The virtual director unit can analyze the gaze and detect the heads of the meeting participants, and provide spacing in the tiles such that the framing size in each tile is equal (equal fair viewfinder, as discussed further below). The spacing can be added asymmetrically in front of the meeting participants' gaze direction (e.g., the direction they are looking).
[0344] Continuing with the example, after a short discussion in the room, one of the remote (e.g., remote / virtual) participants can start speaking. The vision pipeline can detect that the remote participant is speaking and maintain the previous layout. As each (in-person) meeting participant looks at the screen, the vision pipeline can continue to evaluate the gaze of each meeting participant. As an example, if all the meeting participants look at the screen above the center camera, the layout can transition to show the framing from the center camera stream in two tiles.
[0345] Then, the discussion can return to the (physical) room, and the meeting participants can walk to the whiteboard at the back of the room. The vision pipeline can detect the movement of the meeting participants by detecting the speed of the trajectories associated with the meeting participants. The virtual director unit can follow the meeting participants. The tiles corresponding to the meeting participants can follow the movement of the meeting participants. Additionally, the virtual director unit can indicate to the layout engine a 2x2 grid, where each representation of each meeting participant is displayed on a corresponding tile. The meeting participants can reach the whiteboard and start writing on the whiteboard. The vision pipeline can detect the activity on the whiteboard, and the virtual director unit can instruct the layout engine to change to a 3x3 layout, where the stream from the whiteboard camera occupies the top two rows of the first column. The meeting participant writing on the whiteboard can be framed by the camera that best captures their face, and the representation of the meeting participant can be displayed in the top two rows of the last column. The representation of each of the other meeting participants can occupy the tiles in the bottom column using the camera that best captures their face.
[0346] The meeting participant at the whiteboard can talk for a few minutes, presenting ideas. A second meeting participant can talk to discuss their comments. The second meeting participant can raise their hand so that they do not disturb the first meeting participant. The vision pipeline can detect that the second meeting participant has raised their hand, and the virtual director unit can maintain the same framing until a specified / specific duration (e.g., 1 second, 5 seconds, 10 seconds, 30 seconds, 1 minute) has passed. Once that duration has passed (or the time threshold has been reached), the virtual director unit can instruct the layout engine such that the second meeting participant (who raised their hand) is displayed on two tiles in the top two rows of the last column, while the representation of the first meeting participant is displayed in the bottom row.
[0347] Then, the group of meeting participants can start discussing the presented comment(s), and the first meeting participant at the whiteboard can stop writing. When the vision pipeline detects that there are no more changes on the whiteboard after a specified / specific duration, the virtual director unit can instruct the layout engine to return to displaying the 2x2 layout. The representation of the first meeting participant at the whiteboard and another meeting participant on the left side of the table can each occupy one tile in the first column. The representations of the two other meeting participants sitting on the right side of the table can be displayed together in the second column based on proximity.
[0348] Continuing with the example, when the discussion has ended, the vision pipeline can detect that the meeting participants are leaving, getting up, or waving. The vision pipeline can return to an overview framing from the center camera, showing all the meeting participants in the room and them waving goodbye.
[0349] Embodiments of the present disclosure may include features and techniques for providing an equal fair viewfinder. The equal fair viewfinder may provide an experience where everyone can contribute and participate at the same level in a meeting. This may be achieved by visually including everyone in the session.
[0350] In scenarios where one or a few people do most of the talking in a session, it may be desirable to provide remote participants with the context they need to follow the session and know what is going on in the room. To facilitate collaboration, it may be important that all meeting participants feel connected and have equal opportunities to contribute. This may be easier if participants (remote and in-person) can see the reactions and participation from everyone.
[0351] For example, if there are people in the meeting room who would be seen by remote participants if they were also in the physical meeting room, then the remote participants should also be able to see them via video conferencing. Occasionally switching from the speaker to the listener can improve the video conferencing experience and make it more engaging for remote participants. Additionally, in some embodiments, meeting participants who react to the speaker or events taking place in the meeting environment may be captured in, for example, reaction shots. Non-limiting examples of indicators of reactions include smiling, frowning, nodding, laughing, clapping, pointing, and raising his hand. It is envisioned that the views, displays, tiles, and other streams of displays may switch between the speaker, listener, and reacting person in any order / combination.
[0352] Rules and / or machine learning (ML) system training may indicate when it is appropriate to switch from the speaker shot to the listener shot. If there are more than two people in the room, there may also be several possibilities for the listener shot. An equal fair metric may be obtained, which may include a fairness score, to determine what the next listener shot will be.
[0353] The fairness score may be rated from 0 - 1. A fairness score closer to 1 may indicate a more even distribution of screen time among everyone in the room. The equal fair viewfinder may select the next listener shot based on who has the least amount of screen time.
[0354] Embodiments of the present disclosure may relate to the framing using a multi-camera system. To frame people (e.g., meeting participants), rules for lenses designed to produce good frames of the participants may be used. A larger number of people present in the room may result in greater difficulty in synthesizing good individual frames. In some embodiments, the disclosed systems and methods may focus on the head and upper body when synthesizing frames to capture body language and create a space for connection and understanding.
[0355] When a person is alone in a room or not sitting within a predetermined distance near other participants, the camera can frame them individually. The person's head can appear in the top third of the image to give the person as much space as possible and reduce the amount of unnecessary space on the screen. By positioning the head closer to the top of the image, the disclosed systems and methods can emphasize the visibility and presence of each person and can ensure that each person has enough space to move and act naturally.
[0356] If two or more people are sitting within a predetermined distance of each other, two or more people can appear in the same shot, regardless of whether any of them is a speaker or a listener. Group framing can ensure that the people in a group are shown in the best way, such as by creating a shot that includes the heads and upper bodies of all participants. If people move, the group framing shown during the session may change. As an example, if people are sitting such that there is only one group in the room, the shot shown or displayed can include that one group, unless someone moves.
[0357] Embodiments of the present disclosure may relate to methods of group framing in single-camera systems and multi-camera systems. In scenarios where participants are sitting close to each other in a meeting, challenges may arise in finding a suitable frame for a single speaker or listener. For example, if the framed view of a subject is cropped, the user experience may be diminished. Even if the framed shot includes one or more suitably framed subjects but also includes one or more cropped subjects, the user experience may also be negatively affected. Embodiments of the present disclosure generate a framed shot of the subjects that omits the cropped subjects.
[0358] In some embodiments, the identification of potential subjects for framing can be performed based on an analysis of the originally captured video frame (original frame). One or more sub-frames can be generated based on the original frame, and each sub-frame can include one or more subjects. For example, a subject framer can generate Frame A that includes two subjects, Frame B that includes one subject, and Frame C that includes one subject. The subject framer can avoid selecting and / or generating Frame D, where Frame D can include two subjects, with one of the subjects being cropped at the edge of Frame D.
[0359] Embodiments of the present disclosure can generate a framed shot according to a method in which participants (e.g., participants sitting closely together in a meeting) are grouped together to produce a shot / frame that focuses on a particular participant of interest while including adjacent participants in a visually pleasing shot.
[0360] For example, in some embodiments, the subject framing system can capture a video frame, detect the subject represented in the video frame, and generate a bounding box or other potential frame indicator for each detected subject. In some embodiments, the bounding box or frame indicator can be shaped as a rectangle centered on the head of the detected subject (e.g., horizontally centered, vertically centered, or both in the frame). The bounding box or frame indicator can also include a representation of at least a portion of the body of each detected subject. It is contemplated that the subject framing system can perform the operations discussed above in any order. As an example, the subject framing system can detect the subject starting from the left side of the meeting environment or room, and compare (one or more) detected subjects with the nearest subject on the right. The process can continue to the right. In other embodiments, the detection of the subject can start from the seating side of the meeting environment or room.
[0361] Detected subjects that do not overlap in potential framed shots can be shown separately in the framed shot. In other embodiments, detected subjects that are at least partially overlapping (particularly within the potential framed shot) can be grouped together in the same framed shot. As an example, if two subjects (people) overlap visually relative to the camera viewpoint (for example, one subject / person sits slightly behind the other subject / person), they can be grouped. As another example, a subject (person) visible in the (person's) character viewfinder (for example, the potential subject picture) of another subject, the two subjects can still be grouped together even if they do not overlap. Such grouped pictures can be selected, for example, where separating non-overlapping subjects into individual pictures may result in an undesirable picture representation of at least one of the objects. Non-limiting examples of such undesirable representations can include when the height of the head box located at the maximum vertical and horizontal edges of the subject's head is greater than a predetermined percentage of the vertical size of the video picture.
[0362] In some embodiments, certain situations may be encountered in which two subjects do not overlap in the overview screen, and a separate sub-screen may be placed relative to each subject so that none of the sub-screens displays a cropped subject. The subject shown in one or more of the sub-screens may still provide an undesirable screen representation. For example, the position or relative size of one or more of the subjects in the corresponding sub-screens may appear unbalanced, lacking sufficient buffer space around the subject, over-enlarged, etc. In the presence of the above conditions and other undesirable screen representation conditions, embodiments of the present disclosure may group non-overlapping subjects together in a single sub-screen. This situation may occur, for example, in a conference room including two or more people.
[0363] Some embodiments of the present disclosure may employ exceptions to the methods outlined above. For example, in the case of different head sizes (e.g., when one subject is sitting further away from the camera relative to another subject), the system may separate the subjects into more than one frame. This method of separating an object into individual frames may be based on a predetermined threshold relative to head size, shoulder width, distance to the camera, or other spatial metrics. As an example, if the system determines that a first generated head box associated with a first subject is at least three times the width of a second head box generated relative to a second subject, the first and second subjects may be separated into different frame shots. Separating the subjects in this way can achieve a closer view of the subject located further from the camera, which can significantly improve the user experience (compared to a frame in which both the near object and the far object are shown together in a common frame). In some embodiments, the representations of the first and second meeting participants in the first and second focused video streams may be sized respectively by a video processing unit such that the corresponding sizes or dimensions associated with the representations of the first and second meeting participants in the first and second focused video streams satisfy a size similarity condition. For example, the size may be the head size (e.g., head diameter), and the similarity condition may be a head size similarity condition (e.g., the head sizes must be within a predetermined ratio or rate). In other words, the video processing unit may adjust the zoom levels of the two focused video streams to reduce the size difference between the representations of the first and second meeting participants, which may be caused by the meeting participants sitting at different distances from the camera. In other embodiments, the video processing unit may analyze the video output stream from the camera to determine the distances to the camera associated with the first and second meeting participants. The video processing unit may then generate focused streams and cause the display to show the focused streams in corresponding display tiles, the focused streams being a first focused video stream characterized by the representation of the first meeting participant and a second focused video stream characterized by the representation of the second meeting participant, where the size of the representation of the meeting participant is set to at least partially compensate for the difference in the distances of the first and second meeting participants from the camera. The above resizing has the effect that the same or a similar level of detail and information about each meeting participant present in the meeting environment can be provided to remote meeting participants.
[0364] Various systems may require a transition from a frame in which one object is shown separately to a frame in which two (or more) subjects are shown together. For example, as Fig.33As shown, a first subject (e.g., a speaker on the left side of the main screens 3310a, 3320a, 3330a) can be featured alone in a sub - screen, as shown in sub - screen 3310b. If a second subject moves closer to the first subject, the disclosed embodiments can generate a new sub - screen in which the first and second subjects are represented together. An example of a transitional sub - screen when the second subject moves closer to the first subject is shown as sub - screen 3320b derived from main screen 3320a. A new sub - screen 3330b can be generated and displayed, showing both the first and second subjects, and sub - screen 3330b can be derived from main screen 3330a.
[0365] Similarly, in some embodiments, the disclosed systems and methods can transition from a sub - screen including two or more subjects to a sub - screen including one (single) subject. Non - limiting examples of this can include when the subjects move away from each other or when one subject moves closer to the camera relative to another subject. In some embodiments, the transition from one sub - screen shot to another can be associated with a dwell time (e.g., 2.5 seconds), such that the dwell time must elapse at least before transitioning to a new sub - screen shot. The dwell - time constraint can be changed / altered or ignored in cases where the subject movement may require a new sub - screen shot.
[0366] As described above, a person viewfinder can refer to a viewfinder (or screen) focused on one person (or meeting participant). A person viewfinder can be created for each person detected within the field of view of a camera (or cameras). As Fig.34 shown, a meeting room with one person viewfinder is disclosed, as Fig.34 shown in output 3410. Additionally, in some embodiments, a meeting room can include four meeting participants, and each meeting participant can correspond to (or have) a person viewfinder, as shown in output 3420. Consistent with the disclosed embodiments, a person viewfinder output 3430 can be generated and displayed on a display, for example, within a tile in a layout.
[0367] Furthermore, in some embodiments, most or all of the people in a room can be viewed in a single shot (e.g., an overview view). For example, when a remote participant is speaking, all participants in a physical meeting environment can be shown such that they are all visible and can be seen in context.
[0368] In addition, in some embodiments, the stream can start with an overview shot for 20 seconds (or any other suitable duration). This can ensure that remote participants can orient themselves in the scene, see everyone in the meeting, and receive context. If there is any large movement within the field of view of the camera, such as someone entering the room, or someone getting up from a chair and / or moving around, the virtual director can select the overview shot as the next shot. This can happen regardless of any other framing selection (e.g., if someone else is still talking or it is time for a listening shot). This can ensure that remote participants receive context about the meeting environment.
[0369] If no one within the field of view of the camera is speaking (e.g., if one of the remote participants is speaking), the system can output an overview shot for display. This can allow remote participants to keep track of everyone they are speaking to and see the faces and reactions of the participants who are listening. This makes for a more comfortable setting for remote participants to speak. In some embodiments, the overview shot includes everyone within the field of view of the camera.
[0370] Full framing can involve a shot that utilizes the entire field of view of the camera and can be used to establish context when there is a lot of movement. This can ensure that remote participants can follow what is happening in the room when someone enters or gets up from their chair.
[0371] More generally, the system can be configured to display an overview shot of the meeting environment in response to detecting an overview trigger event. The overview shot can be shown for at least a predetermined time interval, such as 5 to 30 seconds, 10 to 30 seconds, or about 20 seconds. As described above, the overview trigger event can be the start of a video conference conducted using the video conferencing system. This provides context to remote participants about the physical meeting environment. Similarly, the overview trigger event can be a meeting participant who is located remotely with respect to the meeting environment (i.e., a remote meeting participant) joining a video conference conducted using the video conferencing system. In such a case, the overview shot can be shown on the display used by the remotely located meeting participant to participate in the video conference, e.g., on the display screen of the user device used by the remote meeting participant to participate in the video conference. The overview trigger event can be a speech by at least one meeting participant who is located remotely with respect to the meeting environment. For example, if the system determines that a remote meeting participant is the speaker, e.g., based on an audio signal from a microphone associated with the remote meeting participant, the overview shot can be shown on the display used by the remotely located meeting participant to participate in the video conference so that the remote meeting participant can understand the context and reactions of all the meeting participants present in the meeting environment.
[0372] In some embodiments, the overview trigger event can be a movement trigger event associated with the movement of at least one individual (e.g., a meeting participant) present in the meeting environment. Triggering the display of an overview shot when specific movement conditions are met ensures that important information about the physical meeting environment and the physical movement of meeting participants present in the meeting environment is provided to remote meeting participants. The movement trigger event can be an individual within the meeting environment moving from one location within the meeting environment to another location within the meeting environment, transitioning between a sitting and a standing position, entering or leaving the meeting environment. Alternatively, the movement trigger event can be that an aggregated movement level, score, or value associated with multiple individuals present in the meeting environment exceeds a predetermined movement threshold. Again, the overview shot can be shown on a display used by a remote meeting participant to participate in a video conference, e.g., on a display screen of a user device used by a remote meeting participant to participate in a video conference. In fact, throughout this disclosure generally, the display on which various video outputs are shown can be the display used by a remote meeting participant to participate in a video conference. The overview video stream is preferably characterized by an individual whose movement is associated with the trigger event such that remote participants can observe the movement.
[0373] Accordingly, the video processing unit can be configured to analyze the video output from one or more video conference cameras to determine whether an overview trigger event has occurred. In other words, the video processing unit can be configured to analyze the video output from one or more video conference cameras to identify such a trigger event and then show an overview shot if such a trigger event is detected.
[0374] The rules for when the speaker view switches from one shot to another can be based on multiple objectives, such as but not limited to, wanting to capture the flow of the conversation, while creating an engaging experience for remote participants and ensuring that the in-room experience is as inclusive as possible. The speaker view can select the best shot type based on different conditions.
[0375] For example, whenever the speaker view is certain that a person is speaking, that person can be considered the speaker and can be framed individually or as part of a group. This can be applied to any scenario where there is a new or existing speaker.
[0376] If a person has been speaking for more than a given number of seconds (e.g., approximately 7 - 15 seconds), a listening shot can be shown or displayed. The next listening shot can be for someone who has been given the least amount of screen time. If no one is speaking, the camera can output a context shot (e.g., an overview shot).
[0377] In addition, in some embodiments, speaker framing can deliver a more dynamic meeting experience that feels closer to being together in the same room to allow remote participants to see who is speaking and feel more included, to help remote participants follow the conversation and know what is going on in the meeting environment, to make it easier for remote participants to be an active part of the conversation by providing a better view of the speaker and listener for a greater sense of context, and to create a more inclusive experience by ensuring that everyone in the room is visually present in the conversation regardless of whether they are speaking.
[0378] In some embodiments, when a task starts speaking and the virtual director has sufficient confidence that it is the speaker, the person can be framed as the speaker. Additionally or alternatively, as long as the person continues to speak, they can be considered the speaker. The speaker shot can be at least 3 seconds long to allow the virtual director sufficient time to have sufficient confidence that someone else (another person) is the new speaker and to provide a more stable and enjoyable experience for any remote participants. The speaker shot can include a person, i.e., the speaker, or a group of people including the speaker.
[0379] In some embodiments, the frame can be updated three times per second. The virtual director can check the audio input, detections, and the rules for which frame it should select as the next shot. This information can be stored over time and a history created so that future decisions are based on that history. Reframing lock can involve the minimum duration for which a frame can be shown or presented. For example, the reframing lock can be 2.5 seconds, meaning that any new frame must be shown for 2.5 seconds. The virtual director can also check the movement of the participants' heads and bodies.
[0380] In some embodiments, if a speaker has been speaking for more than a given number of seconds (e.g., 8 seconds, 9 seconds, or 10 seconds), the virtual director can look for the next listening shot. The next listening shot can include the person who has received the least amount of screen time to ensure that everyone is visually included, expressions, and to create an understanding of the reactions for remote participants. The listening shot can include one or more people who are not speaking. In some embodiments, the listening shot can be shown for 3 seconds.
[0381] Embodiments of the present disclosure may include features and techniques for improving transition and framing methods in a video conferencing system. Framing methods (e.g., speaker framing, listener framing, etc.) may use two or more types of transitions between shots: smooth transitions and hard cuts. The type of transition used may depend on the shot to transition to and the degree of difference from the shot being transitioned from. If there are only minor differences between two shots, a smooth transition may be used. In some embodiments, a smooth transition may be used when transitioning towards an overview shot. A hard cut may be used when there are significant differences between the shots (such as when the view switches from one side of a table to the other).
[0382] Additionally or alternatively, various types of transition types may be employed when transitioning from an initial frame (e.g., any one of a speaker frame, a subject / object framed by a genius view, a video tile framed by a gallery view, etc.) to a target frame.
[0383] For example, a sudden change between an initial frame and a target frame separated by a large difference in camera parameter values may distract the user. In some embodiments, a multi-camera system may provide a smooth transition from the initial frame to the target frame by non-linearly changing at least one camera parameter value (e.g., zoom, pan, etc.) across three or more frames (e.g., of a primary video stream).
[0384] The number of frames (e.g., transition time) included between the initial frame and the target frame may vary based on the characteristics of the initial frame and the target frame. For example, the transition time may vary based on the direction of the planned camera parameter change (e.g., zoom out vs. zoom in) or based on the magnitude of the planned change (e.g., a small change may be associated with a longer transition time).
[0385] The disclosed embodiments may identify a new target frame (e.g., in response to a newly detected condition or target frame trigger, such as a person entering the room, etc.) before completing an ongoing transition. Instead of completing the ongoing transition before transitioning to the new target frame, the ongoing transition may be changed or adjusted. For example, the deceleration phase of the current transition may be omitted, the acceleration phase of the next transition may be omitted, and / or the current transition rate of the ongoing transition may be matched to the initial transition rate of the planned transition.
[0386] Some of the above features have been described with respect to a single camera system for an exemplary system. However, it should be noted that the same principles can also be applied to multi-camera systems and setups. For example, any of the framing methods, transition methods, speaker views, gallery views, overview shots, group shots, and other features and techniques discussed herein can be employed in a multi-camera system. Additionally, any of the framing methods, transition methods, speaker views, gallery views, overview shots, group shots, and other features and techniques described with respect to a multi-camera system can be employed in a single camera system.
[0387] Additionally or alternatively, group framing / group shots can be performed with or without speaker framing / AI. For example, in some embodiments, the speaker framing techniques described above can be employed by a system including two, three, or more cameras. Additionally, it can be shown with a hard cut demonstration, gallery view, or dynamic layout similar to speaker framing. Thus, it is possible on both single-camera and multi-camera setups.
[0388] As an example, the subject / person framing step can be performed by two or more of the cameras included in a multi-camera group. Since each camera can have a unique field of view, pointing direction, zoom level, and / or focal length, a sub-frame of a particular subject generated or enabled based on the output of a first camera can be more preferable than a sub-frame of the same subject generated or enabled by a second camera. Due to the difference in the field of view between the two cameras, a subject can be framed separately based on the output of the first camera, while the same subject can be framed together with a second subject based on the output of the second camera. In some embodiments, a frame showing a single subject can be preferred over a frame showing multiple subjects (e.g., in the case of determining that a single subject is speaking, or in the case of desiring to focus on or highlight a single subject without showing other subjects). In some embodiments, such as in the case of a conversation occurring between subjects, it may be preferred to show multiple subjects together rather than splitting the subjects between sub-frames.
[0389] Furthermore, in some embodiments, the disclosed embodiments can transition between a sub-frame from a first camera showing a single first subject and a sub-frame from a second camera showing the first subject and at least one other subject (and vice versa).
[0390] The disclosed systems and methods can also transition between a sub-frame from a first camera showing a single first subject and a sub-frame from a second camera showing a single second subject (and vice versa). Such a system can be useful in cases where the first camera cannot capture the face of the second subject and / or the second camera cannot capture the face of the first subject.
[0391] In some embodiments, a multi-camera system can include dedicated cameras (e.g., together with dedicated microphones) for each seat position in a venue, where each dedicated camera can be used to generate a sub-frame representing a single subject. These dedicated cameras can be combined with one or more overview cameras capable of generating sub-frames showing multiple subjects together.
[0392] Fig.15 Examples of different lens types are provided. As Figure 15 shown, the total lens frame 1510 can include most or all of the meeting participants and most of the room. The total lens frame 1510 can provide context for the scene (e.g., the meeting environment) and an understanding of the location of the room as well as the people and objects within it. This lens can frame the heads of most of the participants in the top third of the lens and add an equal amount of space between the leftmost and rightmost people at the edges of the frame to fit the requested aspect ratio from the video stream, to give sufficient visual space and to fit the aspect ratio of the video stream.
[0393] The medium lens frame 1520 can include the representation of two to three participants and focus on one participant. The medium lens frame 1520 can be used when focusing on a conversation, session, or speaker. The system can frame the person who is typically speaking in the foreground of the lens and align the head of the speaker with the heads of the other participants in the top third of the lens. Padding can be added in the direction the speaker is looking to bring the lens to the correct aspect ratio and provide sufficient visual space.
[0394] The close-up lens frame 1530 can frame only one person. The close-up lens frame 1530 can be used to focus on one participant who is conversing over a long period or duration. The close-up lens frame 1530 can be employed after the medium lens frame 1520, where the same person is included in both lens frames 1520 and 1530. In the close-up lens frame 1530, the eyes of the participant can be aligned with the top third line, and the close-up lens frame 1530 can show the upper body shoulders / chest of the participant. In some embodiments, the participant can be framed slightly off-center based on the viewing direction rather than in the center of the frame. For example, a person looking to the right can be placed off-center to the left of the close-up lens frame 1530. Additional space can be added in the direction the person is looking and in the area behind the person to bring the frame to the desired aspect ratio.
[0395] Figures 16A - 16B Examples of lens frames including other meeting participants and auxiliary items are provided. The virtual director unit can continuously evaluate all possible lenses for each participant and select clean lenses. The selection can be determined by a set of rules, including not selecting lenses where a person is partially visible, as Figure 16AAs shown. Such as Figure 16A As shown, the virtual director unit may prefer a clean medium shot 1610a in which both Person A and Person B are fully visible over a close-up shot 1610b of Person A in which Person B is partially visible. In some embodiments, the virtual director unit may add additional padding in the direction the person is looking, thereby providing a representation of the participant space to be viewed. This may also be the case for medium and full shots where there may not be enough participants to fully fill the shot. Such as Figure 16B As shown, in some embodiments, items of interest (such as a whiteboard) may also be considered in the evaluation. Such as Figure 15 As shown in B, the virtual director unit may prefer a shot 1620a that includes the entire auxiliary item (e.g., whiteboard) over a shot 1620b that includes a portion of the auxiliary item.
[0396] The disclosed embodiments may include interesting shot frames. The interesting shot frames may include items or people that the vision pipeline determines are of interest in the context of the scene. This may be the item / person that everyone is looking at, or a classified item identified from the sound and video. The item may be framed according to any of the principles disclosed herein. Additionally, in some embodiments, a close shot may be used to frame the item / person.
[0397] The disclosed embodiments may include listening shots. The listening shots may include shots that frame one or more of the participants who are not speaking. This type of shot may be a medium shot or a close shot, depending on whether one or more participants are to be framed. The system may use these shots in different situations, such as when the active speaker has spoken for a predetermined time or duration. Additionally, in some embodiments, if the vision pipeline determines that someone is looking away, looking bored, looking at the table, or has not been framed for a long time, the shot may switch to a listening shot. This may give the participants viewing the video stream from the room an equal opportunity to understand the participation of the other participants in the meeting room.
[0398] The listener shot can be shown together with one or more speaker shot video streams (as well as gallery shots and / or overview video streams) such that remote meeting participants are provided with information about the reactions of non-speaking (listening) meeting participants to the speaker. Selection of non-speaking meeting participants to be included in the listener shot video stream can be based on one or more listener participation characteristics of the non-speaking meeting participants, such as determined according to an analysis of the video and / or audio output streams provided to the video processing unit by the video processing unit. The listener participation characteristics can be as previously described and can include one or more of the following: reactions to the speaking meeting participant, posture changes, gaze direction (e.g., gaze direction towards the speaking meeting participant), facial expressions or changes in facial expressions, head movements (e.g., nodding or shaking of the head), applause, or raised arms or hands. Other body movements can also be used as listener participation characteristics.
[0399] The listener shot video stream featuring non-speaking meeting participants can be displayed side-by-side or simultaneously with one or more speaker shot video streams, e.g., where each video stream is displayed in a corresponding display tile. Alternatively, the listener shot and the speaker shot can be shown at different times, e.g., the listener shot can be shown alternately with one or more speaker shot video streams. The speaker shot can be shown for a longer duration than the listener shot. For example, the speaker shot can be shown for a duration of about 7 - 10 seconds, and the listener shot can be shown for a duration of about 1 - 3 seconds. Providing the remote meeting participants with the non-speaking meeting participants characterized by these characteristics gives information about the relevant reactions of the listeners that they would otherwise lack.
[0400] The disclosed embodiments can include a presenter shot. The presenter shot can be focused on the presenter and optionally on one or more objects with which the presenter is interacting, such as a classroom presenter or a presenter in a boardroom scenario. The speaker shot can be characterized by the speaker without other meeting participants, i.e., only the speaker. For example, in some embodiments, in a case where one participant is talking for most of the meeting, the system can add a presenter shot and a listener shot. These shots can be variations of close-up or medium shots that only show the presenter, but use different camera angles and composites to give a variation in the video and prevent it from feeling static.
[0401] Embodiments of the present disclosure may relate to features and techniques for providing different types of user experiences based on the type of meeting environment. For example, in a conference room, a virtual director unit may use the overall shot of the most central camera at the start of the stream to create an understanding of the context, the room, and the visual relationships among the people in the room. After a predetermined time, the virtual director unit may switch to a camera with the best view showing a group of people in the room (using a medium shot). In some embodiments, the best view may include the left and / or right side of the table. Then, the virtual director unit may frame the person who is speaking through the camera, and the camera may best see the speaker from in front of their face (using a medium shot). If the speaking person talks for longer than a predetermined time, the virtual director unit may switch to using a camera that can best see the selected listening or reacting person from in front of their face to frame the other people in the room who are listening (using a listening shot) or reacting (using a reaction shot). If no one is speaking or the virtual director unit determines that the voice is from an artificial sound source (e.g., a speaker), the virtual director unit may switch to using a camera to frame most or all of the participants in the room, and the camera can best see all the participants from the front (using an overall shot).
[0402] As another example, in a classroom, the parameters in the virtual director unit and the visual pipeline may be adapted to a presenter scenario. The presenter scenario may be employed in a classroom or a lecture meeting environment where one person talks for most of the meeting (e.g., more than half of the meeting duration), and the audience ...
Claims
1. A video conferencing system, comprising: At least one camera, each camera being configured to generate an overview video output stream representing a conference environment; And At least one video processing unit, configured to: Automatically analyze at least one overview video output stream to identify at least a first conference participant and a second conference participant represented in the at least one overview video output stream; Based on one or more determined characteristics associated with at least one of the first conference participant or the second conference participant, assign a first role name to the first conference participant and a second role name to the second conference participant; Based on the priorities associated with the first role name and the second role name, determine a relative framing priority associated with the first conference participant and the second conference participant; Based on the relative framing priority, generate one or more main video streams featuring the first conference participant and / or the second conference participant as output.
2. The system according to claim 1, wherein Generating one or more main video streams featuring the first conference participant and / or the second conference participant as output based on the relative framing priority includes: Based on the relative framing priority, select a framing representation of the first conference participant or a framing representation of the second conference participant; and Generate a main video stream featuring the selected framing representation of the first conference participant or the second conference participant as output.
3. The system according to claim 2, wherein, The video processing unit is configured to select the framing representation of the conference participant associated with the higher relative framing priority.
4. The system according to claim 2 or 3, wherein, The selected framing representation is the framing representation of the second conference participant, and wherein the at least one video processing unit is configured to: Based on the first role name assigned to the first conference participant, determine the likelihood that the first conference participant will start speaking; In response to the likelihood that the first conference participant will start speaking satisfying a likelihood condition, perform one or more actions to prepare the video processing unit to output a main video stream featuring the framing representation of the first conference participant; Based on the analysis of the at least one overview video output stream, determine whether the first conference participant has started speaking; and In response to determining that the first conference participant has started speaking, output the main video stream including the framing representation of the first conference participant for display.
5. The system according to claim 4, wherein, Performing one or more actions to prepare the video processing unit to output a main video stream featuring the first conference participant includes one or more of the following: Select the framing representation of the first conference participant of the main video stream; and Generate the main video stream featuring the framing representation of the first conference participant.
6. The system according to claim 4 or claim 5, wherein Determining the likelihood that the first conference participant will start speaking is also based on one or more determined quantities characterizing the session dynamics between the first conference participant, featured in one or more video output streams, and one or more other conference participants.
7. The system according to claim 6, wherein, The one or more determined quantities include one or more determined vectors, and the one or more determined vectors represent one or more session exchanges between the first conference participant characterized in the one or more video output streams and one or more other conference participants.
8. The system according to claim 7, wherein Each vector represents a speaking episode during which either the first conference participant or one of the one or more other conference participants is determined to have spoken.
9. The system according to claim 8, wherein, The magnitude of each vector represents the duration of the speaking episode, and the direction of each vector represents the next conference participant to speak.
10. The system according to any one of claims 6 to 9, wherein, The one or more determined quantities include at least one of the following: The identities of the one or more other conference participants; One or more durations characterizing one or more episodes during which the first conference participant or one or more other conference participants are determined to have spoken; One or more frequencies characterizing the frequency of episodes during which the first conference participant or one or more other conference participants are determined to have spoken; One or more time periods characterizing one or more gaps between episodes during which the first conference participant or one or more other conference participants are determined to have spoken.
11. The system according to any one of claims 6 to 10, wherein, Determining the likelihood that the first conference participant will start speaking is also based on one or more determined properties associated with the first conference participant.
12. The system according to claim 11, wherein, The one or more determined properties include at least one of the following: the time since the participant last spoke, the frequency of the participant's speech, head turning in the direction of the speaker, gestures, facial expressions, changes in facial expressions, nods, gaze direction, non-verbal sounds, and inhalations.
13. The system according to claim 1, wherein, Generating one or more primary video streams characterized by the first conference participant and / or the second conference participant as output based on the relative framing priorities includes: Generating a first primary video stream and a second primary video stream as output, the first primary video stream including a framed representation of the first conference participant during a first time period, and the second primary video stream including a framed representation of the second conference participant during a second time period; wherein the relative durations of the first time period and the second time period are determined based on the relative framing priorities associated with the first conference participant and the second conference participant.
14. The system according to claim 13, wherein The framing priority associated with the first conference participant is higher than the framing priority associated with the second conference participant; and wherein the duration of the first time period is longer than the duration of the second time period.
15. The system according to claim 13 or claim 14, wherein, The first time period and / or the second time period is / are discontinuous.
16. The system according to any one of the preceding claims, wherein, One or more detected characteristics include one or more of the following: whether a participant is speaking, the length of time a participant speaks, the percentage of the duration of a participant's speech, whether a participant is standing, the viewing direction of a participant, changes in the viewing direction of a participant, gestures performed by a participant, or reactions performed by a participant.
17. The system according to any one of the preceding claims, wherein, The at least one video processing unit is further configured to: assign observed attributes to each of the first conference participant and the second conference participant based on an analysis of the at least one overview video output stream.
18. The system according to claim 17, wherein The observed attributes include demeanor; optionally, where the demeanor includes at least one of serious, humorous, questioning, listening, or disengaged.
19. The system according to claim 17 or claim 18, wherein, The relative framing priority associated with the first conference participant and the second conference participant is determined based on the priority associated with the assigned observed attributes.
20. The system according to claim 19, wherein The priority associated with at least one of the observed attributes is user-assignable.
21. The system according to any one of the preceding claims, wherein, The at least one video processing unit is further configured to assign observed actions to each of the first conference participant and the second conference participant based on an analysis of the at least one overview video output stream; optionally, where the observed actions include at least one of standing, sitting, raising a hand, interacting with a whiteboard, applauding, speaking, sleeping, writing, or laughing.
22. The system according to claim 21, wherein, The relative framing priority associated with the first conference participant and the second conference participant is determined based on the priority associated with the assigned observed actions.
23. The system according to claim 22, wherein, The priority associated with at least one of the observed actions is user-assignable.
24. The system according to any one of the preceding claims, wherein, The at least one video processing unit is configured to: receive a user-assigned priority associated with the first conference participant; and determine the relative framing priority based on the user-assigned priority associated with the first conference participant.
25. The system according to any one of claims 1 to 23, wherein, The at least one video processing unit is configured to: receive a user-assigned role name associated with the first conference participant and a user-assigned priority associated with the user-assigned role name; assign the user-assigned role name to the first conference participant; and determine the framing priority of the first conference participant based on the user-assigned priority associated with the user-assigned role name.
26. The system according to any one of the preceding claims, wherein, The at least one camera includes a plurality of cameras; optionally, where the at least one video processing unit is further configured to automatically analyze the overview video output stream from each of the plurality of cameras and track and correlate the representations of conference participants across the overview video output streams based on at least one identifier.
27. The system according to claim 26, wherein, The at least one video processing unit is included in one of the plurality of cameras.
28. The system according to any one of the preceding claims, wherein, The at least one video processing unit includes one or more microprocessors on-board the camera.
29. The system according to any one of the preceding claims, wherein, The role assigned to each of the first conference participant and the second conference participant based on an analysis of the at least one overview video output stream is updated periodically.
30. The system according to any one of the preceding claims, wherein, Role names include presenter role names, contributor role names, audience role names, and / or user-assigned role names.
31. The system according to any one of the preceding claims, wherein, The at least one video processing unit is configured to implement at least one trained network, and the at least one trained network is configured to output the role names of the identified conference participants based on an input including one or more captured video frames from the at least one overview video output stream.
32. The system according to any one of the preceding claims, wherein, The at least one video processing unit is configured to cause at least one display to show the one or more main video streams.
33. The system according to any one of the preceding claims, further comprising at least one directional audio unit.
34. The system according to any one of the preceding claims, wherein, The main video stream output includes a speaker shot, a presenter shot, an overview shot, a reaction shot, an over-the-shoulder shot, or a group shot.
35. The system according to any one of the preceding claims, wherein, The video processing unit is configured to transition between main video stream outputs using hard cut transitions and / or smooth cut transitions.
36. A video processing method performed by at least one video processing unit, the method comprising: automatically analyzing at least one overview video output stream representing a conference environment generated by at least one camera to identify at least a first conference participant and a second conference participant represented in the at least one overview video output stream; assigning a first role name to the first conference participant and a second role name to the second conference participant based on one or more detected characteristics associated with at least one of the first conference participant or the second conference participant; determining a relative framing priority associated with the first conference participant and the second conference participant based on a priority associated with the first role name and the second role name; and generating, based on the relative framing priority, one or more main video streams featuring the first conference participant and / or the second conference participant as an output.
37. A video conferencing system, comprising: at least one camera, each camera being configured to generate an overview video output stream representing a conference environment; and at least one video processing unit, configured to: automatically analyze the at least one overview video output stream to identify at least a first conference participant represented in the at least one overview video output stream; assign a first role name to the first conference participant based on one or more detected characteristics associated with the first conference participant; determine the likelihood that the first conference participant will start speaking based on the first role name assigned to the first conference participant; in response to the likelihood that the first conference participant will start speaking satisfying a likelihood condition, perform one or more actions to cause the video processing unit to prepare to output a main video stream featuring a framed representation of the first conference participant for display; determine whether the first conference participant has started speaking based on an analysis of the at least one overview video output stream; and in response to determining that the first conference participant has started speaking, output the main video stream featuring the framed representation of the first conference participant for display.
38. A video processing method performed by at least one video processing unit, the method comprising: Automatically analyze at least one overview video output stream representing a meeting environment generated by at least one camera to identify at least a first meeting participant represented in the at least one overview video output stream; Based on one or more detected characteristics associated with the first meeting participant, assign a first role name to the first meeting participant; Based on the first role name assigned to the first meeting participant, determine the likelihood that the first meeting participant will start speaking; In response to the likelihood that the first meeting participant will start speaking satisfying a likelihood condition, perform one or more actions to cause the video processing unit to prepare to output a main video stream characterized by a framed representation of the first meeting participant for display; Based on the analysis of the at least one overview video output stream, determine whether the first meeting participant has started speaking; and In response to determining that the first meeting participant has started speaking, output the main video stream characterized by the framed representation of the first meeting participant for display.
39. A video conferencing system, comprising: At least one camera, each camera being configured to generate an overview video output stream representing a meeting environment; And At least one video processing unit, configured to: Automatically analyze the at least one overview video output stream to identify a plurality of meeting participants characterized in the at least one overview video output stream; Identify a plurality of non-speaking meeting participants among the identified meeting participants; Based on one or more detected characteristics associated with each non-speaking meeting participant among the non-speaking meeting participants, assign a corresponding role name to each non-speaking meeting participant among the identified non-speaking meeting participants; and Generate a main video stream characterized by one or more non-speaking meeting participants among the non-speaking meeting participants as an output for display, wherein the selection of the one or more non-speaking meeting participants to be included in the main video stream is based on the roles assigned to each non-speaking meeting participant among the non-speaking meeting participants.
40. The system according to claim 39, wherein, The at least one video processing unit is configured to identify speaking meeting participants among the identified meeting participants; and wherein the main video stream is characterized by the speaking meeting participants.
41. A video processing method performed by at least one video processing unit, the method comprising: Automatically analyze at least one overview video output stream representing a meeting environment generated by at least one camera to identify a plurality of meeting participants characterized in the at least one overview video output stream; Identify a plurality of non-speaking meeting participants among the identified meeting participants; Based on one or more detected characteristics associated with each non-speaking meeting participant among the non-speaking meeting participants, assign a corresponding role name to each non-speaking meeting participant among the identified non-speaking meeting participants; and Generate a primary video stream characterized by one or more of the non-speaking conference participants as output for display, wherein the selection of the one or more non-speaking conference participants to be included in the primary video stream is based on the roles assigned to each of the non-speaking conference participants.
42. At least one video processing unit configured to perform the method according to any one of claims 36, 38, and 41.
43. One or more computer-readable storage media storing instructions that, when executed by at least one video processing unit, cause the at least one video processing unit to perform the method according to any one of claims 36, 38, and 41.
44. The system according to any one of claims 1 to 35, 37, 39 or 40, wherein, The camera system is a multi-camera video conferencing camera system, and wherein the conference environment is a meeting room, office space, classroom, and / or lecture hall.