Automatic video conference framing using multiple cameras

By using multiple cameras for real-time calibration and dynamic mapping in the video conferencing system, the problem of poor frame formation by a single camera was solved, enabling flexible and efficient frame formation for multiple participants and improving the user experience.

CN121002840APending Publication Date: 2025-11-21HEWLETT PACKARD DEVELOPMENT COMPANY LP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380096589.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-03-30
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

In existing video conferencing systems, a single camera struggles to effectively frame multiple participants, especially since distant participants appear tiny in the image, and manually adjusting camera pan, tilt, and zoom is inconvenient.

Method used

Employing a multi-camera system, including at least one main camera and at least one auxiliary camera, the system optimizes framing of participants of interest through real-time or offline calibration, dynamic mapping, and automatic switching of video streams.

Benefits of technology

It enables effective framing of multiple participants in video conferencing, especially those at a distance, improving the user experience and providing flexible video layout options and automated framing adjustments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121002840A_ABST
    Figure CN121002840A_ABST
Patent Text Reader

Abstract

A method is described in which a first video stream of a scene is received (602) from a first camera (102) and a second video stream of the scene is received (604) from a second camera (104a, 104b). A participant of interest is identified (606) in a first video stream and mapped (608) from the first video stream to a second video stream. A first gesture of the participant of interest relative to the first camera is determined (610), and a second gesture of the participant of interest relative to the second camera is determined (612). The first video stream and the second video stream are rated (614) based on comparing the first gesture and the second gesture. The processor automatically switches between sending one of the first video stream or the second video stream to the display system based on a respective one of the first video stream or the second video stream having a higher rating (616).
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Video conferencing technology allows participants to communicate with each other from remote locations. As an example, a video conferencing system may include: a camera that generates audio and video streams conveying the voice and appearance of one or more participants; speakers that output audio received from the audio streams of remote participants; and a display that outputs video received from the video streams of remote participants.

[0002] Different types of cameras can be used in video conferencing. Mechanical pan-tilt-zoom (“PTZ”) cameras have physical components that allow them to pan, tilt, and zoom. Alternatively, electronic PTZ cameras are still cameras that use image processing techniques to simulate the effect of a mechanical PTZ camera. For example, an electronic PTZ camera can capture a wide-angle image, then digitally magnify and crop the image to create the illusion of panning, tilting, and zooming. In both cases, the camera's pan, tilt, and zoom controls can be adjusted in real time based on the seating positions of the video conference participants and who is speaking to achieve framing of the participants. Attached Figure Description

[0003] The accompanying drawings are provided to help illustrate various features of the examples of this disclosure, and are not intended to limit the scope of this disclosure or exclude alternative implementations.

[0004] Figure 1 This is a block diagram illustrating an exemplary system with two video conferencing endpoints connected via a network.

[0005] Figure 2 This is a block diagram of exemplary components that enable a video conferencing system.

[0006] Figure 3 The illustration shows an exemplary video conferencing environment and system, the video conferencing system including a static main camera and two pan-tilt-zoom auxiliary cameras.

[0007] Figure 4A Exemplary image frames of the same video conferencing environment scene captured from a static main camera (left) and a secondary camera (right).

[0008] Figure 4B The diagram illustrates the feature mapping between features extracted from the main camera image frame and the corresponding features extracted from the auxiliary camera image frame.

[0009] Figure 5A The illustration shows an exemplary scene of a video conferencing environment captured by a static wide-angle camera, wherein the video conferencing environment includes two pan-tilt-zoom auxiliary cameras.

[0010] Figure 5B Illustration Figure 5A Image frames from a static wide-angle camera of the scene, where the locations of the participants have been identified and distinguished as bounding areas.

[0011] Figure 5C The illustration shows a frame of a participant in a video conferencing environment using only a static wide-angle camera.

[0012] Figure 5D The illustration shows participant framing in a video conferencing environment where optimal framing of participants is used between a static wide-angle field-of-view camera and two pan-tilt-zoom auxiliary cameras.

[0013] Figure 6 This is a flowchart illustrating the steps of an exemplary method for automatically selecting video streams captured from different cameras in a multi-camera video conferencing system based on optimally framed video streams for participants of interest.

[0014] Figure 7 This is a flowchart illustrating the steps of an exemplary method for automatically switching between video streams captured by the main camera and the auxiliary camera using real-time calibration between the main camera and the auxiliary camera to optimally frame video conference participants of interest. Detailed Implementation

[0015] As described above, video conferencing technology enables users (i.e., video conference participants) to communicate with each other from remote locations. In a non-limiting example, a video conferencing system may include multiple cameras, each producing audio and video streams that convey the voice and appearance of the participants. Also as described above, video conferencing systems may use different types of cameras, including mechanical PTZ and electronic PTZ cameras. During a video conference, the pan, tilt, and zoom controls of the cameras can be adjusted in real time based on the seating positions of the participants and who is speaking to achieve frame-by-frame capture of the participants.

[0016] In some existing video conferencing systems, participants perform a series of actions to pan, tilt, and zoom the MPTZ camera to capture better frames of participants (such as active speakers). Manually guiding the camera using a remote control is challenging and inconvenient. For these reasons, many video conferencing systems use a fixed wide-angle field of view covering the entire room. In these instances, participants at the far end of the room will appear tiny in the image frame and be difficult to observe on the video stream.

[0017] The advantage of this disclosure is that it utilizes multiple cameras (e.g., at least one main camera and at least one auxiliary camera) in a video conferencing system and provides automatic switching between video streams captured in real time by those cameras to optimally frame participants of interest (e.g., active speakers). Participants of interest can be identified in the main camera video stream, which can be captured using a main camera with a wide field of view. Using calibration among the multiple cameras, the location of the identified participants of interest can then be mapped to the auxiliary camera video stream; this calibration can be performed in real time or based on offline calibration data. In some examples, camera parameters (e.g., pan, tilt, and / or zoom factors) can be determined to frame participants of interest using the auxiliary camera, and the video stream that best frames the participants (e.g., the video stream that best frames the faces of the participants of interest) can be selected.

[0018] By using multiple cameras in a video conferencing system as described above, users are given flexibility in choosing the video conferencing layout. This overcomes the limitations of single-camera or multi-camera systems, which do not allow for real-time calibration between cameras, dynamic mapping between video streams, and / or automatic switching between video streams to optimally frame participants of interest. In the example described in this disclosure, optimal framing output allows users to have an improved hybrid working experience.

[0019] Therefore, a system and method are described for associating positions within a scene from video streams from one or more main cameras with scenes from one or more auxiliary cameras, for example, by extracting features and mapping each captured scene to one another. Using this mapping, a participant of interest identified in a first video stream can be identified in another video stream from an auxiliary camera. A first pose of the participant of interest relative to the main camera is determined, and a second pose of the participant of interest relative to the auxiliary camera is determined. Based on a comparison of the first and second poses, the first and second video streams are rated. Based on the corresponding video stream in the first or second video stream with the higher rating, the processor can then automatically switch between sending one of the first or second video streams to a display system.

[0020] Figure 1 The diagram illustrates an exemplary system 100 for enabling communication between video conferencing participants at one endpoint (such as a video conferencing room) and remote participants at one or more remote endpoints. Figure 1In the illustrated example, system 100 includes a first video conferencing system 120a and a second video conferencing system 120b (collectively referred to as "video conferencing system 120"), and a network 140. The first video conferencing system 120a may include multiple cameras, each producing audio and video streams conveying the voice and appearance of participants at the first video conferencing endpoint 130a. Similarly, in the illustrated example, the second video conferencing system 120b may include multiple cameras, each producing audio and video streams conveying the voice and appearance of participants at the second video conferencing endpoint 130b. From the perspective of the first video conferencing endpoint 130a, the second video conferencing endpoint 130b is a remote endpoint, and similarly, from the perspective of the second video conferencing endpoint 130b, the first video conferencing endpoint 130a is a remote endpoint.

[0021] and Figure 1 Compared to the video conferencing system illustrated in the diagram, and according to various configurations, system 100 can include additional, fewer, or different video conferencing systems 120. Each video conferencing system 120 can be associated with one user or with multiple users (e.g., multiple participants located in a single venue). For example, a first video conferencing system 120a can be associated with a first user or a first group of users, and a second video conferencing system 120b can be associated with a second user or a second group of users.

[0022] The first video conferencing system 120a and the second video conferencing system 120b can communicate via network 140. Network 140 can be a remote wireless network, such as the Internet, a local area network (“LAN”), a wide area network (“WAN”), or a combination thereof. In other examples, network 140 can be a short-range wireless communication network, and in other examples, network 140 can be a wired network using, for example, Ethernet cables, USB cables, etc. Alternatively or additionally, network 140 can include a combination of remote, short-range, and / or wired connections. In some examples, network 140 may include both wired devices and connections and wireless devices and connections. Alternatively or additionally, in some examples, via… Figure 1 One or more intermediate devices not shown in the figure enable communication between two or more components of system 100.

[0023] Reference Figure 2 An exemplary video conferencing system 120 is illustrated. As described above, the video conferencing system 120 can communicate with at least one remote endpoint 135 via network 140 or via a direct communication link. As shown, the remote endpoint 135 may include the remote video conferencing system 125, which in some examples may have components arranged and configured similarly to those of the video conferencing system 120, or may have components arranged differently, configured differently, or both.

[0024] The video conferencing system 120 may include components for capturing audio and video streams and presenting audio and video streams received from a remote endpoint 135. For example, the video conferencing system 120 may include at least one main camera 102, at least one auxiliary camera 104, at least one microphone 106, at least one speaker 108, and at least one display 110. Additionally, the video conferencing system 120 may include an electronic processor 150, a memory 152, and a network interface 154, all of which may be coupled to and communicate via one or more control buses, data buses, etc., which may include a device communication bus 156. Control and / or data buses are generally shown in... Figure 2 For illustrative purposes only. In some examples, one or more control and / or data buses can be used for interconnection between various modules, circuits and components of the video conferencing system 120 and for communication between various modules, circuits and components of the video conferencing system 120.

[0025] Electronic processor 150 communicates with memory 152 to store data, retrieve stored data, process data (e.g., audio stream data, video stream data), etc. Electronic processor 150 can receive instructions 158 and data from memory 152 and execute (among other things) instructions 158. Specifically, electronic processor 150 can execute instructions 158 stored in memory 152. Therefore, electronic processor 150 and memory 152 can perform the methods described herein (e.g., ...). Figure 6 The methods illustrated in the diagram are in various aspects. Figure 7 (The various aspects of the method illustrated in the diagram).

[0026] Memory 152 may include read-only memory (“ROM”), random access memory (“RAM”), other non-transitory computer-readable media, or combinations thereof. Memory 152 may include instructions 158 (e.g., machine-readable instructions) for execution by electronic processor 150. Instructions 158 may include software executable by electronic processor 150 to enable electronic processor 150 to (among other things) receive data and / or commands, transmit data, control the operation of connected main camera 102 and / or auxiliary camera 104, etc. Software may include, for example, firmware, one or more applications, program data, filters, rules, one or more program modules, and other executable instructions or otherwise machine-readable instructions.

[0027] The electronic processor 150 retrieves instructions 158 from memory 152 and executes (among other things) instructions relating to the control processes and methods described herein. The electronic processor 150 is also configured to store data in memory 152, including audio stream data, video stream data, camera calibration data, camera control parameters, etc.

[0028] Additionally, the electronic processor 150 may store other data on the memory 152, including: information identifying participants in the video conferencing environment; the location of participants in the video conferencing environment, which may include borders or other areas identified in the video streams captured by one or more main cameras 102 and / or one or more auxiliary cameras 104; estimated poses of the identified participants (e.g., estimated yaw values ​​of the head pose of each participant in each video stream, estimated pitch values ​​of the head pose of each participant in each video stream, estimated direction of eye gaze of each participant in each video stream); rating values, rating the pose of each participant in each video stream; and so on.

[0029] Network interface 154 provides communication between video conferencing system 120 and one or more remote endpoints 135. As described above, video conferencing system 120 can communicate with one or more remote endpoints 135 via network 140, or alternatively can interface with each other to provide a direct communication link (e.g., via network interface 154 of video conferencing system 120 and the corresponding network interface of one or more remote endpoints 135).

[0030] In some examples, network interface 154 can communicate using wireless communication protocols, such as... Cellular protocols, proprietary protocols, etc. For example, via... Network interface 154 can communicate via a wide area network (such as the Internet) or a local area network. Communication via network interface 154 can be encrypted to protect data exchanged between video conferencing system 120 and one or more remote endpoints 135 from third-party interference.

[0031] The video conferencing system 120 may include at least one speaker 108. The speaker 108 can be used to play video conferencing audio received from one or more remote endpoints 135 to participants in the local video conferencing environment.

[0032] The video conferencing system 120 may also include at least one display 110. The display 110 can be used to display video conferencing video received from one or more remote endpoints 135 to participants in the local video conferencing environment. In some examples, the display 110 may include a flat panel display, such as a liquid crystal display (“LCD”) panel, an LED display panel, etc. The display 110 can also present additional information to the user. For example, the display 110 may provide a graphical user interface (“GUI”) for controlling the operation of one or more main cameras 102, one or more auxiliary cameras 104, or both. Alternatively, the GUI may be presented to the user for controlling the operation of the video conferencing system 120 and its interaction with one or more remote endpoints 135, such as controlling participant interaction, controlling video conferencing audio and video settings, etc.

[0033] During a video conference, one or more main cameras 102 and one or more auxiliary cameras 104 capture video streams and provide those video streams to an electronic processor 150 for processing. As described above, the video conferencing system 120 uses one or more main cameras 102 and one or more auxiliary cameras 104 to dynamically switch between different fields of view in the video conferencing environment so that optimal frames for interested participants can be provided for display by one or more remote endpoints 135.

[0034] One or more main cameras 102 may include a static or fixed camera with a wide field of view. Using one or more main cameras 102, for example, video conferencing system 120 captures video of a room or at least a wide or narrowed field of view of the room, which may include all video conferencing participants and some of the surrounding objects in the video conferencing environment. In one example, one or more main cameras 102 may be controllable to adjust the camera's pan, tilt, and zoom to control its field of view and frame the environment. In some examples, video conferencing system 120 may include a single main camera 102. In other examples, video conferencing system 120 may include two or more main cameras 102. Typically, the one or more main cameras 102 may be referred to as the first camera system of video conferencing system 120. In most examples, the main camera 102 may be a static camera with a wide field of view, as described above. In some other examples, the main camera 102 may include a PTZ camera.

[0035] When more than one main camera 102 is used, a single main camera 102 can be selected to calibrate an auxiliary camera 104, or alternatively, different main cameras 102 can be used to calibrate different auxiliary cameras 104. For example, a main camera 102 with a field of view that significantly overlaps with the field of view of an auxiliary camera 104 can be used to calibrate that auxiliary camera 104, compared to another main camera 102 whose field of view does not significantly overlap with that of the auxiliary camera 104. Typically, a main camera 102 that provides most of the spatial overlap with the auxiliary camera 104 in the scene can be used to calibrate that auxiliary camera 104 because more spatial information is shared among those cameras.

[0036] One or more auxiliary cameras 104 may be controllable cameras with a wide field of view or a narrower field of view compared to one or more main cameras 102. In some examples, the video conferencing system 120 may include a single auxiliary camera 104. In other examples, the video conferencing system 120 may include two or more auxiliary cameras 104. Typically, the one or more auxiliary cameras 104 may be referred to as the second camera system of the video conferencing system 120. As described below, the one or more auxiliary cameras 104 may include PTZ cameras, still cameras, or a combination of both.

[0037] In some examples, one or more auxiliary cameras 104 may face the same direction as one or more main cameras 102. In other examples, one or more auxiliary cameras may face a different direction than one or more main cameras 102. For example, in Figure 3 In the exemplary configuration shown in the figure, the main camera 102 faces a first direction; the first auxiliary camera 104a faces a second direction, which is perpendicular to the first direction or at an angle to the first direction; and the second auxiliary camera 104b faces a third direction, which is also perpendicular to the first direction or at an angle to the first direction. Figure 3The second and third orientations illustrated in the diagram may be opposite to each other, so that different sides of the conference table can be observed by the first and auxiliary cameras 104a, 104b. One or more main cameras 102 and one or more auxiliary cameras 104 may have overlapping fields of view, or some of the cameras may have fields of view that do not overlap with the others. As a non-limiting example, a single main camera 102 may have a field of view that overlaps with the fields of view of two auxiliary cameras 104, while those two auxiliary cameras 104 may have fields of view that do not overlap with each other. Similarly, the video conferencing system 120 may include two main cameras 102 and a plurality of auxiliary cameras 104, wherein one main camera 102 has a field of view that overlaps with the fields of view of only some of the auxiliary cameras 104, and a second main camera 102 has a field of view that overlaps with the fields of view of the other auxiliary cameras 104.

[0038] In some examples, one or more auxiliary cameras may be PTZ cameras. The PTZ camera may be a mechanical PTZ camera (“MPTZ”) or an electronic PTZ camera (“EPTZ”). In an MPTZ, the camera's pan, tilt, and zoom functions are achieved by mechanically controlling the optics within the camera; in an EPTZ, these functions are achieved digitally by panning, tilting, and zooming within a larger field of view to frame a smaller field of view. In some other examples, one or more auxiliary cameras 104 may be still cameras with a wide or narrow field of view.

[0039] Video conferencing system 120 uses one or more auxiliary cameras 104 to capture video of one or more participants (such as an active speaker or another interested participant) in a video conferencing environment within a narrow or wide field of view. In some examples, the auxiliary cameras 104 may include PTZ cameras. The auxiliary cameras 104 may be mechanical PTZ cameras, or in some instances, electronic PTZ cameras. The auxiliary cameras 104 may have narrow or wide fields of view. In some instances, the auxiliary cameras 104 may have a narrower field of view than the main camera 102, or when multiple main cameras 102 are used, a narrower field of view than at least one of the main cameras 102.

[0040] Additionally, at least one microphone 106 can capture audio streams and provide those audio streams to the electronic processor 150 for processing. As a non-limiting example, the microphone(s) 106 may be a desktop microphone, a microphone integrated with one or more main cameras 102, a microphone integrated with one or more auxiliary cameras, etc. The video conferencing system 120 uses the audio streams captured by the microphone(s) 106 for video conferencing audio. In some examples, additional microphones or microphone arrays may also be used to capture audio streams for camera tracking purposes. For example, additional audio streams can be captured and processed by the electronic processor 150 to determine the location of audio sources during a video conference, such as the location of a participant actively speaking.

[0041] Electronic processor 150 sends the captured audio and video streams to one or more remote endpoints 135. In some examples, electronic processor 150 encodes the video stream using an encoding standard. For example, the video stream can be encoded using encoding standards such as MPEG-1, MPEG-2, MPEG-4, H.261, H.263, H.264, etc. Electronic processor 150 may also encode the audio stream using a suitable codec. Electronic processor 150 outputs the encoded audio and video streams to network interface 154, which then transmits the encoded audio and video streams to one or more remote endpoints 135 via network 140. Similarly, network interface 154 receives audio and video streams transmitted from one or more remote endpoints 125 via network 140. Audio and video streams received from one or more remote endpoints 125 are transmitted from network interface 154 to electronic processor 150. The electronic processor 150 sends the received audio stream(s) to the speaker 108 to output video conference audio captured from one or more remote endpoints, and sends the received video stream(s) to the display 110 to output video conference video captured from one or more remote endpoints 135.

[0042] As described above, an advantage of this disclosure is that the electronic processor 150 is capable of processing video streams captured from one or more main cameras 102 and one or more auxiliary cameras 104 to automatically frame a participant by selecting the video stream that best frames the face of a participant of interest (e.g., the current speaker) in the video conference. As an example, the electronic processor 150 can process the video stream captured from the main camera 102 to identify a participant of interest in the main camera video stream. The location of the participant of interest can be mapped from the main camera video stream to one or more auxiliary camera video streams. In some examples, dynamic feature mapping is performed by the electronic processor 150 to map the location of the participant of interest from the main camera video stream to one or more auxiliary camera video streams.

[0043] As the participant of interest is identified in each video stream, the electronic processor 150 processes the video streams to estimate or otherwise determine the pose of the participant of interest in each video stream. The pose may be a head pose, body pose, etc. The estimated pose in each video stream is then rated by the electronic processor 150 to determine which video stream best frames the participant of interest. Based on these ratings, the electronic processor 150 selects the video stream that best frames the participant of interest and sends the selected video stream to the network interface 154 for transmission and display at one or more remote endpoints 135.

[0044] In some examples, the video conferencing system 120 outputs a video stream from one or more main cameras 102 or one or more auxiliary cameras 104 at any given time. As described above, during a video conference, the output video stream from the video conferencing system 120 automatically switches between one or more main cameras 102 and one or more auxiliary cameras 104 to select the best framed video view of the interested participants (e.g., active speakers).

[0045] When the selected video stream is captured from one or more auxiliary cameras 104, the video conferencing system 120 may periodically switch between the selected video stream and the video stream captured by one or more main cameras 102, so that one or more remote endpoints 135 can understand the magnified view of the active speaker or other interested participants, while still periodically observing the wide view of the video conferencing environment (e.g., seeing other participants).

[0046] In some other examples, the video conferencing system 120 is capable of simultaneously transmitting video streams from both a main camera 102 and an auxiliary camera 104, while switching between different video streams captured from different auxiliary cameras 104 as the video conference progresses, based on which auxiliary camera 104 best frames the participants of interest. In these instances, the electronic processor 150 is capable of sending video streams to one or more remote endpoints 135, such that one video stream can be composited with another video stream into a picture-in-picture format. For example, via the electronic processor 150, the auxiliary camera video stream can be composited with the main camera video stream, and the composited video stream can be sent to one or more remote endpoints 135 for display in picture-in-picture format. Alternatively, the video stream can be sent by the electronic processor 150 to one or more remote endpoints 135, where the video stream can be composited by the electronic processors of one or more remote endpoints 135.

[0047] As described above, a video conferencing system can use multiple cameras that can be calibrated against each other based on dynamic feature mapping or other techniques used to correlate the coordinate systems of the multiple cameras. In some examples, the calibration process based on dynamic feature mapping can be performed in real time without the use of external references (e.g., patterns, including chessboards, bitmaps, etc.). Dynamic feature mapping can alternatively generate calibration data based on real-time scene data (e.g., conference room data) from video streams captured by the multiple cameras.

[0048] Reference Figure 3 An exemplary video conferencing system 120 may include a main camera 102, a first auxiliary camera 104a, and a second auxiliary camera 104b. The main camera 102 may be coupled to the display 110, such as by clipping it to an edge or surface of the display 110 or otherwise attaching it to an edge or surface of the display 110. Alternatively, the main camera 102 may be integrated with the display 110.

[0049] In the illustrated example, the main camera 102, the first auxiliary camera 104a, and the second auxiliary camera 104b capture video streams, each depicting a scene in the video conferencing environment 160. For example, the video conferencing environment 160 may include a conference room where multiple participants 161, 162, 163, and 164 are seated around a conference table.

[0050] In the illustrated example, the field of view (“FOV”) 122 of the main camera 102 may have a wide field of view angle that facilitates observation of the video conferencing environment 160 (e.g., a conference room), while the FOV 124a of the first auxiliary camera 104a and / or the FOV 124b of the second auxiliary camera 104b may have a narrower field of view angle that facilitates better framing of participants 161, 162, 163, 164.

[0051] As will be described in more detail below, the main camera 102 can be used to capture a video stream depicting a scene including all participants 161, 162, 163, and 164 in the video conferencing environment 160. When participant 161 is an active speaker, the video stream from the main camera 102 can provide optimal framing for participant 161. However, when participant 162 is an active speaker, the video stream from the main camera 102 may not optimally frame participant 162. For example, participant 162 may be located at a far end of the video conferencing environment 160 relative to the main camera 102 and will therefore appear small in the video stream. In such an instance, participant 162 can be identified in the video stream captured by the main camera 102, and location data associated with participant 162 can be mapped to the video stream captured by the second auxiliary camera 104b, which provides a better field of view for participant 162. As described above, the calibration between the main camera 102 and the second auxiliary camera 104b can be based on the dynamic feature mapping between the two cameras, facilitating this mapping of participant location data.

[0052] Now refer to Figure 4A and 4B This document describes a non-limiting example of such a calibration process based on dynamic feature mapping. In the illustrated example, a first video stream is captured using a primary camera (which is a static wide-angle camera), and a second video stream is captured using an auxiliary camera (which is an MPTZ camera). In this example, both cameras are facing forward on a video conferencing environment 160, which includes three participants 161, 162, and 163. The first and second video streams are received by an electronic processor 150 of a video conferencing system 120, where they are processed to generate calibration data that correlates the coordinate systems of the primary and auxiliary cameras.

[0053] A first image frame 172 is selected by the electronic processor 150 from a first video stream captured by the main camera. Similarly, a second image frame 174 is selected by the electronic processor 150 from a second video stream captured by an auxiliary camera, wherein the first and second image frames 172 and 174 are captured simultaneously by their respective cameras. Alternatively, multiple images may be captured by the auxiliary camera and stitched together by the electronic processor 150 to produce a composite image frame (or composite video stream) that matches the wide-angle field of view of the main camera.

[0054] Feature detection and matching are performed on first and second image frames 172 and 174 using electronic processor 150, and feature 176 is identified in the first image frame 172 and the second image frame 174. As a non-limiting example, feature 176 can be extracted from the first and second image frames 172 and 174 using a scale-invariant feature transform (“SIFT”) operation. The extracted feature 176 can then be matched by electronic processor 150 using a feature matching process. As a non-limiting example, feature 176 extracted from the first image frame 172 can be matched with a corresponding feature 176 extracted from the second image frame 174 using a k-nearest neighbor (“k-NN”) process. Alternatively, other artificial intelligence or machine learning processes or models can be used to match features extracted from the first and second image frames. The matched feature 176 in Figure 4B The diagram is illustrated by line 178, which connects the matched feature 176 from the first image frame 172 to the corresponding feature 176 in the second image frame 174.

[0055] Using these matched features, the transformation matrix is ​​constructed by the electronic processor 150. In a non-limiting example, when generating the transformation matrix, the electronic processor 150 considers features matched at a threshold confidence level (e.g., features matched at a high level of accuracy). The transformation matrix can be a homography matrix. Alternatively, the transformation matrix can take into account both local homography and global similarity between the first and second image frames. By using both local homography and global similarity when constructing the transformation matrix, the functionality of the auxiliary camera used for participant framing can be improved.

[0056] Then, using the main camera image, by processing the first video stream using an electronic processor, video conference participants 162, 164, and 166 are detected and their locations are determined. Using the detected participants, a face detection process can be performed on the first video stream by the electronic processor 150 to detect the head position of each participant 161, 162, and 163 in the first video stream. Participant locations and head positions can be identified by borders 180 and 182, or alternatively by other indicators by identifying groups of pixels associated with the participant's location and / or head position, etc.

[0057] The participant location data (e.g., participant location, head position) is then mapped from the first video stream to the second video stream by the electronic processor 150 using a constructed transformation matrix.

[0058] In some examples, the mapped participant location data can then be used to determine the camera parameters of the auxiliary camera that will provide the best frame for each participant. As a non-limiting example, for each detected bounding box mapped to the second video stream, the distance from the center of the bounding box to the center of the second video stream image frame is calculated. Using this distance, the electronic processor 150 estimates translation and tilt parameters for the auxiliary camera. Additionally, using the bounding box size, the electronic processor 150 can estimate the zoom factor of the auxiliary camera. If a depth map is available, it is used to estimate the zoom factor.

[0059] or

[0060] ZF∝HD

[0061] ZF is the zoom factor, HBB is the size of the head bezel, and HD is the head depth.

[0062] As a non-limiting example, using a regression model, the equations for both the translation and tilt factors can be estimated. Exemplary equations for the translation and tilt factors are as follows.

[0063] Translation = (-x ± 19.083) / 27.99 ("+" is used for right translation and "-" is used for left translation)

[0064] Tilt = (y ± 3.2) / 24.42 ("+" is used for upward translation and "-" is used for downward translation)

[0065] Where x and y are the distances estimated from the center.

[0066] By using sound source localization (“SSL”) from audio and speakers as described above, the participants of interest and their positions in the second video stream image frames can also be determined. This location data can be used by electronic processor 150 to estimate the translation, tilt, and zoom factors of the auxiliary camera, which can then be sent by electronic processor 150 to the auxiliary camera 104 to move the auxiliary camera 104 to frame the participants’ heads.

[0067] Once the auxiliary camera is moved to frame the participant of interest, the electronic processor 150 estimates the participant's head pose in both the first and second video streams. As a non-limiting example, the head pose (or other body pose) can be estimated relative to a corresponding camera used to capture the video stream processed by the electronic processor 150, based on yaw and pitch (e.g., tilt). As another example, the head pose may include the participant's eye gaze direction (i.e., the direction the participant is looking). In these instances, the eye gaze direction can be represented as coordinates in three-dimensional space. The eye gaze direction can be estimated by the electronic processor 150 in both the first and second video streams. As a non-limiting example, the eye gaze direction can be estimated relative to a corresponding camera used to capture the video stream processed by the electronic processor 150, based on yaw.

[0068] Head pose data from a first and second video stream can then be rated by an electronic processor 150, for example, by comparing individual head pose datasets with reference head pose data, which may include thresholds for different head pose parameters (e.g., yaw threshold, pitch threshold, and eye gaze direction threshold). Head poses with larger yaw values ​​(i.e., where the face orientation is turned away from the camera to a greater extent, and the eye gaze direction is turned away from the camera to a greater extent) receive lower ratings, and similarly, head poses with larger pitch values ​​(i.e., where the face orientation is tilted away from the camera to a greater extent) receive lower ratings. For example, if the estimated yaw value of the head pose is less than 3 degrees and / or the estimated pitch value is less than 3 degrees, the video stream can be rated higher compared to a video stream with larger yaw and / or pitch values. As a non-limiting example, different ranges of yaw or pitch values ​​can be assigned different grades of scores, such as yaw or pitch values ​​between 0 and 5 degrees receiving grade 1, yaw or pitch values ​​between 5 and 10 degrees receiving grade 2, yaw or pitch values ​​between 10 and 25 degrees receiving grade 3, and so on. In this example, the grade score is assigned based on a 5-degree range of yaw and pitch. In other examples, different ranges of yaw and pitch values ​​can be used (e.g., smaller or larger ranges). Alternatively, such as by using estimated yaw or pitch values ​​as grade score values, grade scores can be assigned continuously. As mentioned above, yaw values ​​can be estimated for head, eye gaze direction, or both. Similarly, different thresholds can be established for head yaw and eye gaze direction.

[0069] The combined yaw and pitch scores of the head posture can be used as the overall head posture rating, with higher scores corresponding to lower ratings. Alternatively, estimated yaw and pitch values, or other posture data, can be assigned to any arbitrary rating scale. For example, first and second video streams can be rated from 1-5, 1-10, etc., with 1 being the worst and 5 (or 10) being the best. The higher-rated video stream is then sent by electronic processor 150 to remote endpoint 135 for observation by a remote participant.

[0070] Reference Figures 5A-5D An exemplary process for rating the postures of video conference participants is illustrated. Figure 5A An exemplary video stream image from a main camera is shown, which has a static wide-angle field of view of the video conferencing environment 160. In the illustrated example, six participants (161, 162, 163, 164, 165, 166) are seated around a conference table. A first auxiliary camera 104a is positioned on one side of the video conferencing environment 160 opposite to a second auxiliary camera 104b, similar to... Figure 3 The layout is shown in the diagram.

[0071] The participants' locations are identified in the main camera video stream, such as Figure 5B As indicated in the document. For example, as described above, participant location data may include regions identified as containing each participant's body, each participant's head, or both. Participant location data may include borders (e.g., body border 180, head border 182), other indicators, pixel groups, etc.

[0072] Figure 5CThe illustration shows participant framing using a prior art technique that employs only a single wide-angle camera, such as the main camera of a video conferencing system described in this disclosure. In this illustrated example, the framing of participants 162 and 163 may be good, but the framing of participants 161, 164, 165, and 166 only shows the portion of their faces oriented towards the camera. This suboptimal participant framing is overcome using the technique described in this disclosure. By mapping participant location data from the main camera video stream to video streams captured by the first and second auxiliary cameras 104a, 104b, participant pose data can be estimated for each participant 161, 162, 163, 164, 165, 166 in each video stream, and the pose data can be rated as discussed above. Based on these ratings, the optimal field of view for framing participant 161 is determined by electronic processor 150 to be from the first auxiliary camera 104a, and the optimal field of view for framing participants 164, 165, and 166 is determined by electronic processor 150 to be from the second auxiliary camera 104b. Therefore, when the participants of interest change during a video conference, electronic processor 150 can send different video streams to a display system (e.g., the display system of remote endpoint 135) based on which video stream is best for framing the currently interested participant.

[0073] Now refer to Figure 6 The flowchart is illustrated as a step-by-step guide to an exemplary method for automatically selecting video streams captured from different cameras in a multi-camera video conferencing system based on optimally framing the video streams of participants of interest. As described above, calibration can be performed between the multiple cameras to facilitate the identification of video conferencing participants in multiple video streams without duplication.

[0074] The first video stream is received by the electronic processor 150 from the first camera, as indicated in step 602. For example, the first video stream can be received by the electronic processor 150 from the main camera 102. As described above, the main camera 102 has a wide-angle field of view of the video conferencing environment (e.g., a conference room). Therefore, the first video stream depicts the scene corresponding to the video conferencing environment and any participants within the video conferencing environment.

[0075] The second video stream is also received by the electronic processor 150 from the second camera, as indicated in step 604. For example, the second video stream can be received by the electronic processor 150 from an auxiliary camera 104. The auxiliary camera 104 has a different field of view of the video conferencing environment (e.g., a conference room). Therefore, the second video stream depicts the same scene corresponding to the video conferencing environment and any participants within the video conferencing environment, but from a different perspective than the main camera. Different perspectives may correspond to different field of view angles (e.g., narrower field of view), different viewing directions, or a combination of both.

[0076] In some examples, the second video stream may include a composite video stream generated by the electronic processor 150 by synthesizing different image data captured by the second camera to match the FOV of the first camera. For example, by capturing image data by moving the second camera with different pan and tilt values, image data can be captured by the second camera at a wider field of view. This image data can be synthesized to produce a composite video stream or image data that can be processed by the electronic processor 150 during the calibration of the first and second cameras, as described in more detail below.

[0077] The first video stream is processed by electronic processor 150 to identify participants of interest within the first video stream, as indicated in step 606. Alternatively, by utilizing electronic processor 150 to process the first video stream, other participants can be similarly identified within the first video stream. Advantageously, as participants move from off-frame locations into the video conferencing environment, the locations of new participants can be identified and updated in real time.

[0078] In some examples, identifying a participant in the first video stream may include processing the first video stream using a face detection process to determine or otherwise identify regions in the first video stream that correspond to or otherwise contain the participant's face. This process can be repeated for each participant depicted in the first video stream. As an example, the electronic processor 150 can output borders or other indicators for identifying spatial regions in the first video stream corresponding to a specific participant. Alternatively, regions (e.g., groups of pixels) in the first video stream identified as corresponding to or otherwise containing the participant's face can be identified by the electronic processor 150, and their locations are stored. In some other examples, identifying a participant in the first video stream may include identifying the participant's body in addition to, or as an alternative to, the participant's face. The location of the body can be similarly identified and output as borders, other indicators, groups of pixels, etc.

[0079] Additionally, the electronic processor can also receive audio stream data from the microphone 106 of the video conferencing system 120 to facilitate the localization and identification of participants in the video conferencing environment. For example, audio-based localization can be used to identify different participants in a video conferencing environment.

[0080] The locations of interested participants and / or other identified video conference participants are then mapped from the first video stream to the second video stream by the electronic processor 150, as indicated in step 608. Thus, based on the processed first video stream, interested participants or other participants can be located in the second video stream. As a non-limiting example, mapping participant location data from the first video stream to the second video stream can be facilitated using calibration data corresponding to the camera used to capture the first video stream (e.g., the main camera) and the camera used to capture the second video stream (e.g., the auxiliary camera).

[0081] In some examples, the calibration procedure may be executed by the electronic processor 150 to generate calibration data used when mapping participant location data from a first video stream to a second video stream. The calibration procedure may use dynamic feature mapping to estimate camera parameters of the first and second cameras to match the FOV and position of the second camera to the first camera. For example, the calibration data may include data that correlates the coordinate system of the second camera with the coordinate system of the first camera.

[0082] Advantageously, by identifying participants in a first video stream (which can be a wide-angle view of the video conferencing environment so that all video conferencing participants can be observed in the first video stream) and then mapping the location data of those participants to a second camera (or other cameras in the video conferencing system, such as other auxiliary cameras), participant identification can be achieved across multiple cameras to prevent duplication of users in the synthetic video. Because this mapping effectively allows pixel-to-pixel mapping across these different cameras, participant identification can be performed without having to match facial features or other aspects of the participants in different video streams.

[0083] As described above, the calibration data may include a transformation matrix that correlates the coordinate system of the first camera with the coordinate system of the second camera. The calibration data can be generated by processing the first and second video streams using the electronic processor 150 to extract or otherwise detect features from either the first or second video streams. As a non-limiting example, the SIFT operation can be used by the electronic processor 150 to extract features from the first and second video streams. The features extracted from the first video stream can then be matched with features extracted from the second video stream using the electronic processor 150. As a non-limiting example, the KNN process or other machine learning models can be applied to the extracted feature data using the electronic processor 150. Based on the matched features, the electronic processor 150 can construct a transformation matrix that takes into account local homography, global similarity, or both. The constructed transformation matrix can be stored by the electronic processor 150 as calibration data (e.g., by storing the transformation matrix in memory 152).

[0084] The electronic processor 150 estimates the pose of the participant of interest from the first video stream, thereby generating first pose data as output, as indicated in step 610. For example, a region containing the participant of interest is extracted from the first video stream, and the first pose data is estimated from the extracted region. In one example, the region may include a bounding box containing the identified participant of interest. The region may correspond to the participant's head, their body, or both. Therefore, the pose data may be head pose data (or facial orientation data), body pose data, or both. Typically, the pose data includes estimates of pitch, rotation, and / or tilt relative to the first camera. For example, the pose data may include estimates of yaw and pitch values ​​relative to the image plane of the first camera.

[0085] Alternatively, the first pose data may include pose data estimated for each participant identified in the first video stream.

[0086] Similarly, the pose of the participant of interest is estimated by the electronic processor 150 from the second video stream, as indicated in step 612. For example, a region containing the participant of interest (whose pose has been mapped from the first video stream to the second video stream, for example) is extracted from the second video stream, and second pose data is estimated from the extracted region. In one example, the region may include a bounding box containing the identified participant of interest. The region may correspond to the participant's head, their body, or both. Therefore, the pose data may be head pose data (or facial orientation data), body pose data, or both. Typically, the pose data includes estimates of pitch, rotation, and / or tilt relative to the first camera. For example, the pose data may include estimates of yaw and pitch values ​​relative to the image plane of the first camera.

[0087] As described above, in some implementations, camera parameters are estimated for the second camera based on the region of interest contained in the second video stream. For example, translation, tilt, and / or zoom factors can be estimated to control the second camera to frame the interest participants.

[0088] Alternatively, the second pose data may include pose data estimated for each participant identified in the second video stream. In these instances, camera parameters may also be estimated for framing each identified participant.

[0089] Based on the comparison of the first and second pose data, the first and second video streams can then be rated by the electronic processor 150, as indicated in step 614. As described above, a rating score can be assigned to the first and second pose data by comparing the pose data with a reference value or threshold. Based on the estimated pose, a rating from best to worst is given to the first and second video streams.

[0090] In some examples, the first and second pose data can be head pose data. As an example, head pose data may include estimates of head pose values ​​(such as head pitch, head rotation, and / or head tilt). Based on these head pose values, the first and second head pose data can be rated, for example, by comparing these values ​​to reference values ​​or thresholds indicating the optimal orientation of the face toward the camera's imaging plane. Alternatively, the first and second head pose data can be rated based on the percentage of a participant's face oriented toward the corresponding camera. Thus, head pose data may include estimates of the percentage of a participant's face oriented toward the camera (used to capture the video stream used to estimate the head pose data). This percentage value can be estimated based on the head pose values, or using other processes, such as by feeding the head pose data, the corresponding video stream data, or both into a machine learning model trained on suitable training data to estimate the percentage of a face oriented toward the camera (e.g., the percentage of a face oriented toward the camera's imaging plane).

[0091] The video stream with the highest rating is then selected by the electronic processor 150, and the selected video stream is sent to a display system, such as the display system of the remote video conferencing system 125 at the remote endpoint 135, as indicated in step 616. Thus, while the video conference is in progress, the electronic processor 150 can automatically switch between video streams from different cameras based on the identified participants of interest and the optimal framed video streams for those participants.

[0092] When the selected video stream corresponds to a video stream captured using a PTZ camera (e.g., a second video stream), the estimated camera parameters for framing the participants of interest using that camera can also be selected by the electronic processor 150. In these instances, when switching to the video stream optimally framed for the participants of interest, the selected camera parameters are then used by the electronic processor 150 to control the corresponding camera to automatically frame the participants of interest. Therefore, the selected video stream will include better framing of the participants of interest, without the user having to manually readjust the pan, tilt, and zoom of the PTZ camera.

[0093] In some implementations, the electronic processor 150 may not automatically switch between video streams, but instead may control the timing of transitions between video streams, as frequent camera switching can distract video conference participants. For example, if the participant of interest is a frequent speaker in the video conference, the electronic processor 150 may be able to select the best framed video stream for the participant of interest when speaking begins, but may avoid or delay switching to another speaker who may only respond with a brief answer or comment.

[0094] Now refer to Figure 7 This document illustrates a non-limiting exemplary method for automatically switching between video streams captured by a primary camera and an auxiliary camera to optimally frame video conference participants of interest. The method typically includes: an automatic runtime calibration process for correlating the coordinate systems of the primary and auxiliary cameras; a participant identification process for identifying video conference participants in the video stream and its image frames; and a production rule process for automatically framing participants of interest and automatically switching to the video stream that provides the optimal frame for those participants.

[0095] The video stream is captured by the main camera in step 702 and by the auxiliary camera in step 704. In some examples, the video streams are captured in parallel so that they depict the same scene of the video conferencing environment. An automatic runtime calibration process is then performed by an electronic processor in step 706. The automatic runtime calibration process may be performed once during the video conference or may be repeated (e.g., at certain intervals). The automatic runtime calibration process includes extracting features from the main camera video stream or selected image frames from the main camera video stream, as indicated in step 708. As described above, a SIFT operation or other suitable feature extraction process may be used to extract features. Image frames from the auxiliary video stream are stitched together to create a composite image frame having an FOV that matches the FOV of the main camera, and features are extracted from this composite image frame, as indicated in step 710. As described above, a SIFT operation or other suitable feature extraction process may be used to extract features from the composite image frame. The extracted features are then matched in step 712. For example, the KNN process can be used to match features. Alternatively, other suitable feature matching processes can be implemented, including those for implementing artificial intelligence and / or machine learning models. Using the matched features, a transformation matrix is ​​generated in step 714. The transformation matrix can be a homography matrix. In some examples, the transformation matrix may take into account both local homography and global similarity between the matched features.

[0096] Face detection is then performed on the main camera video stream or image frames(s) from it, as indicated in step 716. If no face is detected, as determined in decision block 718, additional video stream data is acquired from the main camera, and the preceding steps are repeated. Alternatively, using the preceding steps, different image frames from the main camera video stream can be selected and processed. When the face of at least one participant is detected, a face bounding box is generated for each detected face, and the method continues as indicated in step 720 by transforming each face bounding box to the auxiliary camera video stream (or image frames(s) from it). For example, as described above, face bounding boxes can be transformed from the main camera video stream to the auxiliary camera video stream by applying a transformation matrix.

[0097] The participant of interest is then identified in step 722. If no participant of interest is identified, additional auxiliary camera video stream data (or image frames from it) can be acquired and processed using the preceding steps. Otherwise, the method continues as indicated in step 724 by estimating the camera parameters (e.g., pan / tilt / zoom factors) of the auxiliary camera that will frame the participant of interest in the auxiliary camera video stream. As described above, these camera parameters can be estimated based on the facial bounding boxes of the participants of interest converted from the main camera video stream to the auxiliary camera video stream. The camera parameters can be stored and otherwise sent to the auxiliary camera to control the framing of the auxiliary camera.

[0098] The head poses of the participants of interest and other participants in the video conference are then estimated in step 726. As described above, head poses can be estimated by estimating head yaw, head pitch, or both. Alternatively, head poses can be estimated by estimating the percentage of the face oriented toward the imaging plane of the respective camera. The best-fit video streams of the participants of interest based on their head poses are then selected and displayed (e.g., by sending the selected video streams to a remote endpoint), as indicated in step 728. If a participant does not have the best head pose in the secondary camera video stream, the primary camera video stream can be selected and displayed, as indicated in step 730.

[0099] This disclosure has described one or more examples, and it should be understood that many equivalents, substitutes, variations and modifications are possible and within the scope of this disclosure, in addition to those expressly stated.

Claims

1. A method comprising: The processor receives a first video stream of the scene from the first camera. The processor receives a second video stream of the scene from the second camera; Using the processor, identify participants of interest in the first video stream; Using the processor, the identified participants of interest are identified in the second video stream by mapping them from the first video stream to the second video stream; Using the processor and from the first video stream, determine the first pose of the interested participant relative to the first camera; Using the processor and from the second video stream, a second pose of the interested participant relative to the second camera is determined; Using the processor, the first video stream and the second video stream are rated based on a comparison of the first pose and the second pose; and Based on the corresponding video stream in the first or second video stream with the higher rating, the system automatically switches between sending one of the first or second video streams to the display system by the processor.

2. The method of claim 1, wherein identifying the interested participants comprises: Identify the pixel group in the first video stream associated with the participant of interest.

3. The method of claim 2, wherein mapping the identified participants of interest from the first video stream to the second video stream comprises: Map the pixel group in the first video stream to the pixel group in the second video stream.

4. The method of claim 1, wherein mapping the identified participants of interest from the first video stream to the second video stream comprises: The electronic processor receives calibration data, wherein the calibration data correlates the coordinate system of the first camera with the coordinate system of the second camera. and Using the calibration data, the identified participants of interest are mapped from the first video stream to the second video stream.

5. The method of claim 4, wherein receiving the calibration data using the electronic processor comprises: The electronic processor is used to perform dynamic feature mapping between the first video stream and the second video stream, thereby generating the calibration data as output.

6. The method of claim 5, wherein the electronic processor performs the dynamic feature mapping by: Extract features from the first video stream; Extract features from the second video stream; The features from the first video stream are matched with the features from the second video stream; and A transformation matrix is ​​generated that maps the features from the first video stream to matching features in the second video stream.

7. The method of claim 1, wherein rating the first video stream and the second video stream comprises: The first and second postures are compared with reference posture data, which associates the range of posture values ​​with the grade score.

8. The method of claim 7, wherein determining the first attitude includes determining first attitude data, the first attitude data including at least one of a first yaw value or a first pitch value, determining the second attitude includes determining second attitude data, the second attitude data including at least one of a second yaw value or a second pitch value, and comparing the first attitude and the second attitude includes comparing the first attitude data and the second attitude data with a range of the attitude values.

9. The method of claim 1, wherein the first camera comprises a still camera and the second camera comprises a pan-tilt-zoom camera.

10. A non-transitory computer-readable storage medium having instructions stored thereon, which, when executed by a processor, cause the processor to: Receive the first video stream of the scene from the first camera; Receive a second video stream of the scene from the second camera; Based on the first video stream, identify interested participants in the scene; From the first video stream, determine the first facial orientation of the participant of interest relative to the first camera; From the second video stream, determine the second facial orientation of the participant of interest relative to the second camera; The first face orientation and the second face orientation are rated based on the percentage of the face of the interested participant depicted in each of the first and second video streams. and Based on the levels of the first and second face orientations, one of the first video streams or the second video stream is sent to the display system.

11. The non-transitory computer-readable storage medium of claim 9, wherein the participant of interest is identified by determining a region in the first video stream containing the participant of interest.

12. The non-transitory computer-readable storage medium of claim 11, wherein, using calibration data that correlates the coordinate system of the first camera with the coordinate system of the second camera, the region containing the participant of interest is mapped from the first video stream to the second video stream.

13. The non-transitory computer-readable storage medium of claim 12, wherein the calibration data comprises a transformation matrix generated by the following operations: Extract features from the first video stream; Extract features from the second video stream; The features from the first video stream are matched with the features from the second video stream; and The transformation matrix is ​​generated based on features from the first video stream that match the features in the second video stream.

14. A system comprising: The first camera; The second camera system includes at least one camera; Electronic processor, used for: Receive a first video stream from the first camera; Receive at least one second video stream, each second video stream being received from one of the at least one cameras in the second camera system; Receive calibration data, which correlates the coordinate system of the first camera with the coordinate system of each camera in the second camera system; Identify pixels in the first video stream corresponding to the participants; Using the calibration data, the pixels in the first video stream are mapped to pixels in each of the second video streams; Determine the participant's first head pose in the first video stream; Determine the second head pose of the participant in each second video stream; Rating data is generated by rating the first head pose based on the orientation relative to the first camera and each second head pose based on the orientation relative to each corresponding camera in the second camera system. Select one of the first video streams that has a higher rating score in the rating data or one of the at least one second video streams as the selected video stream; and Send the selected video stream to the remote endpoint.

15. The system of claim 14, wherein the first camera comprises a still camera, and the second camera system comprises a single pan-tilt-zoom camera.

16. The system of claim 14, wherein the second camera system comprises at least two pan-tilt-zoom cameras.

17. The system of claim 14, wherein the second camera system comprises at least two still cameras.

18. The system of claim 14, wherein the first camera comprises a first camera system, the first camera system comprising at least two still cameras.

19. The system of claim 18, wherein the processor receives the first video stream from one of the at least two still cameras in the first camera system.

Citation Information

Cited By

  • Synchronous display method and device of video conference and storage medium

    CN121309766A