Autonomous video conferencing system using virtual director support

The autonomous video conferencing system addresses the challenge of capturing and framing all objects and interactions in real-time using smart cameras and sensors, enhancing engagement by applying television studio production principles.

JP7852187B2Active Publication Date: 2026-04-28HUDDLY AS
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HUDDLY AS
Filing Date
2022-04-19
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Conventional video conferencing systems struggle to capture subtle changes in expressions and gestures of participants positioned variously within the meeting space, leading to a less engaging experience for remote participants, especially in large spaces where objects far from the camera are difficult to see or read.

Method used

An autonomous video conferencing system utilizing multiple smart cameras and sensors, equipped with sub-symbolic and symbolic artificial intelligence, detects objects and their interactions, applies television studio production principles through a virtual director unit to create automated productions, and streams these in real-time to remote users.

Benefits of technology

Enhances video conferencing by providing a more engaging and comprehensive experience for remote participants, capturing and framing all objects and interactions in real-time, similar to a television studio production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007852187000001
    Figure 0007852187000001
  • Figure 0007852187000002
    Figure 0007852187000002
  • Figure 0007852187000003
    Figure 0007852187000003
Patent Text Reader

Abstract

A system and method are provided for enhancing video conferencing and remote collaboration using sub-symbolic and symbolic artificial intelligence. The autonomous video conferencing system of the present disclosure comprises one main smart camera and multiple peripheral smart cameras, optionally coupled to one or more smart sensors. Each smart camera is equipped with a vision pipeline supported by machine learning to detect objects and their interactions and associated changes in gestures and postures, and a virtual director adapted to apply a predefined set of rules consistent with TV studio production principles. The main camera is adapted to select and update an attention video stream in real time under the direction of the virtual director, and stream the updated attention stream to a user computer. A method is provided for creating automated TV studio productions for various conference spaces and dedicated scenarios with virtual director assistance.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The present disclosure generally relates to video conferencing and remote collaboration technologies. Specifically, the present disclosure relates to enhancing video conferencing and remote collaboration using smart cameras and smart sensors that use artificial intelligence. More specifically, the present disclosure relates to an autonomous video conferencing system using virtual director assistance, and a method for creating an automated television studio production for a video conferencing space using virtual director assistance.

[0002] Remote collaboration and video conferencing have become a mainstay in the modern business world and society as a whole. In conventional video conferencing, the experience for participants is mainly static. Cameras in meeting rooms generally do not notice social or spatial cues, such as speaker-listener dynamics, participant reactions, body language, direction of attention, and degree of engagement. Thus, as remote participants, the experience often lacks engagement. Existing hardware and software video conferencing solutions generally rely on a single camera system. Such systems are generally limited to displaying objects within a video conferencing space from one direction or one angle. Thus, it is difficult for the system to capture subtle changes in expressions, and gestures and postures among objects positioned variously within the video conferencing space. As a result, remote participants are presented with a second-best and less engaging experience. In large video conferencing spaces, it is often even more difficult to capture and frame objects positioned far from the camera, and thus, it may be difficult, if not impossible, for remote participants to see or read and engage with those objects.

[0003] More recent remote collaboration and video conferencing solutions have seen improvements that allow remote users to adjust settings regarding their preferences, among other things, screen selection, image filtering, recording, ambient audio control, file sharing, and privacy controls. However, the inability of remote users to see or read the room and fully engage with all objects in the meeting space in real time remains a drawback. Thus, due to these limitations, despite the dramatic growth of video conferencing solutions in recent years, face-to-face meetings remain preferred in various contexts when it is impossible or undesirable for users to compromise their ability to actively participate in a particular meeting or conference program.

[0004] Therefore, there is a need for smart video conferencing solutions that can capture and detect all objects and their interactions, as well as related changes in gestures and postures, throughout the entire meeting space. More specifically, there is a need to create a cohesive video conferencing production in real time, similar to a television studio production, for the entire video conference, and to stream such a production to remote users, thereby improving remote user engagement. [Overview of the Initiative]

[0005] Therefore, the purpose of this disclosure is to enhance video conferencing solutions using sub-symbolic artificial intelligence, thereby enabling the detection of objects and their interactions, as well as related changes in gestures and postures. A further purpose of this disclosure is to use symbolic artificial intelligence to develop a set of rules that adhere to television production principles, and to create automated video conferencing productions for real-time streaming to remote users by applying such a set of rules.

[0006] In particular, according to this disclosure, in one embodiment, an autonomous video conferencing system comprising a plurality of smart cameras is provided. Each of the plurality of smart cameras is an image sensor adapted to capture video in a video conferencing space and output an overview video stream and a focus video stream, wherein the focus video stream includes sub-images framing detected objects in the overview video stream; an audio device adapted to capture audio signals in the video conferencing space; a visual pipeline unit adapted to receive the overview video stream and the audio signals and to detect objects and poses using machine learning, wherein the objects include people and non-person items, and the poses include the posture, orientation, gestures, and orientation of the detected objects; and connected to the visual pipeline unit and the audio device The system comprises a virtual director unit adapted to frame target objects according to a predetermined set of rules, thereby enabling the focus video stream to be updated in real time, wherein the predetermined set of rules is adapted to create automated television studio productions from the video conferencing space; a stream selector connected to the virtual director unit and adapted to transition the focus video stream to an updated focus video stream; and an application programming interface connected to the virtual director unit and adapted to provide at least one connection to another smart camera among the plurality of smart cameras and one connection to a user computer.

[0007] According to one embodiment, one of the plurality of smart cameras is configured as a main camera, and each of the remaining smart cameras is configured as a peripheral camera. The application programming interface of each peripheral camera is connected to the application programming interface of the main camera. The stream selector of each peripheral camera is further configured to send its updated attention stream to the stream selector of the main camera, and the stream selector of the main camera is further configured to select one of the updated attention streams from the plurality of smart cameras as the updated main attention stream and output the updated main attention stream to the user computer.

[0008] In another embodiment, the autonomous video conferencing system further comprises a plurality of smart sensors located within the video conferencing space and adapted to capture non-image signals within the video conferencing space. Each of the plurality of smart sensors is connected to the application programming interface of the main camera, which includes an application programming interface adapted to provide input to the virtual director unit of the main camera.

[0009] In yet another embodiment, each of the plurality of smart sensors is selected from the group consisting of a touchpad, microphone, smartphone, GPS tracker, echolocation sensor, thermometer, humidity sensor, and biometric sensor.

[0010] In a further embodiment, the audio device includes an array of microphones adapted to provide the direction of the audio to the captured audio signal.

[0011] In another embodiment, the visual pipeline unit includes at least one hardware-accelerated convolutional neural network. In yet another embodiment, the visual pipeline unit is pre-trained with a training set of video and audio data adapted to a dedicated video conferencing space. In a further embodiment, the dedicated video conferencing space is selected from a group consisting of classrooms, workshops, town halls, newsrooms, executive meeting rooms, courtrooms, interview studios, and voting rooms.

[0012] In another embodiment, the predetermined set of rules includes a first rule for evaluating possible framing for each object in the video conferencing space based on a first set of parameters, thereby determining the best framing. In yet another embodiment, the first set of parameters includes: (i) whether the object is speaking; (ii) the length of time it is speaking; (iii) the direction of the object's line of sight; (iv) the degree of visibility of the object within the frame; (v) the orientation of the object; and (vi) which other objects are visible within the frame.

[0013] In a further embodiment, the predetermined set of rules further includes second rules for detecting changes in the video conferencing space and triggering frame transitions based on a second set of parameters. In another embodiment, the second set of parameters includes: (i) an object starting to speak; (ii) an object moving; (iii) an object standing up; (iv) a change in the direction of the object's line of sight; (v) an object reacting; (vi) an object pointing to a new item in the scene; (vii) an object speaking for a predetermined length of time; and (viii) a lack of meaningful reaction from other objects for a predetermined length of time.

[0014] In yet another embodiment, the predetermined set of rules further includes a third set of rules for applying an appropriate shot type to each frame based on a third set of parameters that are consistent with the principles of television studio production. In yet another embodiment, the third set of parameters includes: (i) a wide shot that frames substantially all objects and most of the video conferencing space, thereby providing overall context to the video conferencing space; (ii) an intermediate shot that frames a predetermined number of objects, focuses on the person speaking, thereby highlighting the active dialogue; and (iii) a close-up shot that frames a single object speaking for a predetermined length of time, thereby highlighting the presenter.

[0015] In another embodiment, the third set of parameters further includes: (i) a target shot that frames the target object based on a cue of the scene in the video conferencing space, including the object at the center of the line of sight from all objects in the video conferencing space, and an item raised by the object; (ii) a listener shot that frames at least one object that is not speaking, thereby highlighting the involvement of the non-speaking object in the video conferencing space; and (iii) a presenter shot that frames an object that is speaking for a predetermined longer duration compared to other objects, thereby highlighting the presenter from various camera angles and compositions in the video conferencing space. According to various embodiments, the target shot is adapted as a close-up shot, the listener shot is adapted as one of a close-up and an intermediate shot, and the presenter shot is adapted as one of a close-up and an intermediate shot.

[0016] In another embodiment, the predetermined set of rules further includes a fourth set of rules for applying a virtual director's cut to the video conferencing space based on a fourth set of parameters, thereby aligning it to a dedicated television production scenario. In yet another embodiment, the fourth set of parameters is: (i) a classroom production scenario, beginning with showing the presenter and audience using a wide shot, then transitioning to framing the presenter in a presenter shot for a predetermined length of time, followed by switching between an audience shot showing the audience and a presenter shot showing the presenter; (ii) a workshop production scenario, beginning with showing all participating objects using a wide shot, then transitioning to showing active objects using an intermediate shot, followed by showing objects moving within the video conferencing space using an intermediate shot, and after a predetermined length of time, rotating back to showing active objects within the video conferencing space; and (iii) all visual cues. The meeting room production scenario includes beginning with the overall shot that provides an understanding of the entire video conferencing space with active objects, transitioning after a predetermined length of time to framing a group of objects using an intermediate shot at a sub-position of the video conferencing space that focuses on an active object, then framing an object speaking at the sub-position using an intermediate shot that best displays the front of the object's face, then after another predetermined length of time to framing other objects in the video conferencing space using a listener-side shot that best displays the front of the object's face, and if there are no objects speaking in the video conferencing space, rotating back to an overall shot that captures all objects.

[0017] According to various embodiments, the active object includes a speaking object, a whiteboard on which its contents are being drawn, and items held up by the object during the performance, and the sub-positions of the video conferencing space include a tableside, a stage, benches, a podium, and rows of chairs. In further embodiments, the meeting room production scenario is further adapted to represent a production scenario selected from a group consisting of news broadcasts or podcasts, bilateral negotiations, court proceedings, panel discussions, and polling rallies.

[0018] In another embodiment, the predetermined set of rules further includes a fifth rule for framing clean shots of objects in the virtual meeting space. The fifth rule includes: not selecting any shots with partially visible objects; and aligning the eyes of the active object in the upper third of the frame. In yet another embodiment, the fifth rule further includes: adding spatial padding in the direction of the object's line of sight; and using intermediate shots to frame nearby active objects together. According to various embodiments, nearby active objects include objects, whiteboards, display monitors, lecterns, podiums, and items being demonstrated.

[0019] In a further embodiment, the plurality of smart cameras further include at least one peripheral smart camera located in a separate video conferencing space. The predetermined set of rules in the main camera is further adapted to create automated television studio productions relating to the combined virtual conferencing space.

[0020] In another embodiment, the autonomous video conferencing system further comprises at least one smart sensor located within the separate video conferencing space. The smart sensor is adapted to capture non-image signals within the separate video conferencing space and provide input to the virtual director unit of the main camera.

[0021] In yet another embodiment, the application programming interface of the main camera is adapted to receive input from the user computer and thereby modify the predetermined set of rules relating to the virtual director unit of the main camera.

[0022] According to the present disclosure, in another embodiment, a method is provided for creating an automated television studio production relating to a video conferencing space using virtual director assistance. The method comprises: capturing video within the video conferencing space using a plurality of image sensors; capturing audio signals within the video conferencing space using a plurality of audio devices; generating an overview video stream and a focus video stream, wherein the focus video stream includes sub-videos framing detected objects in the overview video stream; detecting objects and poses from the overview stream and the audio signals using machine learning, wherein the objects include people and non-personal items, and the poses include the posture, orientation, gestures, and direction of the detected objects; implementing a virtual director including a predetermined set of rules that conform to television studio production principles; applying the predetermined set of rules to the detected objects and poses, thereby framing the target objects and updating the focus video stream in real time; and outputting the updated focus video stream to a user computer.

[0023] In yet another embodiment, the machine learning is performed on at least one hardware-accelerated convolutional neural network. In yet another embodiment, the neural network is pre-trained with a training set of video and audio data adapted to a dedicated video conferencing space. In yet another embodiment, the dedicated video conferencing space is selected from a group consisting of classrooms, workshops, town halls, newsrooms, executive meeting rooms, courtrooms, interview studios, and voting rooms.

[0024] According to various embodiments, the predetermined set of rules includes a first rule for evaluating possible framing for each object in the video conferencing space based on a first set of parameters, thereby determining the best framing; a second rule for detecting changes in the video conferencing space and triggering frame transitions based on a second set of parameters; a third rule for applying an appropriate shot type to each frame based on a third set of parameters consistent with television studio production principles; a fourth rule for applying a virtual director's cut to the video conferencing space based on a fourth set of parameters, thereby aligning it with a dedicated television production scenario; and a fifth rule for framing clean shots of objects in the virtual conferencing space based on a fifth set of parameters.

[0025] In another embodiment, the shot type is selected from the group consisting of a wide shot, an intermediate shot, a close-up shot, a subject shot, an audience shot, and a presenter shot. In yet another embodiment, the dedicated television production scenario includes a classroom, a workshop, a meeting room, a broadcast, a bilateral negotiation, a courtroom trial, a panel discussion, and a voting rally.

[0026] In a further embodiment, the method for creating an automated television studio production further comprises capturing non-image signals within the video conferencing space using a plurality of smart sensors. Each of the plurality of smart sensors includes an application program interface connected to the virtual director, thereby providing an input to the virtual director. In various embodiments, each of the plurality of smart sensors is selected from the group consisting of a touch pad, a microphone, a smartphone, a GPS tracker, an echolocation sensor, a thermometer, a humidity sensor, and a biometric sensor.

Brief Description of the Drawings

[0027] [Figure 1] FIG. 1 is a diagram of an autonomous video conferencing system according to one embodiment.

[0028] [Figure 2] FIG. 14 is a diagram of an overview video stream and a spotlight video stream according to one embodiment.

[0029] [Figure 3] FIG. 20 is a diagram of an overall shot, a medium shot, and a close-up shot according to one of the rules of a predetermined rule set in one embodiment.

[0030] [Figure 4] FIG. 26 is a diagram of the framing of a clean shot according to one of the rules of a predetermined rule set in one embodiment.

[0031] [Figure 5] FIG. 32 is a diagram of alignment for the framing of a clean shot according to one of the rules of a predetermined rule set in one embodiment.

[0032] [Figure 6]The diagram shows two examples of small video conferencing spaces for dedicated scenarios in which an autonomous video conferencing system is deployed, according to one embodiment.

[0033] [Figure 7] The diagram shows two examples of a medium-sized video conferencing space for a dedicated scenario in which an autonomous video conferencing system is deployed, according to one embodiment.

[0034] [Figure 8] The diagram shows two examples of large video conferencing spaces for dedicated scenarios in which an autonomous video conferencing system is deployed, according to one embodiment. [Modes for carrying out the invention]

[0035] The video conferencing systems and methods of this disclosure are enhanced by multiple smart cameras and smart sensors using sub-symbolic and symbolic artificial intelligence. In various embodiments, the autonomous video conferencing system of this disclosure includes multiple smart cameras optionally coupled with multiple smart sensors positioned within a video conferencing space. Each smart camera is equipped with a visual pipeline supported by machine learning to detect objects and poses and identify speaker-listener dynamics, and a virtual director adapted to apply a predetermined set of rules consistent with television studio production principles. Simultaneously, a method is provided for creating automated television studio productions for various conferencing spaces and dedicated scenarios using virtual director assistance. Multiple smart cameras and sensors

[0036] Referring to Figure 1, an autonomous video conferencing (AVC) system according to one embodiment comprises a plurality of smart cameras (100, 101) and optionally one or more smart sensors (102). One of the plurality of smart cameras is configured as the main camera (100), and the remaining smart cameras are configured as peripheral cameras (101). Each of the plurality of smart cameras comprises an image sensor (201, 301), an audio device (204, 304), a visual pipeline unit or visual pipeline (202, 302), a virtual director unit or virtual director (203, 303), a stream selector (205, 305), and an application programming interface (API) (206, 306). Each of the one or more smart sensors comprises an API (401) that can connect to the API of a smart camera. According to one embodiment, the API of the main camera (100) is connected to the APIs of each peripheral camera (101) and each smart sensor (102). The API of the main camera is further adapted to provide a connection to a user computer (103).

[0037] Multiple smart cameras and one or more smart sensors within an AVC system are connected via Ethernet®, other local area networks, or wireless networks in various embodiments. According to one embodiment, the main camera and peripheral cameras are positioned in various ways within the video conferencing space to provide effective coverage of the video conferencing space. In another embodiment, multiple smart sensors are activated and strategically positioned within the video conferencing space to capture non-image signals and provide input to the main camera of the AVC system. Smart sensors of this disclosure include, in various embodiments, touchpads, microphones, smartphones, GPS trackers, echolocation sensors, thermometers, humidity sensors, and biometric sensors.

[0038] In an alternative embodiment, one or more peripheral cameras and smart sensors of the AVC system are located in a separate video conferencing space, which serves as a secondary space for video conferencing. These peripheral cameras and smart sensors are networked with the main camera and adapted to provide image and non-image inputs from the secondary space to the main camera. Thus, the AVC system in these alternative embodiments is further adapted to generate automated television studio productions related to the combined video conferencing space based on inputs from all cameras and smart sensors in both spaces.

[0039] The smart cameras in an AVC system are adapted to have various field-of-view angles in various embodiments. For example, if the video conferencing space is small and the AVC system has a small number of cameras, the smart cameras may have a wide field of view, for example, extending to approximately 150 degrees. On the other hand, if the video conferencing space is large and the AVC system has a large number of cameras, the smart cameras may have a narrower field of view, for example, extending to approximately 90 degrees. In another embodiment, the AVC system is equipped with smart cameras having various field-of-view angles, thereby enabling optimal coverage of the video conferencing space. In a further embodiment, the image sensors (201, 301) are adapted to zoom up to 10x, enabling close-up images of objects at the back of the video conferencing space. In an alternative embodiment, one or more smart cameras in the AVC system are adapted to capture content on or around non-person items in the video conferencing space, such as whiteboards, TV displays, posters, and demonstration stands. These cameras may be smaller than the other smart cameras in the AVC system and may be positioned differently, and may be mounted on the ceiling to provide effective coverage of the target content.

[0040] In a further embodiment, the audio device (204, 304) within the smart camera is a microphone array adapted to capture audio signals from various locations around the camera. By using signals from various microphones, the smart camera can determine the direction of audio (DOA) and identify whether silence is present at a particular location or direction. This information is also available to the AVC system's visual pipeline (202, 302) and virtual director (203, 303). In a further alternative embodiment, a high-performance computing device is connected to the AVC system via an Ethernet switch and adapted to provide additional computing power to the AVC system. It has one or more high-performance CPUs and GPUs and operates a portion of the visual pipeline for the main camera and any designated peripheral cameras.

[0041] Therefore, by including multiple smart cameras and smart sensors within the AVC system, effective and robust coverage of various video conferencing spaces and scenarios is possible. By positioning multiple smart cameras within the video conferencing space to collaborate in framing objects within the meeting from various camera angles and zoom levels, the AVC system of this disclosure creates a more comprehensive, natural, and engaging experience for all participants, including remote users. Visual pipeline; overview stream & focus stream

[0042] Referring to Figure 2, for each smart camera in the AVC system, there are two video streams internally: an overview stream and a focus stream. The overview stream displays the entire scene and is consumed by the visual pipeline (202, 302) as shown in Figure 1. The focus stream is a high-resolution stream that frames the target object in response to activity within the video conferencing space over time. Here, video settings are applied under the direction of the virtual director (203, 303), which transitions the focus stream to an updated focus stream, as will be discussed in detail below.

[0043] The visual pipeline of this disclosure is adapted to process incoming overview streams and audio signals and to detect objects and poses using machine learning. In various embodiments, objects include people and non-person items, and poses include posture, orientation, gestures, and direction. In one embodiment, the visual pipeline includes one or more hardware-accelerated programmable convolutional neural networks that employ pre-trained weights to enable the detection of specific characteristics of objects in the field of view. For example, the visual pipeline detects where objects are in the field of view of a smart camera, the degree of their visibility in the field of view, whether they are speaking or not, their facial expressions, their body posture, and their head posture. The visual pipeline also tracks each object over time to determine where the objects were previously in the field of view, whether they are moving or not, and in which direction they are facing.

[0044] One advantage of the visual pipeline, in various embodiments, leveraging subsymbolic artificial intelligence to detect objects and their activities and interactions is that these convolutional neural networks are trained to be unbiased based on characteristics such as gender, age, race, scene, lighting, and size. This allows the AVC system to create more accurate and natural video stream productions for the entire scene and all objects within the video conferencing space.

[0045] The visual pipeline in various embodiments is adapted to operate on a GPU or other dedicated chipset having hardware accelerators for the relevant mathematical operations in its convolutional neural network. In alternative embodiments, the visual pipeline operates on the CPU capabilities available within the AVC system. Further optimization of the visual pipeline to adapt to its hardware chipset in a particular embodiment is achieved by replacing the mathematical operations within its convolutional neural network architecture with uniform mathematical operations supported by the chipset. Thus, dedicated hardware support for the visual pipeline enables it to perform object and pose detection at high frequencies and fast processing times, which further allows the AVC system to respond quickly and on demand to changes in the smart camera's field of view.

[0046] In a particular embodiment, the visual pipeline is pre-trained by running thousands of images and videos related to scenes and detection targets within a video conferencing space. During the above training, the visual pipeline is evaluated using a loss function that measures how well it performs a particular detection. Feedback from the loss function is then used to adjust the weights and parameters of the visual pipeline until it performs a particular detection at a predetermined level of achievement. In one embodiment, the visual pipeline is further pre-trained and fine-tuned with a training set of video and audio data adapted to a dedicated video conferencing space for a dedicated scenario. For example, the visual pipeline may be fine-tuned for a classroom, workshop, town hall, newsroom, executive boardroom, courtroom, interview studio, or polling room to support a dedicated scenario, such as a lecture, interview, news broadcast or podcast, court proceedings, workshop, bilateral negotiation, or polling rally.

[0047] The visual pipeline is further adapted to collect and process audio signals from microphones or other audio devices within the AVC system in various embodiments. It is possible to distinguish speech, including whether an object has increased or decreased its volume, depending on what is happening in the video conferencing space. In one embodiment, the visual pipeline is adapted to classify the topic of conversation based on the audio signal. Speech that does not belong to an object is classified as artificial sound and may be attributed to other sources, such as a loudspeaker. The speech classification and features are combined by the visual pipeline with other information and knowledge it detects and collects from image data about the relevant objects and their activities and interactions, thereby generating a comprehensive understanding of the video conferencing space and all detected objects. The visual pipeline makes this corpus of comprehensive understanding available to a virtual director, who is further responsible for selecting the best shots and creating automated television studio productions from the video conferencing space.

[0048] Virtual director; a predetermined set of rules. As discussed above, the AVC system's virtual director (203, 303), as shown in Figure 1, is connected to the visual pipeline (202, 302) and audio devices (204, 304) and is adapted to frame target objects according to a predetermined set of rules, thereby enabling the focus stream to be updated in real time. The predetermined set of rules is formulated and implemented for the virtual director in accordance with the principles of television studio production, as will be discussed in detail below. The transition from the focus stream to the updated focus stream is performed by the stream selector (205, 305) at the direction of the virtual director within each smart camera.

[0049] Multiple smart cameras in the AVC system work seamlessly together to generate and update attention streams for various locations and target objects within the video conferencing space. The main camera's virtual director (203) is connected to the virtual directors (303) of each peripheral camera via their respective APIs (206, 306), as shown in Figure 1. The main camera's stream selector (205) is connected to the stream selectors (305) of each peripheral camera and is adapted to consume attention streams from the main camera and all peripheral cameras. The main camera's stream selector (205), at the direction of the main camera's virtual director (203), is further responsible for selecting one of the updated attention streams from all cameras in the AVC system as the updated main attention stream. This updated main attention stream is the stream output that becomes available to the user computer, as shown in Figure 1. Therefore, the main camera's virtual director (203) is the brain or command center of the AVC system, responsible for creating automated television studio productions for the video conferencing space based on inputs from all cameras and any smart sensors deployed within the AVC system.

[0050] The virtual director of this disclosure is a software component that leverages rule-based symbolic artificial intelligence to optimize decision-making in various embodiments. The virtual director takes input from the visual pipeline and determines which frames from which cameras in the AVC system should be selected and ultimately streamed to the user. In certain embodiments, the virtual director achieves this by evaluating possible framing for each object in the video conferencing space and their activities over time. For each object captured by a particular smart camera in the AVC system, for example, the virtual director evaluates various crops of the image in which the object is visible in order to discover the best frame for that object in the context of their activities and interactions. The virtual director then determines appropriate video settings to transition the attention stream of a particular camera to the selected best frame. The stream selector then performs the transition and update of the attention stream using these video settings, at the direction of the virtual director.

[0051] In certain embodiments, the main camera's virtual director further acquires input from one or more smart sensors deployed within the AVC system, through their respective APIs. Input from the smart sensors includes non-image signals or cues in general about the video conferencing space and all objects within it, such as the position, movement, and physiological or biometric characteristics of objects.

[0052] The AVC system's API is adapted to send messages between various components via an internal network bus. These messages include information about the status of each smart camera, such as whether it is connected, what type of software it is running, and its current health status. The API also communicates what the camera has detected from the image data, such as where objects are detected in the image, where they are located within the conference space, and other information detected by the visual pipeline. The API further communicates video settings applied to the attention stream, such as image characteristics, color, and brightness. In addition, the API is adapted to communicate virtual director parameters, allowing the AVC system to automatically set and adjust virtual director rule sets and related parameters for all of its component cameras, and allowing AVC system users to personalize the virtual director experience by modifying specific parameters and rules with respect to its predefined rule set.

[0053] As discussed above, the virtual director of this disclosure utilizes rule-based decision-making that leverages symbolic artificial intelligence. A predetermined set of rules is formulated in various embodiments according to television studio production principles. This enables the AVC system to generate automated television studio productions of the conference experience that resemble television productions directed by real-world professionals. In certain embodiments, the predetermined set of rules for each camera in the AVC system may be differently adapted for its location and target object or content within the video conferencing space. In one embodiment, as shown in Figure 1, the predetermined set of rules applied by the virtual director (203) of the main camera contains more rules than the predetermined set of rules applied by the virtual director (303) of the peripheral cameras. In an alternative embodiment, the user is provided with the option to modify or tune specific rules and parameters of the predetermined set of rules in the AVC system through the main camera's API connected to the main camera's virtual director unit (203), as shown in Figure 1.

[0054] According to one embodiment, a predetermined set of rules includes a first set of rules for evaluating possible framing for each object in a video conferencing space based on a first set of parameters. The first set of rules determines the best frame for each object. The first set of parameters in one embodiment include (i) whether the object is speaking; (ii) the length of time it is speaking; (iii) the direction of the object's line of sight; (iv) the degree of visibility of the object in the frame; (v) the orientation of the object; and (vi) which other objects are visible in the frame.

[0055] A predetermined set of rules includes, in another embodiment, a second set of rules for detecting changes in the video conferencing space based on a second set of parameters. The second set of rules triggers frame transitions. In one embodiment, the second set of parameters includes: (i) an object starting to speak; (ii) an object moving; (iii) an object standing up; (iv) an object changing its line of sight; (v) an object reacting; (vi) an object pointing to a new item in the scene; (vii) an object speaking for a predetermined length of time; and (viii) a lack of meaningful reaction from other objects for a predetermined length of time.

[0056] A predetermined set of rules includes a third set of rules in another embodiment for applying an appropriate shot type to each frame based on a third set of parameters. In one embodiment, the third set of parameters includes: (i) a wide shot that frames substantially all objects and most of the video conferencing space, thereby providing overall context to the video conferencing space (see the top frame in Figure 3); (ii) an intermediate shot that frames a predetermined number of objects, focuses on the person speaking, thereby highlighting the active dialogue (see the middle frame in Figure 3); and (iii) a close-up shot that frames a single object speaking for a predetermined length of time, thereby highlighting the presenter (see the bottom frame in Figure 3).

[0057] Multiple parameters of the third rule in another embodiment include: (i) a target shot that frames the target object based on a scene cue in the video conferencing space, including the object at the center of the line of sight from all objects in the video conferencing space, and items raised by the object; (ii) a listener shot that frames at least one object that is not speaking, thereby highlighting the involvement of non-speaking objects in the video conferencing space; and (iii) a presenter shot that frames the object that speaks for the longest duration compared to other objects, thereby highlighting the presenter from various camera angles and compositions in the video conferencing space. In various embodiments, the target shot is adapted as a close-up shot, the listener shot is adapted as a close-up or mid-range shot, and the presenter shot is adapted as a close-up or mid-range shot.

[0058] A predetermined set of rules includes a fourth rule in another embodiment for applying a virtual director's cut to a video conferencing space based on a fourth set of parameters. The fourth rule aligns the video settings to a dedicated television production scenario. In one embodiment, the fourth set of parameters includes classroom production scenarios, workshop production scenarios, and meeting room production scenarios. The meeting room production scenario is further adapted in certain embodiments for news broadcasts or podcasts, bilateral negotiations, court proceedings, panel discussions, and polling rallies.

[0059] A predetermined set of rules includes a fifth rule in further embodiments for framing clean shots of objects within the virtual meeting space. In one embodiment, the fifth rule includes not selecting any shots that have partially visible objects. Referring to Figure 4, for example, the left frame is a clean shot, while the right frame is not a clean shot and is not selected in the AVC system under the fifth rule. In another embodiment, the fifth rule includes aligning the eye of the active object to the top third of the frame. Referring to Figure 5, for example, each of the images (top, middle, bottom) is optimally aligned so that the eye of the active object is in the top third of the frame. These shots are clean shots that are framed and selected in the AVC system under the fifth rule.

[0060] A fifth rule in yet another embodiment includes adding spatial padding in the line of sight of an object and using a mid-shot to frame nearby active objects together. Nearby active objects include, in various embodiments, objects, whiteboards, display monitors, lecterns, podiums, and items being demonstrated. Exclusive scenario

[0061] The AVC system of this disclosure is adapted in various embodiments for various dedicated video conferencing spaces for dedicated video conferencing scenarios. The visual pipeline unit of the smart camera within the AVC system is pre-trained with training sets of video and audio data adapted to dedicated video conferencing spaces, including classrooms, workshops, town halls, newsrooms, executive conference rooms, courtrooms, interview studios, and voting rooms, according to a particular embodiment.

[0062] As discussed above, among the predetermined set of rules implemented and applied by the virtual director unit of the AVC system, there is a fourth rule for creating a virtual director's cut based on a fourth set of parameters. This fourth rule for the virtual director's cut is designed to adapt the video conferencing space to a dedicated television production scenario. In one embodiment, the fourth set of parameters includes classroom production scenarios, workshop production scenarios, and meeting room production scenarios.

[0063] In one embodiment, a classroom scenario begins by showing the presenter and audience using a whole shot, then transitions to framing the presenter in a presenter shot for a predetermined length of time, and then changes to switching between an audience shot showing the audience and a presenter shot showing the presenter.

[0064] In one embodiment, the workshop deliverable scenario begins by using a whole shot to show all participating objects, then transitions to using intermediate shots to show active objects, then changes to using intermediate shots to show objects moving within the video conferencing space, and finally rotates back to showing active objects within the video conferencing space after a predetermined length of time.

[0065] In one embodiment, the meeting room production scenario begins with a wide shot that provides an understanding of the entire video conferencing space with all visible objects, transitions after a predetermined length of time to framing a group of objects using intermediate shots at a sub-position of the video conferencing space that focus on the active object, then changes to framing the object speaking at the sub-position using intermediate shots that best show the front of the object's face, then after another predetermined length of time, switches to framing other objects in the video conferencing space using listener-side shots that best show the front of the object's face, and finally rotates back to a wide shot that captures all objects if there is no object speaking in the video conferencing space. In various embodiments, the active object can represent the object speaking, a whiteboard on which its contents are being drawn, or an item being held up by the object during a demonstration. In various embodiments, the sub-position of the video conferencing space can represent the side of a table, a stage, a bench, a podium, or a row of chairs.

[0066] Meeting room production scenarios in specific embodiments have been further adapted to represent news broadcasts or podcasts, interviews, executive meetings, bilateral negotiations, court proceedings, panel discussions, and polling rallies.

[0067] For various dedicated scenarios, the main camera and peripheral cameras are strategically and variedly positioned within the video conferencing space to provide effective spatial coverage. One or more optional smart sensors are additionally distributed within the video conferencing space, having connections to the main camera and providing input to the main camera's virtual director unit. Examples of the AVC system of this disclosure deployed for dedicated scenarios are shown in Figures 6 to 8. In these drawings, each point represents a smart camera or smart sensor of the AVC system, and each circle represents an object, including a person or a non-person item, such as a chair or demonstration material. Small rectangles along each side of these drawings represent a TV, whiteboard, poster, or projection display (601, 602, 701, 702, 703, 704, 801, 802). A square or rectangle in the middle of the drawing represents a table.

[0068] Referring to Figure 6, both the top and bottom diagrams depict a small video conferencing space where the AVC system is deployed. The top configuration represents a small meeting and panel discussion scenario in one embodiment. The bottom configuration represents a news broadcast or podcast and interview scenario in one embodiment.

[0069] Referring to Figure 7, both the top and bottom diagrams depict a medium-sized video conferencing space where the AVC system is deployed. The top configuration represents an executive meeting scenario in one embodiment. The bottom configuration represents a bilateral negotiation in one embodiment.

[0070] Referring to Figure 8, both the top and bottom diagrams depict a large video conferencing space where the AVC system is deployed. The top configuration represents a workshop scenario in one embodiment. The bottom configuration represents a classroom scenario in one embodiment.

[0071] The descriptions of various embodiments, including the drawings and examples, are for illustrative purposes only and do not limit the present invention or its various embodiments.

Claims

1. An autonomous video conferencing system equipped with multiple smart cameras, wherein each of the multiple smart cameras is: An image sensor adapted to capture video within a video conferencing space and output an overview video stream and a focus video stream, wherein the focus video stream includes sub-images framing detected objects within the overview video stream; An audio device adapted to capture audio signals within the aforementioned video conferencing space; A visual pipeline unit receiving the overview video stream and the audio signal, and adapted to detect objects and poses using machine learning, wherein the objects include people and non-person items, and the poses include the posture, orientation, gestures, and direction of the detected objects; A virtual director unit connected to the visual pipeline unit and the audio device, adapted to frame target objects according to a predetermined set of rules, thereby enabling the focus video stream to be updated in real time, wherein the predetermined set of rules is adapted to create automated television studio productions from the video conferencing space; A stream selector connected to the virtual director unit and adapted to transition the featured video stream to an updated featured video stream; and An application programming interface connected to the virtual director unit and adapted to provide at least one connection to another smart camera among the plurality of smart cameras and one connection to a user computer. It has, One of the plurality of smart cameras is configured as the main camera, each of the remaining smart cameras is configured as a peripheral camera, the application programming interface of each peripheral camera is connected to the application programming interface of the main camera, the stream selector of each peripheral camera is further configured to send its updated attention stream to the stream selector of the main camera, and the stream selector of the main camera is further configured to select one of the updated attention streams from the plurality of smart cameras as the updated main attention stream and output the updated main attention stream to the user computer. Autonomous video conferencing system.

2. The autonomous video conferencing system according to claim 1, further comprising a plurality of smart sensors arranged within the video conferencing space and adapted to capture non-image signals within the video conferencing space, each of the plurality of smart sensors including an application program interface, the application program interface being connected to the application program interface of the main camera and thus adapted to provide input to the virtual director unit of the main camera.

3. The autonomous video conferencing system according to claim 2, wherein each of the plurality of smart sensors is selected from the group consisting of a touchpad, a microphone, a smartphone, a GPS tracker, an echolocation sensor, a thermometer, a humidity sensor, and a biometric sensor.

4. The autonomous video conferencing system according to any one of claims 1 to 3, wherein the audio device includes an array of microphones adapted to provide the direction of audio to the captured audio signal.

5. The autonomous video conferencing system according to any one of claims 1 to 3, wherein the visual pipeline unit includes at least one hardware-accelerated convolutional neural network.

6. The autonomous video conferencing system according to claim 5, wherein the visual pipeline unit is pre-trained using a training set of video and audio data adapted to a dedicated video conferencing space.

7. The autonomous video conferencing system according to claim 6, wherein the dedicated video conferencing space is selected from a group consisting of classrooms, workshops, town halls, newsrooms, executive conference rooms, courtrooms, interview studios, and voting rooms.

8. The autonomous video conferencing system according to any one of claims 1 to 3, wherein the predetermined set of rules includes a first rule for evaluating possible framing for each object in the video conferencing space based on a first set of parameters, thereby determining the best framing.

9. The autonomous video conferencing system according to claim 8, wherein the first set of parameters includes: (i) whether the object is speaking; (ii) the length of time the object is speaking; (iii) the direction of the object's line of sight; (iv) the degree of visibility of the object within the frame; (v) the orientation of the object; and (vi) which other objects are visible within the frame.

10. The autonomous video conferencing system according to claim 8, wherein the predetermined set of rules further includes a second set of rules for detecting changes in the video conferencing space based on a second set of parameters and for triggering frame transitions.

11. The second set of parameters mentioned above are: (i) the object begins to speak; (ii) the object moves; (iii) the object stands up; (iv) the direction of the object's line of sight changes; (v) the object reacts; (vi) the object points to a new item in the scene; (vii) the object has spoken for a predetermined length of time; and (viiii) the absence of a meaningful reaction from other objects for a predetermined length of time. The autonomous video conferencing system according to claim 10, including the above.

12. The autonomous video conferencing system according to claim 10, wherein the predetermined set of rules further includes a third set of rules for applying an appropriate shot type to each frame based on a third set of parameters consistent with the principles of television studio production.

13. The multiple parameters of the third set described above are: (i) a wide shot that frames substantially all objects and most of the video conferencing space, thereby providing overall context to the video conferencing space; (ii) an intermediate shot that frames a predetermined number of objects and focuses on the person speaking, thereby highlighting the active dialogue; and (iii) a close-up shot that frames one object speaking for a predetermined length of time, thereby highlighting the presenter. The autonomous video conferencing system according to claim 12.

14. The autonomous video conferencing system according to claim 13, wherein the third set of parameters further includes: (i) a target shot that frames a target object based on a scene cue in the video conferencing space, including an object that is at the center of the line of sight from all objects in the video conferencing space, and an item raised by the object; (ii) a listener shot that frames at least one object that is not speaking, thereby highlighting the involvement of the non-speaking object in the video conferencing space; and (iii) a presenter shot that frames the object that is speaking for the longest duration compared to other objects, thereby highlighting the presenter from various camera angles and compositions in the video conferencing space, wherein the target shot is adapted as a close-up shot, the listener shot is adapted as one of a close-up and an intermediate shot, and the presenter shot is adapted as one of a close-up and an intermediate shot.

15. The autonomous video conferencing system according to claim 12, wherein the predetermined set of rules further includes a fourth set of rules for applying a virtual director's cut to the video conferencing space based on a fourth set of parameters, thereby aligning it to a dedicated television production scenario.

16. The fourth set of parameters described above are: (i) A classroom production scenario that begins with a wide shot showing the presenter and audience, then transitions to framing the presenter in a presenter shot for a predetermined length of time, followed by switching between an audience shot showing the audience and a presenter shot showing the presenter; (ii) A workshop production scenario that begins with a wide shot showing all participating objects, then transitions to showing active objects using intermediate shots, followed by showing objects moving within the video conferencing space using intermediate shots, and after a predetermined length of time, rotates back to showing active objects within the video conferencing space; and (iii) The same The meeting room production scenario includes beginning with the overall shot that provides an understanding of the entire video conferencing space, transitioning after a predetermined length of time to framing a group of objects using an intermediate shot at a sub-position of the video conferencing space that focuses on an active object, then framing an object speaking at the sub-position using an intermediate shot that best displays the front of the object's face, then after another predetermined length of time to framing other objects in the video conferencing space using a listener-side shot that best displays the front of the object's face, and if there are no objects speaking in the video conferencing space, rotating back to an overall shot that captures all objects. The active object includes the speaking object, the whiteboard on which its contents are being drawn, and the items being held up by the object during the demonstration, and the sub-positions of the video conferencing space include the side of the table, the stage, the bench, the podium, and the rows of chairs. The autonomous video conferencing system according to claim 15.

17. The autonomous video conferencing system according to claim 16, wherein the meeting room production scenario is further adapted to represent a production scenario selected from the group consisting of news broadcasts or podcasts, bilateral negotiations, court proceedings, panel discussions, and voting rallies.

18. The autonomous video conferencing system according to claim 15, wherein the predetermined set of rules further includes a fifth rule for framing clean shots of objects in a virtual meeting space, the fifth rule comprising: not selecting any shots having partially visible objects; and aligning the eyes of an active object in the upper third of the frame.

19. The fifth rule further comprises adding spatial padding in the direction of the object's line of sight; and using intermediate shots to frame nearby active objects together, wherein the nearby active objects include objects, whiteboards, display monitors, lecterns, podiums, and items being demonstrated, according to the autonomous video conferencing system of claim 18.

20. The autonomous video conferencing system according to any one of claims 1 to 3, wherein the plurality of smart cameras further include at least one peripheral smart camera located in a separate video conferencing space, and the predetermined set of rules in the main camera is further adapted to create an automated television studio production relating to the combined virtual conferencing space.

21. The autonomous video conferencing system according to claim 20, further comprising at least one smart sensor located within the separate video conferencing space, wherein the smart sensor is adapted to capture non-image signals within the separate video conferencing space and provide input to the virtual director unit of the main camera.

22. The autonomous video conferencing system according to any one of claims 1 to 3, wherein the application programming interface of the main camera is adapted to receive input from the user computer and thereby modify the predetermined set of rules relating to the virtual director unit of the main camera.

23. A method for creating an automated television studio production relating to a video conferencing space using virtual director support, comprising: capturing video within the video conferencing space using multiple image sensors; capturing audio signals within the video conferencing space using multiple audio devices; generating an overview video stream and a focus video stream, wherein the focus video stream includes sub-videos framing detected objects in the overview video stream; detecting objects and poses from the overview video stream and the audio signals using machine learning, wherein the objects include people and non-personal items, and the poses include the posture, orientation, gestures, and direction of the detected objects; implementing a virtual director including a predetermined set of rules that conform to television studio production principles; applying the predetermined set of rules to the detected objects and poses, thereby framing the target objects and updating the focus video stream in real time; and outputting the updated focus video stream to a user computer.

24. The method according to claim 23, wherein the machine learning is performed on at least one hardware-accelerated convolutional neural network.

25. The method according to claim 24, wherein the neural network is pre-trained using a training set of video and audio data adapted for a dedicated video conferencing space.

26. The method according to claim 25, wherein the dedicated video conferencing space is selected from the group consisting of a classroom, workshop, town hall, newsroom, executive conference room, courtroom, interview studio, and voting room.

27. The method according to any one of claims 23 to 26, wherein the predetermined set of rules includes a first set of rules for evaluating possible framing for each object in the video conferencing space based on a first set of parameters, thereby determining the best framing.

28. The method according to claim 27, wherein the predetermined set of rules further includes a second set of rules for detecting changes in the video conferencing space based on a second set of parameters and for triggering frame transitions.

29. The method according to claim 28, wherein the predetermined set of rules further includes a third set of rules for applying an appropriate shot type to each frame based on a third set of parameters consistent with the principles of television studio production.

30. The method according to claim 29, wherein the shot type is selected from the group consisting of a whole shot, an intermediate shot, a close-up shot, a target shot, a listener-side shot, and a presenter shot.

31. The method according to claim 30, wherein the predetermined set of rules further includes a fourth set of rules for applying a virtual director's cut to the video conferencing space based on a fourth set of parameters, thereby aligning it to a dedicated television production scenario.

32. The method according to claim 31, wherein the dedicated television production scenario includes classrooms, workshops, meeting rooms, broadcasts, bilateral negotiations, court proceedings, panel discussions, and voting rallies.

33. The method according to claim 32, wherein the predetermined set of rules further includes a fifth set of rules for framing clean shots of objects in a virtual meeting space based on a fifth set of parameters.

34. The method according to any one of claims 23 to 26, further comprising the step of capturing non-image signals in the video conferencing space using a plurality of smart sensors, each of the plurality of smart sensors including an application program interface connected to the virtual director, thereby providing input to the virtual director.

35. The method according to claim 34, wherein each of the plurality of smart sensors is selected from the group consisting of a touchpad, a microphone, a smartphone, a GPS tracker, an echolocation sensor, a thermometer, a humidity sensor, and a biometric sensor.

Citation Information

Patent Citations

  • Generating real-time director's cuts of live-streamed events using roles

    US20200267427A1

  • Smart video conferencing system

    US9270941B1

  • Systems and methods for detection and display of whiteboard text and / or an active speaker

    WO2021226821A1

  • Autonomous golf competition systems and methods

    WO2021248050A1