Video conferencing system for hybrid conversations
By simulating local participation of remote participants in the video conferencing system, and using multi-camera and speaker systems for localized rendering of visual and audio, the problem of remote participants being ignored and unable to participate equally in the hybrid talks is solved, achieving a more natural interaction and collaboration effect.
Patent Information
- Application Number
- CN202480004909.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-30
- Filing Date
- 2024-03-29
- Publication Date
- 2025-07-04
AI Technical Summary
Existing video conferencing systems In hybrid talks, interactions between remote participants and local participants are unnatural and difficult to form effective social connections. Remote participants are often overlooked and cannot participate in talks equally, and collaboration and brainstorming tasks are challenging.
Using hardware and software configured to simulate local attendees for remote attendees, the remote attendees feels present by displaying video feeds with segmented visual consistency among human attendees, ensuring that remote attendees are seen in a consistent order, and localized rendering of audio and video using multi-camera and speaker systems, enhancing the sense of presence of remote attendees.
It improves the sense of participation and equality of remote participants, enhances the understanding of social signals, improves the reading of body language and facial expressions of remote participants, and promotes a more natural dialogue flow and collaboration effect.
Smart Images

Figure CN120266471A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 493,307, filed on March 30, 2023, the disclosure of which is hereby incorporated by reference in its entirety. Technical Field
[0002] Implementations relate to video conferencing systems. Background Art
[0003] Hybrid meetings (e.g., video conference meetings) can have a mix of participants. For example, some participants can be in an office in a conference room, while other participants can be remote participants connecting to the meeting from a remote location such as their home or another office. Summary of the Invention
[0004] Implementations relate to a video conferencing system that uses hardware and software configured to simulate local participation of remote participants. Simulating local participation of remote participants is sometimes referred to as a technique for implementing a presence enhancement feature. The video conferencing system can be configured to display a video feed with visual consistency among segmented human participants, or in other words, display a video feed with visual consistency among video conference participants at different physical locations (e.g., local and remote participants).
[0005] In a general aspect, an apparatus, system, non - transitory computer - readable medium (having computer - executable program code stored thereon that is executable on a computer system), and / or method can perform a method execution process including the following operations: receiving, by a first remote device associated with a first remote participant, a remote video stream from a second remote device that includes content associated with a second remote participant of a video conference; receiving, by the first remote device, a local video stream from a local device associated with a local participant of the video conference, the local video stream including an indication of a display order of the first remote participant, the second remote participant, and the local participant; and rendering, based on the display order, the remote video stream and the local video stream on the first device.
[0006] In another general aspect, an apparatus, system, non - transitory computer - readable medium (having computer - executable program code stored thereon that is executable on a computer system), and / or method can perform a method execution process including the following operations: generating a user interface (UI) that includes: a first display portion corresponding to a first camera and a second display portion corresponding to a second camera; rendering the UI on a display of the video conferencing system; receiving a video stream corresponding to a remote participant of the video conference; and rendering the remote participant in one of the first display portion or the second display portion, the rendering including standardizing the presentation of the remote participant relative to the local participant.
[0007] In another general aspect, an apparatus, a system, a non-transitory computer-readable medium (having computer-executable program code stored thereon that can be executed on a computer system), and / or a method can perform a process using a method that includes the following operations: generating a user interface (UI) that includes: a first display portion corresponding to a first camera and a second display portion corresponding to a second camera; rendering the UI on a display of a video conferencing system; receiving a video stream corresponding to a remote participant in the video conference; and generating a video stream of a local participant that includes an indication of a display order of the local participant relative to the remote participant. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Example implementations will be more fully understood from the detailed description given below and the accompanying drawings, in which like elements are represented by like reference numerals, which are given by way of illustration only and thus do not limit example implementations.
[0009] Figure 1A and Figure 1B shows a front view block diagram of a video conferencing system according to at least one example implementation.
[0010] Figure 1C shows a front view block diagram of a video conferencing system according to at least one example implementation.
[0011] Figure 2 shows a top view graphical representation of a video conferencing system according to at least one example implementation.
[0012] Figure 3 shows a front view block diagram of a display of a video conferencing system according to at least one example implementation.
[0013] Figure 4 shows a block diagram of a video conferencing system according to at least one example implementation.
[0014] Figure 5 shows a rear view block diagram of a display according to at least one example implementation.
[0015] Figure 6A 、 Figure 6B and Figure 6C shows a block diagram of a multi-camera display arrangement according to at least one example implementation.
[0016] Figure 6D and Figure 6E shows a block diagram of a display order according to at least one example implementation.
[0017] Figure 7 is a block diagram of a method of operating a video conferencing system according to an example implementation.
[0018] Figure 8 It is a block diagram of a method for operating a video conferencing system according to an example implementation.
[0019] Figure 9 It is a block diagram of a method for operating a video conferencing system according to an example implementation.
[0020] It should be noted that these figures are intended to illustrate the general characteristics of the methods and / or structures utilized in certain example implementations and supplement the written description provided below. However, these drawings are not necessarily to scale and may not accurately reflect the exact structure or performance characteristics of any given implementation, and should not be construed as limiting or restricting the scope of values or properties covered by the example implementations. For example, for clarity, the positioning of modules and / or structural elements may be reduced or enlarged. The use of like or identical reference numerals in the various drawings is intended to indicate the presence of like or identical elements or features. Detailed Description
[0021] Hybrid meetings (e.g., video conference meetings) can have a mix of participants. For example, some participants may be in an office in a conference room (e.g., local participants), while other participants can be remote participants connecting to the meeting from their homes or other offices. At least one problem may be that hybrid meetings using some video conferencing systems are less effective than in-person meetings. For example, in a hybrid meeting, the conversation flow can be awkward and unnatural, it is difficult to read body language, other important social cues (such as gaze and attention) may be missed and / or misinterpreted, people feel distant and have difficulty forming strong social connections, remote participants are often overlooked and unable to participate equally in the meeting, and / or collaborative and brainstorming tasks are challenging.
[0022] At least one technical solution to this problem is to use a video conferencing system with hardware and software configured to simulate the local participation of remote participants (sometimes referred to as presence enhancement features). The video conferencing system can be configured to display a video feed with visual consistency (e.g., consistent display order) among segmented human participants (e.g., local and remote participants). For example, local participants see remote participants in a consistent order, and remote participants see other remote participants in the same order, where local participants are seen within the same order as if they were inserted (see the counterclockwise order described below). The video conferencing system can interoperate with existing systems and operate on existing network systems.
[0023] Figure 1A Shows a front view block diagram of a video conferencing system according to at least one example implementation. As Figure 1AAs shown, the video conferencing system 100 includes displays 105A, 105B, 105C and peripheral device bars 110A, 110B, 110C. Figure 1B A front view block diagram of a video conferencing system according to at least one example implementation is shown. As Figure 1B shown, the video conferencing system 100 includes a display 105D and peripheral device bars 110A, 110B, 110C. Figure 1A The video conferencing system 100 using multiple displays (e.g., three (3)) is shown, and Figure 1B a video conferencing system 100 having a single wide display is shown. The wide display may be a wide aspect ratio display. For example, the wide display may have an aspect ratio of 16:9.
[0024] As Figure 1A and Figure 1B shown, the remote participants 115A, 115B, 115C and local participants 120A, 120B displayed on the displays 105A, 105B, 105C may be locally located, for example, in a conference room. The local participants 120A, 120B may, for example, be sitting at a table (not shown). The video conferencing system 100 may be configured to implement a presence enhancement feature when the remote participants 115A, 115B, 115C are displayed on the displays 105A, 105B, 105C. The peripheral device bars 110A, 110B, 110C may be enclosures in which non-display devices may be held together as a unit. The peripheral device bars 110A, 110B, 110C may include, for example, cameras and speakers. In some implementations, the peripheral device bars 110A, 110B, 110C may include, for example, three (3) cameras and two (2) speakers. The remote participants may be associated with remote devices (e.g., computing devices) operating in the video conference and local devices operating in the video conference. Thus, the local devices may correspond to the video conferencing system 100, and the remote devices may be any other computing devices operating in the video conference. The remote devices may be configured to stream video associated with the remote participants. The local devices may be configured to stream video associated with the local participants. The video streams may be associated with the video conference.
[0025] For example, remote participants 115A, 115B, and 115C can be rendered on displays 105A, 105B, and 105C using natural scale rendering. For example, remote participants 115A, 115B, and 115C can be rendered on displays 105A, 105B, and 105C using optimized display space allocation. For example, remote participants 115A, 115B, and 115C can be rendered on displays 105A, 105B, and 105C using a common background. For example, the video streams transmitted to remote participants 115A, 115B, and 115C can be streamed with a spatially consistent mapping. In other words, the display order or rendering position of remote participants 115A, 115B, and 115C can be consistent across local and remote video conferencing systems.
[0026] In some implementations, a video conference can include a local video conferencing system (e.g., video conferencing system 100) and at least one remote video conferencing system (e.g., video conferencing system 150). As Figure 1A and Figure 1B shown, the mapping, display order, or rendering position of remote participants from left to right are remote participant 115A, remote participant 115B, and remote participant 115C. Additionally, the positions (or physical locations) of local participants are local participant 120A between remote participant 115A and remote participant 115B, and local participant 120B between remote participant 115B and remote participant 115C. The mapping, spatial mapping, display positioning, display order, display position, rendering positioning, rendering order, rendering position, and / or others of the participants can be from one of many perspectives. For example, the mapping, spatial mapping, display positioning, display order, display position, rendering positioning, rendering order, rendering position, and / or others of the participants can be viewed from above in a clockwise or counterclockwise circular rotation. For example, the mapping, spatial mapping, display positioning, display order, display position, rendering positioning, rendering order, rendering position, and / or others of the participants can be viewed linearly from left to right or from right to left from the perspective of a remote participant. For example, the mapping, spatial mapping, display positioning, display order, display position, rendering positioning, rendering order, rendering position, and / or others of the participants can be viewed linearly from left to right or from right to left from the perspective of a local participant.
[0027] Thus, the mapping of participants, spatial mapping, display positioning, display order, display location, rendering positioning, rendering order, rendering location, and / or others can be the remote participants 115A, 115B, 115C, local participant 120B, and local participant 120A (e.g., clockwise), or the remote participants 115C, 115B, 115A, local participant 120A, and local participant 120B (e.g., counterclockwise). Thus, the mapping of participants, spatial mapping, display positioning, display order, display location, rendering positioning, rendering order, rendering location, and / or others can be the remote participant 115A, local participant 120A, remote participant 115B, local participant 120B, and remote participant 115C (e.g., from left to right), or the remote participant 115C, local participant 120B, remote participant 115B, local participant 120A, and remote participant 115A (e.g., from right to left). Such mapping or display order should be consistent across all video conferencing systems in a video conference. Note that on a remote video conferencing system, remote participants may not be rendered, but the remaining participants should be rendered in the above mapping or display order. For example, the remote participants 115A, 115B, 115C can be rendered on the displays 105A, 105B, 105C using localized audio. For example, the remote participants 115A, 115B, 115C can be rendered on the displays 105A, 105B, 105C using standardized lines of sight. In some implementations, the video conferencing system 100 can have fewer than three (3) cameras (e.g., one (1) or two (2)). In this case, remote participants can be associated with the nearest camera and / or the default camera.
[0028] Figure 1C Shows a front view block diagram of a video conferencing system according to at least one example implementation. As Figure 1C shown, the video conferencing system 150 can be used by the remote participant 115A to participate in a video conference (e.g., a hybrid video conference). The video conferencing system 150 can represent a remote video conferencing system communicatively coupled to the video conferencing system 100. In some implementations, the video conferencing system 150 can include a display 155, a camera 160, and a speaker 165. The display 155 can be configured to render the participants of the video conference. For example, as Figure 1C shown, the local participant 120A, local participant 120B, remote participant 115B, and remote participant 115C are rendered on the display 155.
[0029] In some implementations, the video conferencing system 150 can be a home or home office computing system (e.g., a laptop or desktop computer). Thus, the video conferencing system 150 may not be configured to perform all of the hybrid video conferencing features described herein. For example, the video conferencing system 150 may not include all of the hardware associated with the video conferencing system 100. For example, the video conferencing system 150 may not include multiple cameras and / or multiple displays. However, the video conferencing system 150 can be configured to perform the features described herein for which the video conferencing system 150 includes sufficient software and / or hardware to perform.
[0030] For example, the video conferencing system 150 can be configured to receive at least one video stream that includes (e.g., each video stream includes) content associated with at least one participant. For example, the video conferencing system 150 can be configured to receive an indication of the display order of the participants in the video conference. For example, the video conferencing system 150 can be configured to render the at least one received video stream based on the display order.
[0031] For example, as Figure 1A and 1B shown, the video conference participants are in the display order of remote participant 115A, remote participant 115B, remote participant 115C, local participant 120B, and local participant 120A (e.g., clockwise) or remote participant 115B, local participant 120A, remote participant 115B, local participant 120B, and remote participant 115C (e.g., from left to right). Remote participants 115B, 115B, and 115C are rendered on displays 105A, 105B, 105C, 105D (as shown), and local participants 120A and 120B view (e.g., and sit at the conference table) displays 105A, 105B, 105C, 105D (as shown).
[0032] Figure 2 shows a top - view graphical representation of a video conferencing system according to at least one example implementation. As Figure 2As shown, the video conferencing system 100 further includes a table 205 and lights 210A, 210B, 210C (e.g., light emitting diode (LED) lights). Further, display 105A shows remote participant 115A and remote participant 115D, and display 105C shows remote participant 115C and remote participant 115E. In some implementations, local participants 120A, 120B may be positioned on one side of the table 205, and the video conferencing system 100 may be positioned on the other side of the table 205. Remote participants 115A, 115B, 115C, 115D, 115E may be displayed on displays 105A, 105B, 105C. As Figure 2 shown, two (2) or more of the remote participants 115A, 115B, 115C, 115D, 115E may be displayed on displays 105A, 105B, 105C.
[0033] In some implementations, the display of remote participants 115A, 115B, 115C, 115D, 115E may be modified to simulate local presence. For example, remote participants 115A, 115B, 115C, 115D, 115E may be displayed such that their faces are at the same level as local participants 120A, 120B (e.g., appearing at the eye level of local participants 120A, 120B). For example, remote participants 115A, 115B, 115C, 115D, 115E may be displayed such that a portion of their bodies is above the table 205 (e.g., appearing to be sitting next to the table 205). In other words, features associated with remote participants 115A, 115B, 115C, 115D, 115E (sometimes referred to as presence enhancement features) may be modified to simulate local participation in the meeting (e.g., conversation). These features (e.g., presence enhancement features) may include, for example, at least one of the following: ● Rendering of the natural proportions of the participants: Remote participants may be displayed life-size to enhance the sense of presence. In other words, people subconsciously associate size with importance. Thus, displaying remote participants at the same size as local participants may draw the attention of all participants. The increased proportions of remote participants relative to a standard video conferencing system may more easily facilitate reading body language and facial expressions, which improves social signaling. ● Normalized eye line of sight: In addition to accurately scaling remote participants, the video feed can be dynamically positioned up or down to ensure that their eye line of sight matches that of local participants. An elevated head position can be a dominance / submission cue for humans. Therefore, showing local participants and / or remote participants at a lower position around the table may be inadvertently perceived as a lack of equality. The eye line of sight normalization effect can be applied to video feeds containing a single person or, in a conference room video feed, by vertically shifting the portions corresponding to each person in each video. ● Gaze cues require careful attention to the system's geometry and rendering: Research has shown that people can be sensitive to incorrect horizontal angles when detecting gaze cues. However, a downward vertical gaze angle of 5 to 10 degrees is acceptable. Therefore, the peripheral device bars 110A, 110B, 110C including the cameras can be positioned directly above the edges of the displays to minimize the vertical angle. Additionally, the displays 105A, 105B, 105C can be lowered as much as possible (relative to the table 205) to keep the head-to-camera distance small. Additionally, the depth of the table 205 can be large enough to minimize the camera gaze angle. In some implementations, machine learning algorithms can be used to position remote participants on the display to keep them finely centered under the corresponding cameras even when the remote participants move horizontally during the conversation. ● Segmentation: Image segmentation can be used to remove the background from the video feed to render remote participants on a shared background. Rendering remote participants on a shared background can enhance the sense of presence because the remote participants share the background rather than having their own separate backgrounds. Rendering remote participants on a shared background can reduce the visual clutter of different room backgrounds and focus the visual attention on the participants themselves and the conversation they are having. This allows the meeting to accommodate more participants without causing too much visual stress. In addition to segmenting individual remote participants, it is also valuable to remove the backgrounds of other conference rooms in the call. Additionally, the background color, pattern, and brightness can be selected to be neutral to warm colors with moderate visual detail. Given the large display area, a too-bright background may create an overly visually stressful experience, while those that are too dark may make glare more prominent and distracting depending on the display technology. The display brightness and color temperature are calibrated to match the appearance of in-room participants under the current room lighting conditions. ● Segmentation helps with continuous auto-zooming and auto-centering: If the original background is included in a continuously dynamically zooming video feed, distracting zoom effects may occur. Similarly, horizontal panning to keep the subject centered under its respective camera will create false background movement. Using segmentation enables faster and more responsive positioning of remote participants. This improves the gaze cues and allows people to be closely spaced on the display without the risk of adjacent participants overlapping as they move around. ● Auto-zooming video feeds use various different remote hardware to improve presence: Facial detection can be used to measure the inter-pupillary distance of remote participants in each video feed and back-calculate a scale factor to show remote participants at true scale within the room. Some implementations may support any remote camera and setup. Thus, relying on optical settings or seat distance to create accurate sizing may not be an efficient use of available resources. In some implementations, variations in IPD between participants may result in false differences in scale (e.g., a participant with a smaller than average IPD may be inaccurately magnified to a size beyond what is desired). Thus, a machine learning (ML) model can be trained to evaluate the scale (e.g., using IPD and other sizing cues) to modify the zooming to a more desired level. Additionally, face and body detection can be used to detect each remote user within the video feed from the remote conference room and apply individual zooming and repositioning (e.g., using an ML model) to achieve natural scale rendering. ● Edge handling of body truncation and other video shearing artifacts: Remote participants may not be fully captured by their local cameras. The line of sight can be normalized, which may result in a gap between the bottom edge of their video feed and the desktop. This edge will move dynamically depending on the distance of the remote participant from the local camera and other factors. The movement of the distinct boundary can be distracting (e.g., showing a cropped video feed). The amount of movement can depend on how much the remote participant moves during the conversation and the tuning of the filtering for auto-zooming and auto-positioning of the video feed. Some implementations can be configured to fade out the bottom edge of the segmented remote participant, which makes such changes substantially less noticeable. In some implementations, the fade-out process can have a curved shape relative to the positioning of the user's face and body, which makes the truncation look more intentional. The size of the faded-out area can be a configurable parameter. For example, if it is too small, the edge movement may be noticeable. If too large, the rendered remote participant may look ghostly, which may reduce presence. Similarly, fade-out handling of truncation on the left and right edges can be less distracting and more intentional. For shearing along the top of the video feed, a smaller fade-out distance may be beneficial as the top being close to the face will make a larger fade-in look intrusive. ● Multi-view video routing infrastructure: Different architectures can meet the requirements for each remote participant to obtain a specific video feed. For example, a peer-to-peer full-mesh topology implemented using WebRTC in a web application can be used. In this design, when joining a session, each client makes a direct connection to every other client in the session and uses that connection to directly send the video stream from the appropriate camera between the participants. Another architecture is to use a cloud-hosted media router, where each client sends all video feeds along with stream identifiers to the router, and then each client subscribes to the appropriate video stream by requesting the stream with the correct stream ID. ● Multiple video cameras with a spatially consistent mapping to the participants: Each display can have multiple cameras along the top edge (e.g., in the peripheral device bars 110A, 110B, 110C). In some implementations, each display can have three (3) cameras, for a total of nine cameras in the entire system (using three (3) displays or one (1) wide display). Each remote participant can receive a video feed from the camera located above their rendered position. This can create accurate first-party gaze cues. For example, when a local participant looks at the location of a remote participant on the display to talk to that remote participant, the remote participant can see from their video feed that the local participant is looking at (e.g., directly at) that remote participant. ● Display mounting: The display can be flush-mounted to the table, with the lower edge below the tabletop. This mounting configuration can create the illusion that remote and local participants are sitting around the table, rather than appearing on a TV as seen in a standard wall-mounted video conferencing setup. The table may obscure the edges of the display. Thus, there may be a visual illusion that causes local participants to mentally fill in, for example, the legs of remote participants, just as if they were looking across the table at participants in a room. ● Allocating display space: The display space can be allocated based on the number of remote participants in the video feed. For hybrid sessions, the work-from-home participants typically have only one person in front of the camera. Using endpoint metadata or face or body detection, some implementations can be configured to allocate a smaller portion (one-half or one-third of the display) of the display to the remote participant when one (1) remote participant is detected. In some implementations, a smaller horizontal portion of the remote participant's display can be used while maintaining the natural proportions and equally positioned line of sight of that participant. This allows the system to scale to more people rather than simply showing a wide 16:9 aspect ratio image. A more compact layout also feels more natural because it matches how people are typically positioned around a real table. Using less horizontal space also minimizes head turning in smaller sessions, which makes it easier to gauge social reactions. ● Localized audio rendering: Each display may have two (2) speakers placed in the upper left and upper right corners, for a total of six (6) speakers in a three (3)-display system. The audio for each remote participant can be panned proportionally between the two nearest speakers above them to create the illusion that their voice is coming from the location where they are rendered. Vertical audio localization in human perception is malleable, so the visual stimulus of the video causes the sound to lock to the location where they are rendered. This enhances the sense of presence, making remote participants feel as if they are sitting at the table with local participants. Additionally, the localization of the audio makes it easier for local participants to handle overlapping conversations that often occur in fast-paced natural conversations. ● Multiple cameras and localized audio enable larger and more immersive display setups: Some implementations may include three large displays that are physically closer together than typical in other systems. This results in larger head turns. If one camera is used (e.g., a camera positioned in the center), the participants shown on the far left and far right will receive inconsistent gaze cues while turning their heads significantly, even when speaking directly to them. Also, without localized audio, the large field of view of our multi-display setup may require a lot of distracting and tiring head turns to "find" the current speaker. ● Conference table: The table shape can be selected for social dynamics and camera placement. For example, some implementations use multiple cameras. Thus, some implementations may have the constraint that each of the cameras should be able to see all participants. The displays can be table-mounted; thus, a curved shape can enable all local participants to see all remote participants well without tilting. Example tables may include curves on the back edge that can achieve these two design goals. In some implementations, the table depth can be configured to position remote participants within the social space and outside the personal space that may cause social discomfort and fatigue. Further, the front curve where local participants are seated can be configured to maintain the indoor social dynamics such that participants can see each other without intermediate participants blocking the line of sight. This creates a balanced hybrid conversation experience where both remote and local participants have equal ability to see all participants and be seen by all participants. ● Interoperability: Some implementations may operate with existing remote video conferencing hardware. Some implementations may not require remote participants to install special equipment. Some implementations may operate with existing webcams and computers and correct for different fields of view and distances to the camera by using scaling and positioning techniques on the video feed. ● Different numbers of displays: Some implementations are described as using three (3) displays. However, some implementations can scale in size and support one (1) or two (2), four (4), five (5) or more displays. ● Remote systems: Remote systems with ultra-wide curved displays (see Figure 1B ) and multiple cameras can be used. Some implementations can use ultra-wide monitors to allow remote users to have an experience very similar to a local experience. The additional width of the display can enable remote participants to be positioned in a virtual circle as seen in a room. This implementation can also have multiple video cameras with a uniform or staggered layout. This allows remote participants to also send gaze cues to in-room participants and other remote participants.
[0034] In some implementations, the video streams sent to remote participants 115A, 115B, 115C, 115D, 115E can be streamed with a spatially consistent mapping (e.g., including the display order or the order of displaying participants). In other words, the mapping, display order, or rendering position of remote participants 115A, 115B, 115C, 115D, 115E can be consistent across local and remote video conferencing systems. In some implementations, a video conference can include a local video conferencing system and at least one remote video conferencing system. As Figure 2 shown, in a counterclockwise display order, the display order can be remote participant 115E, remote participant 115C, remote participant 115B, remote participant 115D, remote participant 115A, local participant 120A, local participant 120B, and local participant 120C. As Figure 2 shown, the mapping, display order, or rendering position of remote participants from left to right is remote participant 115A, remote participant 115D, remote participant 115B, remote participant 115C, and remote participant 115E. Further, the positions (or physical locations) of local participants are local participant 120A between remote participants 115A and 115D, local participant 120B between remote participants 115D and 115C, and local participant 120C between remote participants 115C and 115E.
[0035] Thus, the mapping, display order, or rendering position of the participants can be (or can be in the following order): remote participant 115A, local participant 120A, remote participant 115D, local participant 120B, remote participant 115C, local participant 120C, and remote participant 115E. Such mapping, display order, or rendering position should be consistent across all video conferencing systems in a video conference. Note that on a remote video conferencing system, remote participants may not be rendered, but the remaining participants should be rendered at the above positions (or in the above order).
[0036] As mentioned above regarding Figure 1C Video conferencing system 150 can be configured to perform features that video conferencing system 150 includes sufficient hardware to perform. Thus, video conferencing system 150 can be configured to render with a spatially consistent mapping (e.g., display order); render with natural scaling, normalized line of sight, and segmented rendering; render with a shared background; render with body truncation and edge handling, consistent display spacing (e.g., tiling as described below), localized audio, and / or others.
[0037] Figure 3 A front view block diagram of a display of a video conferencing system according to at least one example implementation is shown. In some implementations, a user interface (UI) can be rendered on the display of the video conferencing system. In some implementations, the video conferencing system can include at least two displays. In some implementations, one (1) UI can be rendered across at least two displays. In some implementations, at least two UIs can be rendered on associated displays. The UI can be divided into at least two display portions. In some implementations, a first display portion and a second display portion can be rendered on one display. In some implementations, the first display portion can be rendered on a first display, and the second display portion can be rendered on a second display. Hereinafter, a display portion can be referred to as a tile.
[0038] As Figure 3 shown, video conferencing system 100 includes displays 105A, 105B, 105C, and each display 105A, 105B, 105C includes tiles 305, 310, 315, 320, 325, 330, 335, 340, 345. As Figure 3 shown, video conferencing system 100 includes peripheral device bars 110A, 110B, 110C, and each peripheral device bar 110A, 110B, 110C includes three (3) cameras 350. The cameras 350 can be centered above each of the tiles 305, 310, 315, 320, 325, 330, 335, 340, 345.
[0039] The tiles 305, 310, 315, 320, 325, 330, 335, 340, 345 can be configured to display remote participants in a hybrid video conference. For example, the remote participant 115A can be rendered in the tile 305, and the remote participant 115D can be rendered in the tile 310. In some implementations, during a hybrid video conference, the tiles 305, 310, 315, 320, 325, 330, 335, 340, 345 may not display remote participants. For example, in some implementations, during a hybrid video conference, the tiles 305 and 310 may include remote participants, and the tiles 315 and 320 may not display remote participants. In other words, the tiles 315 and 320 can be empty or not render anything. In some implementations, during a hybrid video conference, the tiles 305, 310, 315, 320, 325, 330, 335, 340, 345 can be configured to display content. For example, during a hybrid video conference, the tiles 305, 310, 315, 320, 325, 330, 335, 340, 345 can display the content as the content displayed on the computing device or a screenshot from the computing device. For example, during a hybrid video conference, the tiles 305, 310, 315, 320, 325, 330, 335, 340, 345 can display a whiteboard that displays slides with text and / or images.
[0040] In some implementations, during a hybrid video conference, at least two of the tiles 305, 310, 315, 320, 325, 330, 335, 340, 345 can be configured to display content. For example, during a hybrid video conference, the tiles 320 and 325 can display content. In some implementations, during a hybrid video conference, at least two of the tiles 305, 310, 315, 320, 325, 330, 335, 340, 345 of the displays 105A, 105B, 105C can be configured to display content, and another tile of the displays 105A, 105B, 105C can display remote participants.
[0041] As an example, remote participant 115A can be rendered in tile 305, and remote participant 115D can be rendered in tile 310. In some implementations, rendering remote participant 115A and / or remote participant 115D may include standardizing the presentation of remote participant 115A and / or remote participant 115D relative to local participants (e.g., local participant 120A and / or local participant 120B). Standardizing the presentation of a remote participant may include scaling the presentation of the remote participant. For example, rendering remote participant 115A and / or remote participant 115D may include increasing the size of the presentation to match the size and line of sight of the participants within and / or between tiles 305 and / or 310. For example, rendering remote participant 115A and / or remote participant 115D may include increasing the size of the presentation to approximate the size of the local participants. In other words, the rendered remote participant 115A and / or remote participant 115D may appear as if remote participant 115A and / or remote participant 115D and the local participants were in the same room.
[0042] In some implementations, rendering remote participant 115A and / or remote participant 115D may include standardizing the presentation of remote participant 115A and / or remote participant 115D, including segmenting the remote participant and displaying the remote participant on a common background. For example, segmenting remote participant 115A and / or remote participant 115D may include receiving frames of a streamed video and removing the background associated with remote participant 115A and / or remote participant 115D. Segmenting an image or frame of a video may include partitioning the image or frame into at least two regions, partitions, sections, or others. In some implementations, a trained machine learning model may be used to segment the image or frame.
[0043] In some implementations, segmenting the remote participant 115A and / or the remote participant 115D may include partitioning a frame including the remote participant 115A and / or the remote participant 115D into a region including the remote participant 115A and / or the remote participant 115D and a region including the background. Segmenting the remote participant 115A and / or the remote participant 115D may include removing the background, leaving only the remote participant 115A and / or the remote participant 115D. In some implementations, a common background may be added to the segmented remote participant 115A and / or the remote participant 115D to generate a frame including the remote participant 115A and / or the remote participant 115D on the common background. The common background may be a shared background that is configured to enhance the sense of presence, as the remote participants share the background rather than having their own individual backgrounds. Rendering the remote participants on the shared background may reduce the visual clutter of different room backgrounds and place the visual focus on the participants themselves and the conversation taking place between the local and remote participants.
[0044] In some implementations, rendering the remote participant 115A and / or the remote participant 115D may include standardizing the presentation of the remote participant 115A and / or the remote participant 115D, including shifting the presentation of the remote participant 115A and / or the remote participant 115D based on the eye positioning of the local participant. For example, shifting the presentation of the remote participant 115A and / or the remote participant 115D may include moving the presentation of the remote participant 115A and / or the remote participant 115D up or down such that the local participant looks at the eyes of the remote participant 115A and / or the remote participant 115D when rendered. In some implementations, shifting the presentation of the remote participant 115A and / or the remote participant 115D may include moving the presentation of the remote participant 115A and / or the remote participant 115D left or right to center the presentation within the associated tile. In some implementations, shifting the presentation of the remote participant 115A and / or the remote participant 115D may include fading the boundaries of the presentation of the remote participant. For example, moving the remote participant up may result in truncation of the body and / or other video artifacts due to there being no pixels to render. Fading may include extending the pixels and lightening the pixels to extend missing portions such as the body of the remote participant.
[0045] In some implementations, the video stream transmitted to remote participants can be streamed with a spatially consistent mapping. In other words, the display order or rendering position of remote participants can be consistent across local and remote video conferencing systems. Further, tiles of content, empty tiles, and others can be included in the mapping, display order, or rendering position. The mapping, display order, rendering position, and / or rendering order can be determined by the local video conferencing system. In some implementations, local participants can be grouped together in the video feed, where the video feed can be captured by a camera associated with the remote participant. In other words, the camera 350 above the tile presenting the remote participant can be used to capture local participants for transmission to the video feed of that remote participant. Thus, local participants can be rendered at the same position on the remote participant's display. If there are more than one local participant, the group of local participants can be displayed together and at the same position on the remote participant's display. In some implementations, local participants can be segmented (e.g., using an ML model) and transmitted individually in the video stream. In this implementation, the camera associated with the remote participant is used to capture local participants, and the local participants are segmented after being captured. In this implementation, segmentation can include identifying local participants and removing the background and other local participants from the video feed.
[0046] In some implementations, a hybrid video conference can include and / or add excess remote participants. For example, in some implementations, the number of participants may exceed the space on the display required to render the participants. Thus, the number of tiles may not be sufficient to present all participants. Therefore, some participants can be identified as excess participants. Excess participants (e.g., late-joining participants) can be identified when the video conference is initiated and / or added to the video conference after initiation. Further, if the excess participant is the current speaker, the excess participant can be repositioned. For example, if the excess participant is an active speaker, the excess participant can be swapped with a remote participant (swap tile positions). Thus, an excess remote participant is a remote participant not associated with a tile, and a tile cannot be added to the display. In some embodiments, a tile can be divided into at least two tiles. In some implementations, a tile can be divided vertically. In some implementations, a tile can be divided horizontally. For example, as Figure 3 shown, tile 340 can be divided to generate tile 340' and tile 340''.
[0047] The UI mentioned above may include additional features useful for a video conferencing system. For example, the UI may include a scheduling feature for arranging video conferences. For example, the UI may include an indication for waiting for participants to join. For example, when a participant joins or leaves, the UI may include user animations (e.g., fade in, fade out, and others). For example, the UI may include a join request feature. For example, the UI may include a non-video (e.g., phone) participant view (e.g., tile). For example, the UI may include an identification (e.g., participant name, location, email, and others) feature. For example, the UI may include an attention (e.g., raise hand) feature. For example, the UI may include a mute and / or camera off, avatar, and other features. For example, the UI may include control icons for the included features. For example, the UI may include an indication of the current speaker. For example, the UI may include presentation and recording utilities. Other video conferencing features and / or utilities are within the scope of the present disclosure.
[0048] Figure 4 A block diagram of a video conferencing system according to at least one example implementation is shown. As Figure 4 shown, the video conferencing system 100 includes displays 105A, 105B, 105C and cameras 350A, 350B, 350C. The video conferencing system 100 further includes hubs 405A, 405B, 405C; speakers 410A, 410B, 410C; a microphone 415; an audio amplifier 420; an audio processor 425; a computing device 430; a power supply 435; and a server 440. In some implementations, the audio amplifier 420, the audio processor 425, the computing device 430, and the power supply 435 may be installed as components of an equipment rack 445.
[0049] As mentioned above, the displays 105A, 105B, 105C may be configured to display a user interface UI associated with a hybrid video conference. The UI may be configured to display tiles on the displays 105A, 105B, 105C. The tiles may be used to render remote participants of the video conference. In some implementations, the UI may include three (3) tiles for each display 105A, 105B, 105C. For example, display 105A may display three (3) tiles, display 105B may display three (3) tiles, and display 105C may display three (3) tiles. Thus, in some implementations, the UI may be configured to display a total of nine (9) tiles.
[0050] Also as mentioned above, each tile may be used to render at least one remote participant, may be empty, render content, and / or others. Further, each tile may be associated with cameras 350A, 350B, 350C. In Figure 4Among them, cameras 350A, 350B, and 350C each include three (3) cameras. In some implementations, cameras 350A, 350B, and 350C may be included in peripheral device bars 110A, 110B, and 110C. As an example, display 105A may display three (3) tiles, where the left tile may be associated with the left camera in camera 350A, the center tile may be associated with the center camera in camera 350A, and the right tile may be associated with the right camera in camera 350A. If a remote participant is rendered in a tile associated with display 105A, speaker 410A may be used to generate audio associated with the remote participant. If a remote participant is rendered in a tile associated with display 105B, speaker 410B may be used to generate audio associated with the remote participant. If a remote participant is rendered in a tile associated with display 105C, speaker 410C may be used to generate audio associated with the remote participant. By generating audio in speakers 410A, 410B, and 410C associated with displays 105A, 105B, and 105C, the audio can be localized to the remote participants rendered in the tiles of displays 105A, 105B, and 105C. In some implementations, speakers 410A, 410B, and 410C may be included in peripheral device bars 110A, 110B, and 110C. For example, localizing the audio may include performing stereo panning on the audio feed of the remote participant based on the rendering position of the remote participant and the positions of speakers 410A, 410B, and 410C. Audio localization may depend on the geometry of the video conferencing system. For example, if the left speaker is directly above the left remote participant and the right speaker is above the right participant, the panning may be, for example, 100% left speaker for the left remote participant, 100% right speaker for the right remote participant, and 50% of each speaker for the middle remote participant.
[0051] Server 440 may be configured to transmit and receive video streams associated with a hybrid video conference. Server 440 may be a multi-input multi-output computing device. Server 440 may be configured to receive video streams from remote participants. Further, server 440 may be configured to transmit video streams to remote participants. Audio processor 425 may be configured to process the audio received from microphone 415 to transmit the audio to remote participants. Audio processor 425 may be configured to process the audio received from remote participants via server 440, and audio amplifier 420 may be configured to amplify the audio for playback on speakers 410A, 410B, and 410C. By playing the audio on the associated speaker among speakers 410A, 410B, and 410C, the audio can be localized to the remote participants rendered in the tiles of displays 105A, 105B, and 105C.
[0052] In some implementations, cameras 350A, 350B, 350C can be used to generate videos of local participants for a video stream to be transmitted to remote participants. In some implementations, cameras 350A, 350B, 350C can generate videos and transmit the videos to computing device 430 via hubs 405A, 405B, 405C. Hubs 405A, 405B, 405C can be three (3)-to-one (1) hubs to combine the video feeds of the three (3) individual cameras of cameras 350A, 350B, 350C. Power supply 435 can be configured to provide power to at least computing device 430.
[0053] Computing device 430 can be configured to provide processing resources for the hybrid video conferencing features described herein. For example, computing device 430 can be configured to generate a video stream of local participants for transmission to remote participants. For example, computing device 430 can be configured to render the videos of remote participants in tiles displayed on displays 105A, 105B, 105C. In some implementations, the video stream transmitted to remote participants can be streamed with a spatially consistent mapping. In other words, the display order or rendering position of remote participants can be consistent across local and remote video conferencing systems. Further, tiles including content, empty tiles, and others can be included in the mapping, display order, or rendering position. The mapping, display order, rendering position, and / or rendering order can be determined by computing device 430 of the local video conferencing system.
[0054] Figure 5 A rear view block diagram of a display according to at least one example implementation is shown. As Figure 5 shown, multiple components of the video conferencing system can be attached to display 105. For example, peripheral device bar 110A including camera 350 and speaker 410 can be attached to display 105. Glass cover sheet 505 can be attached to display 105. Glass cover sheet 505 can be attached to and / or overlap the display (e.g., overlap the top edge). Glass cover sheet 505 can be configured to protect camera 350 and speaker 410. Glass cover sheet 505 can be transparent or substantially transparent. Glass cover sheet 505 can be configured to be hidden, less noticeable, and / or less distracting to participants. For example, light 210 can be attached to display 105. In some implementations, light 210 can be an LED.
[0055] Figure 6A , Figure 6B and Figure 6C A block diagram of a multi-camera display arrangement according to at least one example implementation is shown. As Figures 6A to 6CAs shown, monitors A+1, ……, X-T, ……, X, …… X+T, ……, A-1 each have an associated camera 605-1, ……, 605-n. In some implementations, the multi-camera display arrangement can maintain representative remote participant gaze cues for n participants in a video conferencing system. In some implementations, the remote participants can be rendered in the same order on all participant monitors (or tiles) having associated cameras 605-1, ……, 605-n. In some implementations, local participant A can be in a hybrid video conference where n remote participants are rendered on monitors A+1, ……, X-T, ……, X, …… X+T, ……, A-1. In some implementations, monitors A+1, ……, X-T, ……, X, …… X+T, ……, A-1 can be implemented as tiles as discussed above.
[0056] When local participant A looks at the remote participant rendered on monitor X, the remaining remote participants in the hybrid video conference can see the direction in which local participant A is looking at the remote participant rendered on monitor X on the remote monitors of the respective remote participants.
[0057] For example, remote participants on monitors to the left of monitor X (e.g., Figure 6B monitors A+1, ……, X-T, …… identified as section 1 in
[0058] For example, remote participants on monitors to the right of monitor X (e.g., Figure 6CThe display identified as section 2 in... The remote participants on X+1,..., A-1) can view that the local participant A is looking right from the cameras 605-1,..., 605-n associated with the corresponding remote participants. The remote participants can be rendered based on the same order as the displays A+1,..., X-T,..., X,..., X+T,..., A-1 at each location of the video conference. Thus, on the local display of any remote participant (e.g., the remote participant rendered on display X+1), the remote participant X will be shown to the right of the local participant A because when rendered on the remote participant's display, the remote participant X will be rendered after the local participant A. Thus, each remote participant can see that the local participant A is looking in the direction of the remote participant X.
[0059] As mentioned above, the local participant can be associated with each video feed, and the mapping or display order can be linear (e.g., from left to right or from right to left). In this implementation, the local participant (e.g., local participant A) can move. In some implementations, the mapping, display order, or rendering position can be modified. For example, if the local participant A moves to the left of the remote participant X-T, the mapping, display order, or rendering position can be modified based on this movement. As Figure 6B and Figure 6C depicted in, section 2 and section 3 can change based on the movement of the local participant A.
[0060] Figure 6D and Figure 6E shows a block diagram of the display order according to at least one example implementation. The mapping of participants, spatial mapping, display positioning, display order, display position, rendering positioning, rendering order, rendering position, and / or others can be from one of many perspectives. For example, the mapping of participants, spatial mapping, display positioning, display order, display position, rendering positioning, rendering order, rendering position, and / or others can be viewed from above in a clockwise or counterclockwise circular rotation manner. Figure 6D and Figure 6E shows the participants viewed from above and displayed in a counterclockwise circular rotation.
[0061] Referring to Figure 6D , circle 1 can represent the local participant or the participant system, and circles 2, 3, 4, and 5 can represent the remote participants or the participant systems. The display order 610-1 shows the display order from the perspective of the local participant (circle 1), where the display order 610-1 is circle 5, 4, 3, and 2 or the remote participants 5, 4, 3, and 2. The display order for all participants should be the same to ensure that some of the above features (e.g., eye gaze direction) are consistent across all devices associated with the video conference. Thus, referring toFigure 6E , Display order 610-2 shows the display order from the perspective of the remote participant (circle 2), where the display order 610-2 is circle 3, 4, 5, and 1 or remote participants 3, 4, 5, and local participant 1. Display order 610-3 shows the display order from the perspective of the remote participant (circle 3), where the display order 610-3 is circle 4, 5, 1, and 2 or remote participants 4, 5, local participant 1, and remote participant 2. Display order 610-4 shows the display order from the perspective of the remote participant (circle 4), where the display order 610-4 is circle 5, 1, 2, and 3 or remote participant 5, local participant 1, and remote participants 2, 3. Display order 610-5 shows the display order from the perspective of the remote participant (circle 5), where the display order 610-5 is circle 1, 2, 3, and 4 or local participant 1 and remote participants 2, 3, 4.
[0062] Example 1. Figure 7 is a block diagram of a method of operating a video conferencing system according to an example implementation. As Figure 7 shown, in step S705, a first remote device associated with a first remote participant receives a remote video stream from a second remote device, the remote video stream including content associated with a second remote participant of the video conference. In step S710, the first remote device receives a local video stream from a local device associated with a local participant of the video conference, the local video stream including an indication of the display order of the first remote participant, the second remote participant, and the local participant. In step S715, the remote video stream and the local video stream are rendered on the first remote device based on the display order.
[0063] Example 2. The method according to Example 1, wherein the rendering of the remote video stream and the local video stream includes: standardizing the presentation of the second remote participant and the local participant with respect to the first remote participant.
[0064] Example 3. The method according to Example 2, wherein the standardization of the presentation of the second remote participant and the local participant may include: scaling the presentation of the second remote participant and the local participant.
[0065] Example 4. The method according to Example 2, wherein the standardization of the presentation of the second remote participant and the local participant may include: segmenting the presentation of the second remote participant and the local participant, and rendering (or displaying) the second remote participant and the local participant on a common background.
[0066] Example 5. The method according to Example 2, wherein the normalization of the presentations of the second remote participant and the local participant may include: shifting the presentations of the second remote participant and the local participant based on the eye positioning of the first remote participant.
[0067] Example 6. The method according to Example 5, wherein the shifting of the presentations of the second remote participant and the local participant may include: fading the boundaries of the presentations of the second remote participant and the local participant.
[0068] Example 7. The method according to Example 1, wherein the rendering of the remote video stream and the local video stream may include: generating a user interface that includes a first display portion and a second display portion, the first display portion including the second remote participant, and the second display portion including the local participant.
[0069] Example 8. The method according to Example 7, wherein the user interface may include a third display portion that includes content associated with the video conference.
[0070] Example 9. The method according to Example 7, wherein the user interface may include an empty third display portion.
[0071] Example 10. The method according to Example 1, wherein the rendering of the first video stream and the second video stream may include: generating a user interface that includes a plurality of display portions based on a plurality of video streams associated with the video conference.
[0072] Example 11. Figure 8 is a block diagram of a method for operating a video conferencing system according to an example implementation. As Figure 8 shown, in step S805, a user interface is generated that includes a first display portion corresponding to a first camera and a second display portion corresponding to a second camera. In step S810, the UI is rendered on a display of the video conferencing system. In step S815, a video stream corresponding to a remote participant of the video conference is received. In step S820, the remote participant is rendered in one of the first display portion or the second display portion, the rendering including: normalizing the presentation of the remote participant relative to the local participant. In some implementations, the video captured by the first camera may be streamed to the first remote participant. Further, the video associated with and / or received from the first remote participant may be rendered on the first display portion. In some implementations, the video captured by the second camera may be streamed to the second remote participant. Further, the video associated with and / or received from the second remote participant may be rendered on the second display portion.
[0073] Example 12. The method according to Example 11, wherein standardizing the presentation of the remote participant(s) may include scaling the display of the remote participant(s).
[0074] Example 13. The method according to Example 11, wherein standardizing the presentation of the remote participant(s) may include splitting the display of the remote participant(s) and displaying the remote participant(s) on a common background.
[0075] Example 14. The method according to Example 11, wherein standardizing the presentation of the remote participant(s) may include shifting the display of the remote participant(s) based on the eye positioning of the local participant(s).
[0076] Example 15. The method according to Example 14, wherein shifting the display of the remote participant(s) may include fading the boundaries of the display of the remote participant(s).
[0077] Example 16. The generation of the first display portion may include centering the first display portion based on the positioning of the first camera, and the generation of the second display portion may include centering the second display portion based on the positioning of the second camera, according to the method of Example 11.
[0078] Example 17. The method according to Example 11, wherein the generation of the first and second display portions on the display may include generating a third display portion on the display. Alternatively (or additionally), the user interface may include the third display portion on the display.
[0079] Example 18. The method according to Example 17, wherein the first display portion may include the remote participant(s), and the second and third display portions may be empty.
[0080] Example 19. The method according to Example 17, wherein the first display portion may include the remote participant(s), and the second and third display portions may include content associated with the video conference (e.g., content other than the remote participant(s) and / or local participant(s)).
[0081] Example 20. The method according to Example 11, wherein the first camera may be configured to capture video of multiple local participants of the video conference, and the second camera may be configured to capture video of the multiple local participants of the video conference.
[0082] Example 21. The method according to Example 11, wherein the generation of the first and second display portions on the display may include generating multiple display portions based on multiple video streams associated with the video conference. Alternatively (or additionally), the user interface includes multiple display portions based on multiple video streams associated with the video conference.
[0083] Example 22. The method according to Example 11, wherein the generation of the first display portion and the second display portion on the display may include generating a display portion including the excess remote participants.
[0084] Example 23. The method according to Example 11, wherein the method may further include localizing the audio of the remote participants to the speakers.
[0085] Example 24. The method according to Example 11, wherein the video stream may include at least two remote participants, and the generation of the first display portion and the second display portion on the display may include generating a plurality of display portions based on the at least two remote participants. Alternatively (or additionally), the user interface includes a plurality of display portions based on the at least two remote participants.
[0086] Example 25. The method according to Example 11, wherein the method may further include: determining that the video stream includes excess remote participants; and in response to determining that the video stream includes excess remote participants, generating a divided display portion by dividing one of the first display portion and the second display portion, and displaying the excess remote participants in the divided display portion.
[0087] Example 26. The method according to Example 25, wherein one of the first display portion and the second display portion is divided vertically.
[0088] Example 27. The method according to Example 25, wherein one of the first display portion and the second display portion is divided horizontally.
[0089] Example 28. The method according to Example 11, wherein the display may be (or include) a first display including three display portions, a second display including three display portions, and a third display including three display portions.
[0090] Example 29. The method according to Example 11, wherein the display may be a wide aspect ratio display including a first display portion and a second display portion.
[0091] Example 30. Figure 9 is a block diagram of a method of operating a video conferencing system according to an example implementation. As Figure 9As shown, in step S905, a user interface is generated, which includes a first display part corresponding to the first camera and a second display part corresponding to the second camera. In step S910, the UI is rendered on the display of the video conferencing system. In step S915, a first video stream corresponding to a remote participant in the video conference is received. In step S890, a second video stream including local participants is generated, and the video stream includes an indication of the display order of the local participants relative to the remote participants.
[0092] Example 31. The method according to Example 30, wherein the video stream includes a plurality of local participants, and the display order includes a plurality of local participants positioned together in one location.
[0093] Example 32. The method according to Example 30, wherein the rendering may include standardizing the presentation of the remote participants relative to the local participants.
[0094] Example 33. The method according to Example 30, wherein the standardization of the presentation of the remote participants may include scaling the display of the remote participants.
[0095] Example 34. The method according to Example 30, wherein the standardization of the presentation of the remote participants may include splitting the display of the remote participants and displaying the remote participants on a common background.
[0096] Example 35. The method according to Example 30 may further include shifting the presentation of the remote participants by moving the presentation of the remote participants up or down on the display.
[0097] Example 36. The method according to Example 35, wherein the shifting of the presentation of the remote participants may include fading the boundaries of the display of the remote participants.
[0098] Example 37. The method according to Example 30, wherein the generation of the first display part may include centering the first display part based on the positioning of the first camera, and the generation of the second display part may include centering the second display part based on the positioning of the second camera.
[0099] Example 38. The method according to Example 30, wherein the generation of the first display part and the second display part on the display may include generating a third display part on the display. Alternatively (or additionally), the user interface may include a third display part on the display.
[0100] Example 39. The method according to Example 38, wherein the first display part may include remote participants, and the second display part and the third display part may be empty.
[0101] Example 40. The method according to Example 38, wherein the first display portion may include remote participants, and the second and third display portions may include content associated with the video conference (e.g., content other than remote and / or local participants).
[0102] Example 41. The method according to Example 30, wherein generating the first and second display portions on the display may include: generating a plurality of display portions based on a plurality of video streams associated with the video conference. Alternatively (or additionally), the user interface includes a plurality of display portions based on a plurality of video streams associated with the video conference.
[0103] Example 42. The method according to Example 30, wherein generating the first and second display portions on the display may include generating a display portion including an excess number of remote participants.
[0104] Example 43. The method according to Example 30 may further include tracking the eye gaze of local participants.
[0105] Example 44. The method according to Example 30 may further include tracking the eye gaze of remote participants.
[0106] Example 45. The method according to Example 30 may further include localizing the audio of remote participants to speakers.
[0107] Example 46. The method according to Example 30, wherein the video stream may include at least two remote participants, and generating the first and second display portions on the display may include generating a plurality of display portions based on at least two remote participants. Alternatively (or additionally), the user interface includes a plurality of display portions based on at least two remote participants.
[0108] Example 47. The method according to Example 30 may further include: determining that the video stream includes an excess number of remote participants; and in response to determining that the video stream includes an excess number of remote participants, generating a divided display portion by dividing one of the first and second display portions, and displaying the excess remote participants in the divided display portion.
[0109] Example 48. The method according to Example 47, wherein one of the first and second display portions is divided vertically.
[0110] Example 49. The method according to Example 47, wherein one of the first and second display portions may be divided horizontally.
[0111] Example 50. The method according to Example 30, wherein the display may be a first display including three display portions, a second display including three display portions, and a third display including three display portions.
[0112] Example 51. The method according to Example 30, wherein the display may be a wide aspect ratio display including a first display portion and a second display portion.
[0113] Example 52. A method may include any combination of one or more of Examples 1 to 54.
[0114] Example 53. A non-transitory computer-readable storage medium including instructions stored thereon that, when executed by at least one processor, are configured to cause a computing system to perform the method according to any one of Examples 1 to 52.
[0115] Example 54. An apparatus including means for performing the method according to any one of Examples 1 to 52.
[0116] Example 55. An apparatus including at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the apparatus to perform at least the method according to any one of Examples 1 to 52.
[0117] Example implementations may include a non-transitory computer-readable storage medium including instructions stored thereon that are configured to, when executed by at least one processor, cause a computing system to perform any of the methods described above. Example implementations may include an apparatus including means for performing any of the methods described above. Example implementations may include an apparatus including at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the apparatus to perform at least any of the methods described above.
[0118] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuit systems, specially designed ASICs (Application Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be special purpose or general purpose and coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0119] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. As used herein, the term “machine-readable medium,” “computer-readable medium” refers to any computer program product, apparatus, and / or device (e.g., a disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0120] For providing interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., an LED (Light Emitting Diode), or OLED (Organic LED), or LCD (Liquid Crystal Display) monitor / screen) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0121] The systems and techniques described herein can be implemented in a computing system that includes a backend component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a frontend component (e.g., a client computer having a graphical user interface or a web browser through which users can interact with an implementation of the systems and techniques described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (“LAN”), a wide area network (“WAN”), and the Internet.
[0122] A computing system may include a client and a server. The client and the server are generally remote from each other and typically interact via a communication network. The relationship between the client and the server is generated by computer programs that execute on respective computers and cause the client-server relationship to exist between them.
[0123] A variety of implementations have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of this specification.
[0124] In addition, the logical flow depicted in the figures does not require a particular order or sequential order as shown to achieve the desired result. Additionally, other steps may be provided, or steps may be removed from the described flow, and other components may be added to or removed from the described system. Accordingly, other implementations are within the scope of the appended claims.
[0125] Although certain features of the described implementations have been illustrated as described herein, many modifications, substitutions, changes, and equivalents will now occur to those skilled in the art. Accordingly, it should be understood that the appended claims are intended to cover all such modifications and changes that fall within the scope of the implementations. It should be understood that they are presented by way of example only and not of limitation, and that various changes in form and detail can be made. Except for mutually exclusive combinations, any part of the apparatus and / or method described herein can be combined in any combination. The implementations described herein can include various combinations and / or sub-combinations of the functions, components, and / or features of the different implementations described.
[0126] Although the example implementations may include various modifications and alternative forms, their implementations are shown by way of example in the figures and will be described in detail herein. However, it should be understood that the example implementations are not intended to be limited to the specific forms disclosed, but rather, the example implementations will cover all modifications, equivalents, and alternatives that fall within the scope of the claims. Like numerals refer to like elements in the figure descriptions.
[0127] Some of the above example implementations are described as processes and methods depicted as flowcharts. Although the flowcharts depict operations as sequential processes, many operations can be performed in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The processes can terminate when their operations are completed, but can also have additional steps not included in the figures. The processes can correspond to methods, functions, procedures, subroutines, subprograms, etc.
[0128] The methods discussed above, some of which are illustrated by flow diagrams, can be implemented by hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the necessary tasks can be stored in a machine or computer-readable medium, such as a storage medium. One or more processors can perform the necessary tasks.
[0129] The specific structural and functional details disclosed herein are merely representative for the purpose of describing example implementations. However, the example implementations are embodied in many alternative forms and should not be construed as limited to the implementations set forth herein.
[0130] It should be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the example implementations, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element. As used herein, the term and / or includes any and all combinations of one or more of the associated listed items.
[0131] It should be understood that when an element is referred to as being connected or coupled to another element, the element can be directly connected or coupled to the other element, or intervening elements may be present. In contrast, when an element is referred to as being directly connected or directly coupled to another element, no intervening elements are present. Other words used to describe the relationship between elements should be interpreted in a similar manner (e.g., between and directly between, adjacent and directly adjacent, etc.).
[0132] The terms used herein are for the purpose of describing particular implementations only and are not intended to limit the example implementations. As used herein, unless the context clearly indicates otherwise, the singular forms a, an, and the are also intended to include the plural forms. It should be further understood that the terms include, comprises, containing, and / or having, when used herein, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0133] It should also be noted that in some alternative implementations, the indicated functions / actions may not occur in the order indicated in the figures. For example, two figures shown in succession may in fact be executed simultaneously or sometimes in the reverse order, depending on the functionality / behavior involved.
[0134] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which the example implementations belong. Further, it should be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and should not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0135] The above example implementations and corresponding detailed descriptions are presented in terms of software or algorithms and symbolic representations of operations on data bits within a computer memory. These descriptions and representations are the means by which a person of ordinary skill in the art effectively conveys the substance of their work to other persons of ordinary skill in the art. An algorithm, as the term is used herein and as it is commonly used, is considered to be a self-consistent sequence of steps leading to a desired result. The steps are those requiring physical manipulation of physical quantities. Usually, though not necessarily, these quantities take the form of optical, electrical, or magnetic signals capable of being stored, transmitted, combined, compared, and otherwise manipulated. For general reasons chiefly, these signals are referred to as bits, values, elements, symbols, characters, terms, numbers, or others that have sometimes proven convenient.
[0136] In the above illustrative implementations, references to symbolic representations (e.g., in the form of flowcharts) of actions and operations that may be implemented as program modules or functional procedures include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types and may be described and / or implemented using existing hardware at existing structural elements. Such existing hardware may include one or more central processing units (CPUs), digital signal processors (DSPs), application specific integrated circuits, field programmable gate arrays (FPGAs), computers, and the like.
[0137] However, it should be borne in mind that all such and similar terms are to be associated with appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise specifically stated or as is apparent from the discussion, terms such as processing or computing or calculating or determining or displaying or otherwise refer to the actions and processes of a computer system or similar electronic computing device that manipulates data represented as physical electronic quantities within the registers and memories of the computer system and transforms that data into other data similarly represented as physical quantities within the computer system memory or registers or other such information storage, transmission, or display devices.
[0138] It should also be noted that the software implementation aspects of the example implementations are typically encoded on some form of non-transitory program storage medium or implemented over some type of transmission medium. The program storage medium can be magnetic (e.g., a floppy disk or hard disk) or optical (e.g., a compact disc read-only memory or CD ROM), and can be read-only or random access. Similarly, the transmission medium can be a pair of twisted wires, coaxial cable, fiber optic, or some other suitable transmission medium known in the art. The example implementations are not limited by these aspects of any given implementation.
[0139] Finally, it should also be noted that although the appended claims set forth particular combinations of the features described herein, the scope of the present disclosure is not limited to the particular combinations recited hereinafter, but extends to cover any combination of the features or implementations disclosed herein, regardless of whether that particular combination is specifically recited in the appended claims at this time.
Claims
1. A method, comprising: Receiving, by a first remote device associated with a first remote participant, a remote video stream from a second remote device, the remote video stream including content associated with a second remote participant of a video conference; Receiving, by the first remote device, a local video stream from a local device associated with a local participant of the video conference, the local video stream including an indication of a display order of the first remote participant, the second remote participant, and the local participant; And Rendering, based on the display order, the remote video stream and the local video stream on the first remote device.
2. The method according to claim 1, wherein The rendering of the remote video stream and the local video stream includes: standardizing the presentation of the second remote participant and the local participant relative to the first remote participant.
3. The method according to claim 2, wherein, The standardizing of the presentation of the second remote participant and the local participant includes: scaling the presentation of the second remote participant and the local participant.
4. The method according to any one of claims 2 and 3, wherein The standardizing of the presentation of the second remote participant and the local participant includes: segmenting the presentation of the second remote participant and the local participant and rendering the second remote participant and the local participant on a common background.
5. The method according to any one of claims 2 to 4, wherein, The standardizing of the presentation of the second remote participant and the local participant includes: shifting the presentation of the second remote participant and the local participant based on the eye positioning of the first remote participant.
6. The method according to claim 5, wherein The shifting of the presentation of the second remote participant and the local participant includes: fading the boundaries of the presentation of the second remote participant and the local participant.
7. The method according to any one of claims 1 to 6, wherein, The rendering of the remote video stream and the local video stream includes: generating a user interface, the user interface including a first display portion and a second display portion, the first display portion including the second remote participant, and the second display portion including the local participant.
8. The method according to claim 7, wherein, The user interface includes a third display portion, the third display portion including content associated with the video conference.
9. The method according to claim 7, wherein The user interface includes an empty third display portion.
10. The method according to any one of claims 1 to 6, wherein The rendering of the remote video stream and the local video stream includes: generating a user interface including a plurality of display portions based on a plurality of video streams associated with the video conference.
11. A video conferencing system, comprising: A display; A first camera configured to capture video of a local participant of a video conference; A second camera configured to capture video of the local participant of the video conference; And A processor configured to: Generate a user interface UI including: A first display portion corresponding to the first camera; and A second display portion corresponding to the second camera; Receive a video stream corresponding to a remote participant of the video conference; and Render the remote participant in one of the first display portion or the second display portion, the rendering including: standardizing the presentation of the remote participant relative to the local participant.
12. The video conferencing system according to claim 11, wherein, The standardization of the presentation of the remote participant includes scaling the display of the remote participant.
13. The video conferencing system according to any one of claims 11 and 12, wherein The standardization of the presentation of the remote participant includes segmenting the presentation of the remote participant and displaying the remote participant on a common background.
14. The video conferencing system according to any one of claims 11 to 13, wherein, The standardization of the presentation of the remote participant includes shifting the presentation of the remote participant based on the eye positioning of the local participant.
15. The video conferencing system according to claim 14, wherein, The shifting of the display of the remote participant includes fading the boundaries of the display of the remote participant.
16. The video conferencing system according to any one of claims 11 to 15, wherein the generation of the first display portion includes centering the first display portion based on the positioning of the first camera, and the generation of the second display portion includes centering the second display portion based on the positioning of the second camera.
17. The video conferencing system according to any one of claims 11 to 16, wherein, The generation of the first display portion and the second display portion on the display includes generating a third display portion on the display.
18. The video conferencing system according to claim 17, wherein the first display portion includes the remote participant, and the second display portion and the third display portion are empty.
19. The video conferencing system according to claim 17, wherein the first display portion includes the remote participant, and the second display portion and the third display portion include content.
20. The video conferencing system according to any one of claims 11 to 19, wherein the first camera is configured to capture video of a plurality of local participants of the video conference; and the second camera is configured to capture video of the plurality of local participants of the video conference.
21. The video conferencing system according to any one of claims 11 to 20, wherein, The generation of the first display portion and the second display portion on the display includes generating a plurality of display portions based on a plurality of video streams associated with the video conference.
22. The video conferencing system according to any one of claims 11 to 21, wherein, The generation of the first display portion and the second display portion on the display includes generating a display portion including an excess of remote participants.
23. The video conferencing system according to any one of claims 11 to 22, wherein, The processor is further configured to localize the audio of the remote participant to a speaker.
24. The video conferencing system according to any one of claims 11 to 23, wherein the video stream includes at least two remote participants, and the generation of the first display portion and the second display portion on the display includes generating a plurality of display portions based on the at least two remote participants.
25. The video conferencing system according to any one of claims 11 to 24, wherein, The processor is further configured to: determine that the video stream includes an excess of remote participants; and in response to determining that the video stream includes an excess of remote participants, generate a divided display portion by dividing one of the first display portion and the second display portion, and display the excess remote participants in the divided display portion.
26. The video conferencing system according to claim 25, wherein, One of the first display portion and the second display portion is divided vertically.
27. The video conferencing system according to claim 25, wherein, One of the first display portion and the second display portion is horizontally divided.
28. The video conferencing system according to any one of claims 11 to 27, wherein, The display is a first display including three display portions, a second display including three display portions, and a third display including three display portions.
29. The video conferencing system according to any one of claims 11 to 28, wherein, The display is a wide aspect ratio display including the first display portion and the second display portion.
30. A video conferencing system, comprising: A display; A first camera configured to capture video of local participants in a video conference; A second camera configured to capture video of the local participants in the video conference; And A processor configured to: Generate a user interface UI including: A first display portion corresponding to the first camera; and A second display portion corresponding to the second camera; Receive a first video stream corresponding to remote participants in the video conference; Render the remote participants in one of the first display portion or the second display portion; and Generate a second video stream including the local participants, the second video stream including an indication of a display order of the local participants relative to the remote participants.
31. The video conferencing system according to claim 30, wherein The second video stream includes a plurality of local participants, and The display order includes the plurality of local participants positioned together in one location.
32. The video conferencing system according to any one of claims 30 and 31, wherein The rendering includes standardizing the presentation of the remote participants relative to the local participants.
33. The video conferencing system according to claim 32, wherein, The standardizing of the presentation of the remote participants includes: scaling the display of the remote participants.
34. The video conferencing system according to claim 32, wherein, The standardizing of the presentation of the remote participants includes: segmenting the display of the remote participants and displaying the remote participants on a common background.
35. The video conferencing system according to claim 32, further comprising: Shift the presentation of the remote participants by moving the presentation of the remote participants up or down on the display.
36. The video conferencing system according to claim 35, wherein, The shifting of the presentation of the remote participants includes: fading the boundaries of the display of the remote participants.
37. The video conferencing system according to any one of claims 30 to 36, wherein The generation of the first display portion includes centering the first display portion based on the positioning of the first camera, and The generation of the second display portion includes centering the second display portion based on the positioning of the second camera.
38. The video conferencing system according to any one of claims 30 to 37, wherein The generation of the first display portion and the second display portion on the display includes generating a third display portion on the display.
39. The video conferencing system according to claim 38, wherein The first display portion includes the remote participants, and The second display portion and the third display portion are empty.
40. The video conferencing system according to claim 38, wherein The first display portion includes the remote participants, and The second display portion and the third display portion include content.
41. The video conferencing system according to any one of claims 30 to 40, wherein, The generation of the first display portion and the second display portion on the display includes: generating a plurality of display portions based on a plurality of video streams associated with the video conference.
42. The video conferencing system according to any one of claims 30 to 41, wherein, The generation of the first display portion and the second display portion on the display includes generating a display portion including an excess number of remote participants.
43. The video conferencing system according to any one of claims 30 to 42, wherein, The processor is further configured to localize the audio of the remote participants to a speaker.
44. The video conferencing system according to any one of claims 30 to 43, wherein the first video stream includes at least two remote participants, and the generation of the first display portion and the second display portion on the display includes generating a plurality of display portions based on the at least two remote participants.
45. The video conferencing system according to any one of claims 30 to 44, wherein, The processor is further configured to: determine that the video stream includes an excess number of remote participants; and in response to determining that the video stream includes an excess number of remote participants, generate a divided display portion by dividing one of the first display portion and the second display portion, and display the excess remote participants in the divided display portion.
46. The video conferencing system according to claim 45, wherein, One of the first display portion and the second display portion is divided vertically.
47. The video conferencing system according to claim 45, wherein, One of the first display portion and the second display portion is divided horizontally.
48. The video conferencing system according to any one of claims 30 to 47, wherein, The display is a first display including three display portions, a second display including three display portions, and a third display including three display portions.
49. The video conferencing system according to any one of claims 30 to 48, wherein, The display is a wide aspect ratio display including the first display portion and the second display portion.