Virtual conferencing

By modifying video streams to simulate eye contact or offset gaze based on participant attention, the virtual conferencing system improves attention perception and enables private communication, addressing the limitations of conventional systems.

WO2025103574A1PCT designated stage expired Publication Date: 2025-05-22TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2023/081717
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-14
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Conventional virtual conferencing systems fail to accurately convey attention and engagement between participants, making it difficult for presenters to assess audience engagement and for participants to share private gestures or signs.

Method used

The system determines the gaze of participants and modifies video streams to create the illusion of eye contact or offset gaze, allowing for private gestures to be shared between specific participants while maintaining a neutral appearance for others.

Benefits of technology

This solution enhances the perception of attention and engagement in virtual meetings, allowing presenters to better assess audience engagement and enabling private communication between participants without distracting other meeting attendees.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2023081717_22052025_PF_FP_ABST
    Figure EP2023081717_22052025_PF_FP_ABST
Patent Text Reader

Abstract

A method (700) in a virtual conferencing system comprising i) a first user device, UD, used by a first participant and comprising a first display and a first camera and ii) a second UD used by a second participant and comprising a second display and a second camera. The method includes determining (s702) that the first participant is looking at a particular portion of the first display, wherein the particular portion of the first display is associated with the second participant. The method also includes receiving (s / 04) a first video stream captured by the first camera. The method also includes, based on the determination, modifying (s706) the first video stream to produce a modified video stream for causing the second UD to display a first video representation of the first participant, wherein the first video representation of the first participant is configured for causing the second participant to believe that the first participant is not looking directly at the camera but rather looking in a first direction that is offset from a second direction, wherein the second direction is a direction from the first participant to the camera.
Need to check novelty before this filing date? Find Prior Art

Description

VIRTUAL CONFERENCINGTECHNICAL FIELD

[0001] Disclosed are embodiments related to virtual conferencing (e.g., video conferencing) systems and methods.BACKGROUND

[0002] Virtual conferences (e.g., video meetings, a.k.a., video calls) are common. In a typical video meeting, each participant has a camera and a network connection so that each participant can see each other participant - or the most active ones if user’s screen is not large enough to show all participants. Typically, a person in a video meeting looks at their screen and not at their camera, which is typically mounted above the screen, making it impossible for others to see the person’s eyes in a meaningful way (i.e., normal gaze direction is not applicable).

[0003] In one prior art, it was suggested to change the layout of the “gallery” on the screen of the showed video feeds so that the person talking was closer to the camera to make it more likely that each participant could see the others’ eyes. There are also suggestions for how the user can make gestures to adapt the gallery mode layout and mirroring (see, e.g., reference [3]).

[0004] Systems described in the prior art dynamically change the eyes of the user in the video feed so it appears as if the user is looking at the camera, that is, towards other users in the same video meeting, and not on a different place of the screen, which increases the perception of presence (see, e.g., reference [1]).

[0005] In another prior art, the gaze detection system identifies who a first user is looking at, that is, which of the video feeds of other users the first user is looking at, and then only that selected other user gets the first user’s eye-adapted feed, wherein the first user's eye- adapted feed appears as if the first user looks at the camera; whereas all others get the first user’s normal feed (see, e.g., reference [2]). This enables all users to see and understand if the first user looks at them or not even though only one camera feed is used.SUMMARY

[0006] Certain challenges presently exist. For instance, in a convention virtual conference system, if a first user looks at material presented by second user (e.g., the second user is sharing his / her screen), the second user may not get the sense that the first user is focused on what the second user is saying. This is different from a live audience where such attentiondirection interactions are very clear and useful to, e.g., estimate how engaged the audience is. Also, in conventional systems, if the first user is looking at the video of the second user, and the second user is likewise looking at the video of the first user (i.e., the two users have mutual eye contact), there is no way for the two users to share signs or gestures between them without other participants being able to see those gestures or signs as well. Similarly, in conventional systems, it is not possible for one user to make a gesture that is only seen by an intended other user to get that other user’s attention.

[0007] Accordingly, in one aspect there is provided a method in a virtual conferencing system comprising i) a first user device (UD) used by a first participant and comprising a first display and a first camera and ii) a second UD used by a second participant and comprising a second display and a second camera. The method includes determining that the first participant is looking at a particular portion of the first display, wherein the particular portion of the first display is associated with the second participant. The method also includes receiving a first video stream captured by the first camera. The method further includes, based on the determination, modifying the first video stream to produce a modified video stream for causing the second UD to display a first video representation of the first participant, wherein the first video representation of the first participant is configured for causing the second participant to believe that the first participant is not looking directly at the camera but rather looking in a first direction that is offset from a second direction, wherein the second direction is a direction from the first participant to the camera.

[0008] In another aspect there is provided a method in a video conferencing system comprising a first UD used by a first participant and comprising a first display, a second UD used by a second participant and comprising a second display, and a third UD used by a third participant. The method includes determining that the first participant is looking at a particular portion of the first display, wherein the particular portion of the first display is associated with the second participant. The method also includes determining that the second participant is looking at a particular portion of the second display, wherein the particular portion of the seconddisplay is associated with the first participant. The method also includes detecting that the first participant has performed a private gesture during a particular period of time. The method also includes, after detecting that the first participant has made the private gesture and based on the determinations, i) providing to the second UD a first video stream corresponding to the particular period of time for enabling the second UD to produce a first video representation of the first participant during the particular period of time, wherein the first video representation of the first participant shows the private gesture performed by the first participant and ii) generating a second video stream corresponding to the particular period of time for enabling the third UD to produce a second video representation of the first participant during the particular period of time, wherein the second video representation of the first participant does not show the private gesture performed by the first participant.

[0009] In another aspect there is provided a computer program comprising instructions which when executed by processing circuitry of a network node causes the network node to perform any of the methods disclosed herein. In one embodiment, there is provided a carrier containing the computer program wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium. In another aspect there is provided a network node that is configured to perform the methods disclosed herein. The network node may include memory and processing circuitry coupled to the memory.

[0010] An advantage of the embodiments disclosed herein is that they significantly improve the perception of attention in virtual meetings, and bring virtual meetings closer to the utility and function of nonverbal human-to-human communication in physical meetings. As a result, a presenter can more easily assess the engagement of the audience and how they look at the presenter as well as the presented material. There is also an opportunity to naturally create mutual attentions between any two out of many participants in meetings, which is today not possible without more complex and less intuitive, less natural procedures.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate various embodiments.

[0012] FIG. 1 illustrates a system according to an embodiment.

[0013] FIG. 2A illustrates a first user device according to an embodiment.

[0014] FIG. 2B illustrates a first user device according to an embodiment.

[0015] FIG. 3 is a flowchart illustrating a process according to an embodiment.

[0016] FIG. 4 is a flowchart illustrating a process according to an embodiment.

[0017] FIG. 5 illustrates various function of a user device according to an embodiment.

[0018] FIG. 6 illustrates various function of a user device according to an embodiment.

[0019] FIG. 7 is a flowchart illustrating a process according to an embodiment.

[0020] FIG. 8 is a flowchart illustrating a process according to an embodiment.

[0021] FIG. 9 illustrates a network node (e.g., user device or server) according to an embodiment.DETAILED DESCRIPTION

[0022] FIG. 1 illustrates a virtual conferencing system 100 according to an embodiment. System 100 includes a virtual conferencing server 104, which may be a node of network 110 (e.g., the internet). Virtual conferencing server 104 enables a user to engage in a virtual conference (a.k.a., virtual call) with one or more other users. In the example shown in FIG. 1, three user devices (UDs) (a.k.a., “user equipments (UEs)) are “connected” to server 104 and each UD (i.e., UD 101, UD 102, and UD 103) is providing a feed e.g., a video feed to server 104, which then forwards the feed to each other user. Hence, each UD, as shown in FIG. 1, is receiving, in this example, two feeds from server 104. More specifically, UD 101 receives the feed transmitted by UD 102 and also receives the feed transmitted by UD 103. Based on the received feed transmitted by UD 102, UD 101 is operable to produce a video representation of the user (a.k.a., participant) of UD 102 and output the video representation using a display (a.k.a., screen) of UD 101. Likewise, based on the received feed transmitted by UD 103, UD 101 is operable to produce a video representation of the user of UD 103 and output the video representation using the display of UD 101. The feed provided by a UD may comprise encoded images captured by the UD’s camera. The feed provided by a UD may also include metadata, such as data concerning the gaze of the user using the UD and / or data concerning a gesture made by the user of the UD. FIG. 1 is just an example of a video conferencing architecture. In anotherexample architecture, such as a peer-to-peer architecture, there may be no need for server 104 as each UD can transmit directly to each other UD. In such an embodiment, each UD may use its metadata directly to decide to send one feed (e.g., a modified feed as described herein) to one other UD and to send another feed (e.g., a non-modified feed) to another one or more of the other UDs.

[0023] FIG. 2 A further illustrates UD 101 according to an embodiment. In the embodiment shown, UD 101 includes a camera 202 mounted on a display 204 of UD 101. In this example, UD 101 is running a video conferencing application, and a window 208 produced by the application is displayed on display 204. As illustrated, window 208 includes sub-windows (referred to as “panels”) 211, 212, and 213, one panel for each participant. For example, panel 211 includes a video representation of the user of UD 101 (hereafter “the first user”) which enables the first user to see themself, panel 212 includes a video representation of the user of UD 102 (hereafter “the second user”), and panel 213 includes a video representation of the user of UD 103 (hereafter “the third user”). The users of UD 102 and UD 103 would see a similar screen because like UD 101, UD 102 and UD 103 typically also have cameras. Window 208 may also have another panel 214 for displaying information presented by one of the users. For instance, during the virtual conference, the second user may want the other users to see a document on the second user’s screen, and this information can be displayed in panel 214.

[0024] FIG. 2B further illustrates UD 102 according to an embodiment. In the embodiment shown, UD 102 includes a camera 252 mounted on a display 254 of UD 102. In this example, UD 102 is running the same video conferencing application as UD 101, and a window 258 produced by the application is displayed on display 254. As illustrated, window 258 includes panels 261, 262, and 263, one panel for each participant. For example, panel 261 includes a video representation of the user of UD 101, panel 262 includes a video representation of the user of UD 102, and panel 263 includes a video representation of the user of UD 103.Window 258 may also have another panel 264 for displaying information presented by one of the users. For instance, during the virtual conference, the second user may want the other users to see a document on the second user’s screen, and this information can be displayed in panel 264.

[0025] In one embodiment, the gaze of the first user is determined. This determination can be made by software running on UD 101, server 104, UD 102, and / or UD 103. Based on thedetermined gaze, the system can determine if the first user is looking at a particular one of the panels of window 208 (e.g., can determine if the first user is looking at the video representation of the second user or the video representation of the third user).

[0026] In one embodiment, if it is determined, that the first user is looking at the video representation of the second user, then the system causes the video representation of the first user that is displayed on the second user’s screen to make it appear as if the first user is looking directly into the lens of camera 202, thereby giving the second user the sensation that the first user is looking directly at the second user, but the video representation of the first user that is displayed on the third user’s screen does not cause the third user to perceive that the first user is looking directly at the third user (i.e., does not cause the third user to perceive that the first user is looking directly at the lens of camera 202, but rather looking elsewhere).

[0027] In another embodiment, if the first user looks at what is presented by a second user (e.g., the first user is looking at panel 214 at a time when the second user has control of the contents of panel 214), the feed corresponding to UD 101 and sent to UD 102 (e.g., the video stream based on images captured by camera 202) is adjusted so that it appears to the second user as if the first user is “almost” looking directly into the eyes of the second user, where “almost” means the second user can distinguish between not looking in the eyes at all, fully looking in the eyes, or only almost. As another example, if presentation panel 214 contains a chat from the third user, then video stream of the first user to the third user shall be so that it appears as if the first user “almost” looks in the eyes of the third user. In this way, the second or third user can determine whether the first user is focused on their input. A temporary gaze indicating the user is thinking shall not change this adapted video stream gaze, e.g., up towards left or saccades. Other users shall see the video stream from the first user as being further away from looking at them.

[0028] In yet another embodiment, if two users look at each other, this is recognized by the system (own gaze plus eye direction of the other user), then a “mutual attention” (MA) mode is created between those two users allowing them to also send gestures or other information without the other users being aware.

[0029] In still another embodiment, if a first user looks at a second user’s video representation, not only will the second user see a video representation of the first user where the first user’s eyes appear to be looking into the second user’s eyes, but the first user can showgestures (e.g., hand raised) which are only visible in the video representation of the first user that is displayed to the second user but not to anyone else (e.g., the third user).

[0030] In short, when the first user gazes at material of a second user in a video meeting, the feed (e.g., video flow) of the first user to the second user is adapted so it appears as if the first user looks close to the eyes of the second user but not directly in the eyes. When the first user and the second user gaze at each other in a video meeting, this “eye contact” is detected and it is indicated to the first user and second user, after which they can share gestures to each other via their feeds, whereas their feed to all other users are adjusted so that the other user will not see the gestures. And when a first user gazes directly at a second user in a video call, the video feed of the first user to the second user is adapted so that it appears as if the first user looks directly into the eyes of the second user, and this mode is indicated to the first user, after which the first user can make gestures which will be included in the feed to the second user, while the video feed from the first user to all other users is adapted so that it does not contain said gestures and it appears as if the first user is looking elsewhere than into the camera. In this way, the first user can communicate with the second user privately using gesture capture by camera 202.

[0031] Additional Details

[0032] This disclosure provides a system that detects the attention of a user and how the user's attention relates to other users in a virtual conference (e.g., the system, as noted above, can detect whether the first user is looking at the video representation of the second user). The produced output is, possibly multiple, video streams with different characteristics for different participants in the virtual conference.

[0033] A user's attention towards another user, which is an input to the system, is classified as one of: 1) Looking directly at the another user, 2) Looking at some resources associated with the another user (non-limiting examples includes chat-session, presentation or screen-share); and 3) Not looking at the another or any resources associated with the another user.

[0034] The system (e.g., a UD or server) may produce multiple output feeds based on a single input feed, to be sent to different receivers depending on the first users’ attention towards the receivers. The set of output feeds may contain an original camera feed, if no adjustments were needed to achieve the intended attention level in the stream.

[0035] An output feed is (possibly) adjusted to match one of the following attention levels: (1) No Close Gaze (NCG) - The user is looking elsewhere; (2) Close Gaze (CG) - The user appears to be almost looking into the camera, or at an object close to the camera; and (3) Direct Gaze (DG) - The user is looking directly into the camera giving the impression of direct eye contact.

[0036] Video feed of face / eyes

[0037] FIG. 3 is a flowchart illustrating a process, according to an embodiment, that is performed during a video call (or other virtual conference) with multiple (2 or more) participants.

[0038] When the video call has started, the system 100 (e.g., UD 101) detects the gaze of the first user, and relates that to what is shown on screen 204, e.g., placement of video frames, shared slides, or chat windows (step s301).

[0039] The system (e.g., UD 101) then determines if the first user is looking directly at a certain person (e.g., the video representation of the second user shown in panel 212) or looking at information provided by a certain person (e.g., panel 214) (step s302).

[0040] If the system determines that the first user is looking at the second user, then the process proceeds to step s305, else if the system determines that the first user is looking at content associated with the second user (e.g., the first user is looking at panel 214 while the second user has control of the content of panel 214), then the process proceeds to step s307, else the process proceeds to step s309.

[0041] In step s305, the system (e.g., UD 101) obtains two feeds: 1) a direct gaze (DG) feed and a non-close gaze (NCG) feed. The NCG feed may simply be the feed produced by camera 202 (or the conventional video encoder that receives the images from camera) while the DG feed is a processed version of the NCG feed such that when the DG feed is displayed on a screen the video representation of the first user appears to be looking into camera 220. In step s306, the DG feed is provided to UD 102 (i.e., the UD used by the second user) so that the second user will perceive that the first user is looking directly at the second user, while the NCG feed is provided to all of the other participants (e.g.,. the third user). For example, UD 101 sendsboth the DG feed and NCG feed to server 104 together with metadata to cause server 104 to forward the DG feed to UD 102 and forward the NCG feed to UD 103.

[0042] In step s307, the system (e.g., UD 101) obtains, such as creates, two feeds: 1) a close gaze (CG) feed and the NCG feed. As noted above, the NCG feed may simply be the feed produced by camera 202, a feed produced by the conventional video encoder that receives the images from camera, a processed version of the feed produced by the camera, or a processed version of the feed produced by the video encoder, while the CG feed is a modified version of the NCG feed such that when the CG feed is displayed on a screen the video representation of the first user appears to be looking almost a camera 220. In step s306, the CG feed is provided to UD 102 so that the second user will perceive that the first user is looking almost directly at the second user, while the NCG feed is provided to all of the other participants (e.g., the third user).

[0043] In steps s309, the system (e.g., UD 101) obtains a single feed - the NCG feed. In step s310, the NCG feed is provided to all participants.

[0044] Mutual eye contact - Enable Private Gestures

[0045] FIG. 4 is a flowchart illustrating a process, according to an embodiment, that is performed during a video call (or other virtual conference) with multiple (2 or more) participants and that is for sharing private gestures during the video call.

[0046] In step s401, the system determines whether two participants (e.g., the first user and the second user) are looking “at each other”, more precisely, whether the first user is looking at the video representation of the second user and whether the second user is looking at the video representation of the first user.

[0047] If the system determines that the first and second users are looking at each other (i.e., DG in both directions), then the system indicates this fact to each user (e.g., by changing the frame color of panel of the respective video representation to a certain color, e.g., blue) (step s402). This is referred to as the first user and second use entering the “mutual attention” (MA) mode.

[0048] In step s403, the system determines if a user in MA mode with another user has made a gesture (any gesture or one of a set of predefined gestures). If so, the process proceeds to step s404, otherwise the process proceeds to step s405.

[0049] In step s404, assuming the first user made a gesture while in MA mode with the second user, the system produces a modified first user NCG feed that omits the gesture and this modified NCG feed is sent to all participants other than the second user. The second user will get the first user DG feed so that the second user will see the gesture and perceive that the first user is looking into camera 202. The other user will not see the gesture because they receive the modified NCG feed in which the gesture has been removed e.g., a video editing process edits the images of the original first user feed to remove the gesture. In some embodiments, a gesture will be removed from the NCG only if it is a gesture comprised in a predefined set of gestures (e.g., thumbs-up, thumbs-down, wave, smirk, etc.). Accordingly, in some embodiments, hand gestures or facial expression gestures are privately communicated only in MA mode and remain private between those users, whereas others only see the neutral phase without hand gestures with NCG eyes.

[0050] In some embodiments, the gestures being detected while the first and second users are in the MA mode and then determined by the system as being private gestures are indicated to the first user, e.g., by a changing the color of a frame 215 of panel 211 (i.e., the panel containing the video representation of the first user) or by a segmentation of the gesture, allowing the first user to immediately act (e.g., cancel the gesture) if that gesture was not intended to be shared or not intended to be private. In one embodiment, with segmentation of the gesture there is a separation of the video feed into two layers (or two feeds) where one layer holds only the gesture and the other layer holds the person (either person - gesture-pixels, or person appearance modified as if they had not made a gesture). The segmented gesture (i.e. layer or feed containing only the gesture) could be indicated to the first user in many ways (icon, small inset video feed window, colour highlight or other change over the segmented gesture, etc. etc.).

[0051] In some embodiments, when a private gesture is included in the feed for the first user that is sent to the second user, an indication (e.g., visual indication) is provided to the first user (e.g., the frame 216 of panel 212 changes to a certain color) indicating that the gesture being shown is private to the second user and not visible to all. Anyone skilled in the art knows that there are various different types of feedback that can be used for such purposes, and that it need not be limited to visual indications.

[0052] In one embodiment, the first user and the second user leave the MA mode when the system determines that one of the users moves the gaze (step s405 and s406). For practical reasons, there can be a time hysteresis so that there has to be a mutual gaze for e.g., 1 second before MA mode is activated, and a brief change of gaze below e.g., 1 second does not stop MA mode. In some embodiments, MA mode can be turned off by a user, and a user might be able to request MA mode towards a specific other user, and an MA mode can be stopped by pressing a certain key (e.g., escape (Esc) key) or other specific command.

[0053] In one embodiment, when the first user and second user enter MA mode, they can privately talk (e.g., whisper) to each other and not only privately share gestures. In one embodiment, the private talking mode can be initiated either the first user or the second user by some user action, such as, for example, a key press, selection of an on-screen icon, or whispering to identify talking that is meant to be private between the first and second user vs talking that is meant to be towards everyone. With respect to a user whispering to activate the private talking mode, a whisper can be detected by the system using a certain volume threshold (whisper vs. speaking-normally).

[0054] Use cases

[0055] The feeds in above examples can be (A) camera-captured video feeds, (B) partly synthesized feeds (e.g., by morphing of one camera's feed, by using RGB+D cameras), or (C) wholly artificially generated feeds (e.g., virtual synthetic-human avatars, preferably of high- enough quality to feel real, mimicking the user and synchronized to the user's movements determined from a camera or X sensor). The embodiments are not restricted to any specific camera setup.

[0056] For example, in one embodiment, at least UD 101 is configured as shown in FIG. 5. That is, in the embodiment shown in FIG. 5, UD 101 includes camera 202, which produces a sequence of images (a.k.a., “video stream” or “video” for short) 501; a gaze detector 512 for determining whether the first user is looking at information related to one of the other participants in the virtual conference, e.g., whether the first user is looking at panel 214, and for outputting metadata 502 comprising gaze information, such as, for example, information indicating the participant associated with the information on which the first user is focused; an image processor 514 that receives the images (a.k.a., frames) 501 produced by camera 202 andoperable to, based on gaze information received from gaze detector 512, modify the images to produce a modified video stream 503 so that, for example, a visual representation produced when the modified video stream is displayed makes it appear that the first user is looking almost directly at the lens of camera 202 even when the first user is looking at, for example, panel 214; and a multiplexor 516 that receives original video stream 501, modified video stream 503, and metadata 502 and transmits these to server 104. That is, for example, in some embodiments the gaze information from gaze detector 512 to image processor 514 informs image processor 514 of which type of processing, if any, it should do to the images from camera 501. Hence, image processor 514 can produce any type of modified stream based on information from gaze detector 512. Server 104, in this embodiment, determines where to send streams 501 and 503 based on metadata 502. In this example, metadata 502 indicates that the first user is focused on information provided by the second user, and, therefore, server 104 forwards stream 503 to UD 102 and forwards stream 501 to all other participants (i.e., UD 103 in this example).

[0057] In another embodiment, image processor 514 is a component of server 104 rather than UD 101, and UD 101 provides stream 501 and metadata 502 to server 104. In this embodiment, if metadata 502 indicates that the first user is not looking at any information provided by one of the other participants, then server merely forwards stream 501 to all participants (i.e., to UD 102 and UD 103 in this example); but, if metadata 502 indicates that the first user is looking at information provided by a particular participant, then server employs image processor to create modified stream 503, forwards modified stream 503 to the UD used by the participant providing the information on which the first user is focused, and forwards stream 501 to all other participants.

[0058] In another embodiment, image processor (IP) 514 is a component of one or more of the other UDs (e.g. UD 102). In this embodiment, IP 514 receives original video stream 501 and metadata 502, and, if metadata 502 indicates that the user associated with stream 501 is looking at information provided by the user associated with the UD on which IP 514 is running, then IP 514 will produce video stream 503 and the images of video stream 503 will be displayed to the user of the UD such that the user of the UD will perceive the other user as looking almost directly into his / her camera lens.

[0059] In another embodiment, at least UD 101 is configured as shown in FIG. 6. That is, in the embodiment shown in FIG. 6, UD 101 includes camera 202, which produces video stream 501; a gaze detector 612 for determining whether the first user is looking at one of the other participants in the virtual conference (e.g., whether the first user is looking at panel 212 or panel 213) and for outputting metadata 602 comprising information indicating the participant the first user is looking at; an image processor 614 that receives the images (a.k.a., frames) 501 produced by camera 202 and is operable to modify the images to produce a modified video stream 603 so that a visual representation produced when the modified video stream is displayed makes it appear that the first user is looking directly at the lens of camera 202 even when the first user is looking at, for example, panel 212; and a multiplexor 516 that receives original video stream 501, modified video stream 603, and metadata 602 and transmits these to server 104. Server 104, in this embodiment, determines where to send streams 501 and 603 based on metadata 602. In this example, metadata 602 indicates that the first user is focused on the second user, and, therefore, server forwards stream 603 to UD 102 and forwards stream 501 to all other participants (i.e., UD 103 in this example).

[0060] In this embodiment, image processor (IP) 614 includes a gesture detector (GD) 632. In this embodiment, in addition to generating video stream 603 as described above, IP 614 will either output original video stream 501 or video stream 601, which is another modified version of stream 501. More specifically, if GD 632 determines that the first user is in a MA mode with another user (e.g., GD 632 receives information from the another user’s user device indicating that the another user is looking at the first user and GD 632 receives information from gaze detector 612 indicating that the first user is looking at the another user, thereby enabling GD 632 to detect that the first user and the other user are in an MA mode) and GD 632 detects a gesture, then GD 632 causes IP 614 to produce video stream 601 such that the images of stream 601 do not show the gesture. Hence, when the first user and the second user are in an MA mode and the first user makes a gesture, stream 603 will be sent by server 104 to UD 102 and stream 601 will be sent by server 104 to UD 103 (server knows which UD gets which stream based on the metadata 602). In some embodiments, IP 614 with GD 632 may instead of running on UD 101 run on server 104 or on both UD 102 and UD 103. In the embodiment in which IP 614 runs on UD 102 and UD 103, when UD 102 receives stream 501, the IP running on UD 102 will produce and cause to be outputted stream 603, thereby causing the second user to perceive thefirst user looking directly at camera 202, and when UD 103 receives stream 501, the IP running on UD 103 will produce and cause to be outputted stream 601, thereby preventing the third user from seeing the private gesture between the first and second users.

[0061] FIG. 7 is a flow chart illustrating a process 700 according to an embodiment. Process 700 may begin in step s702.

[0062] Step s702 comprises determining that the first participant is looking at portion 214 of display 204.

[0063] Step s704 comprises receiving a first video stream captured camera 202.

[0064] Step s706 comprises, based on the determination, modifying the first video stream to produce a modified video stream for causing UD 102 to display a first video representation of the first participant, wherein the first video representation of the first participant is configured for causing the second participant to believe that the first participant is not looking directly at camera 202 but rather looking in a first direction that is offset from a second direction, wherein the second direction is a direction from the first participant to the camera.

[0065] In one embodiment, the particular portion of the first display is a portion of the first display that contains content presented by or generated by the second participant.

[0066] In one embodiment, the first direction is offset from the second direction by at least a first degree, such as, for example, 1 degree, and no more than a second degree, such as, for example, 45 degrees. In another embodiment, the first degree is about 5 or about 10 degrees and the second degrees is about 25 degrees or about 30 degrees.

[0067] In one embodiment, the first video representation of the first participant is configured for causing the second participant to believe that the first participant is looking at a point x units of distance to the left of the camera or is looking at a point x units of distance to the right of the camera, where x is greater than 0. In one embodiment, x is greater than or equal to 1 centimeter and less than or equal to 20 centimeters.

[0068] In one embodiment, the method is performed by the first UD.

[0069] In one embodiment, the method further comprises: transmitting to a virtual conferencing server the first video stream, the modified video stream, and metadata indicatingthat the modified video stream should be forwarded by the server to the second UD in place of the first video stream.

[0070] In one embodiment, the method is performed by a virtual conferencing server. In one embodiment determining that the first participant is looking at the particular portion of the first display comprises the virtual conferencing server receiving metadata transmitted by the first UD, wherein the metadata indicates that the first participant is looking at the particular portion of the first display. In one embodiment the method further comprises: based on the metadata, transmitting the first video stream to a third UD used by a third participant and transmitting the modified video stream to the second UD.

[0071] In one embodiment, the method is performed by the second UD.

[0072] In one embodiment, determining that the first participant is looking at the particular portion of the first display comprises the second UD receiving metadata transmitted by the first UD, wherein the metadata indicates that the first participant is looking at the particular portion of the first display.

[0073] FIG. 8 is a flow chart illustrating a process 800 according to an embodiment. Process 800 may begin in step s802.

[0074] Step s802 comprises determining that the first participant is looking at portion 212.

[0075] Step s804 comprises determining that the second participant is looking at portion 261 of display 254.

[0076] Step s806 comprises detecting that the first participant has performed a private gesture during a particular period of time.

[0077] Step s808 comprises, after detecting that the first participant has made the private gesture and based on the determinations: (a) providing to the UD 102 a first video stream corresponding to the particular period of time for enabling the UD 102 to produce a first video representation of the first participant during the particular period of time, wherein the first video representation of the first participant shows the private gesture performed by the first participant; and (b) generating a second video stream corresponding to the particular period of time for enabling UD 103 to produce a second video representation of the first participant during theparticular period of time, wherein the second video representation of the first participant does not show the private gesture performed by the first participant.

[0078] In one embodiment, the method is performed by the first UD.

[0079] In one embodiment, providing the first video stream to the second UD comprises transmitting the first video stream to a virtual conferencing server configured to forward the first video stream to the second UD. In one embodiment, the method further comprises transmitting to the virtual conferencing server the second video stream and metadata indicating that the second video stream should be forwarded by the server to the third UD.

[0080] In one embodiment, the method is performed by a virtual conferencing server.

[0081] In one embodiment, determining that the first participant is looking at the particular portion of the first display comprises receiving from the first UD first metadata indicating that the first participant is looking at the particular portion of the first display, and determining that the second participant is looking at the particular portion of the second display comprises receiving from the second UD second metadata indicating that the second participant is looking at the particular portion of the second display. In one embodiment, the method further comprises receiving the first video stream from first UD.

[0082] FIG. 9 is a block diagram of a network node 900, according to some embodiments, which can be used to implement any of the UDs described herein and which can be used to implement server 104. As shown in FIG. 9, network node 900 may comprise: processing circuitry (PC) 902, which may include one or more processors (P) 955 (e.g., one or more general purpose microprocessors and / or one or more other processors, such as an application specific integrated circuit (ASIC), field-programmable gate arrays (FPGAs), and the like), which processors may be co-located in a single housing or in a single data center or may be geographically distributed (e.g., network node 900 may be a distributed computing apparatus comprising two or more computers or a monolithic computing apparatus consisting of a single computer); at least one network interface 948 (e.g., a physical interface or air interface) comprising a transmitter (Tx) 945 and a receiver (Rx) 947 for enabling network node 900 to transmit data to and receive data from other nodes connected to network 110 (e.g., an Internet Protocol (IP) network) to which network interface 948 is connected (physically or wirelessly) (e.g., network interface 948 may be coupled to an antenna arrangement comprising one or moreantennas for enabling network node 900 to wirelessly transmit / receive data); and a storage unit (a.k.a., “data storage system”) 908, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments where PC 902 includes a programmable processor, a computer readable storage medium (CRSM) 942 may be provided. CRSM 942 may store a computer program (CP) 943 comprising computer readable instructions (CRI) 944. CRSM 942 may be a non-transitory computer readable medium, such as, magnetic media (e.g., a hard disk), optical media, memory devices (e.g., random access memory, flash memory), and the like. In some embodiments, the CRI 944 of computer program 943 is configured such that when executed by PC 902, the CRI causes network node 900 to perform steps described herein (e.g., steps described herein with reference to the flow charts). In other embodiments, network node 900 may be configured to perform steps described herein without the need for code. That is, for example, PC 902 may consist merely of one or more ASICs. Hence, the features of the embodiments described herein may be implemented in hardware and / or software.

[0083] While various embodiments are described herein, it should be understood that they have been presented by way of example only, and not limitation. Thus, the breadth and scope of this disclosure should not be limited by any of the abovedescribed exemplary embodiments. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.

[0084] As used herein transmitting a message “to” or “toward” an intended recipient encompasses transmitting the message directly to the intended recipient or transmitting the message indirectly to the intended recipient (i.e., one or more other nodes are used to relay the message from the source node to the intended recipient). Likewise, as used herein receiving a message “from” a sender encompasses receiving the message directly from the sender or indirectly from the sender (i.e., one or more nodes are used to relay the message from the sender to the receiving node). Further, as used herein “a” means “at least one” or “one or more.”

[0085] Additionally, while the processes described above and illustrated in the drawings are shown as a sequence of steps, this was done solely for the sake of illustration. Accordingly, itis contemplated that some steps may be added, some steps may be omitted, the order of the steps may be re-arranged, and some steps may be performed in parallel.

[0086] References

[0087] [1] US 9,743,040 Bl (Symantec filed 2015): Systems and methods for facilitating eye contact during video conferences.

[0088] [2] US 9,626,003 B2 (IBM, Filed 2016): Natural gazes during online video conversations.

[0089] [3] US 11,444,982 Bl (Amazon, filed 2022): Method and apparatus for repositioning meeting participants within a gallery view in an online meeting user interface based on gestures made by the meeting participants.

[0090] [4] US 11,606,508 Bl (Motorola, Filed 2021): Enhanced representations based on sensor data.

[0091] [5] US 11,443,560 Bl (Zoom, Filed 2021): View layout configuration for increasing eye contact in video communication.

Claims

CLAIMS1. A method (700) in a virtual conferencing system (100) comprising i) a first user device, UD (101) used by a first participant and comprising a first display (204) and a first camera (202) and ii) a second UD (102) used by a second participant and comprising a second display (254) and a second camera (252), the method comprising: determining (s702) that the first participant is looking at a particular portion (214) of the first display, wherein the particular portion of the first display is associated with the second participant; receiving (s704) a first video stream captured by the first camera; and based on the determination, modifying (s706) the first video stream to produce a modified video stream for causing the second UD to display a first video representation of the first participant, wherein the first video representation of the first participant is configured for causing the second participant to believe that the first participant is not looking directly at the camera (202) but rather looking in a first direction that is offset from a second direction, wherein the second direction is a direction from the first participant to the camera.

2. The method of claim 1, wherein the particular portion of the first display is a portion of the first display that contains content presented by or generated by the second participant.

3. The method of claim 1 or 2, wherein the first direction is offset from the second direction by at least a first degree and no more than a second degree.

4. The method of any one of claims 1-3, wherein the first video representation of the first participant is configured for causing the second participant to believe that the first participant is looking at a point x units of distance to the left of the camera or is looking at a point x units of distance to the right of the camera, where x is greater than 0.

5. The method of claim 4, wherein x is greater than or equal to 1 centimeter and less than or equal to 20 centimeters.

6. The method of any one of claims 1-5, wherein the method is performed by the firstUD.

7. The method of claim 6, wherein the method further comprises: transmitting to a virtual conferencing server (104) the first video stream, the modified video stream, and metadata indicating that the modified video stream should be forwarded by the server to the second UD in place of the first video stream.

8. The method of any one of claims 1-5, wherein the method is performed by a virtual conferencing server (104).

9. The method of claim 8, wherein determining (s702) that the first participant is looking at the particular portion (214) of the first display comprises the virtual conferencing server (104) receiving metadata transmitted by the first UD, wherein the metadata indicates that the first participant is looking at the particular portion (214) of the first display.

10. The method of claim 8 or 9, wherein the method further comprises: based on the metadata, transmitting the first video stream to a third UD (103) used by a third participant and transmitting the modified video stream to the second UD (102).

11. The method of any one of claims 1-5, wherein the method is performed by the second UD.

12. The method of claim 11, wherein determining (s702) that the first participant is looking at the particular portion (214) of the first display comprises the second UD (102) receiving metadata transmitted by the first UD, wherein the metadata indicates that the first participant is looking at the particular portion (214) of the first display.

13. A method (800) in a video conferencing system comprising a first UD, UD (101), used by a first participant and comprising a first display (204), a second UD (102) used by a second participant and comprising a second display (254), and a third UD used by a third participant, the method comprising: determining (s802) that the first participant is looking at a particular portion (212) of the first display (204), wherein the particular portion (212) of the first display (204) is associated with the second participant; determining (s804) that the second participant is looking at a particular portion (261) of the second display (254), wherein the particular portion (261) of the second display (254) is associated with the first participant; detecting (s806) that the first participant has performed a private gesture during a particular period of time; and after detecting that the first participant has made the private gesture and based on the determinations: providing (s808) to the second UD a first video stream corresponding to the particular period of time for enabling the second UD to produce a first video representation of the first participant during the particular period of time, wherein the first video representation of the first participant shows the private gesture performed by the first participant; and generating (s810) a second video stream corresponding to the particular period of time for enabling the third UD to produce a second video representation of the first participant during the particular period of time, wherein the second video representation of the first participant does not show the private gesture performed by the first participant.

14. The method of claim 13, wherein the method is performed by the first UD (101).

15. The method of claim 14, wherein providing the first video stream to the second UD comprises transmitting the first video stream to a virtual conferencing server (104) configured to forward the first video stream to the second UD.

16. The method of claim 15, wherein the method further comprises: transmitting to the virtual conferencing server (104) the second video stream and metadata indicating that the second video stream should be forwarded by the server to the third UD.

17. The method of claim 13, wherein the method is performed by a virtual conferencing server (104).

18. The method of claim 17, wherein determining (s802) that the first participant is looking at the particular portion (212) of the first display (204) comprises receiving from the first UD first metadata indicating that the first participant is looking at the particular portion (212) of the first display (204), and determining (s804) that the second participant is looking at the particular portion (261) of the second display (254) comprises receiving from the second UD second metadata indicating that the second participant is looking at the particular portion (261) of the second display (254).

19. The method of claim 17 or 18, wherein the method further comprises receiving the first video stream from first UD.

20. A computer program (943) comprising instructions (944) which when executed by processing circuitry (902) of a network node (900) causes the network node to perform the method of any one of claims 1-19.

21. A carrier containing the computer program of claim 20, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium (942).

22. A network node (900) in a virtual conferencing system (100), wherein the network node is configured to perform a method comprising:determining (s702) that a first participant using a first user device, UD (101) is looking at a particular portion (214) of a first display (204) of the first UD, wherein the particular portion of the first display is associated with a second participant using a second UD (102); receiving (s704) a first video stream captured by a first camera (202) of the first UD; and based on the determination, modifying (s706) the first video stream to produce a modified video stream for causing the second UD to display a first video representation of the first participant, wherein the first video representation of the first participant is configured for causing the second participant to believe that the first participant is not looking directly at the camera (202) but rather looking in a first direction that is offset from a second direction, wherein the second direction is a direction from the first participant to the camera.

23. The network node of claim 22, wherein the network node is further configured to perform the method of any one of claims 2-12.

24. A network node (900) in a virtual conferencing system (100), wherein the network node is configured to perform a method comprising: determining (s802) that a first participant using a first user device, UD (101) is looking at a particular portion (212) of a first display (204) of the first UD, wherein the particular portion (212) of the first display (204) is associated with a second participant using a second UD; determining (s804) that the second participant is looking at a particular portion (261) of a second display (254) of the second UD, wherein the particular portion (261) of the second display (254) is associated with the first participant; detecting (s806) that the first participant has performed a private gesture during a particular period of time; and after detecting that the first participant has made the private gesture and based on the determinations: providing to the second UD a first video stream corresponding to the particular period of time for enabling the second UD to produce a first video representation of the first participant during the particular period of time, wherein the first video representation of the first participant shows the private gesture performed by the first participant; andgenerating a second video stream corresponding to the particular period of time for enabling a third UD to produce a second video representation of the first participant during the particular period of time, wherein the second video representation of the first participant does not show the private gesture performed by the first participant.

25. The network node of claim 24, wherein the network node is further configured to perform the method of any one of claims 14-19.

Citation Information

Patent Citations

  • View layout configuration for increasing eye contact in video communications

    US11443560B1

  • Method and apparatus for repositioning meeting participants within a gallery view in an online meeting user interface based on gestures made by the meeting participants

    US11444982B1

  • Enhanced representations based on sensor data

    US11606508B1

  • Method for sensing and communicating visual focus of attention in a video conference

    EP4113982A1

  • Private / public gesture security system and method of operation thereof

    US20140006794A1