Method and system for virtual 3D communication - Patents.com

JP2024518888A5Pending Publication Date: 2025-09-17TRUE MEETING INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023564028
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-05-10
Filing Date
2022-05-10
Publication Date
2025-09-17

AI Technical Summary

Technical Problem

Current video teleconferencing systems suffer from issues such as disconnected participant appearances, lack of eye contact, unclear audio direction, and poor image quality, leading to user fatigue and reduced engagement.

Method used

A method and system for conducting 3D video conferencing that uses gaze direction information to update participant representations in a virtual 3D environment, incorporating 3D models and texture maps to enhance visual interaction and audio synchronization, and employs neural networks for real-time facial expression estimation and audio compression.

Benefits of technology

Enhances user engagement by providing a more immersive and interactive virtual conferencing experience, improving visual and audio clarity, and reducing bandwidth requirements while maintaining high-quality communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for conducting three-dimensional (3D) video conferencing among a plurality of participants may be provided, the method may include acquiring visual information by a visual sensing unit associated with a participant, identifying a plurality of persons appearing in the visual information, discovering at least one associated person from the plurality of persons, determining 3D entity representation information for each of the at least one associated person, and generating, for the at least one participant, a representation of a virtual 3D video conferencing environment based on the 3D entity representation information for each of the at least one associated person.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] cross reference This application claims priority from U.S. Provisional Patent Application No. 63 / 201,713, filed May 10, 2021, which is incorporated by reference in its entirety. [Background technology]

[0002] Video teleconferencing has become very popular. They require each participant to have their own computerized system with a camera usually located near the display.

[0003] Typically, several participants in a meeting will attend in separate small tiles, and another tile may be used to share one of the participants' screens.

[0004] Each participant is typically shown with their own office background, or with a virtual background of their choice.

[0005] Participants are shown from different angles and at different sizes.

[0006] As a result, people may feel disconnected and not as if they were all present in the same room.

[0007] Because users typically look at a screen on which the face of the person they are interacting with is displayed, and not at a camera that may be above or below the screen, for example, the image that appears is of a person looking upwards or downwards, respectively, and not at the other person, and eye contact between the participants in the conversation is thus lost, which enhances the sensation of being disconnected.

[0008] Furthermore, on each participant's screen, the images of the other users may be located in different positions and in a variable order, so it is not clear who is seeing who.

[0009] Because all audio streams from all participants are merged into one single mono track audio stream, it is impossible to know which direction the sound is coming from, which can make it difficult to determine who is speaking at any given moment.

[0010] Because most webcams capture images of the face from mid-chest up, participants' hands are not frequently shown, and therefore hand gestures that are a critical part of standard conversation are not conveyed in a typical video conference.

[0011] Furthermore, the quality of the traffic (bitrate, packet loss, and latency) may change over time and the quality of the video conference call may vary accordingly.

[0012] Typically, video conferencing images tend to be blurry due to limited camera resolution (1080x720 pixels for a common laptop camera), motion blur, and video compression. In many cases, the video freezes and the audio sounds metallic or is lost.

[0013] All these limitations result in an effect known as Zoom fatigue (https: / / hbr.org / 2020 / 04 / how-to-combat-zoom-fatigue), which results in participants becoming more exhausted after hours of video conference meetings than they would typically do in a standard meeting in the same room. [Prior art documents] [Non-patent literature]

[0014] [Non-Patent Document 1] https: / / hbr.org / 2020 / 04 / how-to-combat-zoom-fatigue [Non-Patent Document 2] https: / / en.wikipedia.org / wiki / Iterative_closest_point [Non-Patent Document 3] https: / / flame.is.tue.mpg.de / home Summary of the Invention

[0015] There is an increasing need to enhance virtual interaction between participants and overcome various other problems associated with current video teleconferencing services. [Brief description of the drawings]

[0016] [Figure 1] FIG. 1 illustrates an example of a method. [Diagram 2] FIG. 1 illustrates an example computerized environment. [Diagram 3] FIG. 1 illustrates an example computerized environment. [Figure 4] FIG. 2 illustrates an example of a data structure. [Diagram 5] FIG. 13 illustrates an example of a process for correcting the view direction of a 3D model of a participant's part according to the participant's gaze direction. [Figure 6] FIG. 1 includes an example of a method. [Figure 7] FIG. 1 is an illustration of an image and process. [Figure 8] FIG. 13 is an illustration of an example of parallax correction. [Figure 9] FIG. 1 illustrates an example of a 2.5-dimensional illusion. [Figure 10] FIG. 1 illustrates an example of 3D content for a 3D screen or virtual reality headset. [Figure 11] Illustrated examples of a panoramic view of a virtual 3D environment populated by five participants, a partial view of some of the participants within the virtual 3D environment, and a hybrid view. [Figure 12] 1A-1C are diagrams showing example images of different exposures and example images of a face of different shades. [Figure 13] FIG. 2 is an illustration of an example of a face image and image segmentation. [Figure 14] FIG. 1 illustrates an example of a method. [Figure 15] FIG. 2 is an example of a 3D model and UV map. [Figure 16] FIG. 13 is an example of 2D-2D correspondence calculation for the upper and lower lips. [Figure 17] FIG. [Figure 18] FIG. [Figure 19] FIG. [Figure 20] FIG. 1 is a diagram illustrating a facial texture map. [Figure 21] FIG. 1 illustrates an example of a method. [Figure 22] FIG. 1 illustrates example images capturing two people and example avatars representing one or more people or even more participants. [Diagram 23] FIG. 1 illustrates an example of participants' gaze directions. [Figure 24] FIG. 1 illustrates an example of a method. [Diagram 25] A diagram illustrating examples of various signals exchanged between a computerized environment, a shared folder, and a user device. [Figure 26] FIG. 1 illustrates an example of a timing diagram. [Figure 27] FIG. 1 illustrates an example of a method. [Figure 28] FIG. 1 illustrates an image and examples of foreground and background segmentation. [Figure 29] FIG. 1 illustrates an example of a method. [Diagram 30] FIG. 13 illustrates an example of a participant without lipstick. [Diagram 31] FIG. 1 illustrates an example of a method. [Diagram 32] FIG. 1 illustrates an example of a method. [Diagram 33]1 illustrates different parts of a virtual 3D video conference. [Diagram 34] Illustrates an example of a participant having lipstick. [Diagram 35] FIG. 13 illustrates an example of a participant's avatar without lipstick. [Diagram 36] FIG. 1 illustrates examples of lipstick free expression on participants' lips. [Figure 37] FIG. 13 illustrates an example of a participant avatar with lipstick. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0017] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the disclosed embodiments.

[0018] However, it will be understood by those skilled in the art that the embodiments of the disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, and components have not been described in detail so as not to obscure the embodiments of the disclosure.

[0019] The subject matter regarded as embodiments of the disclosure is particularly pointed out and distinctly claimed in the concluding portion of the specification, however, the disclosed embodiments, both as to organization and method of operation, together with their objects, features, and advantages, may best be understood by reference to the following detailed description when read in conjunction with the accompanying drawings.

[0020] It will be appreciated that for simplicity and clarity of illustration, elements shown in the figures have not necessarily been drawn to scale. For example, the dimensions of some of the elements may be exaggerated relative to other elements for clarity. Further, where considered appropriate, reference numerals may be repeated among the figures to indicate corresponding or similar elements.

[0021] Because the illustrated embodiments of the disclosure can, for the most part, be implemented using electronic components and circuits known to those skilled in the art, details have not been described to any greater extent than is considered necessary for an understanding and appreciation of the basic concepts of the disclosed embodiments, and in order not to obfuscate or distract from the teachings of the disclosed embodiments, as exemplified above.

[0022] Any reference in the specification to a method should apply mutatis mutandis to a system capable of carrying out the method, and should apply mutatis mutandis to a computer readable medium that is non-transitory and has instructions for carrying out the method stored thereon.

[0023] Any reference in the specification to a system should be applied mutatis mutandis to methods that may be performed by the system, and should be applied mutatis mutandis to a computer readable medium having stored thereon instructions that are non-transitory and executable by the system.

[0024] References in the specification to a computer-readable medium being non-transitory should be applied mutatis mutandis to the manner in which it may be applied when executing instructions stored on the computer-readable medium, and should be applied mutatis mutandis to a system configured to execute instructions stored on the computer-readable medium.

[0025] The term "and / or" means in addition or alternatively.

[0026] References to "user" should apply mutatis mutandis to the term "participant" and vice versa.

[0027] A video-related method, non-transitory computer-readable medium, and system are provided and may be applicable, for example, to 3D video teleconferencing. At least some of the examples and / or embodiments illustrated in this application may be applied mutatis mutandis for other purposes and / or during other applications.

[0028] For example, consider a 3D video conference involving multiple participants: a first participant is imaged and a second participant wishes to see a first avatar (or any other 3D visual representation) of the first participant within a virtual 3D video conference environment.

[0029] Generation of the first avatar (or any other 3D visual representation) may be performed in various manners, such as, for example, solely by the second participant's device, solely by the first participant's device, partially by the second participant's device, partially by the first participant's device, by coordination between the first participant's device and the second participant's device, by another computerized system (such as, but not limited to, a cloud system or a remote system), and / or by any combination of one or more devices.

[0030] The inclusion of the avatar (or any other 3D visual representation) in the virtual 3D video conferencing environment may be performed in various manners, such as solely by the second participant's device, solely by the first participant's device, partially by the second participant's device, partially by the first participant's device, by coordination between the first participant's device and the second participant's device, by another device (such as a cloud device or remote device), and / or by any combination of one or more devices.

[0031] Any reference to one manner of execution of any step of generating a first avatar and / or any reference to one manner of execution of any step of including an avatar in a virtual 3D videoconferencing environment may apply mutatis mutandis to any other manner of execution.

[0032] Generating and / or including the first avatar may be responsive to information obtained by the first user's device or a camera or sensor associated with the first user's device. Non-limiting examples of the information may include information about the first participant and / or information regarding the acquisition of an image of the first participant (e.g., camera settings, lighting conditions, and / or ambient conditions).

[0033] The system may include multiple user devices and / or intermediate devices such as servers, cloud computers, and the like.

[0034] FIG. 1 illustrates an example of a method 200 .

[0035] The method 200 is for conducting three-dimensional video conferencing among multiple participants.

[0036] The method 200 may include steps 210, 220, and 230.

[0037] Step 210 may include receiving gaze direction information regarding each participant's gaze direction within a representation of the virtual 3D videoconferencing environment associated with the participant.

[0038] The representation of the virtual 3D videoconferencing environment associated with a participant is the representation that is shown to the participant. Different participants may be associated with different representations of the virtual 3D videoconferencing environment.

[0039] The gaze direction information may represent the detected direction of a participant's gaze.

[0040] The gaze direction information may represent an estimated direction of a participant's gaze.

[0041] Step 220 may include determining, for each participant, updated 3D participant representation information within the virtual 3D videoconferencing environment that reflects the participant's gaze direction. Step 220 may include estimating how the virtual 3D videoconferencing environment will be seen from the participant's gaze direction.

[0042] Step 230 may include generating, for at least one participant, an updated representation of the virtual 3D videoconferencing environment, where the updated representation of the virtual 3D videoconferencing environment represents updated 3D participant representation information for at least a portion of the plurality of participants. Step 230 may include rendering an image of the virtual 3D videoconferencing environment for at least a portion of the plurality of participants. Alternatively, step 230 may include generating input information (a 3D model and / or one or more texture maps) to be fed into the rendering process.

[0043] The method 200 may also include displaying 240, by participant devices of the plurality of participants, an updated representation of the virtual 3D videoconferencing environment, where the updated representation may be associated with the participants.

[0044] The method 200 may include transmitting 250 an updated representation of the virtual 3D videoconferencing environment to at least one device of at least one participant.

[0045] A plurality of participants may be associated with a plurality of participant devices, and the receiving and determining may be performed by at least a portion of the plurality of participant devices. Any of the steps of method 200 may be performed by at least a portion of the plurality of participant devices or by another computerized system.

[0046] Multiple participants may be associated with multiple participant devices, and the receiving and determining may be performed by a computerized system different from any of the multiple participant devices.

[0047] The method 200 may include one of the further additional steps collectively designated 290 .

[0048] The one or more additional steps may include at least one of the following: a. Determining the field of view of a third participant within a virtual 3D video conferencing environment. b. Setting a third updated representation of the virtual 3D video conferencing environment, which may be transmitted to a third participant device to reflect the third participant's field of view. Receiving initial 3D participant representation information for generating a 3D representation of the participant under different conditions, which may include at least one of: (a) different image acquisition conditions (different illumination and / or collection conditions), (b) different directions of gaze, and (c) different facial expressions. d. receiving, during run time, situation metadata and modifying, in real time, updated 3D participant representation information based on the situation metadata. e. for each participant, iteratively selecting a selected 3D model from the participant's multiple 3D models; f. Iteratively facilitating transitions from one selected 3D model of a participant to another 3D model of the participant. g. selecting an output of at least one neural network of the plurality of neural networks based on a required resolution. h. Receiving or generating participant appearance information regarding a participant's head pose and facial expressions. i. Determining updated 3D participant representation information to reflect participant appearance information. j. Determining the shape of each of the avatars representing the participants. k. Determining relevance of segments of the updated 3D participant representation information. l. Selecting which segments to send based on relevance and available resources. m. Generating a 3D model and one or more texture maps of the 3D participant representation information of the participant. n. Estimating 3D participant representation information of one or more occluded areas of a participant's face. o. Estimating 3D model occlusion areas and one or more occlusion texture maps. p. Determining the size of the avatar. q. Receive audio and appearance information regarding audio from Participants. r. Synchronization between audio information and 3D participant representation information. s. Estimating a participant's facial expression based on audio from the participant. t. Estimating participants' movements.

[0049] Receiving the 3D participant representation information may occur during an initialization step.

[0050] The initial 3D participant representation information may include an initial 3D model and one or more initial texture maps.

[0051] The 3D participant representation information may include a 3D model and one or more texture maps.

[0052] The 3D model may have separate parameters for shape, pose, and expression.

[0053] Each of the one or more texture maps may be selected and / or augmented based on at least one of shape, pose, and facial expression.

[0054] Each of the one or more texture maps may be selected and / or augmented based on at least one of the shape, pose, facial expression, and angular relationship between the participant's face and an optical axis of a camera capturing an image of the participant's face.

[0055] Determining, for each participant, the updated 3D participant representation information may include at least one of the following: a. using one or more neural networks to determine updated 3D participant representation information. b. using a plurality of neural networks for determining the updated 3D participant representation information, where different neural networks of the plurality of neural networks may be associated with different situations. c. using a plurality of neural networks for determining the updated 3D participant representation information, where different neural networks of the plurality of neural networks may be associated with different resolutions.

[0056] The updated representation of the virtual 3D video conferencing environment may include avatars for at least some of the multiple participants.

[0057] The gaze direction of an avatar in a virtual 3D video conferencing environment may represent a spatial relationship between (a) the gaze direction of a participant, which may be represented by the avatar, and (b) a representation of the virtual 3D video conferencing environment displayed to the participant.

[0058] The gaze direction of an avatar in a virtual 3D videoconferencing environment may be agnostic to the optical axis of the camera capturing the participant's head.

[0059] The avatars of the participants in the updated representation of the virtual 3D video conferencing environment may appear in the updated representation of the virtual 3D video conferencing environment as captured by a virtual camera located on a virtual plane across the first participant's eyes, and thus the virtual camera and the eyes may be located at the same height, for example.

[0060] The updated 3D participant representation information may be compressed.

[0061] The updated representation of the virtual 3D videoconferencing environment may be compressed.

[0062] The generation of the 3D model and one or more texture maps may be based on images of the participant acquired under different conditions.

[0063] The different situations may include different viewing directions of the camera capturing the images, different postures of the participants, and different facial expressions of the participants.

[0064] Estimation of 3D participant representation information for one or more occlusion areas may be performed using one or more generative adversarial networks.

[0065] Determining, for each participant, the updated 3D participant representation information may include at least one of the following: a. Applying super-resolution technology. b. Applying noise reduction. c. Changing the irradiation conditions. d. Add or change wearable item information. e. Adding or changing make-up information.

[0066] The updated 3D participant representation information may be encrypted.

[0067] The updated representation of the virtual 3D video conferencing environment may be encrypted.

[0068] The appearance information may relate to the participant's head pose, and may relate to facial expressions and / or lip movements of the participant.

[0069] Estimating a participant's facial expression based on audio from the participant may be performed by a neural network trained to map audio parameters to facial expression parameters.

[0070] FIG. 2 illustrates an example of a computer environment including user devices 4000(1)-4000(R) of users 4010(1)-4010(R). An index r ranges from 1 to R, where R is a positive integer. The rth user device 4000(r) may be any computerized device that may include one or more processing circuits 4001(r), memory 4002(r), a man-machine interface such as a display 4003(r), and one or more sensors such as a camera 4004(r). The rth user 4010(r) is associated with (uses) the rth user device 4000(r). The camera may belong to the man-machine interface.

[0071] The user devices 4000(1)-4000(R) and the remote computerized system 4100 may communicate over one or more networks, such as the network 4050. The one or more networks may be any type of network, such as the Internet, a wired network, a wireless network, a local area network, and a global network.

[0072] The remote computerized system may include one or more processing circuits 4101(1), memory 4101(2), and may include any other components.

[0073] Any one of the user devices 4000(1)-4000(R) and the remote computerized system 4100 may participate in the execution of any of the methods exemplified herein, where participating means performing at least one step of any of the aforementioned methods.

[0074] Any processing circuitry may be used, such as one or more network processors, non-neural network processors, rendering engines, and image processors.

[0075] The one or more neural networks may be located on a user device, on multiple user devices, and in a computerized system outside any of the user devices.

[0076] FIG. 3 illustrates an example of a computing environment including user devices 4000(1)-4000(R) of users 4010(1)-4010(R). An index r ranges from 1 to R, where R is a positive integer. The rth user device 4000(r) may be any computerized device that may include one or more processing circuits 4001(r), memory 4002(r), a man-machine interface such as a display 4003(r), and one or more sensors such as a camera 4004(r). The rth user 4010(r) is associated with (uses) the rth user device 4000(r).

[0077] The user devices 4000(1)-4000(R) may communicate through one or more networks, such as network 4050.

[0078] Any one of the user devices 4000(1)-4000(R) may participate in the execution of any of the methods exemplified herein, where participating means performing at least one step of any of the aforementioned methods.

[0079] 4 illustrates various example data structures, which may include user avatars 4101(1)-4101(j), texture maps 4102(1)-4102(k), 3D models 4103(1)-4103(m), 3D representations of objects 4104(1)-4104(n), and any mappings or other data structures referred to in this application.

[0080] Any user may be associated with one or more data structures of any type, such as avatars, 3D models, and texture maps.

[0081] Some examples refer to a virtual 3D videoconferencing environment, such as a meeting room, restaurant, cafe, concert, party, exterior environment, or imaginary environment, in which a user is set up. Each participant may choose or otherwise be associated with a virtual or real background, and / or may select or otherwise receive any virtual or real background in which an avatar associated with at least some of the participants is displayed. The virtual 3D videoconferencing environment may include one or more avatars representing one or more of the participants. The one or more avatars may be virtually located within the virtual 3D videoconferencing environment. One or more characteristics of the virtual 3D videoconferencing environment (which may or may not be associated with an avatar) may vary from one participant to another.

[0082] Either the user's entire body, a portion of the user's body, or only the user's face may be seen within the environment, and thus an avatar may include the participant's entire body, the upper portion of the participant's body, or only the participant's face.

[0083] Within the virtual 3D video conferencing environment, improved visual interaction between users may be provided that may emulate the visual interaction that exists between real users who are actually located near one another, which may include making or withholding eye contact and facial expressions that are directed toward a particular user.

[0084] In a video conference call between different users, each user may be provided with a view of one or more other users and the system may determine (based on gaze direction and the virtual environment) the user's looking position (e.g., looking at one of the other users, looking at none of the users, looking at a screen showing a presentation, looking at a whiteboard, etc.) and this is reflected by a virtual representation (3D model) of the user in the virtual environment so that other users may determine the user's looking position.

[0085] Figure 5 illustrates an example of the process of modifying the view direction of some of the participants' avatars according to the participants' gaze direction. The top part of Figure 5 is a virtual 3D videoconferencing environment represented by a panoramic view 41 of five participants 51, 52, 53, 54, and 55 sitting near a table 60. All participants face the same direction, the screen.

[0086] In the bottom image, the fifth participant's avatar faces the first participant's avatar as it is detected that the fifth participant is looking at a 3D model of the first participant in the environment as presented to the fifth participant.

[0087] Tracking the user's eyes and gaze direction can also be used to determine the direction the user is looking at, and the person or object the user is looking at. This information can be used to rotate the avatar's head and eyes so that in the virtual space it appears to be looking at the same person or object as the user is in the real world.

[0088] Tracking the user's head pose and eye gaze can also be used to control the appearance of the virtual world on the user's screen: for example, if the user is looking to the right of the screen, the virtual camera's viewpoint can be moved to the right, so that the person or object the user is looking at is located in the center of the user's screen.

[0089] Rendering the user's head, body, and hands from a viewpoint different from the original viewpoint of the camera can be done in different ways, as described below.

[0090] In one embodiment, a 3D model and texture map are created before the start of the meeting, and this model is then animated and rendered at run-time according to the user's pose and facial expression estimated from the video image.

[0091] A texture map is a 2D image in which each color pixel represents the red, green, and blue reflectance coefficients of an area in the 3D model. An example of a texture map is shown in Figure 20. Each color pixel in a texture map corresponds to a coordinate within a particular polygon (e.g., triangle) on the surface of the 3D model.

[0092] An example of a 3D model composed of triangles and the mapping of a texture map onto those triangles is shown in Figure 15.

[0093] Generally speaking, each pixel in a texture map has three coordinates that define the index of the triangle it maps to and its exact position within the triangle.

[0094] 3D models that are composed of a fixed number of triangles and vertices may be deformed as the 3D model changes. For example, a 3D model of a face may be deformed as the face transforms its expression. Nevertheless, pixels in the texture map correspond to the same position within the same triangle, even when the 3D position of the triangle changes as the facial expression changes.

[0095] A texture map can be constant or can vary with time, expression, or viewing angle: in either case, the correspondence between a given pixel in the texture map and a certain coordinate in a triangle in a 3D model does not change.

[0096] In yet another embodiment, a new view is created based on real-time images acquired from a video camera and a new viewpoint (virtual camera) position.

[0097] To achieve the best match between audio and lip and facial expressions, the audio and video created from rendering the 3D model based on pose and expression parameters are synchronized. Synchronization can be done by packaging the 3D model parameters and audio in one packet corresponding to the same time frame, or by adding a timestamp to each of the data sources.

[0098] To further improve the natural appearance of the rendered model, an audio neural network can be trained to estimate facial expression coefficients based on audio. This can be done by training the neural network using a database of videos of people speaking and the corresponding audio of this speech. The videos can be videos of the participant who is to be represented by the avatar, or videos of other people. Given sufficient examples, the network learns the correspondence between the audio (i.e., phonemes) and the corresponding facial movements, especially lip movements. Such a trained network enables the continued rendering of facial expressions, especially lip movements, even when the video quality is low or when parts of the face are obscured relative to the original video camera.

[0099] In yet another embodiment, a neural network can be trained to estimate audio sounds from lip and throat movements, as done by professional lip readers, or from any other facial cues. This enables creating or improving the quality of audio when the audio is interrupted or when there is background noise that reduces its quality.

[0100] In yet another embodiment, a neural network is trained to compress audio by finding latent vectors of parameters that can reconstruct the audio with high quality, such a network serving to compress the audio at a lower bitrate than is possible with standard audio compression methods for a given audio quality, or to obtain a higher audio quality for a given bitrate.

[0101] Such a network can be trained to compress an audio signal into a fixed number of coefficients influenced by speech that is as similar as possible to the original speech under some cost function.

[0102] The transformation of speech into a set of parameters can be a non-linear function, rather than simply a linear transformation as is common in standard speech compression algorithms. One example is that the network needs to learn and define a set of basis vectors that form the span of the spoken audio.

[0103] The parameters are then the vector coefficients of the audio as spanned by this set.

[0104] FIG. 6 illustrates the method 2001.

[0105] The method 2001 is for conducting a 3D video conference among multiple participants, and the method may include steps 2011 and 2021.

[0106] Step 2011 may include determining, for each participant, updated 3D participant representation information within the virtual 3D video conferencing environment that represents the participant, which may be based on audio generated by the participant and appearance information regarding the participant's appearance.

[0107] Step 2021 may include generating an updated representation of the virtual 3D videoconferencing environment for at least one participant, where the updated representation of the virtual 3D videoconferencing environment represents updated 3D participant representation information for at least a portion of the multiple participants. For example, any movement by a participant may reveal or bracket a portion of the environment. Additionally, movement by a participant may affect lighting in the room, such that the movement may modify the exposure to illuminate a different portion of the environment.

[0108] The method may include matching between audio from a participant and appearance information of a participant.

[0109] The appearance information may relate to the participant's head pose and facial expressions.

[0110] The appearance information may relate to the participant's lip movements.

[0111] Creating a 3D model The 3D models and texture maps of users can be created on-the-fly from 2D or 3D video cameras, or can be prepared before the start of a 3D video conference call. They can also be a combination of high-quality models prepared before the meeting and real-time models created during the meeting. Changes in the appearance of participants relative to the high-quality model, for example a newly grown beard, can be adjusted using information from the on-the-fly camera. As another example, a new texture map can be created from the video during the meeting based on what the person is currently seeing. However, this texture map may contain dead zones due to obstruction of areas that cannot currently be seen by the camera. Such dead zones can be filled by using previously created texture maps.

[0112] Filling in those zones is performed by matching landmarks in the two texture maps, using a method known as registration. Once matching is performed, data about the occlusion areas are taken from the previously prepared texture map.

[0113] Illumination corrections between the current and previous texture maps can be calculated based on the areas that can be represented in both maps. Those corrections can be applied to the current texture map so that there is no clear border line between textures captured at different times. In addition, to avoid sharp transitions between textures from different times, a continuous blending of the textures can be applied, for example by using a weighted average of the two texture maps, where the weights vary along the transition zone between the textures. The above mentioned methods can also be used to merge texture maps, material maps, and 3D models.

[0114] If the video camera is a 2D camera, computerized models such as convolutional neural networks can be used to create 3D models from the 2D images. These models can be parametric models whose parameters determine the shape, expression, and posture of the face, body, and hands. Such models can be trained using a set of 2D images and corresponding 3D models. The corresponding 3D models can be created in several ways. In the rendering process, different illuminations can be used to make the models robust to variable illumination.

[0115] Alternatively, many 2D images of a real person can be acquired and then a 3D model can be created from those multiple 2D images by using photogrammetry software. In yet another alternative, a depth camera, including even an RGB camera such as a Kinnect camera or an Intel RealSense camera, can be used to acquire both the 3D depth model and the corresponding 2D images. At run-time, after training the network using the method described above, it can be fed with 2D images as input and the network outputs a 3D model. The 3D model can be output as a point cloud, a mesh, or a set of parameters that describe the 3D model in a given parametric space.

[0116] If the camera is a 3D depth camera, the depth data can be used to make the model more accurate and resolve ambiguities. For example, if one only obtains a forward-facing image of a person's head, it may be impossible to know the exact depth of each point in the image, i.e., the length of the nose. When there is more than one image of the face from different angles, such ambiguities can be resolved. Nevertheless, occluded areas or inaccuracies seen in only one image may remain. The depth data from the depth camera can help generate a 3D model with depth information for each point, which solves the ambiguity problem described above.

[0117] If an offline 3D model creation process can be used, this can be done using a single image, multiple images, a video, or several videos. The user can be asked to rotate their head, hands, and body so that it can be seen from many angles to cover all views and avoid missing areas in the model.

[0118] If such areas still exist, they can be extrapolated or inferred from the modeled areas or by neural networks trained using many examples.

[0119] In particular, a Generative Adversarial Network (GAN) can be trained based on many images of a person, or on many images of multiple people, to generate images of a person from angles that may differ from the angle from which the camera may currently view the person.

[0120] At run-time, such a network receives an image of a person as input and the camera position from which the person should be rendered. The network renders images of the person from different camera positions, including parts that may be obscured in the input image due to being nearly parallel to the camera's line of sight or that may be at low resolution in the input image (i.e., the cheek in a frontal image).

[0121] Figure 7 shows an example of a process 100 that uses a generative adversarial network 109 to complete texture in areas that cannot be seen in the original image. With a GAN, there may be no need to build a complete and accurate 3D model with the entire texture map and render it.

[0122] An image 101 is input to a neural network 103, which outputs image characteristics 105 (which may include texture parameters, expression parameters, and / or shape parameters), for example the neural network may extend the texture parameters into a texture map. The neural network may also receive additional information 102 and generate characteristics 105 based on the additional information as well.

[0123] A differential renderer 107 may render an image from the texture map, facial expressions, and shape parameters. This image may have missing parts due to occlusion of parts of the head that were not seen in the original input image. A generative adversarial network 109 (GAN) may complete the rendered image into a full image 110 without any missing parts.

[0124] For example, in cases where a user's face may not be uniformly lit, such as when there is strong illumination from a window on the side of the face or from a spot projector above the user's head, a generative adversarial network (GAN) can also be used to correct the illumination in the model's texture map.

[0125] GAN networks can also be used to correct 3D models and create, for example, ears that may not be properly seen in an image due to obstruction by cheeks or by hair.

[0126] The user may also be asked to pose and perform different facial expressions so that a comprehensive model of postures and expressions can be created. Examples of such postures and expressions may be smiling, frowning, opening and closing the mouth and eyes.

[0127] A 3D model may have separate parameters for shape, pose, and expression. Shape parameters may depend only on a particular person and may be independent of pose and expression. Thus, they remain constant even when a person moves their head, speaks, or makes various facial expressions. Thus, during the modeling process of a person, the expression and pose of the person being modeled do not need to be static or frozen during the capture of the video or images that may be used to create the 3D model. Since the shape of the 3D model is considered static, there is no need to use a 3D camera or a collection of 2D cameras, which would otherwise be necessary to create the 3D model. This alleviates the requirement to use several multi-view cameras that may be synchronized in time. All models created from multiple images may be merged into one 3D model, or several different models that are variable due to expression or lighting conditions, but all of which may have common shape parameters.

[0128] During the real-time rendering process, the closest model or models in terms of viewing angle or illumination may be chosen as the starting point for the model transformation and rendering process.

[0129] For example, if different models are available that refer to viewing angles of 0, 10, 20, 30, and 40 degrees, and at a given moment the user wants to view the model at an angle of 32 degrees, the model corresponding to an angle of 30 degrees may be chosen as the starting point for model transformation.

[0130] Furthermore, some such models may be interpolated or extrapolated to obtain models in conditions that are not some of the pre-recorded conditions.

[0131] During the process of creating the 3D avatar, 3D model, and 2D texture map, the quality of the 3D model that may be created may be evaluated by projecting it onto a two-dimensional image from different angles using a simple linear geometric projection model or a more complex model of the camera that includes optical distortions. The projection of the 3D model onto the 2D image may be compared to the image captured by the camera or cameras. In doing so, it may be beneficial to model the camera that may be used to capture the image so that the geometric distortions of the camera may be modeled in the projection process. Modeling may include, but is not limited to, modeling the focal length of the camera, pixel size, total field of view, non-linear geometric distortions such as barrel distortion or pincushion distortion, or any other distortion of the optical system, especially for cameras with a wide field of view such as a fisheye camera.

[0132] Modeling may also include modeling blurring due to optical and color distortions. A projection of the 3D model may be compared to the captured 2D image to verify that the 3D geometry may be accurate and that the reflectance map may be accurate.

[0133] To compare the projected and captured images several methods can be used, for example: a. Comparing the positions of facial landmarks, such as the corners of the eyes and lips, the tip and edge of the nose, the edges of the cheeks and chin, that may be found in the image pairs. b.Comparing the positions of the silhouettes. c. Comparing the positions of the corners and lines detected in both images. d. Comparing the grey levels of two images.

[0134] Any differences that may be found may be used to update the 3D model and reflection map in a manner that reduces the differences between the projected and captured images. For example, if it may be found that the corners of the eyes may be located far to the left in the projection of the 3D model compared to their location in the captured 2D image, the model may be revised to shift the location of the corners of the eyes to the right in order to reduce the error between the location of the landmarks in the projection and the captured image.

[0135] This can be done by changing the position of the 3D points in the 3D mesh, or by changing parameters in a parametric model that affect the position of the landmarks.

[0136] This process can be used to reduce errors in the rendered and captured images, and therefore improve the quality of the models that can be created.

[0137] In particular, it may be beneficial to project the image at different angles, such as 0, 45, and 90 degrees, to capture any geometric or gray-level differences between the model and the captured image.

[0138] The quality of the 3D model and the texture maps may be analyzed during or after the process of the creation of the avatar and may be specifically checked to verify that all or some of the following cases may be covered: There can be no obscured areas on the face model, body model, or hand model. b. All relevant facial expressions can be covered. c. Both open and closed eyes can be modeled. d. Teething, closed mouth and open mouth can be covered. e. There are no areas with low resolution due to imaging of facial structures that may be nearly juxtaposed to the line of sight, for example imaging of the cheek from the front. f. The illumination can be adequate, there can be no areas that are too dark or too bright and saturated. g. There cannot be any areas that can be very noisy.

[0139] The model may not differ significantly from the user's current appearance in the video image, for example due to shaving or adding a beard, or changing hairstyle.

[0140] In cases where the inspection process discovers that there may be missing information, the user may be asked to add additional photos or video sequences to complete the missing information.

[0141] Prior to the initiation of a call between the users, but after the user's camera begins capturing the user's image, the 3D model and texture map may be enhanced to reflect the user's new appearance as seen at that moment.

[0142] Information from previously created models and texture maps may be merged with updated information obtained before the start of a meeting or during the meeting. For example, new information about the illumination of a person's body and face, the user's hair, shaving, make-up, clothing, etc. may be used to update the 3D model and texture map. Areas not previously seen, such as the top of the head or the bottom of the chin or other parts of the body that may be seen before or during the session, may also be used to update the 3D model or texture map.

[0143] The new information may be used to replace previous information, may be averaged with the previous information, or may be otherwise merged with the previous information.

[0144] To scale a 3D model, i.e. to know its exact dimensions from a 2D camera whose camera parameters may be unknown and whose range to the modeled object may be unknown, several methods can be used. For example: a. Using an object of known size that can be placed next to the object, for example, to place a credit card on a user's forehead. Such objects can include, but are not limited to, credit cards, driver's licenses, bills, coins, rulers, etc. In such cases, the classification method can classify the object used and determine its size from a database. For example, the method can detect a bill as originating from one of multiple countries and / or units, recognize it, and obtain its size from a database. Similarly, the method can detect a ruler and determine its size from a reading against the ruler. b. Asking the user to define their height. The face height may be approximately 13% of the height of an adult. This may be a sufficiently accurate approximation for many application requirements. In addition, children and babies may be known to have different body proportions. For babies, the face height may be known to be approximately 25% of their height. The face size may be a non-linear function of height, e.g., 25% of the height for a person who may be 60 centimeters tall, 20% of the height for a person who may be 100 centimeters tall, and 13% of the height for a person who may be 150 centimeters tall or more.

[0145] The 3D model of the user may include, but is not limited to, the following: a. Parametric models of the face and body, i.e. shape, expression, and pose. b. A high frequency depth map detailing such fine details as wrinkles, skin moles, etc. c. A reflectance map detailing the color of each part of the face or body. Multiple reflectance maps can be used to model the change in appearance from different angles. d. An optional material map detailing the materials each polygon may be made of, e.g. skin, hair, clothing, plastic, metal, etc. e. An optional semantic map that lists what part of the body each part in the 3D model or reflection map represents. f. The models and maps may be created before the meeting, during the meeting, or may be a combination or models created before and during the meeting.

[0146] A user's model may be stored on the user's computer, phone, or other device, and it may also be transmitted to the cloud or to other users, possibly in an encrypted manner to protect the user's privacy.

[0147] FIG. 6 also illustrates a method 90 for generating and using a parametric model.

[0148] The method 90 may include steps 92 , 94 , 96 , and 98 .

[0149] Step 92 may include generating, by the user device, a 3D model associated with the user, where the 3D model may be a parametric model.

[0150] Step 94 may include transmitting parameters of the 3D model to a computerized system.

[0151] Step 96 may include monitoring, by the participant's user device, each participant during the conference call, updating parameters of each participant's 3D model accordingly, and transmitting the updated parameters (the transmitting may be affected by communication parameters).

[0152] Step 98 may include receiving, by each participant's user device, updated parameters of the 3D models associated with the other participants and updating the display accordingly to reflect the changes to the models.

[0153] FIG. 6 also illustrates a method 1800 for generating a 3D visual representation of a detected object, which may be three-dimensional.

[0154] The method 1800 may include steps 1810, 1820, and 1830.

[0155] Step 1810 may include obtaining at least one 3D visual representation parameter, where the visual representation parameter may be selected from a size parameter, a resolution parameter, and a resource consumption parameter.

[0156] Step 1820 may include obtaining object information representative of the detected object and selecting a neural network for generating a visual representation of the detected object based on at least one parameter, for example, the information representative of the detected object may be a viewing angle of the object.

[0157] Steps 1810 and 1820 may be followed by step 1830 of generating a 3D visual representation of the 3D object by means of the selected neural network.

[0158] Step 1830 may include at least one of the following: a. Generating a 3D model of the 3D object and at least one 2D texture map of the 3D object. b. Further processing the 3D model and the 2D texture map during the rendering step to produce at least one rendered image.

[0159] The generating may be performed by a first computerized unit and the generating may be followed by transmitting the 3D model and the at least one 2D texture map to a second computerized unit, the second computerized unit configured to render the at least one rendered image based on the 3D model and the at least one 2D texture map.

[0160] The 3D object may be a participant in a 3D video conference.

[0161] The method may include outputting a 3D visual representation from a set of selected neural network outputs.

[0162] The 3D object may be a participant in a 3D video conference.

[0163] Performing super-resolution and refinement on 3D models Super-resolution techniques can be used to increase the resolution of a 3D model. Super-resolution techniques can be used to increase the resolution of a 3D model or a deformable texture map of a 3D model. For example, several images of the model with some translation or rotation between them can then be used to create a grid at a higher resolution than can be created from a single image. It will be noted that the color values ​​of the model can relate to polygons in the 3D mesh or pixels in the 2D texture map.

[0164] This process can be done using a recursive process: in the first stage, 3D models and texture maps that are upsampled interpolations of the low-resolution models and texture maps are used as an initial guess. These 3D models and texture maps have more vertices and pixels than are in the original 3D models and texture maps, but do not contain any additional detail. The upsampled models and texture maps are then used to render an image of the textured model from a perspective similar to that of the camera.

[0165] The rendered image is compared to a 2D image captured by a camera.

[0166] The comparison can be performed, but is not limited to, by subtraction of the two images, or by global alignment of the images followed by subtraction, or by subtraction of local aligned areas within the images. The result of the comparison, a difference image obtained by this process, contains details from the original camera images that are not present in the rendered image. The difference can be used as feedback to increase the resolution of the initial 3D model and texture maps.

[0167] Enhancing may be done, without limitation, by adding a difference image to the initial inference to obtain a new inference with more detail. The new 3D model and texture map may be rendered again to obtain a second rendered image, which is compared to the original camera image to create a second difference image, which may be used as feedback to enhance the resolution of the 3D model and texture map.

[0168] This process may be repeated a given number of times or until a certain criterion is met, e.g., the difference between the actual camera image and the rendered image is below a certain threshold. The process is repeated when a comparison of the rendered textured 3D model with several camera images from a collection of images, such as from a video sequence, is performed. For each image, the 3D model and texture map may be sampled by the camera at different positions, since there may be many images in the image set or video.

[0169] Thus, the process can create 3D models and texture maps that are effectively based on a higher sampling rate than is available from a single image. As a result of this process, a 3D model with more vertices and a texture map with more pixels are created that show high resolution details that are not present in the original low resolution 3D model and texture map.

[0170] Multiple images of the face and body may also be acquired from the same or different angles by averaging the images together, thus improving the signal-to-noise ratio, i.e., creating a model with a lower level of pixel noise. This may be particularly beneficial when images are acquired in low illumination conditions, where the resulting images may be noisy.

[0171] Super-resolution techniques based on learning methods may also be applied. In such schemes, machine learning methods such as convolutional neural networks may be trained on pairs of high-resolution and low-resolution images or 3D models, so that correspondences between the low-resolution and high-resolution images or models may be learned. During the rendering process, the method receives as input a low-resolution image or model and outputs a corresponding high-resolution image or model. Those types of methods may be particularly useful for generating sharp edges at transitions between different facial tissues, such as sharp edges along the eyes or eyebrows.

[0172] The transition from low to high resolution can be performed based on a single image or on multiple images, and it can be performed in the process of creating a 3D model, a texture map, or when rendering the final image that can be presented to the user.

[0173] Reducing random noise in the 3D model and 2D texture map may also be performed using denoising methods. Such methods may include linear filtering techniques, but preferably include non-linear edge-preserving techniques, such as bilateral filters, anisotropic diffusion, or convolutional neural networks, that reduce random noise while preserving edges and fine details in the image of the 3D model.

[0174] The user's appearance can be altered and improved by manipulating the resulting 3D model or reflection map: for example, different kinds of touch-ups can be applied, such as removing skin wrinkles, applying make-up, stretching the face, lip filling, or changing eye color.

[0175] The user's body shape may also be modified, and the user's clothes may be changed from the real clothes to other clothes according to the user's desire. Accessories such as earrings, glasses, hats, etc. may also be added to the user's model.

[0176] Alternatively, objects such as glasses or headphones may be removed from the user's model.

[0177] Communication systems based on 3D models During a communication session, i.e., a 3D video conference between several users, a 2D or 3D camera (or several cameras) captures videos of the users, from which a 3D model (e.g., a best-fit 3D model) of the user can be created at a high frequency, e.g., 15-120 frames per second.

[0178] Temporal filters or temporal constraints within the neural network can be used to ensure smooth transitions between model parameters corresponding to video frames in order to produce a smooth temporal reconstruction and avoid artificialities in the results.

[0179] Real-time parametric models along with reflectance maps and other maps can be used to render visual representations of faces and bodies that can be very close to the original images of the faces and bodies in the video.

[0180] Since this can be a parametric model, it can be represented by a small number of parameters. Typically, fewer than 300 parameters can be used to create a high-quality model of a face, including the shape, expression, and pose of each person.

[0181] Those parameters can be further compressed using quantization and entropy coding, such as Huffman or arithmetic coders.

[0182] The parameters may be ordered according to their importance, and the number of parameters and the number of bits per parameter that may be transmitted may be variable according to the available bandwidth.

[0183] In addition, instead of coding the values ​​of the parameters, the differences in their values ​​between successive video frames may be coded.

[0184] The parameters of the model can be transmitted directly to all other user devices or to a central server. This can save a lot of bandwidth as instead of transmitting the entire model of the actual high quality image during the entire conference call, much fewer bits representing the parameters can be transmitted. This can ensure high quality video conference calls even when the current available bandwidth is low.

[0185] Transmitting model parameters directly to other users instead of through a central server may reduce latency by approximately 50%.

[0186] Other user devices may reconstruct the appearance of other users from the 3D model parameters and corresponding reflection maps. Because reflection maps representing such things as a person's skin color change very slowly, they may be transmitted only once at the beginning of a session or at a low update frequency according to the changes that occur in their reflection maps.

[0187] Additionally, the reflection map and other maps may only be partially updated, for example according to changed areas or according to semantic maps representing body parts: for example, the face may be updated, but the hair or body, which may be less important for reconstructing the emotion, may not be updated or may be updated less frequently.

[0188] In some cases, the bandwidth available for transmission may be limited. Under such conditions, it may be beneficial to order the parameters for transmission according to some priority, and then transmit the parameters in this order as the available bandwidth allows. This ordering may be done according to their contribution to the visual perception of the realistic video. For example, parameters related to the eyes and lips may have a higher perceived importance than those related to the cheeks or hair. This approach allows for a high degree of degradation of the reconstructed video.

[0189] The model parameters, the unmodeled video pixels, and the audio can all be synchronized.

[0190] As a result, the total bandwidth consumed by the transmission of 3D model parameters can be a few hundred bits per second, much smaller than the 100 kbits per second to 3 Mbits per second that get typically used for video compression.

[0191] Parametric models of the user's speech may also be used to compress the user's speech beyond what may be possible with generic speech compression methods. This further reduces the necessary bandwidth required for video and audio conferencing. For example, a neural network may be used to compress the speech into a limited set of parameters from which the speech can be reconstructed. The neural network is trained such that the resulting decompressed speech is closest to the original speech under a particular cost function. The neural network may be a nonlinear function, unlike the linear transformations used in generic speech compression algorithms.

[0192] Transmission of bits for reconstructing video and audio at the receiving end may be prioritized so that the most important bits may be transmitted or received with a higher quality of service. This may include, but is not limited to, prioritizing audio over video, prioritizing model parameters over texture maps, and prioritizing certain areas of the body or face over others, such as prioritizing information related to the user's lips and eyes.

[0193] The optimization method may determine the allocation of bitrate or quality of service to audio, 3D model parameters, texture maps, or pixels or coefficients that may not be part of the model, to ensure an overall optimal experience. For example, as the bitrate decreases, the optimization algorithm may determine to reduce the resolution of the 3D model or update the frequency of the 3D model to ensure a minimum quality of the audio signal.

[0194] 3D Model Encryption and Security A user's 3D models and corresponding texture maps may be stored on the user's device, on a server in the cloud, or on other users' devices. The models and texture maps may be encrypted to secure the users' personal data. Before a call between several users, a user's device may request access to the other user's 3D models and texture maps so that the device can render the other user's models based on their 3D geometry.

[0195] This process may involve the exchange of encryption keys at a high frequency, e.g., every second, so that after the call is ended, a user is not able to access the other user's 3D models and texture maps or any other personal data.

[0196] The User is able to determine which other Users may have had access to the User's 3D models and texture maps or any other personal data.

[0197] Additionally, a user may be able to delete personal data that may be stored on the user's device, a remote computer, or another user's device.

[0198] A 3D model and texture map of the user, which may be stored on the user's device or on a central computer, may be used to authenticate that the person in front of a 2D or 3D camera may actually be the user, which may eliminate the need to log into a system or service with a password.

[0199] Another security measure may involve protecting access to and use of one or more avatars (e.g., display of the avatar during the 3D video conference) of one or more participants, which can be done by applying digital rights management methods that enable access to and use of the avatar (or avatars), or by using any other authentication method access control to access and / or use of the avatar. Authentication may be done multiple times during the 3D video conference. Authentication may be based on biometrics, may require a password, may include face identification methods based on either 2D images, 2D video (with motion), or based on 3D features.

[0200] Parallax correction, eye contact generation and 3D effects based on 3D models The corrections mentioned below may correct for any deviation between the actual optical axis of the camera and the desired optical axis of the virtual camera. Some of the examples refer to the height of the virtual camera, and any of the following may also refer to the lateral position of the camera, e.g., positioning of the virtual camera at the center of the display (both height and lateral position, positioning of the virtual camera to have a virtual optical axis aimed at the participant's eyes (e.g., via a virtual optical axis that may be perpendicular to the display or may have any other spatial relationship to the display).

[0201] Given that a user can be imaged by a user's camera, other user devices can reconstruct a 3D model of that user from an angle different from the angle at which the original video (of the user) was captured by the camera.

[0202] For example, in many videoconferencing situations, a video camera may be placed above or below the eye level of a user. When a first user looks at the eyes of a second user as they are presented on the first user's screen, the first user is not looking directly into the camera. Thus, the image as captured by the camera and presented to the other user shows the first user's eyes as gazing downward or upward (depending on the camera's position and optical axis).

[0203] By rendering the 3D model from a point directly in front of the user's gaze, the resulting image of the user can be seen as looking directly into the eyes of the other user.

[0204] 8 illustrates an example of parallax correction. Image 21' may be an image captured by camera 162, which is located above display 161 and has a real optical axis 163 (pointed downwards) and a real field of view 163 which may be directed towards fifth participant 55.

[0205] The corrected image 22' may be virtually captured by a virtual camera 162' having a virtual optical axis 163' and a virtual field of view 163', which may be positioned at a point on the screen at eye level and directly in front of the fifth participant 155.

[0206] The face position tracker may track the position of the viewer's face and may change the viewpoint of the rendering accordingly, for example, if the viewer moves to the right, the viewer may see more of the left side of the opposite figure, and if the viewer moves to the left, the viewer may see more of the right side of the opposite figure.

[0207] This creates a 3D sensation of seeing three-dimensional characters or objects even while using a 2D screen.

[0208] 9 illustrates an example of a 3D illusion produced by a 2D device. The image acquired by the camera (and the FOV of the tracker) is denoted as 35, and the various virtual images are denoted as 31, 32, and 33.

[0209] This can be obtained by modifying the rendered image according to the viewer's movements and the viewer's eyes, thus creating a 3D effect. To do this, an image of the viewer is acquired by a camera, such as a webcam.

[0210] A face detection algorithm detects and tracks the face in the image. In addition, the viewer's eyes are detected and tracked within the face. As the viewer's face moves, the algorithm detects the position of the eyes and calculates their position within the 3D world. The 3D environment is rendered from a virtual camera according to the viewer's eye position.

[0211] When rendered images are presented on a 2D screen, only one image is rendered: this image of the 3D environment may be rendered from the perspective of a camera positioned between the viewer's eyes.

[0212] If a viewer uses a 3D display, such as a 3D display or virtual reality (VR) headset or glasses, two images corresponding to the right and left eye perspectives are generated to produce a stereoscopic image.

[0213] FIG. 10 illustrates an example of two stereoscopic images (designated 38 and 39) presented on a 3D screen or VR headset.

[0214] Some displays, such as autostereoscopic displays, do not require glasses to present 3D images. In such 3D displays, different images can be projected at different angles, for example using a lenticular array, so that each eye sees a different image. Some autostereoscopic displays, such as the Alioscopy Glasses-Free 3D Display, project more than two images at different angles, up to eight different images, in the case of some Alioscopy displays. When using such displays, more than two images can be rendered to create a 3D effect on the screen. This represents a significant improvement over conventional 2D videoconferencing systems in creating a more realistic and up-close sensation.

[0215] To enhance the 3D sensation, 3D audio can also be used. For each user, the user's position in the virtual 3D setting relative to all other users can be known. A stereo signal of each user's speech can be generated from the mono audio signal by introducing a delay between the audio signals for the right and left ear according to the relative positions of the audio sources. In such a way, each user gets the sensation of the direction from which the sound comes and therefore of who is speaking.

[0216] Additionally, images of the user's face, and in particular their lips, can be used to perform lip reading.

[0217] Analysis of successive images of lips can detect lip movements. Such movements can be analyzed, for example, by a neural network trained to detect when lip movements are associated with speaking. As input to the training phase, it is possible to tag the input video sequence with a sound analyzer or a human with human sounds. If a person is not speaking, the system can auto-mute the user, thus reducing background noise that may come from the user's environment.

[0218] Lip reading can also be used to know what sounds can be predicted to be produced by the user. This can be used to filter out external noises that do not correlate with those predicted sounds, i.e., are not in the predicted frequency range, and can be used to filter out background noise when the user is speaking.

[0219] Lip reading may also be used to assist in transcription of conversations that may be performed on the system in addition to speech recognition methods that may be based solely on audio.

[0220] This can be done, for example, by a neural network. The network is trained using the person speaking and the associated text that was spoken during the sequence. The neural network can be a recurrent neural network with or without LSTM or any other type of neural network. The method, which can be based on both audio and video, can result in improved speech recognition performance.

[0221] The face, body and hands can be modeled using a limited number of parameters, as explained above.

[0222] However, in real-world video conferencing, not every pixel in the image corresponds to a face, body, and hand model: objects that cannot be part of the body may appear in the image.

[0223] As an illustration, a person speaking in a conference may be holding an object that may or may not be significant to a particular conference call. A speaker may be holding a pen that has no significance to the meeting or a diagram that is very significant to the meeting. To transmit those objects to other viewers, they can be recognized and modeled as 3D objects. The model can be transmitted to other users for reconstruction.

[0224] Some portions of a video image may not be modeled as 3D objects and may be transmitted to other users as pixel values, DCT coefficients, wavelet coefficients, wavelet zero trees, or any other efficient way to transmit those values. Examples include flat objects placed in the background, such as a whiteboard or pictures on a wall.

[0225] The video image and the model of the user may be compared, for example but not limited to, subtracting the rendered image and the video image of the model. This is done by rendering the model as if it were taken from the exact position of the real camera. With a perfect model and rendering, the rendered image and the video image should match. The difference image may be segmented into areas where the model estimates the video image accurately enough and areas where the model may not be accurate enough or does not exist. Any areas that cannot be modeled accurately enough may be transmitted separately as described above.

[0226] In some circumstances, the system may determine that some objects viewed cannot be modeled as mentioned above. In those cases, the system may decide to transmit to the viewer a video stream that includes at least some of the unmodeled parts, and then the existing 3D models are rendered on top of the transmitted video at their respective positions.

[0227] A user may be provided with one or more views of the virtual 3D videoconferencing environment, while the user may or may not select a field of view, e.g., a field of view that includes all of the other users or only one or some of the users, and / or may select or view an object in one or some of the virtual 3D videoconferencing environment, such as a TV screen, a whiteboard, etc.

[0228] When combining video pixels and a rendered 3D model, the areas corresponding to the model, the areas corresponding to the video pixels, or both may be processed so that the combination appears natural and the seams between the different areas are not visible, which may include, but is not limited to, relighting, blurring, sharpening, denoising, or adding noise to one or some of the image components so that the entire image appears to originate from one source.

[0229] Each user may use a curved screen or a combination of physical screens with the effect that the user can view a panoramic image showing a 180 degree or 360 degree view (or any other angle range view) of the virtual 3D video conferencing environment, and / or a narrow field of view image that focuses on a few people, one person, or only a portion of a person, i.e., a person's face, the screen, or a whiteboard or one or more parts of the virtual 3D video conferencing environment.

[0230] A user can control the narrow field of view image or portion or portions of the narrow field of view image(s) by using a mouse, keyboard, touchpad or joystick, or any other device that allows panning and zooming within or from an image.

[0231] A user may be able to focus on a certain area of ​​the virtual 3D videoconferencing environment (eg, a panoramic image of the virtual 3D videoconferencing environment) by clicking on an appropriate portion within the panoramic image.

[0232] Figure 11 illustrates an example of a panoramic view 41 of a virtual 3D video conferencing environment populated by five participants and a partial view 42 of some of the participants within the virtual 3D video conferencing environment. Figure 11 also illustrates a hybrid view 43 that includes a panoramic view (or partial view) and an enlarged image of some of the participants' faces.

[0233] The user may be able to pan or zoom using head, eye, hand or body gestures: for example, by looking at the right or left part of the screen the focal area may move left or right so that it appears in the center of the screen, and by leaning forward or backward the focal area may zoom in or out.

[0234] A 3D model of the person's body may also assist in accurately segmenting the body and the background. In addition to the body model, the segmentation method learns which objects may be connected to the body, for example, the person may be holding a phone, a pen, or a piece of paper in front of the camera. Those objects are segmented together with the person and added to the image in the virtual environment, either by using the model of the object or by transmitting an image of the object based on a pixel-level representation. This is in contrast to existing virtual background methods that may be employed in existing video conferencing solutions that may not show the objects held by the user, so that those objects are not segmented together with the person, but rather are segmented as part of the background that needs to be replaced by the virtual background.

[0235] Segmentation methods typically use some metric that needs to be exceeded in order for pixels to be considered as belonging to the same segment. However, segmentation methods may also use other approaches, such as Fuzzy logic, where the segmentation method only outputs the probability that the pixels belong to the same segment. When the method detects an area of ​​pixels with a probability that makes certain whether the area should be segmented as part of the foreground or background, the user can be queried how to segment this area.

[0236] As part of the segmentation process, objects such as earphones, cables connected to earphones, microphones, 3D glasses, or VR headsets may be detected by the method. Those objects may be removed in the modeling and rendering processes, so that the image viewed by the viewer does not include those objects. Options to show or remove such objects may be selected by the user, or may be determined in any other manner, e.g., based on selections made previously by the user and by other users.

[0237] If the method detects more than one person in the image, it may query the user whether to include the person or people in the foreground and within the virtual 3D video conferencing environment, or whether to segment them out of the image and outside the virtual 3D video conferencing environment.

[0238] In addition to using the shape or geometric features of objects to determine whether they may be part of the foreground or background, the method may also be aided by knowledge of the temporal changes in brightness and color of those objects. Objects that do not move or change have a higher probability of being part of the background, e.g., part of the room in which the user is sitting, and areas where movement or temporal changes may be detected may be considered to have a higher probability of belonging to the foreground. For example, a standing lamp may not be seen to be moving at all, and it is considered to be part of the background. A dog walking around the room is moving and is considered to be part of the foreground. In some cases, periodic repeating changes or movements may be detected, e.g., a fan spinning, and those areas may be considered to have a higher probability of belonging to the background.

[0239] The system learns user preferences and uses feedback regarding which objects, textures, or pixels may be part of the foreground and which may be part of the background, and uses this knowledge to improve subsequent segmentation steps. A learning method, such as a convolutional neural network or other machine learning method, may learn which objects may typically be selected by a user as part of the foreground and which objects may typically be selected by a user as part of the background, and may use this knowledge to improve the segmentation method.

[0240] Automatic exposure control for digital still and video cameras Segmentation of the user's face and body from the background can assist in setting the exposure time of the user's camera so that the exposure can be optimal for the user's face and body and can be influenced by light or dark areas in the background.

[0241] In particular, the exposure may be set according to the brightness of the user's face, so that the face may not be too dark, nor too bright, but may be saturated.

[0242] In determining the correct exposure for a face that may be detected, there may be a challenge in knowing the actual brightness of the person's skin. It may be preferable not to overexpose the skin of people with naturally dark skin (see image 111 in FIG. 12) and turn them into a light face in an overexposed image, see image 112 in FIG. 12.

[0243] In order not to overexpose images of people with dark skin, the auto-exposure method may set the exposure according to the brightness level of the white of the user's eyes or teeth. The camera exposure may be varied slowly, using some temporal filtering, and not change rapidly from frame to frame. Such a method ensures that the resulting video may not have jitter. Furthermore, such a method may allow setting the exposure based on the brightness level of the eyes or teeth, even when the eyes or teeth appear in some frames and not in some others.

[0244] The detection of the face, eyes or teeth may be based on 3D models and texture maps, on methods to detect those parts of the body, or on tracking methods. Such methods may include algorithms such as the Viola Jones algorithm, or neural networks trained to detect faces and specific facial parts. Alternatively, fitting the 2D image to a 3D model of the face may be performed, where the positions of all facial parts in the 3D model are known a priori.

[0245] In another embodiment, the exact darkness of the skin can be estimated in a Hue, Saturation, and Brightness color coordinate system. In such a coordinate system, the Hue and Saturation do not change with exposure, only the Brightness coordinates change. It has been discovered that a correspondence can be found between the Hue and Saturation values ​​of people at the appropriate exposure and respective brightness values ​​of their skin. For example, a pinkish skin tone corresponds to a fair face and a brownish tone corresponds to a dark skin, see for example images 121-126 in FIG. 12.

[0246] In yet another embodiment, a neural network, such as a convolutional neural network or any other network, can be trained to identify correspondences between the shape of the face and other attributes and the skin brightness. Then, at run-time, the face at various exposures can be analyzed independent of the chosen exposure, and the detected attributes can be used to estimate the exact brightness of the skin, which can be used to determine the camera exposure that results in such skin brightness.

[0247] A neural network can be trained to discover this relationship function or correlation between the hue and saturation of skin and the respective brightness in a properly exposed image. In the inference stage, the neural network suggests a proper exposure for a picture based on the hue and saturation of skin in an image that is not necessarily optimally exposed, e.g., too bright or too dark. This calculated exposure can be used to capture a properly exposed image that is neither too dark nor too bright.

[0248] In yet another embodiment, a user of a photographic device, such as a cell phone, professional camera, or webcam, may be asked once to take a photo of themselves or another person with a white paper or other calibration object for reference. This calibration process may be used to determine the exact tone, saturation, and brightness of the person's skin. Then, at run time, the computing device may run a method that recognizes a given person and adjusts the exposure and white balance so that the person's skin corresponds to the exact skin color as discovered in the initial calibration process.

[0249] Running computations on the cloud The processing of this system may be performed on the user's device, such as a computer, phone, or tablet, or on a remote computer, such as a server on the cloud. Computations may also be split and / or shared between the user's device and a remote computer, or they may be performed on the user's device for users with appropriate hardware, and on the cloud for other users (or in any other computing environment).

[0250] The estimation of body and head parameters can be based on compressed or uncompressed images. In particular, they can be performed on compressed video on a remote computer, such as a central computer on the cloud or another user's device. This allows a standard video conferencing system to transmit the compressed video to the cloud or to another user's computer, where modeling, rendering, and processing are performed.

[0251] Use of multiple screens and channels to present information in a videoconferencing application and method to increase meeting efficiency - Patents.com A virtual meeting may appear to take place in any virtual environment, such as a room, in any other closed environment, or in any open environment. Such an environment may include one or more screens, whiteboards, or flip charts for presenting information. Such screens may appear and disappear, be moved, enlarged, and reduced in size according to the desires of the user.

[0252] Multiple participants may share their screens (or any other content) on more than one screen, meaning that multiple sources of information can be viewed simultaneously.

[0253] Materials to be shared or presented may be pre-loaded onto such a screen or repository before the meeting begins for easy access during the meeting.

[0254] One possible way to present different materials is by transmitting them through dedicated streams, one for each material to be presented. In this setting, streams can be allocated to viewers based on many criteria. For example, streams can be specifically allocated to one or more viewers. Alternatively, streams can be allocated according to topic or other considerations. In such a case, the stream to be viewed can be selected by each viewer. This can be done quickly by using a keyboard, mouse, or any other device. Such a selection can be much faster than the common practice of sharing a piece of content, which currently may require requesting permission to share a screen from a manager of the meeting, receiving such permission, clicking a "screen share" button, and selecting the relevant window to share.

[0255] Such a "screen sharing" process can take (for example) up to several minutes. In various applications, the "screen sharing" can be repeated many times with many different participants presenting their material, and a lot of valuable time can be lost. The proposed solution can reduce the duration to a few seconds.

[0256] In some cases, not all of the participants in a meeting or screen, or all of the other objects of interest in the 3D virtual environment, may appear on the viewer's screen at the same time. For example, this may occur if the screen's field of view is smaller than necessary to view all participants. In such cases, it may be necessary to move the viewing user's field of view to the right, left, forward, backward, up, or down to change the viewpoint and see a different participant or object. This can be achieved by different means, including but not limited to: a. Using the keyboard arrows or other keys to pan and tilt the viewpoint, or zoom in and out. b. Using the mouse or other keys to pan and tilt the viewpoint, or zoom in and out. c. Using a method that tracks the user's head position or eye gaze direction, or both, to pan and tilt the viewpoint, or zoom in and out. The input to the method can be video of the user from a webcam, or any other 2D or 3D camera, or any other sensor, such as an eye gaze sensor. d. Using a method that tracks the user's hands to pan and tilt the viewpoint, or zoom in and out. The input to the method can be video from a webcam, or any other 2D or 3D camera, or any other sensor such as an eye-gaze sensor. e. Determining who may be a speaker at any moment and panning, tilting, and zooming in on that speaker at any given moment. If several people may be speaking at the same time, the method can determine who may be the dominant speaker and pan and tilt to that speaker or zoom out to a wide field where several speakers may be shown.

[0257] The computations required to create an avatar within the virtual 3D videoconferencing environment may be performed on the user's computing device, in the cloud, or in any combination of the two. In particular, performing the computations on the user's computing device may be preferable to ensure faster response times and lower latency due to communication with a remote server.

[0258] Two or more 2D or 3D cameras can be placed at different positions around the user's screen, for example integrated into a border or corner of the user's screen, so that there can be simultaneous views of the user from different directions in real time. The 2D or 3D views from different directions can be used to create a 3D textured model that corresponds to the user's appearance in real time.

[0259] If the camera is a 3D camera, the 3D depth acquired by the camera can be merged into a 3D model that is more complete than a model acquired from only one camera, as two or more cameras capture additional areas to what can be captured by only one camera.

[0260] Because the cameras are located in different positions, they obtain slightly different information about the user, and each camera may be able to capture areas that are occluded and not seen by the other cameras. If the cameras are 2D cameras, different methods may be used to estimate a 3D model of the user's face. For example, photogrammetric methods may be used to accomplish this task. Alternatively, neural networks may be used to estimate the 3D model that produces the image as captured by the cameras.

[0261] Color images as captured by the cameras can be used to create a complex texture map. This map then covers more area than can be captured by only one camera. Multiple texture maps as obtained from each camera can be stitched together, averaging overlapping areas to create one more comprehensive texture map. This can also be performed by neural networks.

[0262] This real-time 3D textured model can then be used to render the user's view from various angles and camera positions, and in particular can be used to correct the viewing position of the virtual camera so that the user's screen is virtually located at a virtual position where, for example, the height and / or lateral coordinates are positioned at the position where the participant's eyes are.

[0263] The virtual position may be located in an imaginary plane that virtually intersects the participant's eyes, perpendicular or substantially perpendicular to the display. In this way, a sensation of eye contact may be created for the real-time video of the user. The real-time 3D textured model may also be re-lit differently from the lighting of a real person in a real environment to create a more comfortable illumination, e.g., illumination with less shadows.

[0264] A speech to speech or text to speech method or neural network can be applied to the audio stream to summarize the content of the conversation taking place in the virtual meeting. For example, a neural network can be trained on the full text and their respective summaries. Similarly, a neural network can be trained to generate a list of action items and assignees.

[0265] To expedite the process and assist the neural network in reaching a decision, a human can represent the relevant portion of text on a summary of the task list. This can be done in real time in close proximity to when the relevant text is spoken. The summary and list of action items can be distributed to all meeting attendees or to any other list of recipients. This can be used to enhance meetings and increase their productivity.

[0266] Digital assistants may also help control the application, for example by assisting the recipient with an invitation, presenting information on a screen, or controlling other settings of the application.

[0267] Digital assistants can be used to transcribe meetings in real time and present the transcription on the user's screen. This can be very beneficial when the audio received at the remote participants may be degraded due to echoes or accents that can be difficult to understand, or due to problems with the communication network such as low bandwidth or packet loss.

[0268] Digital assistants can be used to translate speech from one language to another in real time and present the translation on the user's screen. This can be very beneficial when participants speak different utterances. Furthermore, Text To Speech (TTS) engines can be used to generate audio representations of the translated utterances. Neural networks such as generative adversarial networks or recurrent neural networks can be used to generate natural-sounding speech rather than robotic speech. Such networks can also be trained and then used to generate translated speech with the same intonation as is in the original utterance in the original language.

[0269] Another neural network, such as a convolutional neural network, can be used to animate the face and lips of the 3D model to move according to the generated translated speech. Alternatively, a GAN or other network can be used to generate a sequence of 2D images of the face and lips that move according to the generated translated speech. For this, a neural network can be trained to learn the lip movements and facial distortions as they relate to the speech. By combining all the steps described above, a sequence of images and corresponding audio of a person speaking in one language can be translated into a sequence of images and corresponding audio of a person speaking in another language, where the audio sounds natural and the image sequence corresponds to the new audio, i.e., the lip movements can be synchronized with the phonemes of the speech.

[0270] Such a system may be used as described above, but is not limited to video conferencing applications, television interviews, automatic dubbing of movies or e-learning applications.

[0271] A method for accurate 3D tracking of faces via monocular RGB video To track a user's facial pose and expressions, a method for accurate 3D tracking of the face (without depth) via monocular RGB video input would be beneficial. The method needs to detect various facial expressions, e.g., smiling, frowning, and neck pose changes, along with the 3D movement of the face in the video relative to the camera.

[0272] Typically, monocular video face tracking can be done using a sparse set of landmarks (dlib landmarks, HR-net facial landmarks, and Google's Media Pipe landmarks).

[0273] These landmarks can typically be generated using a sparse collection of user-annotated images, or synthetically using parametric 3D models.

[0274] The limitations of these conventional methods are: Absence of landmarks in certain areas (ears, neck). b.Clarity of landmarks. c. Landmark accuracy and stability. D.Temporal coherence. e.Mapping landmarks onto 3D models.

[0275] The input to the proposed method can be a 2D monocular video, an approximation of the tracked parameters of the first frame of the video (specific parameters), approximated deformation parameters (of the person) in the video and an approximate camera model, along with a templated 3D model of the face (global) with a deformation model (person-specific or global) for this 3D template.

[0276] A 3D face template mesh (templated 3D model) may include a coarse triangular mesh of a typical human face. By coarse, I mean with a dimension of 5K or 10K polygons, which may be enough to represent the overall shape but not wrinkles, microstructure, or other fine details.

[0277] 3D facial deformation models for the template may include standard parametric methods that deform the template, changing the overall shape of the 3D mesh (jaw structure, nose length, etc.), facial expressions (smile, frown, etc.), or its exact position and orientation based on the positions and cues found in the image. Users of the method can choose to use statistical 3D meshes as deformation models, such as Basel Face Model / Facewarehouse / Flame models, and / or use prior deformations, such as As-Rigid-As-Possible, elastic or isometric targets.

[0278] Approximate deformation parameters and an approximate camera model for the person in the video can be found by standard 3D MMM fitting techniques, for example by using face landmark detection methods that detect known face part parameters and optimize the camera and pre-annotated landmarks in a least squares sense. The initialization does not need to be exact, but only needs to be approximated, and can generally be generated via known techniques.

[0279] The output of this method can be the geometry (deformation parameters and mesh model) for each frame, as well as a set of approximate camera parameters for each image.

[0280] For each frame, the deformed mesh is called the current 3D face mesh, and its deformation parameters on top of the template can be chosen based on a set of landmarks deduced from the 2D face part segmentation and the pre-annotated segmentation. To that end, the proposed method can use a 2D face part segmentation method together with a classical 2D rigid registration technique that utilizes the ICP (Iterative Closest Point) method to track and deform a 3D face model based on the input RGB monocular video.

[0281] The proposed method builds a common face part facial part segmentation network that annotates each pixel with a given face part.

[0282] Fig. 13 illustrates face segmentation. The input image 131 is a color image acquired by a camera. Image 132 illustrates the segmentation of different face parts, visualized by different colors.

[0283] In addition, the triangular mesh template can be pre-annotated with predefined annotations of facial features (e.g., nose, eyes, ears, neck, etc.). Mesh annotation can assist in finding correspondences between various facial features on the 3D model and facial features on a given target image. Annotation of facial features can be done only once on the 3D template, so that the same annotation can be used automatically for multiple people. Annotation can be defined by listing triangles belonging to each facial feature or by using UV coordinates for the mesh along with a 2D texture map to color the facial features with different colors as in FIG. 12.

[0284] FIG. 14 illustrates a method 1700.

[0285] The method may perform the following steps for each pair of consecutive video frames (denoted as a first image and a second image), including one or more iterations of steps 171-175.

[0286] Step 171 may involve calculating current 2D positions of various facial landmarks in the first image given the current 3D face mesh and camera parameters.

[0287] Step 171 may include using a previous iteration's model of the deformed face mesh and camera screen space projection parameters, where the method uses the camera's extrinsic and intrinsic parameters to perform a perspective projection on the 3D face mesh to obtain 2D screen space pixel locations of each visible annotated face part vertex. Using the 3D pre-annotation (see FIG. 15, 3D model 141 and UV map 142), the method finds the 2D positions of the vertices within each face part by matching the annotations.

[0288] Step 172 may include calculating 2D positions of various facial landmarks in the second image.

[0289] Step 172 may include running a face part segmentation method to annotate each pixel of the image, and if the pixel does not belong to the background, the method saves the defined face part to which it belongs (eyes, nose, ears, lips, eyebrows, etc.) as an annotation.

[0290] Step 173 may include calculating a 2D->2D density correspondence between 2D locations of face features in the first image and 2D locations of face features in the second image.

[0291] Step 173 may include finding, for each face part, a correspondence between the face part points of the first image and one of the second image by running a symmetric ICP method (https: / / en.wikipedia.org / wiki / Iterative_closest_point). The ICP method proceeds iteratively between two steps, the first step finding the correspondence between the shape of the first image and the shape of the second image by aspirationally choosing, for each point in the shape of the first image, the closest point on the shape of the second image. The second step optimizing and finding the rotation and translation that best transforms the points of the first image to the points of the second image in a least squares sense. To find the optimal solution, the process repeats those two steps until convergence occurs when a convergence metric is satisfied.

[0292] Here, the shapes in the first image can be the current 2D positions of the various face parts, and the shapes in the second image can be the 2D positions given by the segmentation map (see above). The rigorous adaptation of the ICP can be done separately for each face part. For example, for each visible projected nose pixel in the first image, find the corresponding pixel on the nose in the second image given by the face part segmentation for the target image.

[0293] Step 174 may include calculating a 3D->2D density correspondence between 3D locations in the first image and 2D locations in the second image.

[0294] Step 175 may include deforming the face mesh to match the correspondence.

[0295] Step 174 may include using a rasterizer and the given camera parameters to backproject the 2D pixels rendered from the 3D face mesh, defined by their barycentric coordinates, and the camera model of the first image back to their 3D positions on the mesh, thus producing a correspondence between face feature points in 3D on the mesh and positions of the second image in 2D under the perspective projection of the camera.

[0296] Step 175 may include using a deformation model (e.g., a 3D MMM as described above) to deform the face mesh and modify camera parameters so that projections of 3D features in the first image match 2D locations in the second image, as in typical sparse landmark and camera fitting.

[0297] Steps 171-175 may be repeated until the convergence metric is met.

[0298] For example, as in the correspondence and matching procedure, the above steps can be repeated until convergence, at each step finding different and better correspondences and optimizing them. Convergence is achieved when a convergence metric is satisfied.

[0299] The method produces a set of landmarks in areas and face parts that may not be covered by traditional landmark methods, such as ears, neck, and forehead, due to the use of a 3D mesh. The method produces a dense set of landmarks, and the density correspondence requires little annotation except for a one-time annotation of the face part in the 3D model template. The method produces a dense set of high quality landmarks that can be temporally coherent due to the regression performed. Coherence in this context means that the landmarks do not have jitter between frames.

[0300] It also allows obtaining landmarks on various face or body parts, e.g., ears and neck, by simply adopting common segmentation / classification methods.

[0301] FIG. 16 may be an illustration of 2D-2D density correspondence calculation for the upper lip (pixels that are identically colored in both images correspond to each other).

[0302] FIG. 17 illustrates a method including a sequence of steps 71, 72, 73, and 74.

[0303] Step 71 may include obtaining a virtual 3D environment, which may include generating or receiving executed once instructions that cause the virtual 3D environment to be displayed to a user. The virtual 3D environment may be a virtual 3D video conferencing environment or may be different from the virtual 3D video conferencing environment.

[0304] Step 72 may include obtaining information about an avatar associated with a participant, the participant avatar including at least a face of the participant in the conference call. The participant avatar may be received once, at least once per period, or at least once per conference call.

[0305] Step 73 may include virtually positioning an avatar associated with the participant within the virtual 3D environment. This may be done in any manner, such as based on the participant's previous sessions, based on metadata such as job title and / or priority, based on role in a conference call, e.g., call initiator, as well as based on participant preferences. Step 73 may include generating a virtual representation of the virtual 3D environment populated with the participant's avatar.

[0306] Step 74 may include receiving information regarding a spatial relationship between a position of the participant's avatar and the participant's gaze direction, and updating at least an orientation of the avatar associated with the participant within the virtual 3D environment.

[0307] FIG. 18 illustrates a method 1600.

[0308] The method 1600 may be for updating a current avatar for a person and may include steps 1601 , 1602 , 1603 , 1604 , and 1605 .

[0309] Step 1601 may include calculating a current location in two-dimensional (2D) space of a current facial landmark point of the person's face, which may be based on a current avatar and one or more current acquisition parameters of the 2D camera, and the person's current avatar may be located in 3D space.

[0310] Step 1602 may include calculating, in 2D space, target positions of facial landmark points of the person's face, where calculating the target positions may be based on one or more images acquired by the 2D camera.

[0311] Step 1603 may include calculating a correspondence between the current position and the target position.

[0312] Step 1604 may include calculating locations of facial landmark points in 3D space based on the correspondences.

[0313] Step 1605 may include modifying the current avatar based on the location of the facial landmark points in 3D space.

[0314] The current facial landmark points can only be edge points of the current facial landmarks.

[0315] The current facial landmark points may include edge points of the current facial landmark and non-edge points of the current facial landmark.

[0316] Calculating the correspondence may involve applying an iterative closest point (ICP) process, where the current position may be considered as the source position.

[0317] The locations of the target facial landmark points in 3D space can be represented by barycentric coordinates.

[0318] The current avatar may include a reference avatar and a current 3D deformed model, and modifying the current avatar may include modifying the current 3D deformed model without substantially modifying the reference avatar.

[0319] The current 3D deformation model can be a 3D morphable model (3DMM).

[0320] The method may include repeating steps 1601-1605 for the current image and until convergence.

[0321] Step 1602 may include segmentation.

[0322] FIG. 18 also illustrates an example method 1650 for conducting 3D video conferencing between multiple participants.

[0323] The method 1650 may include steps 1652, 1654, and 1656.

[0324] Step 1652 may include receiving initial 3D participant representation information for generating 3D representations of participants under different circumstances. This receiving may be based on videos or images of participants acquired specifically for the video conference or for other purposes. The received information may also be retrieved from additional sources such as social networks and the like. The participant information may relate to participants of the conference call, for example, a first participant and a second participant.

[0325] Step 1654 may include receiving, by the first participant's user device, second participant situation metadata indicating one or more current situations regarding the second participant during the 3D video conference call.

[0326] Step 1656 may include updating, by the first participant's user device, the 3D participant representation within the first representation of the virtual 3D videoconferencing environment.

[0327] The different situations may include at least one of different image acquisition conditions, different gaze directions, different viewer viewpoints, and different facial expressions.

[0328] The initial 3D participant representation information may include an initial 3D model and one or more initial texture maps.

[0329] FIG. 18 also illustrates an example method 1900 for conducting 3D video conferencing between multiple participants.

[0330] The method 1900 may include steps 1910 and 1920.

[0331] Step 1910 may include determining, for each participant, updated 3D participant representation information within the virtual 3D video conference environment multiple times during the 3D video conference.

[0332] Step 1920 may include generating, for at least one participant, an updated representation of the virtual 3D video conference environment multiple times during the 3D video conference, where the updated representation of the virtual 3D video conference environment represents updated 3D participant representation information for at least a portion of the multiple participants.

[0333] The 3D participant representation information may include a 3D model and one or more texture maps.

[0334] The 3D model may have separate parameters for shape, pose, and expression.

[0335] Each texture map may be selected and / or augmented based on at least one of shape, pose, and facial expression. Augmenting may include modifying values ​​due to lighting, facial makeup effects (such as lipstick and blush), removing facial hair features (such as beards, mustaches, etc.) and accessories (such as glasses, earphones, etc.), etc.

[0336] Each texture map may be selected and / or augmented based on at least one of the shape, pose, facial expression, and angular relationship between the participant's face and the optical axis of a camera capturing an image of the participant's face.

[0337] The method may include iteratively selecting, for each participant, a selected 3D model from a plurality of 3D models of the participant, and smoothing a transition from one selected 3D model of the participant to another 3D model of the participant.

[0338] Step 1910 may include at least one of the following: a. using one or more neural networks to determine updated 3D participant representation information. b. using a plurality of neural networks for determining the updated 3D participant representation information, where different neural networks of the plurality of neural networks may be associated with different situations. c. using a plurality of neural networks for determining the updated 3D participant representation information, where different neural networks of the plurality of neural networks may be associated with different resolutions.

[0339] The method may include selecting an output of at least one of the multiple neural networks based on a required resolution, the multiple neural networks operating on different output resolutions, and the one having a resolution closest to the required resolution is selected.

[0340] FIG. 18 further illustrates an example method 2000 for 3D video conferencing among multiple participants.

[0341] The method 20 may include steps 2010 and 2020.

[0342] Step 2010 may include determining, for each participant, updated 3D participant representation information within the virtual 3D videoconferencing environment that represents the participant. Determining may include estimating 3D participant representation information of one or more occlusion areas of the participant's face that may be hidden from a camera capturing at least one viewable area of ​​the participant's face.

[0343] Step 2020 may include generating an updated representation of the virtual 3D video conferencing environment for at least one participant, the updated representation of the virtual 3D video conferencing environment representing updated 3D participant representation information for at least a portion of the multiple participants.

[0344] The method may include texture mapping the 3D model occlusion areas and one or more occlusion portions.

[0345] Estimating the 3D participant representation information of one or more occlusion areas may be performed using one or more generative adversarial networks.

[0346] The method may include determining a size of the avatar.

[0347] Multiresolution Neural Networks for Rendering 3D Models of People In 3D virtual meeting applications, there may be a need to present 3D virtual video conference participants with very high quality within a virtual 3D video conference environment. To achieve a high level of realism, neural networks may be used to create 3D models of the head and body of each participant. Neural networks may also be used to create texture maps of the participants, and the 3D models and texture maps may then be rendered to create images of the participants that can be viewed from different angles.

[0348] When there are more than two participants in a meeting, each participant may want to zoom in and out to see the other participants from a close-up, rather than zooming in and out to see more or all of the participants in the meeting.

[0349] Using a neural network to create 3D models and texture maps of participants can typically be a computationally intensive operation. Running a neural network multiple times to render images of many participants may not be scalable or possible using standard computers, as the number of calculations required may be high and computer resources may be wasted without achieving real-time rendering. Alternatively, using a network of computers on the cloud may be very costly.

[0350] According to this embodiment, a collection of networks can be trained to produce 3D models and texture maps at different levels of detail (number of polygons in the 3D model and number of pixels in the texture map).

[0351] For example, a very high resolution network may create a 3D model with 10,000 polygons and a 2D texture map with 2000 x 2000 pixels. A high resolution network may create a 3D model with 2500 polygons and a 2D texture map with 1000 x 1000 pixels.

[0352] The medium resolution network may create a 3D model with 1500 polygons and a 2D texture map with 500 x 500 pixels. The low resolution network may create a 3D model with 625 polygons and a 2D texture map with 250 x 250 pixels.

[0353] In an embodiment, all these networks can be one network with several outputs after a variable number of layers, for example the output of the final network is a texture map with 2000x2000 pixels, and the output of the previous layer is a texture map with 1000x1000 pixels.

[0354] During run time, the software determines what size the image of each participant in the meeting should be, according to the zoom level that the user may be using.

[0355] Depending on the size required following the zoom level, the method determines which network should be used to create the 3D model and 2D texture map with the relevant level of detail. In this way, smaller numbers require lower resolution networks resulting in fewer calculations per network. Thus, the total number of calculations required to render an image of many people is reduced compared to running many full resolution networks.

[0356] According to an embodiment, a texture map of a person's face can be generated based on texture maps of different areas of the face.

[0357] One of the texture maps for an area of ​​the face (e.g., for facial landmarks eyes and mouth) may be of higher resolution (more detail) than the texture map for another area of ​​the face (e.g., the area between the eyes and nose may have higher resolution than the cheeks or forehead). For example, a higher resolution texture map for the eyes may be added to a lower resolution texture map for another area of ​​the face to provide a hybrid texture map 2222, see FIG. 20.

[0358] The texture maps for different areas can be of two or more different resolution levels. The selection of resolution for each texture map can be fixed or can change over time. The selection can be based on the priority of the different areas. The priority can change over time.

[0359] According to another embodiment, the texture maps of different areas of the face may be updated and / or transmitted at different frequencies according to the frequency of changes in those areas. For example, the eyes and lips may change more frequently than the nostrils or eyebrows. Thus, the texture maps of the nostrils and eyebrows may be updated less frequently than for the eyes and lips. In this way, the number of calculations is further reduced compared to a situation in which the texture maps of the nostrils and eyebrows are updated with more frequent updates of the texture maps of the eyes and lips.

[0360] The resolution of the texture maps for different facial areas may be based on additional parameters such as available computational and memory resource conditions.

[0361] Generating the texture map of the face from the texture maps of different areas of the face may be performed in any manner and may include, for example, smoothing boundaries between the different texture maps of the different areas, etc. Any references made to the face may mutatis mutandis be applied to the entire person or to any other body tissue of the person.

[0362] FIG. 18 also illustrates an example method 2100 for generating texture maps for use during a video conference, such as a virtual 3D conference.

[0363] Method 21 may include steps 2110, 2120, and 2130.

[0364] Step 2110 may include obtaining (e.g., generating or receiving in any manner) multiple texture maps for multiple areas of at least a portion of the 3D object, where the multiple texture maps may include a first texture map for a first area and a first resolution and a second texture map for a second area and a second resolution, where the first area is different from the first area and the first resolution is different from the second resolution.

[0365] Step 2120 may include generating a texture map for at least a portion of the 3D object, where the generating may be based on a plurality of texture maps.

[0366] Step 2130 may include utilizing a visual representation of at least a portion of the 3D object based on a texture map of at least a portion of the 3D object during the video conference.

[0367] Multi-view texture maps It can result in generating highly realistic faces, which can be applied to other objects.

[0368] High quality and highly realistic images and videos or faces and bodies can be a common problem in computer graphics.

[0369] This can be applied to the creation of movies or computer games, among other uses.

[0370] It can also be applied to create 3D video conferencing applications where users can sit in a common space, and 3D avatars represent the participants, moving and speaking according to the users' actual movements as captured by a standard webcam.

[0371] To create a realistic looking 3D representation of a face, head or body, 3D models and 2D texture maps can be created offline and then manipulated. Manipulating means creating maneuvers within the 3D model that enable different parts of the model to move, much like muscles do in a real body.

[0372] The 3D models and texture maps should include views of the external parts of the body and face, but also internal parts such as the mouth, teeth, and tongue. They should enable body parts such as eyelids to move and to present open and closed eyes.

[0373] To create highly realistic looking images or videos, very high level 3D models may be used, typically with up to 100,000 in a model of a head.

[0374] In addition, the texture maps should contain a description of all internal and external body / head parts at high resolution.

[0375] In addition to texture maps, material or reflectivity maps may be required to enable the rendering engine to simulate non-uniform (non-Lambertian) reflections of light from the body and face, for example from moist or oily skin, or from glaring eyes.

[0376] Creating such 3D models and 2D texture and material maps typically requires a well-equipped studio with many cameras and controlled lighting, which limits the use of these models to offline and post-production use cases.

[0377] Due to all this, rendering highly realistic bodies and heads can be a complex process requiring many calculations, the amount of which may not be capable of being processed on a standard computer either in real time or at high frame rates (at least 30 frames per second).

[0378] This problem becomes even more severe when many bodies and heads need to be rendered in each image, for example when there can be many participants in a 3D meeting.

[0379] Instead of using 3D models with a very high number of polygons, texture maps with interior and exterior parts and many options for material / reflection maps, an alternative solution is provided that requires much less computation and also enables real-time rendering of many bodies and faces.

[0380] The solution may be based on capturing images or videos of a person from different viewpoints, for example from the front, side, back, top, and bottom.

[0381] This can be done by scanning the head with a handheld cell phone camera, or by turning the head in front of a fixed camera such as a webcam or cell phone camera mounted on a tripod or any other device. Images of people can also be obtained in other ways and from other sources, including using scanned photographs of people, etc., or extracted from social networks or internet resources.

[0382] During the scanning process, the person may be asked to perform different facial expressions and speak. To scan the whole body, the user may be asked to pose in different body postures, move and change posture continuously.

[0383] The images collected in this step can be used to train a neural network or several neural networks that create a 3D model of the head and / or body depending on the pose and facial expression required, as well as depending on the viewpoint.

[0384] In addition, texture-map dependent viewpoints can be generated depending on the required pose and expression, and depending on the viewpoint.

[0385] The 3D model and texture maps can be used to render an image of the head and / or body or person.

[0386] Since the 2D texture map output by the neural network may depend on the viewpoint, pose, and / or expression, it should only contain information that may be relevant to rendering the image from the viewpoint, pose, and / or expression. This enables the 3D model of the head or body to be less precise, so that missing 3D details such as skin wrinkles can be compensated for by the fact that those details appear in the 2D texture image. Similarly, there may be no need to produce manipulated models of open or closed eyelids, so that the texture of the open or closed eyelids can be found in the 2D image and projected onto the 3D model.

[0387] In fact, the 3D model can be highly inaccurate, as it omits many facial details and does not take into account small muscles and their movements. It also does not include the interior, which is not a moving part of the face as mentioned above, and the 2D image presents the exterior from one perspective, not multiple perspectives. This means that inaccuracies in the 3D model are not reflected in rendering the 3D model and texture maps from one perspective.

[0388] As a result, the 3D model used to render the image does not need to be very detailed and does not contain many polygons: typically it may have thousands or hundreds of polygons, compared to tens or hundreds of thousands of polygons in conventional solutions.

[0389] This allows for fast, real-time rendering of the head and / or body in real time on computing devices with inexpensive processing units.

[0390] Furthermore, the 3D model and 2D texture map can be output by different networks depending on the resolution of the desired output image. A low-resolution image is rendered based on a low-resolution polygonal 3D model and a low-resolution texture map that can be output by a neural network with fewer coefficients that require fewer calculations.

[0391] This further allows rendering several heads and / or bodies at once in one image using low-cost, low-power computing devices such as laptops without GPUs, mobile phones, or tablets.

[0392] It will also be noted that the solution does not require a studio and can be based on a single camera, it does not require complex systems with many cameras and illumination sources, and it does not require controlled lighting.

[0393] FIG. 19 illustrates an example of a method 2200 for 3D video conferencing.

[0394] The method 2200 may include steps 2210 and 2220.

[0395] Step 2210 may include determining, for each participant, updated 3D participant representation information within the virtual 3D videoconferencing environment that represents the participant, which may include compensating for differences between an actual optical axis of a camera capturing an image of the participant and a desired optical axis of the virtual camera.

[0396] Step 2220 may include generating an updated representation of the virtual 3D video conferencing environment for at least one participant, where the updated representation of the virtual 3D video conferencing environment represents updated 3D participant representation information for at least a portion of the multiple participants.

[0397] The updated representation of the virtual 3D video conferencing environment may include avatars for at least some of the multiple participants.

[0398] The gaze direction of a first avatar in the virtual 3D video conferencing environment may represent a spatial relationship between (a) the gaze direction of the first participant, which may be represented by the first avatar, and (b) a representation of the virtual 3D video conferencing environment displayed to the first participant.

[0399] The gaze direction of the first avatar in the virtual 3D videoconferencing environment may be agnostic to the actual optical axis of the camera.

[0400] A first avatar of a first participant within the updated representation of the virtual 3D video conferencing environment appears within the updated representation of the virtual 3D video conferencing environment as captured by the virtual camera.

[0401] The virtual camera may be located in a virtual plane virtually across the first participant's eyes.

[0402] The method may include receiving or generating participant appearance information relating to a participant's head pose and facial expression, and determining updated 3D participant representation information to reflect the participant appearance information.

[0403] The method may include determining a shape of each of the avatars.

[0404] FIG. 19 also illustrates an example of a method 2300 for generating an image from the perspective of an object, which may be three-dimensional.

[0405] The method 2300 may include a step 2310 of rendering an image of the object based on at least one two-dimensional (2D) texture map associated with a compact 3D model of the object and a viewpoint.

[0406] Rendering may include virtually placing textures generated from at least one 2D texture map onto the compact 3D model.

[0407] The method may include selecting at least one 2D texture map associated with a viewpoint from a plurality of 2D texture maps, which may be associated with different texture map viewpoints.

[0408] The rendering may also be responsive to the desired appearance of the object.

[0409] The object may be a representation of an acquired object that may be acquired by a sensor.

[0410] The rendering may also be responsive to obtained appearance parameters of the object.

[0411] The captured objects may be participants in a three-dimensional (3D) video conference.

[0412] The method may include receiving at least one 2D texture map from the one or more neural networks.

[0413] FIG. 19 further illustrates an example method 2400 for conducting 3D video conferencing among multiple participants.

[0414] The method 2400 may include steps 2410, 2420, and 2430.

[0415] Step 2410 may include receiving, by a first unit that may be associated with the first participant, second participant metadata and first perspective metadata, where the second participant metadata may indicate a posture of the second participant and a facial expression of the second participant, and the first perspective metadata may indicate a virtual position from which the first participant desires to view the second participant's avatar.

[0416] Step 2420 may include generating, by the first unit, second participant representation information based on the second participant metadata and the first viewpoint metadata, where the second participant representation information may include a compact 3D model of the second participant and a second participant texture map.

[0417] Step 2430 may include determining, for the first participant, a representation of the virtual 3D videoconferencing environment during the 3D videoconference, where the determining may be based on the second participant representation information.

[0418] The method may include generating one each of the compact 3D and second participant texture maps in response to the second participant metadata and the first viewpoint metadata.

[0419] Generating at least one of the compact 3D model and the second participant texture map may include feeding the second participant metadata and the first viewpoint metadata to one or more neural networks trained to output at least one of the compact 3D model and the second participant texture map based on the second participant metadata and the first viewpoint metadata.

[0420] A compact 3D model may contain fewer than 10,000 points.

[0421] A compact 3D model may necessarily consist of 5 thousand points, such as the FLAME model (https: / / flame.is.tue.mpg.de / home).

[0422] Determining the representation of the virtual 3D videoconferencing environment may include determining an estimate of an appearance of the second participant within the virtual 3D videoconferencing environment based on a second participant texture map, and revising the estimate based on a compact 3D model of at least the second participant.

[0423] The correcting may include correcting the estimation based on concealment and illumination effects associated with compact 3D models of one or more participants in the 3D conference video.

[0424] Retention and mood estimation from video. Due to Covid 9, many in-person meetings have been replaced with video conference calls.

[0425] Such calls can be long, participants can lose their attention or focus, and may be tempted to do other things in parallel with the meeting, such as browsing the Internet, reading e-mail, or playing on their phones.

[0426] In many cases, it may be important for some meeting participants to know whether other participants may be attentive (i.e., paying attention to the meeting) and how other participants feel - for example, they may be happy, sad, angry, stressed, agree, or disagree with what other participants are saying.

[0427] Example cases for such video teleconferencing may be associated with, for example, school lectures, university lectures, sales meetings, and team meetings managed by a team manager.

[0428] Solutions can be provided to analyze the video and estimate the memory capacity of participants, especially those who are not actively participating and speaking.

[0429] A database of videos from video conference meetings can be collected.

[0430] For each participant (or at least a portion of participants) appearing in one or more of the videos, the videos can be divided into portions where the user's memory and emotion can be assumed to be constant. In each portion of each video, memory levels and emotions can be estimated by using several possible measures.

[0431] Participants can be queried as to how engaged they were during that portion of the meeting and what their mood was during that time. a. External annotators can be asked to estimate memory and mood based on participants' appearance, including head pose, eye movements, and facial expressions. b. External devices may be used to measure the participant's heartbeat and other biological signals, such as is done by a polygraph machine or other less sophisticated methods. c. Computer software or observers may verify whether a participant was looking at another window on their computer screen that may not be related to the meeting, i.e., not fully focused on the meeting.

[0432] A numerical score for retention may be created for each video segment, or alternatively, participants' retention may be classified into several classes, such as "highly interested," "interested," "uninterested," "boring," "extremely boring," and "lots of tasks."

[0433] In a similar manner, the user's mood can be inferred, for example "happy", "satisfied", "sad", "angry", "stressed".

[0434] Conversely, a numerical value can be assigned to certain sensations such as happiness, relaxation, interest, etc.

[0435] Neural network models can be trained to find correlations between participants' appearances in videos and their levels of memorability and mood.

[0436] At run-time, a video can be fed into the network, which outputs an estimate of the retention level as a function of time.

[0437] This output may be presented to some participants, such as the meeting host or manager (teacher, salesperson, manager), to improve their performance or to help some other participants who may have lost memory.

[0438] In an embodiment, faces detected in a video may be modeled by a neural network that generates a parametric model, including parameters for head pose, eye gaze direction, and facial expression, as described in a previous patent.

[0439] Once a parametric model is discovered, only the parameters can be input into the neural network, which estimates the retention level instead of inputting the raw video.

[0440] The parameters may be input as a series of parameters over time, so that temporal changes in facial expressions, head, and eye movements may be taken into account. For example, if there is no change in parameters coding facial expressions or head and eye direction over an extended period of time, the network may learn that this may be a sign of lack of attention.

[0441] Such a method may be beneficial because it reduces the amount of data that can be input to a network that estimates the level of retention.

[0442] In another embodiment, the output of the video analytics network may be combined with data collected by computer software.

[0443] Such additional data can be: a. Are other windows visible on the screen? b. Can users type or click a mouse during a video conference meeting? c. Using eye gaze tracking, the direction a person can look can be estimated.

[0444] The method may estimate whether a user is looking at someone who is talking or at other people in a video conferencing application, or is just staring around.

[0445] Using eye gaze detection, the method can also estimate whether the user is looking at areas of the screen that are not occupied by the videoconferencing software, such as other open windows.

[0446] Using eye gaze detection, the method can estimate whether a user is reading text during a meeting.

[0447] The combination of all data sources can be used to estimate whether meeting participants are likely to have multiple tasks during the meeting and whether they are likely to pay attention to other tasks instead of the video meeting.

[0448] It will be noted that the above mentioned process is not limited to rendering images of people, but can also be used to render animals or any other object.

[0449] FIG. 19 further illustrates an example method 2500 for determining psychological parameters of participants in a video conference.

[0450] The method 2500 may include applying 2510 a machine learning process to videos of the participants captured during the video conference to determine a mental state of the participants during the video conference, where the mental state may be selected from mood and memory retention. The machine learning process may include training on video segments of one or more persons and training mental state metadata indicative of the mental state of the one or more person participants during each of the training video segments together with which it has been trained by the training process.

[0451] The training mental state metadata may be generated in any manner, for example by at least one of the following: a. Querying one or more people. b. be generated by entities other than one or more persons (such as medical staff and experts); c. Measuring one or more physiological parameters of the one or more persons during acquisition of the training video segments. d. During acquisition of the training video segments, the training video segments are generated based on the interaction of one or more persons with a component other than a display associated with the one or more persons. e. training video segments are generated based on the gaze direction of one or more persons during acquisition;

[0452] One or more persons may be participants.

[0453] The video conference may be a three-dimensional (3D) video conference.

[0454] The method 2500 may include training.

[0455] FIG. 18 further illustrates an example method 2600 for determining the mental state of a participant in a video conference.

[0456] The method 2600 may include steps 2610 and 2620.

[0457] Step 2610 may include obtaining participant appearance parameters during the 3D video conference. An example of such parameters is given in the Flame model (https: / / flame.is.tue.mpg.de / home).

[0458] Step 2620 may include determining the mental state of the participant, which may include analyzing the parameters by a machine learning process.

[0459] The machine learning process may be implemented by a thin neural network.

[0460] The analysis is carried out iteratively during the 3D video conference.

[0461] The analyzing may include following one or more patterns in the values ​​of the appearance parameters.

[0462] The method may include determining, by a machine learning process, the mental state of the participant based on the one or more patterns.

[0463] The method may include determining a memory deficit in which one or more appearance parameters may not change substantially for at least a predetermined period of time.

[0464] The mental state may be the mood of the participant.

[0465] The mental state may be the participant's memory.

[0466] The determining may be further responsive to one or more interaction parameters relating to the participants' interactions in devices other than the display.

[0467] The participant appearance parameters may include the participant's gaze direction.

[0468] FIG. 19 illustrates an example method 2700 for determining psychological parameters of participants in a video conference.

[0469] The method 2700 may include steps 2710 and 2720.

[0470] Step 2710 may include obtaining participant interaction parameters during the 3D video conference.

[0471] Step 2720 may include analyzing the participant interaction parameters to determine psychological parameters of the participants through a machine learning process.

[0472] FIG. 19 also illustrates an example method 2800 for determining the mental state of a participant in a video conference.

[0473] The method 2800 may include steps 2810, 2820, and 2830.

[0474] Step 2810 may include obtaining participant appearance parameters during the 3D video conference.

[0475] Step 2820 may include obtaining participant computer traffic parameters indicative of computer traffic exchanged with a participant computer utilized to participate in the 3D video conference.

[0476] Step 2830 may include determining the mental state of the participant, which may include analyzing participant appearance parameters and participant computer traffic parameters by a machine learning process.

[0477] FIG. 19 also illustrates an example method 2900 for determining the mental state of a participant in a video conference.

[0478] The method 2900 may include steps 2910, 2920, and 2930.

[0479] Step 2910 may include obtaining participant appearance parameters during the 3D video conference.

[0480] Step 2920 may include obtaining participant computer traffic parameters indicative of computer traffic exchanged with a participant computer utilized to participate in the 3D video conference.

[0481] Steps 2910 and 2920 may be followed by step 2930 of determining the mental state of the participant, which may include analyzing participant appearance parameters and participant computer traffic parameters by a machine learning process.

[0482] It should be noted that the total number of calculations that may need to be performed may not be bound by the number of people appearing in the Field Of View (FOV), but rather by the resolution of the view: if the screen resolution remains constant, for example, widening the FOV may result in more participants being shown, but with a smaller size that needs to be captured and rendered.

[0483] Multiple participants in one visual detection unit Existing teleconferencing systems assume one participant per camera, so one tagged name appears per camera even if more than one person uses it. This can lead to a lack of understanding of who the participants are, especially if other participants cannot recognize them.

[0484] Even when multiple participants are captured by a single camera, or by a visual sensing unit that may include more than a single camera, it may be beneficial to provide an accurate representation of each participant captured by the camera.

[0485] Participants may appear in one or more representations of the virtual 3D videoconferencing environment, and each participant may be represented by an avatar.

[0486] It should be noted that non-participants may also appear in one or more representations of the virtual 3D videoconferencing environment. Thus, a person who is to appear in at least one representation of the virtual 3D videoconferencing environment may be considered a relevant person. A relevant person may be a participant or a non-participant.

[0487] The method may begin with visual information analysis, such as detecting the number of people captured by the visual sensing unit and attempting to identify the people. Any identification process may be used, for example face detection and recognition.

[0488] Once a person is detected, the method may determine whether the person is relevant or not and may be ignored.

[0489] Assuming that the people are relevant, the images of the people (portions of the image acquired by the visual sensing unit) may be segmented. Segmentation may include associating different segments with each participant's clothing or other possible accessories (watches, glasses, jewellery, etc.). Optionally, the relevant people may be enabled to identify the different segments (by receiving input from a user).

[0490] In a virtual 3D video conferencing environment where each participant is represented by an avatar, each one of the relevant persons captured by the visual sensing unit may be represented by a different avatar. Without identifying that there are multiple relevant persons, such a system would not function.

[0491] Within this framework, it may happen that one of the related persons makes a gesture or, in some cases, looks at another related person in the same camera. This is then reflected by the behavior of the avatar. As an illustration, if one of the related persons hands an object to another related person, this action may be reflected in the virtual 3D video conferencing environment, showing the first related person and corresponding avatar handing a similar object to the avatar corresponding to the second related person.

[0492] Optionally, the system also has a temporary tracking mechanism with some temporary memory. This allows participants to move in and out of the camera's view and be separately identified. This tracking can be based on face recognition, clothing color tracking, or similar methods.

[0493] Another option is that when more than one person appears in the camera view, the system can be instructed to show only a subset of those people in the video conference. For example, if a video conference is conducted from home, it is quite customary for other household people and animals - children, pets, spouse (considered to be unrelated) to appear occasionally in the camera's view. In this case, the system can be configured to not show the unrelated people or animals in the video conference.

[0494] FIG. 21 illustrates examples of several methods, method 3000, method 3001, method 3003, and method 3200.

[0495] The method 3000 is for conducting a virtual 3D video conference between multiple participants.

[0496] Execution of the virtual 3D video conference may include displaying multiple representations of the virtual 3D video conference environment on multiple participant devices. Computations required for provision of the virtual 3D video conference may be performed by one or more computing systems other than any of the multiple participant devices, may be performed solely (or nearly solely) by the multiple participant devices, or may be performed by a combination of one or more participant devices and one or more other systems.

[0497] Information related to the presence of the relevant person within the field of view of the visual detection unit associated with any of the participants may be transmitted to one or more other participant devices, may be transmitted to one or more other systems, and may be subject to filtering rules, transmission blocking rules, or any other rules related to the processing and / or transmission and / or display of any indications related to the plurality of persons.

[0498] The participant devices may display multiple representations of the virtual 3D videoconferencing environment, and typically the representations of the virtual 3D videoconferencing environment vary from one participant device to another. The presence of one or more relevant persons may be reflected in at least a portion of the multiple representations of the virtual 3D videoconferencing environment.

[0499] The method 3000 may begin by acquiring 3010 visual information by a visual sensing unit associated with a participant.

[0500] Step 3010 may be followed by step 3020 of identifying one or more people appearing in the visual information. In some cases, multiple people may appear in the visual information. In some other cases, only one person appears in the visual information. In some further cases, no people appear in the visual information.

[0501] If a single person appears in the visual information or if no people appear in the visual information, step 3020 may be followed by step 3029, responding to the detection, or responding to the single person, or responding to the absence of any people.

[0502] If multiple persons appear in the visual information, step 3020 may be followed by a step 3030 of finding at least one relevant person from the multiple persons.

[0503] An associated person is a person whose presence can be indicated to at least one participant (or participant device) of the virtual 3D video conference. At the very least, an indication of the associated person's presence can be transmitted outside of a participant device of a participant.

[0504] The presence of an associated person may be represented (or at least be a candidate for being represented) within the virtual 3D video conference environment that is displayed to one or more participants of the virtual 3D video conference. A participant may decide not to receive an indication of the person and / or the display of said presence may be subject to filtering and / or display rules. An associated person may or may not be a participant.

[0505] Step 3030 may include at least one of the following: a. Determining which persons of a plurality of persons are participants in a virtual 3D video conference. b. Determining whether a Participant is a Related Person. c. Determining whether a non-participant in a 3D video conference is a relevant person. d. Applying a facial recognition process. e. Applying any biometric identification process, as well as facial recognition processes. f. storing identification information for at least one associated person for at least a period of time following appearance of a participant and person, which may reduce utilization of computational resources since there is no need to initiate another association determination process; g. Identifying any of the at least one relevant person after the at least one relevant person leaves the field of view of the visual sensing unit and then re-enters the field of view of the visual sensing unit, the identifying being based on the identification information. This may provide a certain "memory" since a person identified as relevant may leave the field of view for up to a predefined amount of time and still be considered relevant. h. Continuing to indicate that the relevant person is within the field of view of the visual sensing unit and is required to regenerate and / or update the virtual 3D videoconferencing environment even during a predefined period when the relevant person leaves the field of view to reduce computational resources and may also reduce utilization of communication resources (no need to transmit information regarding updates to the virtual 3D videoconferencing environment). This may make the virtual 3D videoconferencing environment smoother. The method may use a hysteresis mechanism or any other smoothing mechanism when determining whether to update information related to the presence or absence of the relevant person in the virtual 3D videoconferencing environment.

[0506] Step 3030 may be followed by step 3040, which is responsive to finding at least one associated person from the plurality of persons.

[0507] Step 3040 may include step 4041 of determining 3D entity representation information for each of the at least one associated person, and step 3042 of generating, for the at least one participant, a representation of the virtual 3D videoconferencing environment based on the 3D entity representation information for each of the at least one associated person.

[0508] Steps 3010, 3020, 3030, and 3040 can be performed in connection with either a visual sensing unit or a participant.

[0509] The method 3001 is for conducting a virtual 3D video conference between multiple participants.

[0510] The method 3001 may begin by step 3010 of acquiring visual information by a visual sensing unit associated with a participant.

[0511] Step 3010 may be followed by step 3020 of identifying one or more people appearing in the visual information. In some cases, multiple people may appear in the visual information. In some other cases, only one person appears in the visual information. In some further cases, no people appear in the visual information.

[0512] If a single person appears in the visual information or if no people appear in the visual information, step 3020 may be followed by step 3029, which responds to the detection, or responds to a single person, or responds to the absence of any people.

[0513] If multiple persons appear in the visual information, step 3020 may be followed by a step 3030 of finding at least one relevant person from the multiple persons.

[0514] Step 3030 may be followed by step 3040, which is responsive to finding at least one associated person from the plurality of persons.

[0515] Step 3040 may include step 4041 of determining 3D entity representation information for each of the at least one associated person, and step 3042 of generating, for the at least one participant, a representation of the virtual 3D videoconferencing environment based on the 3D entity representation information for each of the at least one associated person.

[0516] Step 3040 may include a step 3043 of searching for a physical interaction between the relevant persons. Upon finding a physical interaction, step 3040 may also include a step of generating a representation (for at least one participant) of the virtual 3D videoconferencing environment that may respond to the physical interaction.

[0517] The method 3002 is for conducting a virtual 3D video conference between multiple participants.

[0518] The method 3002 may begin by acquiring 3010 visual information by a visual sensing unit associated with a participant.

[0519] Step 3010 may be followed by step 3020 of identifying one or more people appearing in the visual information. In some cases, multiple people may appear in the visual information. In some other cases, only one person appears in the visual information. In some further cases, no people appear in the visual information.

[0520] If a single person appears in the visual information or if no people appear in the visual information, step 3020 may be followed by step 3029, which responds to the detection, or responds to a single person, or responds to the absence of any people.

[0521] If multiple persons appear in the visual information, step 3020 may be followed by a step 3030 of finding at least one relevant person from the multiple persons.

[0522] Step 3030 may be followed by step 3040, which is responsive to finding at least one associated person from the plurality of persons.

[0523] Step 3040 may include step 3041 of determining 3D entity representation information for each of the at least one associated person, and step 3042 of generating, for the at least one participant, a representation of the virtual 3D videoconferencing environment based on the 3D entity representation information for each of the at least one associated person.

[0524] Step 3040 may include step 3045 of generating a same visual detection unit indication that indicates that the associated person is captured by a single visual detection unit, see, for example, same visual detection unit indication 3099 in FIG.

[0525] A same visual detection unit indication may be included in the representation of the virtual 3D videoconferencing environment (for at least one participant). The visual detection unit may include a first camera and a second camera. A same visual detection unit indication may or may not be generated that one of the relevant persons is within the field of view of the first camera and another of the relevant persons is within the field of view of the second camera.

[0526] The method 3003 is for conducting a virtual 3D video conference between multiple participants.

[0527] The method 3003 may begin by acquiring 3010 visual information by a visual sensing unit associated with a participant.

[0528] Step 3010 may be followed by step 3020 of identifying one or more persons appearing in the visual information. In some cases, multiple persons may appear in the visual information. In some other cases, one person appears in the visual information. In some further cases, no persons appear in the visual information.

[0529] If a single person appears in the visual information or if no people appear in the visual information, step 3020 may be followed by step 3029, responding to the detection, or responding to the single person, or responding to the absence of any people.

[0530] If multiple persons appear in the visual information, step 3020 may be followed by a step 3030 of finding at least one relevant person from the multiple persons.

[0531] Step 3030 may be followed by step 3040, which is responsive to finding at least one associated person from the plurality of persons.

[0532] Step 3040 may include step 4041 of determining 3D entity representation information for each of the at least one associated person, and step 3042 of generating, for the at least one participant, a representation of the virtual 3D videoconferencing environment based on the 3D entity representation information for each of the at least one associated person.

[0533] Step 3040 may include, for each associated person captured by the same visual sensing unit, a step 3047 of determining whether the associated person is speaking.

[0534] Step 4047 may be followed by responding to a determination of whether the associated person is speaking within the representation of the virtual 3D videoconferencing environment that is displayed to one or more participants.

[0535] Responding can include enabling a single speaking person to be displayed within the virtual 3D environment.

[0536] The method 3200 is for conducting a virtual 3D video conference between multiple participants.

[0537] The method 3200 may include an initialization step 3202. The initialization step 3202 may include receiving initial 3D participant representation information for generating 3D representations of participants under different conditions. The 3D participant representation information may include a 3D model and one or more texture maps.

[0538] The method 3200 may include receiving 3210 gaze direction information regarding a gaze direction of a participant. The gaze direction information may represent a detected or estimated gaze direction of the participant.

[0539] Step 3210 may be followed by step 3220 of estimating (a) whether a participant's gaze is directed toward a person located within the field of view of a visual detection unit that also captures at least the participant's head, or (b) whether the person's gaze is directed toward a representation of the person within the virtual 3D video conferencing environment.

[0540] Step 3220 may be followed by step 3230 of determining whether (i) a 3D representation of the person should appear within the virtual 3D video conferencing environment and / or whether to update the gaze direction of the participant's representation to indicate that the participant is looking at the person.

[0541] The determining may be responsive to different parameters, for example, whether a participant's gaze was directed at a person, whether the person's gaze was directed at a representation of the person within the virtual 3D video conferencing environment, whether the person is a participant in the current virtual 3D conference, whether the participant participated in any previous virtual 3D conferences, etc.

[0542] Step 3230 may include at least one of the following: a. Determining that a 3D representation of a person should appear within the virtual 3D video conferencing environment when the person is not a participant. b. Enabling non-participants to appear within the virtual 3D video conferencing environment. c. Performing a determination based on rules or definitions provided by one participant, which may also be based on rules provided by other participants, which rules may define which persons should appear in their representation of the virtual 3D video conferencing environment. d. Performing the determining based on at least one of: (a) a size of the person; and (b) an estimated age of the person. For example, children may be excluded from being represented. e. Performing the determining based on communication bandwidth and / or computational resource conditions. For example, when the available bandwidth of a communication link or channel from one participant device to another device or system falls below a certain threshold, the decision may tend to ignore a person, e.g., especially if the person is not a participant, and in another example, even if the person is not associated with an existing avatar. f. Using facial recognition to identify people. g. Using an identification process to identify certain participants and persons. h. Performing the determining based on stored identification information regarding the person and a participant, the information being stored for at least a period of time following appearance of the participant and the person. Identifying a person after the person exits a field of view of a visual sensing unit and then re-enters the field of view of the visual sensing unit. The identifying is based on the identification information.

[0543] Step 3230 may be followed by step 3240 responsive to the determination of step 3230 .

[0544] Step 3240 may include at least one of steps 3240(a) through 3240(n): When it is determined that a 3D representation of a person should appear within the virtual 3D videoconferencing environment, generating person information regarding the appearance of the person. Person information may include information that can be processed by a rendering engine or other image processor to provide a 3D representation of the person, or an avatar or other 3D representation of the person within one or more representations of the virtual 3D videoconferencing environment. The person may or may not be associated with an avatar. When associated with an existing avatar, the person information may be instructions on how to update the avatar (e.g., to provide context information). When not associated with an existing avatar, it may be necessary to generate a new avatar or use an existing avatar even if not initially associated with the person. b. When determining to update the gaze direction, updating the gaze direction of the participant's expression to indicate that the participant is looking at the person, which may include updating context regarding the participant and the like. c. Generating a same visual detection unit indication that the person and a participant are captured by the same visual detection unit, where the visual detection unit may include a first camera and a second camera, where the participant is within a field of view of the first camera and the person is within a field of view of the second camera. d. Searching for physical interactions between a person and a participant (when determining that the person should appear). e. When a physical interaction is discovered, determining whether it should appear in one or more representations of the virtual 3D video conferencing environment, and if so, how it should appear, and generating information in which the physical interaction is represented in the one or more representations. f. Generating 3D person representation information indicating that the person is not a participant. g. Maintaining a participant's gaze direction within the virtual 3D video conferencing environment unchanged during changes in a participant's gaze direction from person to visual representation of a person within the virtual 3D video conferencing environment. h. Generating an updated representation of the virtual 3D video conferencing environment including avatars for at least some of the multiple participants. i. determining the relevance of segments of updated 3D participant representation information and selecting which segments to transmit based on relevance and available resources; j. determining the relevance of segments of the updated representation of the virtual 3D videoconferencing environment information and selecting which segments to transmit based on relevance and available resources. k. Generating a 3D model and one or more texture maps of the 3D participant representation information of the participant. l. Estimating 3D participant representation information for one or more occluded areas of the participant's face that are located outside the field of view of a camera that captures at least one visual area of ​​the participant's face. m. For each participant, determining updated 3D participant representation information by changing illumination conditions. n. For each participant, determining updated 3D participant representation information by adding or modifying wearable item information.

[0545] All of steps 3240(a)-3240(n) may be performed by the same device or system, but one or more of steps 3240(a)-3240(n) may be performed by different devices and / or systems, for example, step (h) may be generated by a rendering engine located in a computerized system or a participant device different from the participant device performing step 3240(a).

[0546] There may be multiple representations of the virtual 3D videoconferencing environment (e.g., one for each participant), and steps 3230 and / or 3240 may be performed for each one of the representations. The updates themselves (including visual information, e.g., the appearance of people) may vary from one representation to another.

[0547] The participants of the virtual 3D conference are associated with multiple participant devices, which may be different computerized systems than any of the multiple participant devices.

[0548] Various steps of the method 3200 may be performed by at least one of the computerized systems and one or more of the multiple participant devices.

[0549] 22 illustrates an image 3009 that is part of a video captured by a visual sensing unit. The image 3009 captures a first person 3004 and a second person 3005. There is a physical interaction between the people as they embrace each other. The physical interaction may be represented within a virtual 3D video conferencing environment.

[0550] In one example, both persons are considered related persons and their avatars 3004' and 3005' appear within a representation 3009' of a virtual 3D videoconferencing environment (only a portion of the environment is shown).

[0551] In another example, only the first person is considered the relevant person, and his avatar 3004' (and not the second person's avatar) appears in the representation 3009" of the virtual 3D videoconferencing environment (only a portion of the environment is shown).

[0552] 22 also illustrates an image 3008 that is part of a video captured by the visual sensing unit. The image 3008 captures a third person 3007 looking at a fourth person 3008.

[0553] In one example, both persons are considered to be associated persons and their avatars 3006' and 3007' appear in a representation 3008" of a 3D videoconferencing environment (only a portion of the environment is shown). Additional avatars of other associated persons 51-53 are also shown.

[0554] 23 illustrates an example of participant gaze directions. The top part of the figure illustrates a fifth participant 85 as looking at a 3D visual representation (51) of a first participant 81 within a virtual 3D video conferencing environment (within a panoramic view 41).

[0555] The second example illustrates a fifth participant 85 as viewing the first participant 81, in this example both participants may be using the same device and are captured by the same visual capture unit.

[0556] In both cases, the virtual 3D video conferencing environment may be updated to indicate that the fifth participant is seeing the first participant (either the actual participant or a 3D participant representation). An indication may be provided as to whether the fifth participant is seeing the actual first participant or a representation of the participant.

[0557] Share content It is important that video conferences are as efficient as possible because they are susceptible to communication problems and lack the benefits of face-to-face meetings. One issue that can limit the efficiency of video conferences is the need to share information that is usually accomplished by sharing files and screens.

[0558] Existing solutions such as Zoom, Webex, and Microsoft Teams allow sharing applications or their entire screen during a meeting. Some of those applications even allow multiple users to share content at the same time. If other participants want to share content before the meeting so that they can prepare for and be notified of the meeting, they have to do so through some additional application. For example, they push the material through an email to other participants. If other participants are interested in the material after the meeting, it needs to be pushed to them.

[0559] While the proposed method is directed to 3D videoconferencing, it can also be beneficial for other systems, especially for 2D videoconferencing environments.

[0560] The proposed method allows each participant to share more than one piece of data during a conference call. Moreover, information can be easily shared with other participants before their meeting and viewed following its conclusion.

[0561] According to the proposed method, when a meeting is planned and an invite is sent out, a shared folder is created, like a folder in Google Drive or Microsoft Teams, and a link to the drive is sent to subsequent participants, which can be the same link that is later used for the meeting itself.

[0562] The meeting host is allowed to set permissions (access control rules) for access to the shared folder. Those permissions may include being able to upload documents, edit them, create subfolders, etc. The following paragraphs detail the possible options, assuming they are enabled for participants.

[0563] Participants can upload word processed documents, presentations, spreadsheets, etc. (collectively referred to as "documents") into folders. They can also create subfolders within folders based on different criteria. Participants may be able to set specific settings for documents they upload to the same folder.

[0564] One possible option is to send notifications to participants when documents are uploaded or when they are modified.

[0565] Participants can upload documents to the shared folder during the meeting itself.

[0566] An additional option is the collaborative creation of one or more documents by one or more of the participants during the meeting (eg, as Google Drive allows).

[0567] During the meeting, participants may decide at a particular time that they will share one or more of the documents in the shared folder during the meeting.

[0568] Having a shared folder allows for the following novel advantages: if a participant cannot join the meeting or has communication problems, their documents can still be viewed by other participants. It is simple for a single participant to share more than one piece of material at a time. As mentioned above, existing solutions only allow one application, one window, or one screen to be shared at a time. Sharing information before the meeting does not require multiple applications, as it takes care to update participants when documents are available.

[0569] Following the end of the meeting, it is possible to remove or delete the shared folder for any defined period of time. One additional possibility is to add a record of the meeting to the same shared folder. This then allows participants who may have missed all or part of the meeting to find all relevant information in one place. It also allows participants who take part in the meeting to review the information at their own pace after the meeting is over.

[0570] The proposed method also enables instant sharing of material with other participants without the need to send the material out after the meeting. If summaries and / or action items are captured for the meeting, they can also be placed in the shared folder.

[0571] FIG. 24 illustrates a method 3400 for sharing content during a virtual 3D video conference.

[0572] The method 3400 may begin with steps 3410, 3420, and 3430.

[0573] Step 3410 may include inviting a number of participants to join the virtual 3D video conference.

[0574] Step 3420 may include creating a dedicated shared folder for storing the shared content items, the shared content being accessible at least during the virtual 3D video conference, the shared content including at least one of text, documents, video units, and audio units.

[0575] Step 3430 may include enabling access to the shared folder for multiple participants, where access is governed by one or more access control rules, which may include adding a link to the invitation of step 3410, or performing any enabling step following step 4310, or regardless of step 3410, in conjunction with step 3410 below.

[0576] Access control rules may determine such things as retrieval of content that is shared and uploading of content to a shared folder.

[0577] One or more access control rules may be responsive to the availability of storage resources in the shared folder, for example, preventing uploads when the size of the content to be uploaded exceeds a first size threshold (which may be determined per participant, per type of participant, per organizer, per participant, etc.), when a participant reaches a second aggregate size of uploaded content from participants.

[0578] One or more access control rules may be responsive to bandwidth availability of a communication link to and / or from a shared folder.

[0579] Access may be enabled to begin prior to the start of the conference call, to begin at the time of the conference call, and the like.

[0580] Access may be terminated at or after the end of the conference call.

[0581] Steps 3410, 3420, and 3430 may be followed by a step 3440 of conducting a virtual 3D video conference, which includes sharing at least one of the content items.

[0582] Step 3440 may include recording the virtual 3D video reference.

[0583] Sharing may be performed based at least in part on one or more sharing rules. For example, all participants may share any content in a shared folder. For yet another example, the sharing rules may impose restrictions on the manner in which sharing may be performed by one or more participants.

[0584] One or more sharing rules may be included in one or more access control rules.

[0585] One or more shared rules may not be included in one or more access control rules.

[0586] Step 3440 may be followed by an additional step 3450, which is performed at or following the end of the virtual 3D conference.

[0587] Step 3450 may include at least one of the following: a. Delete the dedicated shared folder after the virtual 3D video conference is completed. b. Maintaining the shared folder dedicated after completion of the virtual 3D video conference and enabling access to the shared folder after completion of the virtual 3D video conference. c. Maintaining the shared folder dedicated until a predefined period of time after completion of the virtual 3D video conference and enabling access to the shared folder until a predefined period of time after completion of the virtual 3D video conference. d. Maintaining a shared folder dedicated after completion of the virtual 3D video conference, and applying after completion access control rules for accessing the shared folder. e. Maintaining a dedicated shared folder after completion of the virtual 3D video conference and adding a recording of the virtual 3D video conference to the shared folder.

[0588] One, some, or all of steps 3410, 3420, 3430, 3430, and 3450 may be managed by a virtual 3D video conferencing application.

[0589] FIG. 25 illustrates user devices 4000(1)-4000(R) (and 4000(r), where r ranges from 1 to R), a network 4050, a remote computerized system 4100 (which may include a virtual 3D videoconference router 4111), and a shared folder 4105 containing a number M of shared content items 4105(1)-4105(M) (and 4105(m), where m ranges from 1 to M). FIG. 25 also illustrates invitations 4106(1)-4106(R) sent by user device 4000(r) inviting other participants to access the shared folder and join the virtual 3D videoconference. During the virtual 3D videoconference, various signals (VC-related signals) 4108 are exchanged with the user devices.

[0590] The shared folder may be implemented in any manner, for example by the remote computerized system 4100 or any other unit of the system. A recording 4109 of a virtual 3D video conference is illustrated as being stored in the virtual folder.

[0591] FIG. 25 also illustrates various rules 4104(1) through 4104(N) (collectively referred to as 4104), which may include sharing rules, access control rules, and the like.

[0592] A rule may apply to all participants, to a subset of participants, or to only one participant.

[0593] FIG. 26 illustrates two examples of a first timing diagram 3480 and a second timing diagram 3480'.

[0594] The first timing diagram 3480 illustrates the sequence of events: opening a shared folder and / or initiating access to a shared folder 3482, a shared folder 3483, initiating a conference call 3485, ending a conference call 3486, and notifying participants regarding the closing of a shared folder 3487.

[0595] There may be other timing relationships between these events.

[0596] The virtual 3D conference takes place between the start of conference call 3485 and the end of conference call 3486 .

[0597] In the first timing diagram, the conference call may be recorded and available to participants until its conclusion, for example, in a shared folder 3487. The recording may be available in the shared folder or provided in any other manner.

[0598] The second timing diagram 3480 illustrates the sequence of events: (a) opening a shared folder and / or a shared folder 3482 occurring simultaneously with the initiation of access to the shared folder 3483, (a) the initiation of a conference call 3485, and (b) notifying participants regarding the end of the conference call 3486 occurring simultaneously with the closing of the shared folder 3487.

[0599] Foreground and Background It is often important in VC systems to distinguish between foreground and background. In this context, the background is the part of the scene captured by the participant's camera that is less important than other parts of the scene. The less important parts can be modified or removed altogether, since their appearance has no role in the conference or meeting taking place. In fact, existing solutions often allow for modification of the background.

[0600] This is often done to replace a participant's clear background with a more pleasing one, or one chosen for a variety of reasons, such as commercial reasons, to create a particular atmosphere, or for other reasons.

[0601] Due to the increasing importance of videoconferencing systems, it is important that the distinction between foreground and background is as accurate as possible. This task is especially important within upcoming 3D VC environments where only the participants' avatars, and possibly some accessories they may be using, are presented to the other participants.

[0602] Some solutions may distinguish between foreground and background on a frame-by-frame basis. This method works well when a method known as "green screen" is used. When using the method, the background has a known color (green), which is achieved by placing a screen behind the participants. Each pixel captured by the camera is examined. If its color matches the known screen color, the pixel is assumed to be part of the background. This method can be augmented in several ways. Nevertheless, most environments that host participants do not facilitate such screens, and other methods are used.

[0603] Existing methods for this typically search for skin color first. They then try to find some reasonable surrounding shape or color around the skin color before they determine that they identify a person in the captured picture. This often leads to body parts appearing and disappearing in a haphazard manner from the presented picture, because sometimes the body parts are perceived as part of the background and are replaced, and other times the body parts are considered part of the foreground and are not replaced.

[0604] Another drawback of this system is that if a participant wants to add some accessories, such as an easel or whiteboard, they will appear as if they are part of the background and will not be shown when they are replaced.

[0605] While the present method is directed to 3D videoconferencing, it can also be beneficial for other systems.

[0606] According to current methods, background and foreground are differentiated based on temporal tracking and not performed on a frame-by-frame basis.

[0607] According to this method, the captured picture is first segmented to identify so-called blobs. In addition to identifying blobs as part of the foreground or background based on their static characteristics (such as color or surrounding color), as in today's methods, blobs are also classified based on their temporal or dynamic characteristics. Blobs that may move and change their appearance, color, or other characteristics may be classified as belonging to the foreground, or at least as having a high probability of belonging to the foreground. In some cases, blobs with temporal motion (such as fans, fluttering pieces of paper, etc.) can be categorized as belonging to the background, or as having a high probability of belonging to the background.

[0608] An option is to let the user decide, at times, preferably as the user joins the conference, to choose whether a blob belongs to the foreground or background as this progresses. An alternative is to have a machine learning system in place to learn the temporal and spatial behavior of blobs that belong to the foreground and background. This system would learn from the user selection whether to include a blob in the foreground or background. One way to implement this is through a neural network.

[0609] Once the background is known, participants can add accessories and they will appear to other viewers in the conference. In addition, if a whiteboard or similar device is used, the writing on it and the board itself are not classified as part of the background, so the system will continue to show it to other participants.

[0610] FIG. 27 illustrates a method 3500 for foreground and background segmentation in connection with virtual three-dimensional (3D) videoconferencing.

[0611] The method 3500 may begin by segmenting 3510 each image of a plurality of images of a video stream into segments. Each segment may have one or more characteristics that are substantially constant.

[0612] The segmenting may include applying blob analysis, where the segments are blobs. The segmenting may apply a segmentation method different from blob analysis.

[0613] Step 3510 may be followed by step 3520 of determining a temporal quality of the segment.

[0614] Step 3520 may be followed by step 3530 of classifying each segment as a background segment or a foreground segment based at least in part on a temporal characteristic of the segment.

[0615] Step 3530 may include at least one of the following: Classifying static segments as background segments. b. Classifying segments that show periodic changes as background segments. c. Searching for one or more facial segments. d. Classifying each face segment as a foreground segment. e. Classifying segments that show periodic changes, rather than facial segments, as background segments. f. using a machine learning process to classify each segment as a background segment or a foreground segment, the machine learning process being trained to perform the classification based on classification input received from a user. g. Classifying based at least in part on user feedback. h. Displaying at least one user segment of the image and receiving a classification input from the user relating to at least a portion of the segment, the classification being also based on the classification input. i. classifying as a foreground segment one or more items added to the virtual 3D video conferencing environment that are visible to at least one participant of the virtual 3D conference;

[0616] FIG. 27 also illustrates a method 3501 for foreground and background segmentation in connection with virtual three-dimensional (3D) videoconferencing.

[0617] The method 3501 may begin by segmenting 3510 each image of a plurality of images of a video stream into segments. Each segment may have one or more characteristics that are substantially constant.

[0618] The segmenting may include applying blob analysis, where the segments are blobs. The segmenting may apply a segmentation method different from blob analysis.

[0619] Step 3510 may be followed by step 3520 of determining a temporal quality of the segment.

[0620] The method 3520 may be followed by a step 3525 of providing user information and receiving feedback from the user.

[0621] Step 3525 may include at least one of the following: a. Providing temporal information to the user regarding the temporal nature of the segment. b. Receiving feedback from a user, such as a classification input related to at least a portion of the segment. c. Displaying the segments to a user and providing the user with temporal information regarding the temporal nature of the segments. d. Receiving feedback from a user, such as a classification input related to at least a portion of the segments.

[0622] Step 3525 may be followed by step 3535 of classifying each segment as a background segment or a foreground segment based at least in part on the feedback. The feedback may include, for example, a classification input.

[0623] Step 3535 may be responsive to the feedback and to the temporal characteristics of the segment. Step 3535 may include any of the sub-steps of step 3530, each of which may be modified based on the feedback or have its output and possible feedback from the user.

[0624] FIG. 38 shows an example of image segments into foreground and background.

[0625] Image 3490 captures a person 3493, a ventilator 3494, and a grey wall. When in operation, the ventilator may perform periodically changing movements and can be considered to belong to the background 3492. The person itself forms the foreground 3491.

[0626] Touch-up - Noise removal make-up Existing video conferencing systems such as Zoom and Microsoft Teams allow participants to add "filters" that enhance or otherwise modify their appearance. For example, it is possible to add makeup such as lipstick or blush, it is possible to add gadgets such as glasses, it is possible to appear as if adding a moustache and beard, modifying hair color and style, etc.

[0627] Such filters can be used only to not add makeup or gadgets. They can also be utilized to tweak the appearance of the participant (such a feature known as photoshopping) and denoise (to reduce noise added by the camera, lighting conditions).

[0628] A need exists for providing an accurate and efficient (in terms of memory resource utilization and / or computational resource utilization) method for determining the appearance of a participant within a virtual 3D environment.

[0629] Performing segmentation on a frame-by-frame basis to identify face parts to enhance them is very inefficient and suffers from noise introduced in the image. For example, in each frame, lips are identified and then the associated color of lipstick is applied. Similarly, the chin is detected and possibly the inclination of the face and moustache is placed on top of it at the correct angle. This is a costly operation. Especially if a person chooses to add lipstick, blush, glasses, moustache and also retouch the hair color, this requires that detection of all relevant face parts needs to be done 10 times per second (depending on the frame rate, which is typically 30 times per second or more). Once the parts are detected, touch-ups and makeup are added frame by frame.

[0630] This costly action also limits the possibility of denoising a participant's appearance. The main reason for doing it this way is that the system does not maintain a model of a particular participant's face.

[0631] A method is provided for participants to appear in the meeting environment through avatars. Any of the methods described above for generating such a representation may be used.

[0632] A 3D model of the participant, or at least the participant's head and face and / or torso, may be obtained, and this model (and one or more texture maps) may be manipulated or used as a basis to create an avatar of the participant.

[0633] Different facial parts of the participant are an integral part of the 3D model.

[0634] To add touch-ups and makeup, the 3D model can be updated once. For example, a chosen color is added to the lips in lipstick. As with other parts of the 3D model, the voxels corresponding to the lips then have a reflectance associated with it, and as the avatar is rendered, the reflectance allows for a realistic appearance of the lips. Similarly, any chosen color is applied to the cheeks to appear as rouge. To make this appear more realistic, the chosen color can be combined linearly or otherwise, in intensity or spatially, with the original skin color, so that it appears as if the model really does have rouge on its cheeks. Then, when the model is manipulated to create the avatar, all the additions are ready in the appropriate places.

[0635] Moreover, this method allows for easy noise removal and "photoshopping." Noise removal is possible because the model should be insensitive to noise introduced by the camera, lighting, or other sources. As the model's existence is ongoing, it can be easily made clearer of noise introduced during the capture of a single image by the camera by averaging the reflectance values ​​of each point in the model over time.

[0636] Such "photoshopping" of facial features (such as correcting the nose, lifting cheekbones, removing a "double chin") is performed on a 3D model instead of performing those actions over and over again frame by frame. Once the 3D model is created, all the effects are performed on the model. In other words, the model's cheekbones are lifted and its double chin is removed. Those adjustments are then noted, and whenever a new image is captured by the camera, all that is needed to create a new avatar is to understand the person's new location, orientation, and gaze. They are then applied to the adjusted 3D model.

[0637] FIG. 29 illustrates the method 3600.

[0638] Method 3600 refers to a first participant and a second participant for ease of explanation. The first and second participants described above may be any pair of participants. Any step of method 3600 may be applied to any combination of participants.

[0639] The method 3600 may include an initialization step 3602 .

[0640] The initialization step 3602 may include receiving, by a user device of the first participant of the virtual 3D video conference, 3D representation information of a reference second participant for generating a 3D representation of the second participant under different constraints, where the different constraints may include at least one from (a) a touch-up constraint, (b) a make-up constraint, and (c) one or more situation constraints.

[0641] Constraints, eg, makeup and / or touch-ups, may be provided even when the actual participant is not actually wearing the makeup defined in the makeup constraint.

[0642] The at least one other constraint may be determined by other means, such as image analysis, and the like.

[0643] An example of a situation constraint is illustrated in method 3200 .

[0644] For example, different constraints may include different gaze directions of the second participant, different facial expressions of the second participant, different lighting conditions, different fields of view of the camera, and the like.

[0645] The initial 3D participant representation information may include an initial 3D model and one or more initial texture maps.

[0646] The 3D participant representation information may include a 3D model and one or more texture maps.

[0647] The initial second participant 3D representation information may represent a modified representation of the second participant. It is "modified" in the sense that the modified representation differs from the second participant's actual appearance. The modified representation differs from the second participant's actual appearance by at least one of the following: size, shape, and position of facial elements.

[0648] The method 3600 may include a step 3610 of receiving, by the first participant's user device, second participant constraint metadata indicating one or more current constraints with respect to the second participant during the 3D video conference call.

[0649] Step 3610 may be followed by step 3620 of updating, by the user device of the first participant, a 3D participant representation of the second participant within the first representation of the virtual 3D videoconferencing environment based on the constraint metadata of the second participant.

[0650] Step 3620 may be followed by step 3630 of generating an avatar for the second participant based on the 3D participant representation information of the second participant.

[0651] Step 3630 may include step 3632 of generating a makeup version of the facial element based on the makeup-free appearance of the facial element and the selected makeup, such that the selected makeup may be virtually added or placed over or otherwise integrated with the makeup-free facial appearance of the facial element.

[0652] Step 3632 may include generating a makeup version of the facial element by applying a linear function to the makeup-free appearance of the facial element and the selected makeup voxels.

[0653] The make-up free version may be replaced by any referring representation of the second participant, which may be modified according to one or more make-up constraints.

[0654] The method 3600 may include obtaining 3670 updated reference second participant 3D representation information for generating an updated second participant 3D representation under the different constraints. The updated reference second participant 3D representation information may replace the initial reference second participant 3D representation under the different constraints.

[0655] The updated reference second participant's expression information may be generated by performing noise removal.

[0656] There may be multiple representations of the virtual 3D videoconferencing environment (one for each participant), and steps 3630 and / or 3640 may be performed for each one of the representations. The updates themselves (including visual information, e.g., the appearance of people) may differ from one representation to another.

[0657] Participants of a virtual 3D conference are associated with multiple participant devices, which may be different computerized systems than any of the multiple participant devices.

[0658] Various steps of the method 3600 may be performed by at least one of the computerized systems and one or more of the multiple participant devices.

[0659] FIG. 30 illustrates an example of a participant without lipstick, FIG. 34 illustrates an example of a participant with lipstick, FIG. 35 illustrates an example of a participant avatar without lipstick, FIG. 36 illustrates an example of a lipstick-free representation of a participant's lips, and FIG. 37 illustrates an example of a participant avatar with lipstick.

[0660] The absence of lipstick or the required addition of lipstick can be learned from the participant's image and transmitted as a constraint to the other participant's device. Additionally or alternatively, a participant may request that their 3D representation be updated by adding and / or removing lipstick regardless of the actual state of their lips.

[0661] A participant may, for example, request to add any wearable items to the avatar that the participant does not actually wear, remove from the avatar any wearable items that the participant actually wears, and / or introduce any required changes in the participant and his / her surroundings (the actual appearance of wearable items, jewelry, accessories, as well as the participant's avatar (in any manner, from the participant's device or from any other device or system)).

[0662] Improved audio quality in video conferencing It is important that participants can hear each other well and clearly in video conferences because these settings are less natural and typically require more focused attention on the part of the participants than face-to-face meetings. Nonetheless, background noise is often audible during online meetings. In other cases, problems with microphones or other system components reduce the quality and clarity of what is being said, reducing the effectiveness of such meetings.

[0663] Noise cleaning methods exist today. Some solutions, such as Krisp, clean up non-human voices. This particular application is installed on the client side of the video conference. In other words, participants who do not have it installed do not get its benefits. Meanwhile, the noisy or unclear soundtrack is transmitted to all participants.

[0664] The proposed method utilizes image and video processing to enhance the audio within a video conference, which is entirely possible because in a video conference environment, participants typically have cameras that see and capture them.

[0665] In short, the enhancement is performed by visually analyzing the participant's mouth, lip, and tongue movements, or a subset of these that may appear on a camera viewing the speaker.

[0666] Using machine learning techniques, the system is trained to learn how their movements respond to different sounds. This training can be performed by neural networks or other methods.

[0667] Training can be performed on whole words and sentences. Additionally or alternatively, it can be performed on only a subset of "sounds." For example, in the English language, it is universally agreed that there are 44 phonemes or distinct sounds, with some variations based on stress and articulation.

[0668] When such a system sees a speaking video conference participant, it can make educated assumptions about the sounds the speaker is making. These assumptions can then be used in two ways: A. Clarifying background noise by removing sounds that do not appear to come from a speaker. b. Improving the quality of audio transmitted from the system, for example, when a participant's microphone is not working or even if it is muted by mistake (force unmuting may be an optional setting in a video conference and can be set separately by each participant and / or by the meeting host).

[0669] These audio corrections may be performed in the system of speakers or in a central location, depending on available resources or based on other considerations.

[0670] It is also possible to have the participant say some word or some sound in order for the system to calibrate itself to a particular participant. This can be done only once, at the beginning of the meeting, when the participant joins them, or once every number of meetings.

[0671] FIG. 31 illustrates a method 3700 for improving audio quality associated with participants in a virtual three-dimensional (3D) video conference.

[0672] The method 3700 may begin by step 3710 determining, by a machine learning process, audio generated participants based on image analysis of video of the participants captured during the virtual 3D video conference.

[0673] The machine learning process may be trained to convert image analysis output into a participant's generated audio. The machine learning process may be trained to convert video into a participant's generated audio.

[0674] The method may include training a machine learning process or receiving a trained machine learning process.

[0675] Step 3710 may be followed by step 3720 of generating associated audio information for the participant based at least on the generated audio of the participant. The participant's associated audio information, when provided to another participant's computerized system, causes the other participant's computerized system to generate associated audio for the participant that is of higher quality than the participant's audio when the participant's audio is included in detected audio detected by an audio sensor associated with the participant.

[0676] Step 3720 may include at least one of the following: determining one or more audio processing features of an audio processing algorithm and applying the audio processing algorithm to the detected audio, The one or more audio processing features may be any time-domain and / or spectral-domain audio parameters, such as a desired spectral range of the associated audio of the participant; b. Applying an audio processing algorithm, which may include a filtering step Applying an audio processing algorithm may include filtering the detected audio. c. Applying a noise reduction algorithm to the detected audio. d. Applying a speech synthesis algorithm.

[0677] The determining step 3710 may be applied even when the participant's audio sensors (eg, microphone) are muted.

[0678] Step 3710 may be preceded by or may include determining when the audio sensor is muted. The determination regarding the mute state of the audio sensor may be based on a comparison between an output of the audio sensor and an estimated audio output by the user based on an image analysis of the video of the participant.

[0679] When the audio sensor determines that the audio is to be muted, step 3720 may include applying a speech synthesis algorithm.

[0680] Step 3720 may include step 3722 of determining how to generate relevant audio information for the participant based on at least one of the presence and quality of the detected audio.

[0681] Step 3722 may include selecting between (i) applying an audio processing algorithm to the detected audio and (ii) applying a speech synthesis algorithm.

[0682] prediction In a virtual 3D video conference, participants may appear as avatars or have any other 3D representation.

[0683] This may involve creating 3D models of participants. During the meeting, participants sit in front of a camera. They capture their movements and some analysis is performed to discover the participants' posture, orientation, and facial expressions. Then, for each viewer of the meeting, an avatar of the participant is created, so that the avatar's posture, orientation, and facial expressions appear in the viewer's field of view as it would if the participant were physically located in the meeting environment. This real-time processing can be seen as having two components: one that performs the analysis of the participants and the other that performs the rendering.

[0684] The two components may or may not be co-located. For example, the analysis may need to be performed only once per participant, but the rendering may need to be performed multiple times, once per viewer. Thus, one option is to have the analysis performed at the participant's location or at a central location, and the rendering, or parts of it, may be performed at each viewer's location. The analysis component needs to inform the rendering component of changes in posture, orientation, and facial expression, so that the rendering component renders the avatar accurately.

[0685] In order to increase efficiency, reduce the possibility of error, and conserve resources, it is important to reduce the amount of communication between these two components while maintaining a high degree of reliability.

[0686] Any changes in motion or other characteristics can be expected even over short periods of time.

[0687] Consider the following simplified example: Assume a participant in a video conference nods, and assume that an image is captured every 33 milliseconds, which is also the interval between rendering of the participant's avatar in the meeting's viewer system. When the participant's head is moving upwards, this movement is assumed to continue for at least several hundred milliseconds, say 200 milliseconds.

[0688] Under those assumptions, if the rendering component is able to predict that this motion is occurring, it may be able to continuously render this motion without receiving any additional information from the analysis component, for example, as long as the prediction is accurate, at least for a short period of time. If the actual motion differs from the predicted motion, the analysis component may only need to update the rendering unit to the predicted motion with corrections. Those corrections contain much less information than the actual motion information. This therefore allows for a lot of savings in communication.

[0689] For example, assume there is no prediction at the client. The server needs to send all values. For example, every frame the orientation should change by 1 degree upwards. If the client does not have prediction capabilities, the server only needs to send the correction. For example, the client predicted by 1 degree upwards, but in reality the change is 1.0001 degrees, so the client only needs to send a value of 0.0001.

[0690] If the prediction is "generally" good, then a correction, if made, is of a lower order of magnitude than a perfect prediction.

[0691] For example, if one predicts a value of 100, but the actual value turns out to be 101, then the correction is simply 1. Because the corrections typically have a much smaller value than the prediction, they can be coded with fewer bits. If the corrections are large, but they are made infrequently, using Huffman coding or arithmetic coding allows for more communication bits.

[0692] If this is not the case, in other words if the correction is of the same magnitude as the prediction, this actually means that there is no prediction.

[0693] Machine learning systems can be trained to learn how to predict posture, orientation, and facial expressions based on their recent history. Predicting their near future can be performed for each of their histories separately, or based on any combination of their histories. Predictive models can be learned for each participant separately, or for the "whole" participants.

[0694] For example, an RNN neural network or an LSTM neural network may receive the pose, orientation, and expression values ​​at any given time and learn to predict the next values.

[0695] This is very similar to how a NN can be taught how to compose text by studying existing text, or by creating music by studying musical sequences.

[0696] Once the model is learned, it is shared with the analysis and rendering components.

[0697] The third component is that the decider may decide among three options: a. Having the analysis component transmit all data to the rendering component. b. Having the rendering component render based solely on the predictive model. c. Having a rendering component render based on the predictive model along with corrections sent by the analytical model.

[0698] For ease of explanation, it is assumed that the analysis component and the decision maker are in a first computerized unit and the rendering component is in a second computerized unit.

[0699] The determination can be made by setting a threshold for the amount of data that needs to be transmitted, or by setting a threshold for the number of consecutive times that a correction needs to be transmitted, or a combination thereof.

[0700] The analysis component is aware of the predictive model used by the rendering component, so it can evaluate what the rendering component would be doing if it were rendering based on the predictive model.

[0701] FIG. 32 illustrates a method 3800 for predicting changes in participant behavior in a virtual three-dimensional (3D) video conference.

[0702] The prediction may reduce the volume of traffic between computerized units.

[0703] The method 3800 may be an iterative method. Each iteration may use one behavior predictor, one may need to use another behavior predictor, and the next iteration begins. Each iteration is applied to a portion of the virtual 3D video conference.

[0704] It is assumed that a first computerized unit performs the various steps of the method 3800 and can be considered as an analyzer and / or a transmitter.

[0705] The second computerized unit may receive the information generated by the first computerized unit and may display (or cause a display to show) representations of the participants within the virtual 3D videoconferencing environment.

[0706] The second computerized unit may be considered as a receiver.

[0707] A first computerized entity may have access to the participants' videos, which are captured during the virtual 3D video conference, and a second computerized entity may not have access to the videos.

[0708] The first computerized entity may be an image analyzer. The second computerized entity may be a rendering unit.

[0709] Each one of the first computerized unit and the second computerized unit may be a participant device and a computerized system other than any participant device, etc.

[0710] The method 3800 may begin by performing the following steps for each of the multiple portions of the virtual 3D video conference: A step 3810 of determining, by the first computerized unit, a participant behavior predictor to be applied by the second computerized unit during the portion of the virtual 3D video conference. Any determination or selection method may be applied, including any method of finding the best estimator, good estimator, etc. may be used. b. Determining 3820 one or more prediction inaccuracies associated with applying the participant behavior predictors during portions of the virtual 3D video conference. c. Step 3830 of determining whether to generate and transmit to a second computerized unit prediction inaccuracy metadata indicating at least one prediction inaccuracy affecting a representation of the participant within the virtual 3D video conference environment as presented by another participant of the virtual 3D video conference during a portion of the virtual 3D video conference.

[0711] Step 3830 may be followed by a step 3840 of generating and transmitting prediction inaccuracy metadata to the second computerized unit when it is determined to generate and transmit prediction inaccuracy metadata to the second computerized unit.

[0712] Participant behavior predictors may be determined and transmitted at the start of the portion or after the portion has started.

[0713] One or more prediction inaccuracies may be generated in real time and transmitted to a second computerized unit (if so determined) to enable real time corrections to the participant's representation.

[0714] Step 3840 may include transmitting an end of portion indicator and / or an identifier of the next behavior predictor, etc.

[0715] Step 3480 may include sending to the second computerized entity information regarding the participant behavior predictor to be applied by the second computerized unit.

[0716] Step 3810 may be based on the participant's behavior during the previous portion of the virtual 3D video conference.

[0717] Step 3810 may include determining when a portion ends and a new portion begins based on one or more prediction inaccuracies associated with applying the participant's behavior predictors during the portion.

[0718] Step 3810 may include determining when a portion ends and a new portion begins based on one or more prediction inaccuracies associated with applying the participant's behavior predictors during the portion.

[0719] For example, a determination may be made when the size of the transmitted information related to prediction inaccuracy (Spi) exceeds a threshold, when Spi exceeds the size of "direct" behavior information (Sdbi) that directly exemplifies the participant's behavior (without prediction), when the accuracy of the currently used participant's behavior predictor falls below a threshold, etc.

[0720] Step 3830 may be based on an effect of at least one prediction inaccuracy on a participant's expression.

[0721] At least one of steps 3810, 3820, 3830, and 3840 may be performed by a machine learning process.

[0722] Steps 3810, 3820, 3830, and 3840 may be performed by a first computerized unit.

[0723] The method 3800 may include a step 3850 of determining, by the second computerized unit, in each portion, a participant behavior predictor to be applied by the second computerized unit.

[0724] Step 3850 may be followed by step 3860 of applying, by a second computerized unit, a participant behavior predictor in each portion, the application being influenced by prediction inaccuracy information received in real time from the first computerized unit.

[0725] 33 illustrates three periods in a virtual 3D video conference 4201, 4202, and 4203. During the first period 4201, two participants 4211 and 4212 are in a location (as illustrated in image 4215) and both are looking at a computer display. The participants move and change their gaze direction (during the second period 4202) and thus may remain in the latter location during the third period 4203 until they see each other as illustrated in image 4216.

[0726] The first behavior predictor 4241, which was accurate during the first period 4201, is no longer accurate when the participants start moving and thus, at the end of the first period or slightly thereafter (as shown in FIG. 33), the first portion 4231 of the virtual 3D video conference may end (and the second portion 4232 may begin) and the second behavior predictor 4242 may be used.

[0727] The second behavior predictor 4242, which was accurate during the second period 4202, is no longer accurate when the participant stops moving and thus, at the end of the second period or slightly thereafter (as shown in FIG. 33), the second portion 4232 may end (and the third portion 4233 may begin) and the third behavior predictor 4243 may be used.

[0728] At least some of the methods described above may be applicable, subject to modification, to 2D video conferencing.

[0729] In the foregoing specification, the disclosed embodiments have been described with reference to specific examples of the disclosed embodiments. It will, however, be apparent that various modifications and changes can be made thereto without departing from the broader spirit and scope of the disclosed embodiments as set forth in the appended claims.

[0730] Moreover, the terms "front," "back," "top," "bottom," "over," and "under," etc. in the description and claims, when present, are used for descriptive purposes and not necessarily for describing permanent relative positions. Terms so used are interchangeable under appropriate circumstances, such that the disclosed embodiments described herein are capable of operating in orientations other than those illustrated or otherwise described herein, for example.

[0731] A connection as discussed herein may be any type of connection suitable for transferring signals from or to a respective node, unit, or device, for example, via an intermediate device. Thus, unless otherwise implied or stated, a connection may be, for example, a direct connection or an indirect connection. A connection may be illustrated or described in the reference as being a single connection, multiple connections, unidirectional connections, or bidirectional connections. However, different embodiments may vary the implementation of a connection. For example, separate unidirectional connections may be used rather than bidirectional connections, and vice versa. Also, multiple connections may be replaced with a single connection that transfers multiple signals in a serial or time-multiplexed manner. Similarly, a single connection carrying multiple signals may be separated into various different connections carrying subsets of those signals. Thus, there are many options for transferring signals.

[0732] Any arrangement of components to achieve the same functionality is effectively associated such that the desired functionality is achieved. Thus, any two components combined herein to achieve a particular functionality, regardless of architecture or intermediate components, may be viewed as being "associated" with one another such that the desired functionality is achieved. Similarly, two components so associated may also be viewed as being "operably connected" or "operably coupled" with one another such that the desired functionality is achieved.

[0733] Moreover, those skilled in the art will recognize that the boundaries between operations described above are merely exemplary. Multiple operations may be combined into a single operation, a single operation may be distributed into additional operations, and operations may be performed with partial overlap in time. Moreover, alternative embodiments may include multiple instances of a particular operation, and the order of operations may be rearranged in various other embodiments.

[0734] Also for example, in one embodiment, the illustrated examples may be implemented on a single integrated circuit or within the same device. Alternatively, the examples may be implemented as any number of separate integrated circuits or separate devices interconnected with each other in any suitable manner.

[0735] However, other modifications, variations and adaptations are possible, and the specification and drawings are accordingly to be regarded in an illustrative rather than a restrictive sense.

[0736] In the claims, any reference signs placed between parentheses shall not be construed as limiting the scope of the claim. The word "comprising" does not exclude the presence of other elements or steps than those recited in the claim. Moreover, the terms "a" or "an" are defined as one or more than one as used herein. Also, even when the same claim includes the introductory phrase "one or more" or "at least one" and includes an indefinite article such as "a" or "an", the use of introductory phrases such as "at least one" and "one or more" in the claims shall not be construed as implying that the introduction of another claim element by the indefinite article "a" or "an" limits any particular claim that includes the claim element so introduced to disclosed embodiments that include only one such element. The same applies to the use of definite articles. Unless otherwise stated, terms such as "first" and "second" are used to arbitrarily distinguish between the elements that such terms describe. As such, such terms are not necessarily intended to indicate a chronological or other priority of such elements. The rare fact that certain measures are recited in mutually distinct claims does not indicate that a combination of those measures cannot be used to advantage.

[0737] While certain features of the disclosed embodiments have been illustrated and described herein, many modifications, substitutions, changes, and equivalents will occur to those skilled in the art herein, and it is therefore to be understood that the appended claims are intended to cover all such modifications and changes as fall within the spirit of the disclosed embodiments.

Claims

1. 1. A method for conducting a three-dimensional (3D) video conference among a plurality of participants, comprising: acquiring visual information by a visual sensing unit associated with a participant; identifying a plurality of people appearing in the visual information; Finding at least one associated person from the plurality of persons; determining 3D entity representation information for each of the at least one associated person; generating, for at least one participant, a representation of the virtual 3D videoconferencing environment based on the 3D entity representation information for each of the at least one associated person; A method comprising:

2. The method of claim 1 , wherein the discovering comprises determining which of the plurality of people are participants in the virtual 3D video conference.

3. The method of claim 1 , wherein the discovering comprises determining that a non-participant in the 3D video conference is an associated person.

4. The method of claim 1 , wherein the identifying step comprises applying a facial recognition process.

5. The method of claim 1 , comprising storing, for at least a period of time, identifying information about the at least one associated person according to the certain participant and the person's appearance.

6. 6. The method of claim 5, further comprising identifying any of the at least one associated person after the at least one associated person exits a field of view of the visual detection unit and re-enters the field of view of the visual detection unit, the identifying being based on the identification information.

7. The method of claim 1 , wherein a plurality of participants participate in the virtual 3D conference, and the plurality of participants are detected by a plurality of visual detection units.

8. 1. A method for conducting a three-dimensional (3D) video conference among a plurality of participants, comprising: receiving gaze direction information regarding each participant's gaze direction within a representation of a virtual 3D videoconferencing environment associated with the participant; estimating whether a participant's gaze is directed towards a person located within a field of view of a visual detection unit that also captures at least the head of the participant; determining whether a 3D representation of the person should appear within the virtual 3D videoconferencing environment; determining, for each participant, updated 3D participant representation information within the virtual 3D videoconferencing environment that reflects the gaze direction of the participant, wherein for the given participant, determining the updated 3D participant representation information is responsive to results of the estimating and determining; generating an updated representation of the virtual 3D videoconferencing environment for at least one participant, the updated representation of the virtual 3D videoconferencing environment representing the updated 3D participant representation information for at least a portion of the plurality of participants; A method comprising:

9. The method of claim 8 , wherein the determining includes checking whether the person is one of the participants.

10. The method of claim 9 , wherein when determining that the person is one of the participants, searching for physical interactions between the person and a participant.

11. 11. The method of claim 10, wherein (a) determining the updated 3D participant representation information for the one participant and (b) determining the updated 3D participant representation information for the person reflects the physical interaction.

12. 10. The method of claim 9, further comprising: (a) determining that the 3D representation of the person should appear in the virtual 3D video conferencing environment; and (b) generating 3D person representation information when determining that the person is not one of the participants, wherein the updated representation of the virtual 3D video conferencing environment further includes the 3D person representation information.

13. 10. The method of claim 9, comprising maintaining the gaze direction of the one participant within the virtual 3D video conferencing environment unchanged during a change in the gaze direction of the one participant from the person to a visual representation of the person within the virtual 3D video conferencing environment.

14. The method of claim 9 , comprising receiving initial 3D participant representation information for generating the 3D representation of the participant under different conditions.

15. 1. A method for foreground and background segmentation in connection with virtual three-dimensional (3D) video conferencing, comprising: Segmenting each image of a plurality of images of a video stream into segments, each segment having one or more characteristics that are constant; determining a temporal characteristic of the segment; classifying each segment as a background segment or a foreground segment based at least in part on the temporal characteristics of the segment; A method comprising:

16. The method of claim 15, comprising classifying static segments as background segments.

17. 16. The method of claim 15, further comprising: displaying at least one user segment of the image; and receiving classification input from the user relating to at least a portion of the segment, wherein the classification is also based on the classification input.

18. 16. The method of claim 15, further comprising: providing temporal information to a user regarding a temporal characteristic of the segment; and receiving a classification input from the user relating to at least a portion of the segment, wherein the classification is also based on the classification input.

19. 16. The method of claim 15, further comprising: displaying the segment to a user; providing the user with temporal information regarding a temporal aspect of the segment; and receiving a classification input from the user relating to at least a portion of the segment, the classification also being based on the classification input.

20. 16. The method of claim 15, wherein said classifying is followed by classifying one or more items as foreground segments to be added to a virtual 3D video conference environment displayed to at least one participant of said virtual 3D conference.