Generating a 3D representation of a participant's head in a video communication session - Patent Application 20070233633

JP2025513707A5Active Publication Date: 2025-06-09TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024555182
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-03-14
Publication Date
2025-06-09
Estimated Expiration
2042-03-14

AI Technical Summary

Technical Problem

Existing methods for generating a 3D representation of a participant's head in video communication sessions either rely on complex and costly 3D sensors for real-time capture or use animated avatars that may not accurately represent head movements.

Method used

A computing device that collects 3D representations of a participant's head, identifies facial landmarks, determines head posture, and generates an avatar representation for the outer portion of the head using machine learning models, while streaming only the inner face portion for real-time display.

Benefits of technology

This approach reduces bandwidth requirements, improves user experience by maintaining detailed face structure and movement, and eliminates issues with incomplete head captures, all while simplifying the hardware requirements for 3D sensing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computing device for generating a three-dimensional (3D) representation of a head of a participant in a video communication session is provided. The computing device comprises a processing circuit that causes the computing device to be operable to collect a captured 3D representation (300) of the head and to identify positions (1-27) of a set of facial landmarks in the captured 3D representation (300). The set of facial landmarks includes facial landmarks that indicate boundaries of a human face. The computing device is further operable to determine a pose of the head and to determine a boundary (310) between an inner portion (320) and an outer portion (330) of the captured 3D representation (300) based on the identified positions (1-27) of the set of facial landmarks. The inner portion (320) of the captured 3D representation represents the participant's face. The computing device is further operable to generate an avatar representation corresponding to the outer portion (330) of the captured 3D representation (300) using a machine learning (ML) model trained on human heads with the determined pose of the head as input.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a computing device for generating a three-dimensional (3D) representation of the head of a participant in a video communication session, a method for generating a 3D representation of the head of a participant in a video communication session, a corresponding computer program, a corresponding computer-readable data carrier and a corresponding data carrier signal. [Background technology]

[0002] For example, various alternatives are known for generating and displaying to a viewer a three-dimensional (3D) representation of a human head during a video communication session between two or more participants.

[0003] The first type of solution is based on computer-generated 3D avatars. Such avatars are generated using a machine learning (ML) model that is trained using a captured 3D representation of a human head and then adapted to that particular head using a two-dimensional (2D) image representing that head. During use, the head pose of the sending participant's head is continuously detected and used as input to the ML model, which generates a dynamically animated avatar that reflects the actual movement of the sending participant's head. In recent years, considerable improvements in generated 3D avatars have been achieved through the use of generative adversarial networks (GANs), as demonstrated by H. Luo et al. ("Normalized Avatar Synthesis Using StyleGAN and Perceptual Refinement", arXiv:2106.11423, arXiv, 2021).

[0004] A second type of solution relies on capturing the sending participant's head using a 3D sensor, such as a stereo camera, and transmitting the captured 3D representation in real time for display to the receiving participant, e.g., as a point cloud or mesh stream. Solutions based on real time capture and transmission excel in representing the details of the captured head compared to animated 3D avatars. These details, especially the captured facial details, are important for conveying the emotions of the sending participant. However, transmitting the 3D captured representation of the head in real time requires a significantly larger bandwidth of the communication link used to transmit the captured 3D data, e.g., as a point cloud or mesh stream. Furthermore, today's solutions that rely on capturing the sending participant's head in real time generally require relatively complex 3D sensors, such as several cameras, to ensure that the outer part of the head (outside the face) is fully captured as the sending participant's head moves, which increases the complexity and cost. Summary of the Invention

[0005] It is an object of the present invention to provide an improved alternative to the above techniques and to the prior art.

[0006] More specifically, it is an object of the present invention to provide an improved solution for generating a 3D representation of a human head based on capturing the head using a 3D sensor device.

[0007] These and other objects of the invention are achieved by the various aspects of the invention, which are defined by the independent claims.Embodiments of the invention are characterized by the dependent claims.

[0008] According to a first aspect of the present invention, there is provided a computing device for generating a 3D representation of a head of a participant in a video communication session. The computing device comprises a processing circuit that causes the computing device to be operable to collect a captured 3D representation of the head and to identify a position of a set of facial landmarks in the captured 3D representation. The set of facial landmarks includes a facial landmark indicating a boundary of a human face. The computing device is further operable to determine a pose of the head and to determine a boundary between an inner part and an outer part of the captured 3D representation. The boundary is determined based on the identified positions of the set of facial landmarks. The inner part of the captured 3D representation represents the participant's face. The computing device is further operable to generate an avatar representation corresponding to the outer part of the captured 3D representation. The avatar representation is generated using an ML model trained on a human head. The determined pose of the head is used as an input to the ML model.

[0009] According to a second aspect of the present invention, there is provided a method for generating a 3D representation of a head of a participant in a video communication session. The method is implemented by a computing device and includes collecting a captured 3D representation of the head and identifying a position of a set of facial landmarks in the captured 3D representation. The set of facial landmarks includes facial landmarks indicative of a boundary of a human face. The method further includes determining a pose of the head and determining a boundary between an inner portion and an outer portion of the captured 3D representation. The boundary is determined based on the identified positions of the set of facial landmarks. The inner portion of the captured 3D representation represents the participant's face. The method further includes generating an avatar representation corresponding to the outer portion of the captured 3D representation. The avatar representation is generated using an ML model trained on a human head. The determined pose of the head is used as an input for the ML model.

[0010] According to a third aspect of the invention there is provided a computer program comprising instructions which, when executed by a computing device, cause the computing device to perform a method according to an embodiment of the second aspect of the invention.

[0011] According to a fourth aspect of the invention there is provided a computer readable data carrier storing a computer program according to the third aspect of the invention.

[0012] According to a fifth aspect of the present invention there is provided a data carrier signal, the data carrier signal carrying a computer program according to the third aspect of the present invention.

[0013] The present invention exploits the understanding that a 3D representation of a participant's head in a video communication session may be generated based on extracting an inner portion of a captured 3D representation of the head and making the inner portion available for display to a receiving user, e.g., as a real-time stream. The extracted inner portion generally corresponds to the facial region, or face, of the head, including the eyes, nose, ears, and mouth. The remainder of the captured 3D representation, referred to herein as the outer portion, represents the portion of the head outside the face. This outer portion is replaced by an avatar that is generated using an ML model trained on a human head, using the pose of the captured head as input for the ML model. This results in an animated avatar that reflects the actual pose and movement of the captured head.

[0014] The embodiments of the present invention are advantageous in that a receiving user viewing the generated 3D representation of the captured head is not bothered by an incompletely captured 3D representation of the head, which may occur due to limitations in the 3D sensor used to capture the 3D representation. At the same time, the detailed structure and movement of the face and facial parts of the captured head, which are important for conveying emotions in inter-human communication, are preserved. Thereby, the user experience may be improved. The embodiments of the present invention are further advantageous in that the amount of data captured as a 3D representation of the head that needs to be transferred in real time to the receiving user, for example by streaming, is reduced. This is true since only a subset of the captured 3D data, i.e. the inner part of the captured 3D representation representing the face of the captured head, needs to be transferred from the 3D sensor to the display device. Thereby, the bandwidth required to conduct a 3D video communication session is reduced.

[0015] Although advantages of the present invention have been explained in some cases with respect to embodiments of the first aspect of the invention, corresponding reasoning applies to embodiments of the other aspects of the invention.

[0016] Further objectives, features, and advantages of the present invention will become apparent upon study of the following detailed disclosure, drawings, and appended claims. Those skilled in the art will appreciate that different features of the present invention may be combined to produce embodiments other than those described below.

[0017] The above, as well as additional objects, features and advantages of the present invention will be better understood through the following illustrative, non-limiting detailed description of embodiments of the invention, with reference to the accompanying drawings. [Brief description of the drawings]

[0018] [Figure 1] FIG. 1 illustrates a video communication session between two participants according to an embodiment of the present invention. [Diagram 2] FIG. 2 illustrates a captured 3D representation of a human head, according to an embodiment of the present invention. [Diagram 3] FIG. 1 illustrates the use of facial landmarks to determine the boundary between inner and outer portions of a captured 3D representation of a head, according to an embodiment of the present invention. [Figure 4] 2A-2C are schematic diagrams illustrating generating a 3D representation of a participant's head in a video communication session according to an embodiment of the present invention; [Figure 5A] 2 is a sequence diagram illustrating generating a 3D representation of a participant's head in a video communication session according to an embodiment of the present invention. [Figure 5B] 2 is a sequence diagram illustrating generating a 3D representation of a participant's head in a video communication session according to an embodiment of the present invention. [Figure 5C] 2 is a sequence diagram illustrating generating a 3D representation of a participant's head in a video communication session according to an embodiment of the present invention. [Figure 6] FIG. 2 illustrates a schematic diagram of a computing device for generating a 3D representation of a participant's head in a video communication session, according to an embodiment of the present invention. [Figure 7]2 illustrates a method for generating a 3D representation of a participant's head in a video communication session according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0019] All figures are schematic, not necessarily to scale, and generally show only those parts that are necessary to elucidate the invention; other parts may be omitted or merely suggested.

[0020] The present invention will now be described more fully hereinafter with reference to the accompanying drawings, in which several embodiments of the invention are shown. However, the present invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided as examples so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.

[0021] 1 shows a video communication session between two participants 101 and 103, illustrated as a one-way video communication session between a sending computing device 110 and a receiving computing device 130. During the video communication session, sometimes referred to as a 3D video communication session, a 3D representation of the head 102 of the sending participant 101 is captured using a 3D sensor 111, such as a stereo camera, and transmitted (e.g., streamed) over a communication network 140 to the receiving computing device 130 for display to the receiving participant 103 using a display device 131, such as a computer display or a head-mounted display (HMD). It should be noted that embodiments of the present invention are not limited to the one-way video communication session between two participants, as illustrated in FIG. 1. Rather, embodiments of the present invention may be envisioned that enable one-way (e.g., a presentation streamed from a presenter to many viewers) or two-way (e.g., a video call during a virtual conference) video communication sessions between two or more participants. More particularly, in the case of a two-way video communication session, embodiments of the invention also support generating a 3D representation of another participant's head in the opposite direction, the head of participant 103 in Figure 1. Thus, a computing device supporting two-way video communication sessions according to embodiments of the invention comprises both a 3D sensor 111 and a display device 131, or is operatively connected to both.

[0022] A computing device for generating a 3D representation of the head 102 of a participant 101 in a video communication session may be embodied in different forms, for example as a sending computing device 110, as a receiving computing device 130, as an edge computing device 120 provided at the edge of a communication network 140 through which traffic between the sending computing device 110 and the receiving computing device 130 passes, or as a combination thereof. The edge computing device 120 may be provided, for example, near a Radio Access Network (RAN) that is part of the communication network 140, through which the sending computing device 110 and / or the receiving computing device 130 communicate with each other and / or with the edge computing device 120.

[0023] The computing device for generating the 3D representation of the head 102 of the participant 101 in the video communication session, particularly when embodied as the sending computing device 110 or the receiving computing device 130, may be any one of a smartphone, a tablet, a laptop computer, an augmented reality (AR) device, a virtual reality (VR) device, a mixed reality (MR) device, an extended reality (XR) device, and an HMD. Alternatively, the computing device for generating the 3D representation of the head 102 of the participant 101 in the video communication session, particularly when embodied as the edge computing device 120, may be any one of an edge server, an application server, and a cloud computer. It will also be appreciated that the computing device for generating the 3D representation of the head 102 may be embodied in a distributed manner. That is, different operations involved in generating a 3D representation of the head 102, which will be described in further detail below, may be distributed and performed in a collaborative manner among two or more of the sending computing device 110, the edge computing device 120, and the receiving computing device 130. An illustrative example for distributing different operations in a collaborative manner among the sending computing device 110, the edge computing device 120, and the receiving computing device 130 is shown in Figures 5A-5C and will be elucidated in further more detail below.

[0024] Throughout this disclosure, embodiments of the invention are described with respect to generating a 3D representation of a human head 102, including a face. The human face includes the eyes, nose, ears, and mouth. In human-to-human communication, the detailed structure and movement of the face and facial parts are important for conveying emotions, for example, between participants 101 and 103.

[0025] Below and with reference to Figure 6, an embodiment 600 of a computing device (also referred to as a "computing device" for brevity) for generating a 3D representation of a head 102 of a participant 101 in a video communication session is described in more detail. Additional reference is made to Figure 4, which illustrates diagrammatically the flow of data in generating a 3D representation of a head of a participant in a video communication session.

[0026] The computing device 600 comprises a processing circuit 602 that causes the computing device 600 to be operable to collect a captured 3D representation of the head 102. This may be accomplished, for example, by capturing 502 a 3D representation of the head using the 3D sensor 111. Optionally, when the computing device 600 is embodied as a sending computing device 110, the computing device 110 may comprise the 3D sensor 111. Alternatively, the computing device 110 may be operatively connected to the 3D sensor 111. For example, the 3D sensor 111 may be a separate unit connected to the computing device 110 via an interface circuit ("I / O interface" in FIG. 6) using any wired or wireless technology known in the art, such as Universal Serial Bus (USB), Lighting, High Definition Multimedia Interface (HDMI), Bluetooth, etc. Alternatively, if the computing device 600 is embodied as an edge computing device 120 or as a receiving computing device 130, or a combination thereof, the computing device 120 / 130 may be operable to collect the captured 3D representation of the head 102 by receiving 512 the captured 3D representation, for example, as a data stream directly from the 3D sensor 111 (or indirectly via the sending computing device 110) via the communications network 140 using, for example, Real Time Protocol (RTP), Secure Real-time Transport Protocol (SRTP), or any other suitable protocol.

[0027] The 3D sensor 111 may comprise one or more of a 3D camera (also known as a stereo camera), an optical 3D sensor, a LiDAR, and a 2D camera. An optical 3D sensor may be used to capture and reconstruct the 3D depth of a real-world object, such as the head 102. Depending on the source of radiation used, optical 3D sensors may be divided into two categories, passive and active. Stereoscopic sensors, Shape-from-Silhouettes (SfS) sensors, and Shape-from-Texture (SfT) sensors are examples of passive 3D sensors that do not emit any kind of radiation themselves. The 3D sensor collects images of a scene, e.g., the head 102, optionally from different perspectives or with different lighting setups. The images are then analyzed to calculate the 3D depth of points in the captured scene, e.g., points representing the surface of the head 102 and parts of the head 102. Viewed differently, an active 3D sensor emits radiation, e.g., electromagnetic waves such as light, and the interaction between an object, such as the head 102, and the radiation is captured by the sensor. From the analysis of the captured data and based on the properties of the emitted radiation, coordinates of points in the captured scene, e.g., points representing the surface of the head 102 and parts of the head 102, are obtained. Time-of-Flight (ToF) sensors, phase-shift sensors, and active triangulation sensors are examples of active 3D sensors. The output of an optical 3D sensor is generally a depth map image.

[0028] LiDAR (Light Detection and Ranging) can be used to measure distance (aka "ranging") by illuminating a target, such as the head 102, with light and then measuring the reflection by a light sensor. LiDAR sensors can operate in the ultraviolet, visible, or infrared spectrum. Because commonly used laser light is collimated, LiDAR sensors need to scan the scene to generate an image with the desired field of view. The output of a LiDAR sensor is typically a point cloud, which can then be enriched with other sensor data, such as RGB data from a conventional (2D) camera, which may be included in the 3D sensor 111.

[0029] The processing circuit 602 further causes the computing device 600 to be operable to identify 503 the location of a set of facial landmarks (also known as "facial key points" or simply "key points") in the captured 3D representation. The set of facial landmarks includes facial landmarks that indicate the boundaries of a human face. Different sets of facial landmarks are used in the art. As an example, "Fast Facial Landmark Detection and Applications: A Survey" (by K.S.Khabarlak and L.S.Koriashkina, arXiv:2101.10808v2, arXiv, 2021) lists different sets comprising between 21 and 98 facial landmarks. In general, a subset of the facial landmarks of a given set indicates the boundaries of a human face. In Figure 3, an exemplary set of facial landmarks 1-27 that delineate the boundaries of a human face is reproduced as described in "Facial Landmarks for Face Recognition with Dlib" (Retrieved 2022-03-11, https: / / sefiks.com / 2020 / 11 / 20 / facial-landmarks-for-face-recognition-with-dlib / ). For illustrative purposes, facial landmarks 1-27 are overlaid on a sketch of a captured 3D representation 300 of head 102 in a representative position.

[0030] Using an ML model such as a neural network, a set of facial landmarks can be detected in a facial (2D) image and the (3D) locations of the facial landmarks can be determined. This can be accomplished using known facial landmark detection algorithms, for example, as described in "Fast Facial Landmark Detection and Applications: A Survey" using the Dlib library (see "Facial Landmarks for Face Recognition with Dlib"), or the OpenCV library (see, for example, "Head Pose Estimation using Python," https: / / towardsdatascience.com / head-pose-estimation-using-python-d165d3541600, retrieved on 2022-03-11).

[0031] The processing circuit 602 further causes the computing device 600 to be operable to determine 504 a pose of the head 102 (also referred to as a "head pose"). The determined pose of the head 102 may be expressed in terms of Euler angles, e.g., pitch, yaw, and roll, although embodiments of the present invention may also rely on alternative sets of angles. The pose of the head 102 may be determined 504 using a similar approach as described above with respect to identifying 503 the location of a set of facial landmarks, for example. For example, the head pose may be determined based on facial landmarks using the OpenCV library (see "Head Pose Estimation using Python"). This example illustrates how the head pose may be determined using only six facial landmarks, identifying the edges of the eyes, nose, chin, and mouth. The set of facial landmarks used for determining 504 the head pose may be different from the set of facial landmarks that indicate the boundaries of a human face. Alternatively, the set of facial landmarks that indicate the boundaries of a human face may comprise landmarks that indicate the pose of a human head. Different techniques for determining head pose from (2D) images of the head are known in the art, see for example "Fast Facial Landmark Detection and Applications: A Survey".

[0032] The processing circuit 602 further causes the computing device 600 to be operable to determine 505 a boundary between an inner portion and an outer portion of the captured 3D representation. The inner portion of the captured 3D representation represents the face of the participant 101 and may also be referred to as a facial portion of the captured 3D representation. The boundary is determined 505 based on the positions identified in 503 of facial landmarks, in particular a set of facial landmarks that indicate boundaries of a human face. In practice, the boundary between the inner portion and the outer portion of the captured 3D representation may be determined 505 by fitting a 2D shape, such as an oval shape, to the identified positions of a set of facial landmarks that indicate boundaries of a human face. As an example, an oval shape 310 that is fitted to a set of facial landmarks 1-27 is shown in FIG. 3. The boundary 310 separates the inner (facial) portion 320 from the outer portion 330 of the captured 3D representation 300.

[0033] As an alternative to fitting a 2D shape to the identified locations of a set of facial landmarks that indicate the boundaries of the human face, the boundaries between the inner and outer portions of the captured 3D representation may be determined in 505 by first fitting a 3D shape, such as an oval or ellipsoid, to the captured 3D representation of the head 102. The identified locations of the set of facial landmarks that indicate the boundaries of the human face are then projected onto the fitted 3D shape using either a surface normal or a projection to an origin or coordinate system used for the captured 3D representation, advantageously close to the center of the head 102. The projected locations of the facial landmarks are points on the surface of the fitted 3D shape. A 2D shape, such as an oval shape, is then fitted to the points on the surface of the fitted 3D shape.

[0034] It will be appreciated that embodiments of the present invention are not limited to using oval shapes in determining 505 the boundary between the inner and outer portions of the captured 3D representation. In particular, ellipses or circles, which are special cases of oval shapes, may be used. Embodiments of the present invention may also rely on spline shapes.

[0035] The processing circuit 602 further causes the computing device 600 to be operable to generate 507 an avatar representation corresponding to the outer portion 330 of the captured 3D representation 300. The outer portion 330 of the captured 3D representation 300 is defined by the boundary 310 determined in 505 between the inner portion 320 and the outer portion 330 of the captured 3D representation 300. In practice, since the boundary 310 is represented by a 2D shape, such as an oval, the avatar representation corresponding to the outer portion 330 of the captured 3D representation 300 is generated for a portion of the head 102 outside the 2D shape representing the boundary 310 between the inner portion 320 and the outer portion 330 of the captured 3D representation 300. If the generated avatar representation is a point cloud, the generated points are outside the 2D shape representing the boundary 310.

[0036] The avatar representation is generated at 507 using an ML model trained on human heads. The pose determined at 504 of the head is used as an input to the ML model. The avatar representation generated at 507 corresponding to the outer portion 330 of the captured 3D representation 300 is an animated representation of the outer portion of the head 102, i.e., the portion of the head 102 outside the face, defined by the boundary 310 between the inner portion 320 and the outer portion 330 of the captured 3D representation 300. As an example, the avatar representation corresponding to the outer portion 330 of the captured 3D representation 300 may be generated using a GAN, as described in "Normalized Avatar Synthesis Using StyleGAN and Perceptual Refinement". The ML model may be a general ML model trained on human heads in general. Alternatively, the ML model may be a specific ML model trained on one or more of the specific types of human heads, including gender, age or age range, skin color, hair type, etc. It will also be appreciated that embodiments of the present invention may be envisaged for generating a 3D representation of an animal's head.

[0037] The processing circuit 602 optionally further causes the computing device 600 to be operable to extract 506 an inner portion 320 of the captured 3D representation 300 and to merge 508 the extracted inner portion 320 of the captured 3D representation 300 at 506 with the generated avatar representation at 507 into a merged 3D representation of the head 102. In other words, the inner portion 320 of the captured 3D representation 300, which is a subset of the data captured by the 3D sensor 111 representing the face of the head 102, i.e. inside the boundary 310 determined at 505 between the inner portion 320 and the outer portion 330 (wherein the boundary 310 is represented by a 2D shape such as an oval), is merged at 508 with the generated avatar representation at 507, which corresponds to the outer portion 330 of the captured 3D representation 300. In practice, this corresponds to replacing the outer part 330 of the captured 3D representation 300 with a generated avatar representation, i.e. an animated (computer generated) representation of the outer part of the head 102, while preserving the inner part 320 of the captured 3D representation 300, i.e. the one captured in real time representing the face of the head 102. More specifically, if the captured 3D representation and the generated avatar representation are point clouds, extracting the inner part 320 of the captured 3D representation 300 corresponds to selecting, from the point cloud representing the captured 3D representation 300, points whose coordinates are inside the determined boundary 310 between the inner part 320 and the outer part 330 of the captured 3D representation 300, i.e. points that are inside the 2D shape representing the boundary 310. Correspondingly, the points of the point cloud representing the generated avatar representation have coordinates that are outside the determined boundary 310 between the inner portion 320 and the outer portion 330 of the captured 3D representation 300, i.e., they are points that are outside the 2D shape representing boundary 310.Similarly, merging 508 the extracted inner portion 320 of the captured 3D representation 300 and the generated avatar representation into a merged 3D representation of the head corresponds to combining different sets of points (represented by separate point clouds) into a single point cloud. If the captured 3D representation and the generated avatar representation are represented in a format different from a point cloud, such as a mesh och depth map image, the representations may be optionally converted to point clouds before extracting 506 the inner portion 320 of the captured 3D representation 300 and merging 508 the extracted inner portion 320 of the captured 3D representation 300 in 506 and the generated avatar representation in 507 into a merged 3D representation of the head 102. Alternatively, extracting 506 the inner portion 320 of the captured 3D representation 300 and merging 508 the extracted inner portion 320 of the captured 3D representation 300 in 506 and the generated avatar representation in 507 into a merged 3D representation of the head 102 may be performed in the native format of the captured 3D representation and the generated avatar representation, without converting the data to a point cloud.

[0038] The processing circuit 602 optionally further causes the computing device 600 to be operable to display 509 the merged 3D representation of the head 102 using the display device 131. Optionally, the display device 131 may be included in the computing device 600 when the computing device is embodied as a receiving computing device 130. The display device 131 may be any one of a computer display, a television, an AR device, a VR device, an MR device, an XR device, and an HMD device.

[0039] By replacing the outer portion 330 of the captured 3D representation 300 of the head 102, which needs to be transmitted in real time from the sending computing device 110 to the receiving computing device 130 for rendering on the display device 131, with a generated avatar representation, the amount of captured data that needs to be transmitted over the communications network 140 in real time can be reduced.

[0040] A further advantage arises from the fact that the captured 3D representation of the head 102 may be incomplete, especially in the outer part 320 of the head 102, i.e. outside the facial region. This may occur when the head 102 moves and due to limitations in the field of view of the 3D sensor 111. Such a situation is illustrated in FIG. 2, which shows missing patches 211 and 221 in the captured 3D representation at different poses 210 and 220 of the head 102 relative to the 3D sensor 111. By replacing the outer part 330 of the captured 3D representation 300 of the head 102 with the generated avatar representation, the merged 3D representation displayed using the display device 131 does not suffer from missing captured data in the outer part 330 of the captured 3D representation 300. The user experience of the viewing participant 103 is thereby improved, as the displayed merged 3D representation of the head 102 is less likely to suffer from missing captured data.

[0041] Optionally, the ML model used for generating 507 the avatar representation corresponding to the outer portion 330 of the captured 3D representation 300 is trained 510 on the head 102 of the participant 101. In other words, the ML model is trained 510 specifically on the head 102 captured during the video communication session. Thereby, the user experience during the video communication session, and in particular the user experience of the receiving participant 103 viewing the merged 3D representation of the head 102, is improved.

[0042] The processing circuit 602 optionally further causes the computing device 600 to be operable to collect the ML model from a data storage associated with the participant 101. Preferably, the collected ML model is trained on the head 102 of the participant 101 and is also referred to herein as a "participant-specific ML model." For example, the participant-specific ML model may be stored on a user device associated with the participant 101, such as the sending computing device 110, or on a cloud storage. When the participant 101 initiates or joins a video communication session, the participant-specific ML model may be retrieved at 511 / 522 by the computing device 600 and used in generating 507 an avatar representation corresponding to an outer portion of the captured 3D representation. For example, the participant-specific ML model may be transmitted at 511 / 522 from the sending computing device 110, which is a personal device used by the sending participant 101, to the computing device 600 embodied as an edge computing device 120 and / or as a receiving computing device 130. Alternatively, the computing device 600 may retrieve, i.e., request or receive, the participant-specific ML model from cloud storage (not shown in FIGS. 5A-5C) that is associated with and accessible by the computing device 600 to the participant 101. The latter may be the case, for example, if the participant-specific ML model is stored in cloud storage (iCloud, One Drive, etc.) and associated with a user identifier of the participant 101 (participant 101's Apple ID, email address, etc.).

[0043] The processing circuit 602 optionally further causes the computing device 600 to be operable to train 510 the ML model using at least the outer portion 330 of the captured 3D representation 300 and the pose determined at 504 of the head 102. That is, the outer portion 330 of the captured 3D representation 300 is extracted, for example, simultaneously with the extracting 506 of the inner portion 320 of the captured 3D representation 300 and used as an input for training 510 the ML model together with the pose determined at 504 of the head 102. Optionally, the computing device 600 may be operable to train 510 the ML model further based on the inner portion 320 of the captured 3D representation 300, i.e. using a substantially complete captured 3D representation 300 of the head 102. This is advantageous in that the ML model used for generating 507 the avatar representation corresponding to the outer portion 320 of the captured 3D representation 300 may be trained at 510 for the specific head 102 of the participant 101 as soon as the video communication session begins. The ML model may be, for example, a general ML model trained for human heads in general. Alternatively, the ML model may be a specific ML model trained for some type of human head, as described above. As yet another alternative, the ML model may be a participant-specific ML model that has been trained during a previous video communication session or during a dedicated training procedure and stored for later use in a data storage associated with the participant 101, for example a data storage or cloud storage comprised in the sender computing device 110.

[0044] The captured 3D representation, the inner portion of the captured 3D representation, the merged 3D representation, and the avatar representation may be stored and transmitted between the sending computing device 110, the edge computing device 120, and the receiving computing device 130 via the communication network 140 using any suitable data format, in particular 3D immersive media formats and / or protocols. More specifically, the captured 3D representation, the inner and outer portions of the captured 3D representation, the merged 3D representation, and the avatar representation may be stored and transmitted as a point cloud, a mesh, or a depth map image. A point cloud is a set of data points in space that represent a 3D object, such as the head 102. A mesh, also called a polygon mesh, is a collection of vertices, edges, and faces that define the shape of a 3D object, such as the head 102. A depth map image contains information related to the distance of the surface of a 3D object, such as the head 102, from a viewpoint, in particular the viewpoint of the 3D sensor 111. Protocols used to transmit the captured 3D representation, the inner portion of the captured 3D representation, the merged 3D representation, and the avatar representation between the sending computing device 110, the edge computing device 120, and the receiving computing device 130 over the communications network 140 include, but are not limited to, RTP, SRTP, Dynamic Adaptive Streaming over HTTP (DASH), etc.

[0045] Below and with reference to Figures 5A-5C, different embodiments of the present invention are illustrated, with particular focus on whether the operations involved in generating a 3D representation of a human head may be performed on the sending computing device 110, the edge computing device 120, or the receiving computing device 130.

[0046] The embodiment shown in Figure 5A is characterized by edge-centric processing, which is advantageous in that the edge computing device 120 generally has abundant computing resources in terms of computational power, memory, and power supplies, as compared to the sending computing device 110 and the receiving computing device 130, which may be embodied as smartphones, tablets, HMDs, or other types of mobile computing devices that are often battery-powered and less powerful in terms of processing power.

[0047] More specifically, when the computing device 600 for generating a 3D representation of a head 102 of a participant 101 in a video communication session is embodied as an edge computing device 120, the edge computing device 120 is operable to collect a captured 3D representation of the head by receiving 512 the captured 3D representation of the head 102 from the sending computing device 110. The edge computing device 120 is further operable to identify 503 the position of a set of facial landmarks in the captured 3D representation, the set of facial landmarks including facial landmarks indicating boundaries of a human face. The edge computing device 120 is further operable to determine 504 the pose of the head 102. The edge computing device 120 is further operable to determine 505 the boundary 310 between the inner portion 320 and the outer portion 330 of the captured 3D representation 300 based on the positions identified in 503 of the set of facial landmarks. The inner portion 320 of the captured 3D representation 300 represents the face of the participant 101. The edge computing device 120 is further operable to generate 507 an avatar representation corresponding to the outer portion 330 of the captured 3D representation 300 using an ML model trained on human heads with the determined pose 504 of the head 102 as input.

[0048] Optionally, the edge computing device 120 may be further operable to extract 506 the inner portion 320 of the captured 3D representation 300 and merge 508 the extracted inner portion 320 of the captured 3D representation 300 and the generated avatar representation into a merged 3D representation of the head 102.

[0049] The edge computing device 120 may be further operable to transmit 521 the merged 3D representation of the head 102 to the receiving computing device 130, where the merged 3D representation of the head 102 is displayed at 509 using the display device 131.

[0050] Optionally, the edge computing device 120 may be operable to train 510 the ML model using at least the outer portion 330 of the captured 3D representation 300 and the determined pose of the head 504. Further optionally, the edge computing device 120 may be operable to train 510 the ML model further based on the inner portion 320 of the captured 3D representation 300.

[0051] 5A, in the embodiment shown in FIG. 5B and described below, part of the processing has been moved from the edge computing device 120 to the receiving computing device 130. In this case, the edge computing device 120 and the receiving computing device 130 in combination implement an embodiment of the present invention. In other words, the present invention is embodied as a system of computing devices for generating a 3D representation of the heads of participants in a video communication session.

[0052] More specifically, the edge computing device 120 is operable to collect a captured 3D representation of the head by receiving 512 the captured 3D representation of the head 102 from the sending computing device 110. The edge computing device 120 is further operable to identify 503 the location of a set of facial landmarks in the captured 3D representation, the set of facial landmarks including facial landmarks that indicate boundaries of a human face. The edge computing device 120 is further operable to determine 504 a pose of the head 102 and transmit 523 the determined pose of the head 102 to the receiving computing device 130. The edge computing device 120 is further operable to determine 505 a boundary 310 between an inner portion 320 and an outer portion 330 of the captured 3D representation 300 based on the identified 503 locations of the set of facial landmarks. The inner portion 320 of the captured 3D representation 300 represents the face of the participant 101. The edge computing device 120 is further operable, optionally, to extract 506 the inner portion 320 of the captured 3D representation 300 and transmit 524 the extracted inner portion 320 of the captured 3D representation 300 to the receiving computing device 130.

[0053] The receiving computing device 130 is operable to generate 507 an avatar representation corresponding to the outer portion 330 of the captured 3D representation 300 using an ML model trained on a human head with the pose determined at 504 of the head 102 received by the receiving computing device 130 at 523 as input. Optionally, the receiving computing device 130 is operable to merge 508 the inner portion 320 of the captured 3D representation 300 received at 524 and the avatar representation generated at 507 into a merged 3D representation of the head 102. The receiving computing device 130 may further be operable to display 509 the merged 3D representation of the head 102 using the display device 131.

[0054] Optionally, the edge computing device 120 may be further operable to train 510 the ML model using at least the outer portion 330 of the captured 3D representation 300 and the determined pose of the head 504. Further optionally, the edge computing device 120 may be operable to train 510 the ML model further based on the inner 320 portion of the captured 3D representation 300, i.e., using a substantially complete captured 3D representation of the head 102. In this case, the edge computing device 120 may be operable to transmit 525 the updated ML model to the receiving computing device 130.

[0055] A further embodiment of a system of computing devices for generating a 3D representation of the head of a participant in a video communication session is shown in Fig. 5C. Compared to the embodiment shown in Fig. 5B, the additional operations involved in generating the 3D representation of the head 102 have been transferred from the edge computing device 120 to the receiving computing device 130.

[0056] More specifically, the edge computing device 120 is operable to collect a captured 3D representation of the head 102 by receiving 512 the captured 3D representation of the head 102 from the sending computing device 110. The edge computing device 120 is further operable to identify 503 the location of a set of facial landmarks in the captured 3D representation, the set of facial landmarks including facial landmarks that indicate boundaries of a human face. The edge computing device 120 is further operable to determine 504 a pose of the head 102 and transmit 523 the determined pose of the head 102 to the receiving computing device 130. The edge computing device 120 is further operable to determine 505 a boundary 310 between an inner portion 320 and an outer portion 330 of the captured 3D representation 300 based on the identified 503 locations of the set of facial landmarks and transmit 527 the determined boundary between the inner portion and the outer portion to the receiving computing device 130. The inner portion of the captured 3D representation represents the face of the participant 101 .

[0057] The receiving computing device 120 is optionally operable to extract 506 the inner portion 320 of the captured 3D representation 300 using the boundary 310 received at 527 between the inner portion 320 and the outer portion 330 of the captured 3D representation 300.

[0058] The receiving computing device 130 is operable to generate 507 an avatar representation corresponding to the outer portion 330 of the captured 3D representation 300 using an ML model trained on a human head with the pose determined at 504 of the head 102 received by the receiving computing device 130 at 523 as input. Optionally, the receiving computing device 130 is operable to merge 508 the inner portion 320 of the captured 3D representation 300 received at 524 and the avatar representation generated at 507 into a merged 3D representation of the head 102. The receiving computing device 130 may further be operable to display 509 the merged 3D representation of the head 102 using the display device 131.

[0059] Optionally, the receiving computing device 130 may be further operable to train 510 the ML model using at least the outer portion 330 of the captured 3D representation 300 and the determined pose of the head 504. Further optionally, the edge computing device 120 may be operable to train 510 the ML model further based on the inner portion 320 of the captured 3D representation 300, i.e., using a substantially complete captured 3D representation 300 of the head 102.

[0060] Although embodiments of the present invention have been described with respect to a particular distribution of operations involved in generating 3D representations of the heads of participants in a video communication session between the sending computing device 110, the edge computing device 120, and the receiving computing device 130, respectively, as shown in Figures 5A-5C, those skilled in the art may readily envision alternative forms for distributing the operations involved in generating 3D representations of the heads of participants in a video communication session between the sending computing device 110, the edge computing device 120, and the receiving computing device 130.

[0061] In the following, an embodiment of a processing circuit 602 included in a computing device 600 for generating a 3D representation of a head of a participant in a video communication session will be described with reference to Fig. 6. An embodiment of the processing circuit 600 may be included in one or more of the sending computing device 110, the edge computing device 120, and the receiving computing device 130.

[0062] The processing circuit 602 may comprise one or more processors 603, such as a central processing unit (CPU), a microprocessor, an application processor, an application specific processor, a digital signal processor (DSP) including a graphics processing unit (GPU), and an image processor, or a combination thereof, and a memory 604 comprising a computer program 605 comprising instructions. When executed by the processor(s) 603, the instructions cause the computing device 600 to operate according to the embodiments of the invention described herein. When the operations involved in generating a 3D representation of a participant's head in a video communication session are distributed among two or more of the sending computing device 110, the edge computing device 120, and the receiving computing device 130, their respective instructions, when executed by their respective processors 603, cause two or more of the sending computing device 110, the edge computing device 120, and the receiving computing device 130 to operate in a cooperative manner according to the embodiments of the invention described herein. The memory 604 may be, for example, a random access memory (RAM), a read only memory (ROM), a flash memory, etc. The computer program 605 may be downloaded into the memory 604 by the network interface circuitry 601 as a data carrier signal carrying the computer program 605. The processing circuitry 602 may alternatively or additionally comprise one or more application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc., operable to cause the computing device 600 to operate according to embodiments of the invention described herein.

[0063] The network interface circuitry 601 may include one or more of a cellular modem (e.g., GSM, UMTS, LTE, 5G or higher generation), a WLAN / WiFi modem, a Bluetooth modem, an Ethernet interface, an optical interface, etc. for exchanging data between the computing device 600 and other computing devices, particularly between the sending computing device 110, the edge computing device 120, and the receiving computing device 130, and a communications network 140, which may include the Internet and one or more RANs.

[0064] In the following, an embodiment of a method 700 for generating a 3D representation of the head 102 of a participant 101 in a video communication session is described with reference to FIG.

[0065] The method 700 is performed by the computing device 600 and includes collecting 701 a captured 3D representation of the head 102 and identifying 702 a position of a set of facial landmarks in the captured 3D representation. The set of facial landmarks includes facial landmarks indicative of a boundary of a human face. The method 700 further includes determining 703 a pose of the head 102 and determining 704 a boundary between an inner portion and an outer portion of the captured 3D representation. The boundary is determined 704 based on the positions identified in 702 of the set of facial landmarks. The inner portion of the captured 3D representation represents the participant's face. The method 700 further includes generating 705 an avatar representation corresponding to the outer portion of the captured 3D representation. The avatar representation is generated 705 using an ML model trained on a human head with the pose determined in 703 of the head 102 as input. The ML model may optionally be trained on the head 102 of the participant 101.

[0066] The method 700 optionally further includes extracting 707 an inner portion of the captured 3D representation and merging 708 the inner portion of the captured 3D representation extracted in 707 with the avatar representation generated in 705 into a merged 3D representation of the head 102.

[0067] The method 700 optionally further includes displaying 709 the merged 3D representation of the head 102 using a display device 131. The display device 131 may be any one of a computer display, a television, an AR device, a VR device, an MR device, an XR device, and an HMD device.

[0068] Collecting 701 the captured 3D representation of the head 102 may include capturing the 3D representation of the head 102 using the 3D sensor 111. The 3D sensor 111 may comprise one or more of a 3D camera, a LiDAR, and an optical 3D sensor.

[0069] The method 700 optionally further includes collecting the ML model from data storage associated with the participant 101.

[0070] Method 700 optionally further includes training 706 an ML model using at least the outer portion of the captured 3D representation and the pose of the head determined in 703. The ML model is optionally further trained based on the inner portion of the captured 3D representation.

[0071] It will be appreciated that method 700 may include additional, alternative, or modified steps as described throughout this disclosure. The method may also be performed in a collaborative manner by two or more computing devices, for example, two or more of the sending computing device 110, the edge computing device 120, and the receiving computing device 130.

[0072] An embodiment of the method 700 may be implemented as a computer program 605 comprising instructions which, when executed by the computing device 600, cause the computing device 600 to perform the method 700 and to become operable according to embodiments of the invention described herein. The computer program 605 may be stored on a computer readable data carrier such as the memory 604. Alternatively, the computer program 605 may be carried by a data carrier signal via the network interface circuit 601, for example downloaded into the memory 604.

[0073] Those skilled in the art will appreciate that the present invention is in no way limited to the embodiments described above. On the contrary, many modifications and variations are possible within the scope of the appended claims.

Claims

1. An edge computing device (120, 600) for generating a three-dimensional (3D) representation of a head (102) of a participant (101) in a video communication session, the edge computing device comprising a processing circuit (602), the processing circuit (602) being such that the edge computing device receives (512) a captured 3D representation (300) of the head (102) from a sending-side computing device (110); identifies (503) positions (1-27) of a set of face landmarks in the captured 3D representation (300), the set of face landmarks including face landmarks indicating boundaries of a human face, identifying (503) positions (1-27) of the set of face landmarks; determines (504) the pose of the head (102); sends (523) the determined pose of the head (102) to a receiving-side computing device (130); determines (505) a boundary (310) between an inner part (320) and an outer part (330) of the captured 3D representation (300) based on the identified positions (1-27) of the set of face landmarks, the inner part (320) of the captured 3D representation (300) representing the face of the participant (101), determining (505) the boundary (310); extracts (506) the inner part (320) of the captured 3D representation (300); sends (524) the extracted inner part (320) of the captured 3D representation (300) to the receiving-side computing device (130); causes the edge computing device (120, 600) to be operable to perform.

2. At least using the outer part (330) of the captured 3D representation (300) and the determined pose of the head to train (510) an ML model; sending (525) the updated ML model to the receiving-side computing device (130); The edge computing device (120, 600) according to claim 1, which is further operable to perform.

3. The edge computing device (120, 600) according to claim 2, operable to perform training (510) of the ML model further based on the inner part (320) of the captured 3D representation (300).

4. The edge computing device (120, 600) according to any one of claims 1 to 3, wherein the captured 3D representation (300) and the inner part of the captured 3D representation (300) are a point cloud, a mesh, or a depth map image.

5. A receiving-side computing device (130, 600) for generating a three-dimensional (3D) representation of the head (102) of a participant (101) in a video communication session, the receiving-side computing device comprising a processing circuit (602), the processing circuit (602) causing the receiving-side computing device to receive (523) the head pose from the edge computing device (120); receive (524) the inner part (320) of the captured 3D representation (300) of the head (102) from the edge computing device (120), the inner part (320) of the captured 3D representation (300) representing the face of the participant (101), and receive (524) the inner part (320) of the captured 3D representation (300) of the head (102); using a machine learning (ML) model trained on a human head, with the received pose of the head as an input, to generate an avatar representation corresponding to the outer part (330) of the captured 3D representation (300) (507); merge (508) the received inner part (320) of the captured 3D representation (300) and the generated avatar representation into a merged 3D representation of the head (102); A receiving-side computing device (130, 600) causing it to be operable to perform the above.

6. The receiving-side computing device (130, 600) according to claim 5, further operable to display (509) the merged 3D representation of the head (102) using a display device (131).

7. The receiving-side computing device (130, 600) according to claim 6, wherein the display device (131) is any one of a computer display, a television, an augmented reality (AR) device, a virtual reality (VR) device, a mixed reality (MR) device, an extended reality (XR) device, and a head-mounted display (HMD) device.

8. The receiving-side computing device (130, 600) according to any one of claims 5 to 7, wherein the ML model is trained for the head (102) of the participant (101).

9. The receiving-side computing device (130, 600) according to any one of claims 5 to 8, wherein the captured 3D representation (300), the inner part of the captured 3D representation (300), the merged 3D representation, and the avatar representation are a point cloud, a mesh, or a depth map image.

10. A system of a computing device (120, 130, 600) for generating a three-dimensional (3D) representation of the head (102) of a participant (101) in a video communication session, the system comprising the edge computing device (120) according to any one of claims 1 to 4 and the receiving-side computing device (130) according to any one of claims 5 to 9.

11. A method (700) for generating a three-dimensional (3D) representation of the head (102) of a participant (101) in a video communication session, the method being performed by an edge computing device (120, 600), receiving (701) a captured 3D representation (300) of the head (102) from a sending-side computing device (110), identifying (702) positions (1 to 27) of a set of face landmarks in the captured 3D representation (300), the set of face landmarks including face landmarks indicating boundaries of a human face, determining (703) the pose of the head (102), Transmitting the determined pose of the head (102) to the receiving computing device (130), and determining (704) a boundary (310) between an inner portion (320) and an outer portion (330) of the captured 3D representation (300) based on the identified positions (1-17) of the set of facial landmarks, wherein the inner portion (320) of the captured 3D representation (300) represents the face of the participant (101), and determining (704) the boundary (310). Extracting (707) the inner portion (320) of the captured 3D representation (300). Transmitting the extracted inner portion (320) of the captured 3D representation (300) to the receiving computing device (130). A method (700) including the above steps.

12. At least using the outer portion (330) of the captured 3D representation (300) and the determined pose of the head (102) to train an ML model (706), and transmitting the updated ML model to the receiving computing device (130). Transmitting the updated ML model to the receiving computing device (130). The method (700) according to claim 11, further including the above steps.

13. The method (700) according to claim 12, wherein the ML model is further trained based on the inner portion (320) of the captured 3D representation (300) (706).

14. The method (700) according to any one of claims 11 to 13, wherein the captured 3D representation (300) and the inner portion of the captured 3D representation are a point cloud, a mesh, or a depth map image.

15. A method (700) for generating a three-dimensional (3D) representation of a head (102) of a participant (101) in a video communication session, the method being performed by a receiving computing device (130, 600). Receiving the pose of the head from an edge computing device (120). Receiving an inner part (320) of the captured 3D representation (300) of the head (102) from the edge computing device (120), wherein the inner part (320) of the captured 3D representation (300) represents the face of the participant (101), receiving the inner part (320) of the captured 3D representation (300) of the head (102), Using a machine learning (ML) model trained on a human head with the received pose of the head as an input to generate an avatar representation corresponding to an outer part (330) of the captured 3D representation (300), Merging the received inner part (320) of the captured 3D representation (300) and the generated avatar representation into a merged 3D representation of the head (102), A method (700) comprising:

16. The method (700) according to claim 15, further comprising displaying (709) the merged 3D representation of the head (102) using a display device (131).

17. The method (700) according to claim 16, wherein the display device (131) is any one of a computer display, a television, an augmented reality (AR) device, a virtual reality (VR) device, a mixed reality (MR) device, an extended reality (XR) device, and a head-mounted display (HMD) device.

18. The method (700) according to any one of claims 15 to 17, wherein obtaining (701) the captured 3D representation (300) of the head (102) comprises capturing the 3D representation of the head using a 3D sensor (111).

19. The method (700) according to claim 18, wherein the 3D sensor (111) comprises one or more of a 3D camera, a LiDAR, and an optical 3D sensor.

20. The method (700) according to any one of claims 15 to 19, wherein the ML model is trained on the head (102) of the participant (101).

21. The method (700) according to any one of claims 15 to 20, wherein the captured 3D representation (300), the inner portion of the captured 3D representation (300), the merged 3D representation, and the avatar representation are a point cloud, a mesh, or a depth map image. Claim 22 A computer program (605) comprising instructions which, when the computer program (605) is executed by a computing device (600), cause the computing device (600) to perform the method according to any one of claims 15 to 21. Claim 23 A computer-readable data carrier (604) storing the computer program (605) according to claim 22. Claim 24 A data carrier signal carrying the computer program (605) according to claim 22.