Generating a 3D representation of a participant's head in a video communication session
By processing 3D head data to generate an avatar for the outer head parts and transmit only the inner facial portion, the solution addresses bandwidth and complexity issues in 3D video communication, enhancing user experience and reducing data transfer.
Patent Information
- Application Number
- JP2024555182
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-14
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-03-14
AI Technical Summary
Existing 3D video communication solutions face challenges in capturing and transmitting detailed head movements with high bandwidth requirements and complex sensors, leading to incomplete representations and increased costs.
A computing device processes a captured 3D head representation to identify facial landmarks, determine a head pose, and generate an avatar representation for the outer part using an ML model, while transmitting only the inner facial portion for display.
This approach reduces bandwidth requirements and improves user experience by preserving facial details and movements, while minimizing data transfer and sensor complexity.
Smart Images

Figure 0007728472000001 
Figure 0007728472000002 
Figure 0007728472000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a computing device for generating a three-dimensional (3D) representation of the head of a participant in a video communication session, a method for generating a 3D representation of the head of a participant in a video communication session, a corresponding computer program, a corresponding computer-readable data carrier and a corresponding data carrier signal. [Background technology]
[0002] For example, various alternatives are known for generating and displaying to a viewer a three-dimensional (3D) representation of a human head during a video communication session between two or more participants.
[0003] The first type of solution is based on computer-generated 3D avatars. Such avatars are generated using a machine learning (ML) model trained using a captured 3D representation of a human head and then adapted to that specific head using two-dimensional (2D) images representing that head. During use, the head pose of the sending participant's head is continuously detected and used as input to the ML model, which generates a dynamically animated avatar that reflects the sending participant's actual head movements. Recently, significant improvements in generated 3D avatars have been achieved through the use of generative adversarial networks (GANs), as demonstrated by H. Luo et al. ("Normalized Avatar Synthesis Using StyleGAN and Perceptual Refinement," arXiv:2106.11423, arXiv, 2021).
[0004] A second type of solution relies on capturing the sending participant's head using a 3D sensor, such as a stereo camera, and transmitting the captured 3D representation in real time for display to the receiving participant, e.g., as a point cloud stream or mesh stream. Solutions based on real-time capture and transmission excel in representing details of the captured head compared to animated 3D avatars. These details, especially the captured facial details, are important for conveying the sending participant's emotions. However, transmitting the 3D captured representation of the head in real time, e.g., as a point cloud or mesh stream, requires significantly greater bandwidth of the communication link used to transmit the captured 3D data. Furthermore, current solutions that rely on capturing the sending participant's head in real time generally require relatively complex 3D sensors, such as several cameras, to ensure that the outer part of the head (outside the face) is fully captured as the sending participant's head moves, which increases complexity and cost. Summary of the Invention
[0005] It is an object of the present invention to provide an improved alternative to the above techniques and prior art.
[0006] More particularly, it is an object of the present invention to provide an improved solution for generating a 3D representation of a human head based on capturing the head using a 3D sensor device.
[0007] These and other objects of the invention are achieved by various aspects of the invention, which are defined by the independent claims.Embodiments of the invention are characterized by the dependent claims.
[0008] According to a first aspect of the present invention, there is provided a computing device for generating a 3D representation of a head of a participant in a video communication session, the computing device comprising processing circuitry configured to cause the computing device to process a captured 3D representation of the head. acquisition and identifying positions of a set of facial landmarks in the captured 3D representation. The set of facial landmarks includes facial landmarks that indicate boundaries of a human face. The computing device is further operable to determine a head pose and determine a boundary between an inner part and an outer part of the captured 3D representation. The boundary is determined based on the identified positions of the set of facial landmarks. The inner part of the captured 3D representation represents the participant's face. The computing device is further operable to generate an avatar representation corresponding to the outer part of the captured 3D representation. The avatar representation is generated using an ML model trained on a human head. The determined pose of the head is used as input to the ML model.
[0009] According to a second aspect of the present invention, there is provided a method for generating a 3D representation of a head of a participant in a video communication session, the method being performed by a computing device and comprising: acquisitionand identifying positions of a set of facial landmarks on the captured 3D representation. The set of facial landmarks includes facial landmarks that indicate boundaries of a human face. The method further includes determining a head pose and determining a boundary between an inner portion and an outer portion of the captured 3D representation. The boundary is determined based on the identified positions of the set of facial landmarks. The inner portion of the captured 3D representation represents the participant's face. The method further includes generating an avatar representation corresponding to the outer portion of the captured 3D representation. The avatar representation is generated using an ML model trained on a human head. The determined pose of the head is used as input for the ML model.
[0010] According to a third aspect of the present invention, there is provided a computer program comprising instructions which, when executed by a computing device, cause the computing device to: to , a method according to an embodiment of the second aspect of the present invention make it happen .
[0011] According to a fourth aspect of the present invention there is provided a computer readable data carrier storing a computer program according to the third aspect of the present invention.
[0012] According to a fifth aspect of the present invention there is provided a data carrier signal, the data carrier signal carrying a computer program according to the third aspect of the present invention.
[0013] The present invention utilizes the understanding that a 3D representation of a participant's head in a video communication session can be generated based on extracting an interior portion of a captured 3D representation of the head and making the interior portion available for display to a receiving user, e.g., as a real-time stream. The extracted interior portion generally corresponds to the facial region, or face, of the head, including the eyes, nose, ears, and mouth. The remainder of the captured 3D representation, referred to herein as the exterior portion, represents the portion of the head outside the face. This exterior portion is replaced by an avatar generated using an ML model trained on a human head, using the pose of the captured head as input for the ML model. This results in an animated avatar that reflects the actual pose and movement of the captured head.
[0014] Embodiments of the present invention are advantageous in that a receiving user viewing the generated 3D representation of the captured head is not bothered by an incompletely captured 3D representation of the head, which may occur due to limitations in the 3D sensor used to capture the 3D representation. At the same time, the detailed structure and movement of the face and facial parts of the captured head, which are important for conveying emotions in inter-human communication, are preserved, thereby improving the user experience. Embodiments of the present invention are further advantageous in that the amount of data captured as the 3D representation of the head that needs to be transferred to the receiving user in real time, for example, by streaming, is reduced. This is true because only a subset of the captured 3D data, i.e., the inner part of the captured 3D representation representing the face of the captured head, needs to be transferred from the 3D sensor to the display device. This reduces the bandwidth required to conduct a 3D video communication session.
[0015] Although advantages of the present invention have been explained in some cases with respect to embodiments of the first aspect of the invention, corresponding reasoning applies to embodiments of the other aspects of the invention.
[0016] Further objectives, features, and advantages of the present invention will become apparent upon study of the following detailed disclosure, drawings, and appended claims. Those skilled in the art will appreciate that different features of the present invention may be combined to create embodiments other than those described below.
[0017] The above, as well as additional objects, features and advantages of the present invention will be better understood through the following illustrative, non-limiting detailed description of embodiments of the invention, with reference to the accompanying drawings. [Brief explanation of the drawings]
[0018] [Figure 1] FIG. 1 illustrates a video communication session between two participants according to an embodiment of the present invention. [Figure 2] FIG. 1 illustrates a captured 3D representation of a human head, according to an embodiment of the present invention. [Figure 3] FIG. 10 illustrates the use of facial landmarks to determine the boundary between the inner and outer portions of a captured 3D representation of a head, according to an embodiment of the present invention. [Figure 4] 1A-1C are diagrams illustrating schematically generating 3D representations of the heads of participants in a video communication session according to an embodiment of the present invention; [Figure 5A] 1 is a sequence diagram illustrating generating a 3D representation of a participant's head in a video communication session according to an embodiment of the present invention. [Figure 5B] 1 is a sequence diagram illustrating generating a 3D representation of a participant's head in a video communication session according to an embodiment of the present invention. [Figure 5C] 1 is a sequence diagram illustrating generating a 3D representation of a participant's head in a video communication session according to an embodiment of the present invention. [Figure 6] FIG. 1 illustrates a schematic diagram of a computing device for generating a 3D representation of the head of a participant in a video communication session, according to an embodiment of the present invention. [Figure 7]1 illustrates a method for generating a 3D representation of the head of a participant in a video communication session according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0019] All figures are schematic, not necessarily to scale, and generally show only those parts that are necessary to elucidate the invention; other parts may be omitted or merely suggested.
[0020] The present invention now will be described more fully hereinafter with reference to the accompanying drawings, in which several embodiments of the invention are shown. This invention may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided as examples so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0021] 1 illustrates a video communication session between two participants 101 and 103, illustrated as a one-way video communication session between a sending computing device 110 and a receiving computing device 130. During the video communication session, sometimes referred to as a 3D video communication session, a 3D representation of the sending participant's 101 head 102 is captured using a 3D sensor 111, such as a stereo camera, and displayed to the receiving participant 103 using a display device 131, such as a computer display or head-mounted display (HMD). To table Show do For 、The video signal is then transmitted (e.g., streamed) over communications network 140 to receiving computing device 130. Note that embodiments of the present invention are not limited to the one-way video communication session between two participants shown in FIG. 1 . Rather, embodiments of the present invention may be envisioned that enable one-way (e.g., a presentation streamed from a presenter to many viewers) or two-way (e.g., a video call during a virtual conference) video communication sessions between two or more participants. More particularly, in the case of two-way video communication sessions, embodiments of the invention also support generating a 3D representation of another participant's head in the opposite direction, e.g., participant 103's head in FIG. 1 . Accordingly, a computing device supporting a two-way video communication session according to embodiments of the present invention may include both 3D sensor 111 and display device 131 or be operatively connected to both.
[0022] A computing device for generating a 3D representation of the head 102 of a participant 101 in a video communication session may be embodied in different forms, for example, as a sending computing device 110, a receiving computing device 130, an edge computing device 120 provided at the edge of a communication network 140 through which traffic passes between the sending computing device 110 and the receiving computing device 130, or a combination thereof. The edge computing device 120 may be provided, for example, near a radio access network (RAN) that is part of the communication network 140, through which the sending computing device 110 and / or the receiving computing device 130 communicate with each other and / or with the edge computing device 120.
[0023] The computing device for generating the 3D representation of the head 102 of a participant 101 in a video communication session, particularly when embodied as the sending computing device 110 or the receiving computing device 130, may be any one of a smartphone, a tablet, a laptop computer, an augmented reality (AR) device, a virtual reality (VR) device, a mixed reality (MR) device, an extended reality (XR) device, or an HMD. Alternatively, the computing device for generating the 3D representation of the head 102 of a participant 101 in a video communication session, particularly when embodied as the edge computing device 120, may be any one of an edge server, an application server, or a cloud computer. It will also be appreciated that the computing device for generating the 3D representation of the head 102 may be embodied in a distributed manner. That is, different operations involved in generating the 3D representation of the head 102, which will be described in further detail below, may be distributed and performed in a collaborative manner among two or more of the sending computing device 110, the edge computing device 120, and the receiving computing device 130. Illustrative examples for distributing different operations in a collaborative manner among the sending computing device 110, the edge computing device 120, and the receiving computing device 130 are shown in Figures 5A-5C and will be elucidated in further greater detail below.
[0024] Throughout this disclosure, embodiments of the present invention are described with respect to generating a 3D representation of a human head 102, including a face. A human face includes eyes, nose, ears, and mouth. In human-to-human communication, the detailed structure and movement of the face and facial parts is important for conveying emotions, for example, between participant 101 and participant 103.
[0025] Below and with reference to Figure 6, an embodiment 600 of a computing device (also referred to as a "computing device" for brevity) for generating a 3D representation of the head 102 of a participant 101 in a video communication session is described in more detail. Additional reference is made to Figure 4, which shows schematically the flow of data in generating a 3D representation of the head of a participant in a video communication session.
[0026] The computing device 600 may display a captured 3D representation of the head 102. acquisition6 , the computing device 110 may include a processing circuit 602 that causes the computing device 110 to be operable to: capture 502 a 3D representation of the head using the 3D sensor 111; this may be achieved, for example, by capturing 502 a 3D representation of the head using the 3D sensor 111; Optionally, if the computing device 600 is embodied as the sending computing device 110, the computing device 110 may include the 3D sensor 111; alternatively, the computing device 110 may be operatively connected to the 3D sensor 111; for example, the 3D sensor 111 may be a separate unit connected to the computing device 110 via an interface circuit ("I / O interface" in FIG. 6 ) using any wired or wireless technology known in the art, such as Universal Serial Bus (USB), Lighting, High-Definition Multimedia Interface (HDMI), Bluetooth, etc. Alternatively, if the computing device 600 is embodied as an edge computing device 120 or as a receiving computing device 130, or a combination thereof, the computing device 120 / 130 may transmit the captured 3D representation of the head 102 by receiving 512 the captured 3D representation, e.g., as a data stream directly from the 3D sensor 111 (or indirectly via the sending computing device 110) over the communications network 140, e.g., using the Real-Time Protocol (RTP), the Secure Real-Time Transport Protocol (SRTP), or any other suitable protocol. acquisition The device may be operable to:
[0027] The 3D sensor 111 may include one or more of a 3D camera (also known as a stereo camera), an optical 3D sensor, a LiDAR, and a 2D camera. Optical 3D sensors can be used to capture and reconstruct the 3D depth of real-world objects, such as the head 102. Depending on the source of radiation used, optical 3D sensors can be divided into two categories: passive and active. Stereoscopic sensors, Shape-from-Silhouettes (SfS) sensors, and Shape-from-Texture (SfT) sensors are examples of passive 3D sensors that do not emit any type of radiation themselves. The 3D sensor collects images of a scene, e.g., the head 102, optionally from different perspectives or with different lighting setups. The images are then analyzed to calculate the 3D depth of points in the captured scene, e.g., points representing the surface of the head 102 and parts of the head 102. Viewed another way, an active 3D sensor emits radiation, e.g., electromagnetic waves such as light, and the interaction between the radiation and an object, such as the head 102, is captured by the sensor. From an analysis of the captured data and based on properties of the emitted radiation, coordinates of points in the captured scene, e.g., points representing the surface of the head 102 and portions of the head 102, are obtained. Time-of-Flight (ToF) sensors, phase-shift sensors, and active triangulation sensors are examples of active 3D sensors. The output of an optical 3D sensor is typically a depth map image.
[0028] LiDAR (Light Detection and Ranging) can be used to measure distance (also known as "ranging") by illuminating a target, such as the head 102, with light and then measuring the reflection with a light sensor. LiDAR sensors can operate in the ultraviolet, visible, or infrared spectrum. Because commonly used laser light is collimated, the LiDAR sensor must scan the scene to generate an image with the desired field of view. The output of the LiDAR sensor is typically a point cloud, which can then be enriched with other sensor data, such as RGB data from a traditional (2D) camera, which may be included in the 3D sensor 111.
[0029] The processing circuit 602 further causes the computing device 600 to be operable to identify 503 the locations of a set of facial landmarks (also known as "facial keypoints" or simply "keypoints") in the captured 3D representation. The set of facial landmarks includes facial landmarks that delineate boundaries of a human face. Different sets of facial landmarks are used in the art. As an example, "Fast Facial Landmark Detection and Applications: A Survey" (by K.S.Khabarlak and L.S.Koriashkina, arXiv:2101.10808v2, arXiv, 2021) lists different sets comprising between 21 and 98 facial landmarks. Generally, a subset of the facial landmarks in a given set delineates boundaries of a human face. 3 reproduces an exemplary set of facial landmarks 1-27 that demarcate the boundaries of a human face, as described in "Facial Landmarks for Face Recognition with Dlib" (https: / / sefiks.com / 2020 / 11 / 20 / facial-landmarks-for-face-recognition-with-dlib / , retrieved March 11, 2022). For illustrative purposes, facial landmarks 1-27 are overlaid on a sketch of a captured 3D representation 300 of head 102 at a representative position.
[0030] Using an ML model such as a neural network, a set of facial landmarks can be detected in a facial (2D) image and the (3D) locations of the facial landmarks can be determined. This can be achieved using known facial landmark detection algorithms, for example, as described in "Fast Facial Landmark Detection and Applications: A Survey" using the Dlib library (see "Facial Landmarks for Face Recognition with Dlib") or the OpenCV library (see, for example, "Head Pose Estimation using Python," https: / / towardsdatascience.com / head-pose-estimation-using-python-d165d3541600, retrieved March 11, 2022).
[0031] The processing circuit 602 further causes the computing device 600 to be operable to determine 504 a pose of the head 102 (also referred to as “head pose”). The determined pose of the head 102 may be expressed in terms of Euler angles, e.g., pitch, yaw, and roll, although embodiments of the present invention may also rely on alternative sets of angles. The pose of the head 102 may be determined 504 using, for example, a similar approach as described above with respect to identifying 503 the location of a set of facial landmarks. For example, the head pose may be determined based on facial landmarks using the OpenCV library (see “Head Pose Estimation using Python”). This example describes how the head pose may be determined using only six facial landmarks, identifying the edges of the eyes, nose, chin, and mouth. The set of facial landmarks used for determining 504 may differ from the set of facial landmarks that delineate the boundaries of a human face. Alternatively, the set of facial landmarks that delineate the boundaries of a human face may comprise landmarks that delineate the pose of the human head. Different techniques for determining head pose from (2D) images of the head are known in the art, see for example "Fast Facial Landmark Detection and Applications: A Survey".
[0032] The processing circuit 602 further causes the computing device 600 to be operable to determine 505 a boundary between an inner portion and an outer portion of the captured 3D representation. The inner portion of the captured 3D representation represents the face of the participant 101 and may also be referred to as a facial portion of the captured 3D representation. The boundary is determined 505 based on the positions identified 503 of facial landmarks, particularly a set of facial landmarks that delineate the boundaries of a human face. In practice, the boundary between the inner portion and the outer portion of the captured 3D representation may be determined 505 by fitting a 2D shape, such as an oval shape, to the identified positions of a set of facial landmarks that delineate the boundaries of a human face. As an example, an oval shape 310 fitted to a set of facial landmarks 1-27 is shown in FIG. 3 . The boundary 310 separates the inner (facial) portion 320 from the outer portion 330 of the captured 3D representation 300.
[0033] As an alternative to fitting a 2D shape to the identified locations of a set of facial landmarks that demarcate the boundaries of the human face, the boundary between the inner and outer portions of the captured 3D representation may be determined in 505 by first fitting a 3D shape, such as an oval or ellipsoid, to the captured 3D representation of the head 102. The identified locations of the set of facial landmarks that demarcate the boundaries of the human face are then projected onto the fitted 3D shape using either a surface normal or a projection to an origin or coordinate system used for the captured 3D representation, advantageously near the center of the head 102. The projected locations of the facial landmarks are points on the surface of the fitted 3D shape. A 2D shape, such as an oval shape, is then fitted to the points on the surface of the fitted 3D shape.
[0034] It will be appreciated that embodiments of the present invention are not limited to using oval shapes in determining 505 the boundary between the inner and outer portions of the captured 3D representation. In particular, ellipses or circles, which are special cases of oval shapes, may be used. Embodiments of the present invention may also rely on spline shapes.
[0035] The processing circuit 602 further causes the computing device 600 to be operable to generate 507 an avatar representation corresponding to the outer portion 330 of the captured 3D representation 300. The outer portion 330 of the captured 3D representation 300 is defined by the boundary 310 determined in 505 between the inner portion 320 and the outer portion 330 of the captured 3D representation 300. In practice, since the boundary 310 is represented by a 2D shape, such as an oval, the avatar representation corresponding to the outer portion 330 of the captured 3D representation 300 is generated for a portion of the head 102 outside the 2D shape representing the boundary 310 between the inner portion 320 and the outer portion 330 of the captured 3D representation 300. If the generated avatar representation is a point cloud, the generated points are outside the 2D shape representing the boundary 310.
[0036] An avatar representation is generated at 507 using an ML model trained on human heads. The pose of the head determined at 504 is used as input to the ML model. The avatar representation generated at 507 corresponding to the outer portion 330 of the captured 3D representation 300 is an animated representation of the outer portion of the head 102, i.e., the portion of the head 102 outside the face defined by the boundary 310 between the inner portion 320 and the outer portion 330 of the captured 3D representation 300. As an example, the avatar representation corresponding to the outer portion 330 of the captured 3D representation 300 may be generated using a GAN, as described in “Normalized Avatar Synthesis Using StyleGAN and Perceptual Refinement.” The ML model may be a general ML model trained on human heads generally. Alternatively, the ML model may be a specific ML model trained on one or more specific types of human heads, including gender, age or age range, skin color, hair type, etc. It will also be appreciated that embodiments of the present invention may be envisioned for generating 3D representations of animal heads.
[0037] The processing circuit 602 optionally further causes the computing device 600 to be operable to extract 506 the inner portion 320 of the captured 3D representation 300 and to merge 508 the inner portion 320 of the captured 3D representation 300 extracted at 506 with the avatar representation generated at 507 into a merged 3D representation of the head 102. In other words, the inner portion 320 of the captured 3D representation 300, which is a subset of the data captured by the 3D sensor 111 representing the face of the head 102, i.e., which is inside the boundary 310 determined at 505 between the inner portion 320 and the outer portion 330 (the boundary 310 being represented by a 2D shape, such as an oval), is merged at 508 with the avatar representation generated at 507, which corresponds to the outer portion 330 of the captured 3D representation 300. In practice, this amounts to replacing the outer portion 330 of the captured 3D representation 300 with a generated avatar representation, i.e., an animated (computer-generated) representation of the outer portion of the head 102, while preserving the inner portion 320 of the captured 3D representation 300, i.e., the captured one that represents the face of the head 102 in real time. More specifically, if the captured 3D representation and the generated avatar representation are point clouds, extracting the inner portion 320 of the captured 3D representation 300 amounts to selecting points from the point cloud representing the captured 3D representation 300 that have coordinates that are inside the determined boundary 310 between the inner portion 320 and the outer portion 330 of the captured 3D representation 300, i.e., points that are inside the 2D shape representing the boundary 310. Correspondingly, the points of the point cloud representing the generated avatar representation have coordinates that lie outside the determined boundary 310 between the inner portion 320 and the outer portion 330 of the captured 3D representation 300, i.e., they are points that lie outside the 2D shape representing boundary 310.Similarly, merging 508 the extracted interior portion 320 of the captured 3D representation 300 and the generated avatar representation into a merged 3D representation of the head corresponds to combining different sets of points (represented by separate point clouds) into a single point cloud. If the captured 3D representation and the generated avatar representation are represented in a format different from a point cloud, such as a mesh or depth map image, the representations may optionally be converted to point clouds before extracting 506 the interior portion 320 of the captured 3D representation 300 and merging 508 the extracted interior portion 320 of the captured 3D representation 300 at 506 and the generated avatar representation at 507 into a merged 3D representation of the head 102. Alternatively, extracting 506 the inner portion 320 of the captured 3D representation 300 and merging 508 the extracted inner portion 320 of the captured 3D representation 300 in 506 and the generated avatar representation in 507 into a merged 3D representation of the head 102 may be performed in the native formats of the captured 3D representation and the generated avatar representation, without converting the data to a point cloud.
[0038] The processing circuit 602 optionally further causes the computing device 600 to be operable to display 509 the merged 3D representation of the head 102 using the display device 131. Optionally, the display device 131 may be included in the computing device 600 when the computing device is embodied as the receiving computing device 130. The display device 131 may be any one of a computer display, a television, an AR device, a VR device, an MR device, an XR device, and an HMD device.
[0039] By replacing the outer portion 330 of the captured 3D representation 300 of the head 102, which needs to be transmitted in real time from the sending computing device 110 to the receiving computing device 130 for rendering on the display device 131, with the generated avatar representation, the amount of captured data that needs to be transmitted over the communication network 140 in real time can be reduced.
[0040] A further advantage arises from the fact that the captured 3D representation of the head 102 may be incomplete, especially in the outer portion 320 of the head 102, i.e., outside the facial region. This can occur when the head 102 moves and due to limitations in the field of view of the 3D sensor 111. Such a situation is illustrated in FIG. 2, which shows missing patches 211 and 221 in the 3D representation captured at different poses 210 and 220 of the head 102 relative to the 3D sensor 111. By replacing the outer portion 330 of the captured 3D representation 300 of the head 102 with the generated avatar representation, the merged 3D representation displayed using the display device 131 does not suffer from missing captured data in the outer portion 330 of the captured 3D representation 300. The user experience of the viewing participant 103 is thereby improved, as the displayed merged 3D representation of the head 102 is less likely to suffer from missing captured data.
[0041] Optionally, the ML model used to generate 507 the avatar representation corresponding to the outer portion 330 of the captured 3D representation 300 is trained 510 on the head 102 of the participant 101. In other words, the ML model is trained 510 specifically on the head 102 captured during the video communication session. This improves the user experience during the video communication session, and in particular the user experience of the receiving participant 103 viewing the merged 3D representation of the head 102.
[0042] The processing circuitry 602 optionally causes the computing device 600 to retrieve the ML model from data storage associated with the participant 101. acquisition Preferably, the method further causes the device to be operable to: acquisition The generated ML model is trained for the head 102 of the participant 101 and is also referred to herein as a “participant-specific ML model.” For example, the participant-specific ML model may be stored on a user device associated with the participant 101, such as the sending computing device 110, or in cloud storage. When the participant 101 initiates or joins a video communication session, the participant-specific ML model may be retrieved at 511 / 522 by the computing device 600 and used in generating 507 an avatar representation corresponding to an outer portion of the captured 3D representation. For example, the participant-specific ML model may be transmitted at 511 / 522 from the sending computing device 110, which is a personal device used by the sending participant 101, to the computing device 600 embodied as the edge computing device 120 and / or as the receiving computing device 130. Alternatively, the computing device 600 may retrieve, i.e., request and receive, the participant-specific ML model from cloud storage (not shown in FIGS. 5A-5C) that is associated with and accessible by the computing device 600. The latter may be the case, for example, if the participant-specific ML model is stored in cloud storage (iCloud, One Drive, etc.) and associated with the participant 101's user identifier (e.g., the participant's 101's Apple ID, email address, etc.).
[0043] The processing circuit 602 optionally further causes the computing device 600 to be operable to train 510 an ML model using at least the outer portion 330 of the captured 3D representation 300 and the pose determined at 504 of the head 102. That is, the outer portion 330 of the captured 3D representation 300 is extracted, for example, simultaneously with extracting 506 the inner portion 320 of the captured 3D representation 300 and used, together with the pose determined at 504 of the head 102, as input for training 510 the ML model. Optionally, the computing device 600 may be operable to train 510 the ML model further based on the inner portion 320 of the captured 3D representation 300, i.e., using substantially the complete captured 3D representation 300 of the head 102. This is advantageous in that the ML model used to generate 507 the avatar representation corresponding to the outer portion 320 of the captured 3D representation 300 can be trained 510 for the specific head 102 of the participant 101 as soon as the video communication session begins. The ML model can be, for example, a general ML model trained for human heads in general. Alternatively, the ML model can be a specific ML model trained for some type of human head, as described above. As yet another alternative, the ML model can be a participant-specific ML model that was trained during a previous video communication session or during a dedicated training procedure and stored for later use in data storage associated with the participant 101, for example, data storage provided in the sender computing device 110 or cloud storage.
[0044] The captured 3D representation, the inner portion of the captured 3D representation, the merged 3D representation, and the avatar representation are transmitted using any suitable data format, in particular a 3D immersive media format and / or protocol. be remembered, andBetween the sending computing device 110, the edge computing device 120, and the receiving computing device 130 via a communication network 140. Send by The captured 3D representation, the interior and exterior portions of the captured 3D representation, the merged 3D representation, and the avatar representation may be stored and transmitted as a point cloud, a mesh, or a depth map image. A point cloud is a set of data points in space that represent a 3D object, such as the head 102. A mesh, also called a polygonal mesh, is a collection of vertices, edges, and faces that define the shape of a 3D object, such as the head 102. A depth map image contains information related to the distance of the surface of a 3D object, such as the head 102, from a viewpoint, particularly the viewpoint of the 3D sensor 111. Protocols used to transmit the captured 3D representation, the interior portions of the captured 3D representation, the merged 3D representation, and the avatar representation between the sending computing device 110, the edge computing device 120, and the receiving computing device 130 over the communication network 140 include, but are not limited to, RTP, SRTP, Dynamic Adaptive Streaming over HTTP (DASH), etc.
[0045] Below, and with reference to Figures 5A-5C, different embodiments of the present invention are illustrated, with particular focus on whether the operations involved in generating a 3D representation of a human head may be performed on a sending computing device 110, an edge computing device 120, or a receiving computing device 130.
[0046] 5A is characterized by edge-centric processing, which is advantageous in that the edge computing device 120 generally has abundant computing resources in terms of computational power, memory, and power supplies, compared to the sending computing device 110 and the receiving computing device 130, which may be embodied as smartphones, tablets, HMDs, or other types of mobile computing devices that are often battery-powered and less powerful in terms of processing power.
[0047] More specifically, when the computing device 600 for generating a 3D representation of the head 102 of a participant 101 in a video communication session is embodied as an edge computing device 120, the edge computing device 120 receives 512 the captured 3D representation of the head 102 from the sending computing device 110, thereby generating the captured 3D representation of the head. acquisition The edge computing device 120 is further operable to identify 503 the locations of a set of facial landmarks in the captured 3D representation, the set of facial landmarks including facial landmarks that delineate boundaries of a human face. The edge computing device 120 is further operable to determine 504 the pose of the head 102. The edge computing device 120 is further operable to determine 505 the boundary 310 between the inner portion 320 and the outer portion 330 of the captured 3D representation 300 based on the locations of the set of facial landmarks identified in 503. The inner portion 320 of the captured 3D representation 300 represents the face of the participant 101. The edge computing device 120 is further operable to generate 507 an avatar representation corresponding to the outer portion 330 of the captured 3D representation 300 using an ML model trained on human heads with the pose determined in 504 of the head 102 as input.
[0048] Optionally, the edge computing device 120 may be further operable to extract 506 an inner portion 320 of the captured 3D representation 300 and merge 508 the extracted inner portion 320 of the captured 3D representation 300 and the generated avatar representation into a merged 3D representation of the head 102.
[0049] The edge computing device 120 may be further operable to transmit 521 the merged 3D representation of the head 102 to the receiving computing device 130, and the merged 3D representation of the head 102 is displayed at 509 using the display device 131.
[0050] Optionally, the edge computing device 120 may be operable to train 510 the ML model using at least the outer portion 330 of the captured 3D representation 300 and the determined pose of the head 504. Further optionally, the edge computing device 120 may be operable to train 510 the ML model further based on the inner portion 320 of the captured 3D representation 300.
[0051] 5A, in the embodiment shown in FIG. 5B and described below, part of the processing has been moved from the edge computing device 120 to the receiving computing device 130. In this case, the edge computing device 120 and the receiving computing device 130 in combination implement an embodiment of the present invention. In other words, the present invention is embodied as a system of computing devices for generating 3D representations of the heads of participants in a video communication session.
[0052] More specifically, the edge computing device 120 receives 512 the captured 3D representation of the head 102 from the sending computing device 110, thereby generating the captured 3D representation of the head. acquisition The edge computing device 120 is further operable to identify 503 locations of a set of facial landmarks in the captured 3D representation, the set of facial landmarks including facial landmarks that delineate boundaries of a human face. The edge computing device 120 is further operable to determine 504 a pose of the head 102 and transmit 523 the determined pose of the head 102 to the receiving computing device 130. The edge computing device 120 is further operable to determine 505 a boundary 310 between an inner portion 320 and an outer portion 330 of the captured 3D representation 300 based on the identified locations of the set of facial landmarks in 503. The inner portion 320 of the captured 3D representation 300 represents the face of the participant 101. The edge computing device 120 is further operable to optionally extract 506 an inner portion 320 of the captured 3D representation 300 and transmit 524 the extracted inner portion 320 of the captured 3D representation 300 to the receiving computing device 130.
[0053] The receiving computing device 130 is operable to generate 507 an avatar representation corresponding to the outer portion 330 of the captured 3D representation 300 using an ML model trained on a human head, with the pose determined at 504 of the head 102 received by the receiving computing device 130 at 523 as input. Optionally, the receiving computing device 130 is operable to merge 508 the inner portion 320 of the captured 3D representation 300 received at 524 and the avatar representation generated at 507 into a merged 3D representation of the head 102. The receiving computing device 130 may further be operable to display 509 the merged 3D representation of the head 102 using the display device 131.
[0054] Optionally, the edge computing device 120 may be further operable to train 510 the ML model using at least the outer portion 330 of the captured 3D representation 300 and the determined pose of the head 504. Further optionally, the edge computing device 120 may be operable to train 510 the ML model further based on the inner 320 portion of the captured 3D representation 300, i.e., using a substantially complete captured 3D representation of the head 102. In this case, the edge computing device 120 may be operable to send 525 the updated ML model to the receiving computing device 130.
[0055] A further embodiment of a system of computing devices for generating a 3D representation of the head of a participant in a video communication session is shown in Figure 5C. Compared to the embodiment shown in Figure 5B, the additional operations involved in generating the 3D representation of the head 102 have been moved from the edge computing device 120 to the receiving computing device 130.
[0056] More specifically, the edge computing device 120 receives 512 the captured 3D representation of the head 102 from the sending computing device 110, thereby generating the captured 3D representation of the head. acquisition The edge computing device 120 is further operable to identify 503 the location of a set of facial landmarks in the captured 3D representation, the set of facial landmarks including facial landmarks that indicate boundaries of a human face. The edge computing device 120 is further operable to determine 504 a pose of the head 102 and transmit 523 the determined pose of the head 102 to the receiving computing device 130. The edge computing device 120 is further operable to determine 505 a boundary 310 between an inner portion 320 and an outer portion 330 of the captured 3D representation 300 based on the identified locations of the set of facial landmarks in 503 and transmit 527 the determined boundary between the inner portion and the outer portion to the receiving computing device 130. The inner portion of the captured 3D representation represents the face of the participant 101.
[0057] The receiving computing device 120 is optionally operable to extract 506 the inner portion 320 of the captured 3D representation 300 using the boundary 310 received at 527 between the inner portion 320 and the outer portion 330 of the captured 3D representation 300.
[0058] The receiving computing device 130 is operable to generate 507 an avatar representation corresponding to the outer portion 330 of the captured 3D representation 300 using an ML model trained on a human head, with the pose determined at 504 of the head 102 received by the receiving computing device 130 at 523 as input. Optionally, the receiving computing device 130 is operable to merge 508 the inner portion 320 of the captured 3D representation 300 received at 524 and the avatar representation generated at 507 into a merged 3D representation of the head 102. The receiving computing device 130 may further be operable to display 509 the merged 3D representation of the head 102 using the display device 131.
[0059] Optionally, the receiving computing device 130 may be further operable to train 510 the ML model using at least the outer portion 330 of the captured 3D representation 300 and the determined pose of the head 504. Further optionally, the edge computing device 120 may be operable to train 510 the ML model further based on the inner portion 320 of the captured 3D representation 300, i.e., using a substantially complete captured 3D representation 300 of the head 102.
[0060] Although embodiments of the present invention have been described with respect to a particular distribution of operations involved in generating 3D representations of the heads of participants in a video communication session between a sending computing device 110, an edge computing device 120, and a receiving computing device 130, respectively, as shown in Figures 5A-5C, those skilled in the art may readily envision alternative forms for distributing the operations involved in generating 3D representations of the heads of participants in a video communication session between a sending computing device 110, an edge computing device 120, and a receiving computing device 130.
[0061] In the following, an embodiment of a processing circuit 602 included in a computing device 600 for generating a 3D representation of the head of a participant in a video communication session will be described with reference to Figure 6. An embodiment of the processing circuit 600 may be included in one or more of the sending computing device 110, the edge computing device 120, and the receiving computing device 130.
[0062] The processing circuit 602 may comprise one or more processors 603, such as a central processing unit (CPU), a microprocessor, an application processor, an application specific processor, a digital signal processor (DSP) including a graphics processing unit (GPU), and an image processor, or a combination thereof, and a memory 604 comprising a computer program 605 comprising instructions. When executed by the processor(s) 603, the instructions cause the computing device 600 to operate in accordance with embodiments of the present invention described herein. When operations involved in generating 3D representations of the heads of participants in a video communication session are distributed among two or more of the sending computing device 110, the edge computing device 120, and the receiving computing device 130, their respective instructions, when executed by their respective processors 603, cause two or more of the sending computing device 110, the edge computing device 120, and the receiving computing device 130 to operate in a cooperative manner in accordance with embodiments of the present invention described herein. The memory 604 may be, for example, random access memory (RAM), read only memory (ROM), flash memory, etc. The computer program 605 may be downloaded into the memory 604 by the network interface circuitry 601 as a data carrier signal carrying the computer program 605. The processing circuitry 602 may alternatively or additionally comprise one or more application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc., operable to cause the computing device 600 to operate in accordance with embodiments of the present invention as described herein.
[0063] The network interface circuitry 601 may include one or more of a cellular modem (e.g., GSM, UMTS, LTE, 5G or higher generation), a WLAN / WiFi modem, a Bluetooth modem, an Ethernet interface, an optical interface, etc. for exchanging data between the computing device 600 and other computing devices, particularly between the sending computing device 110, the edge computing device 120, and the receiving computing device 130, and a communications network 140 that may include the Internet and one or more RANs.
[0064] In the following, an embodiment of a method 700 for generating a 3D representation of the head 102 of a participant 101 in a video communication session will be described with reference to FIG.
[0065] The method 700 is performed by the computing device 600 to generate a captured 3D representation of the head 102. acquisition The method 700 includes: determining 701 a position of a set of facial landmarks on the captured 3D representation; and identifying 702 a position of a set of facial landmarks on the captured 3D representation. The set of facial landmarks includes facial landmarks that indicate a boundary of a human face. The method 700 further includes determining 703 a pose of the head 102; and determining 704 a boundary between an inner portion and an outer portion of the captured 3D representation. The boundary is determined 704 based on the positions of the set of facial landmarks identified in 702. The inner portion of the captured 3D representation represents the participant's face. The method 700 further includes generating 705 an avatar representation corresponding to the outer portion of the captured 3D representation. The avatar representation is generated 705 using an ML model trained on human heads, with the pose determined in 703 of the head 102 as input. The ML model may optionally be trained on the head 102 of the participant 101.
[0066] The method 700 optionally further includes extracting 707 an interior portion of the captured 3D representation and merging 708 the interior portion of the captured 3D representation extracted in 707 with the avatar representation generated in 705 into a merged 3D representation of the head 102.
[0067] The method 700 optionally further includes displaying 709 the merged 3D representation of the head 102 using a display device 131. The display device 131 may be any one of a computer display, a television, an AR device, a VR device, an MR device, an XR device, and an HMD device.
[0068] The captured 3D representation of the head 102 acquisition Doing 701 may include capturing a 3D representation of the head 102 using the 3D sensor 111. The 3D sensor 111 may comprise one or more of a 3D camera, a LiDAR, and an optical 3D sensor.
[0069] The method 700 optionally retrieves the ML model from data storage associated with the participant 101. acquisition The method further includes:
[0070] Method 700 optionally further includes training 706 an ML model using at least the outer portion of the captured 3D representation and the pose of the head determined in 703. The ML model is optionally further trained based on the inner portion of the captured 3D representation.
[0071] It will be appreciated that method 700 may include additional, alternative, or modified steps in accordance with those described throughout this disclosure. The method may also be performed in a collaborative manner by two or more computing devices, for example, two or more of sending computing device 110, edge computing device 120, and receiving computing device 130.
[0072] An embodiment of method 700 may be implemented as a computer program 605 comprising instructions that, when executed by computing device 600, cause computing device 600 to perform method 700 and become operable according to embodiments of the invention described herein. Computer program 605 may be stored on a computer-readable data carrier such as memory 604. Alternatively, computer program 605 may be carried by a data carrier signal, e.g., downloaded to memory 604, via network interface circuitry 601.
[0073] Those skilled in the art will appreciate that the present invention is in no way limited to the embodiments described above. On the contrary, many modifications and variations are possible within the scope of the appended claims.
Claims
1. 1. An edge computing device (120, 600) for generating a three-dimensional (3D) representation of a head (102) of a participant (101) in a video communication session, the edge computing device comprising a processing circuit (602), the processing circuit (602) configured to: receiving (512) a captured 3D representation (300) of the head (102) from a sending computing device (110); identifying (503) locations (1-27) of a set of facial landmarks in the captured 3D representation (300), the set of facial landmarks including facial landmarks that delineate boundaries of a human face; determining (504) a pose of the head (102); transmitting (523) the determined pose of the head (102) to a receiving computing device (130); determining (505) a boundary (310) between an inner portion (320) and an outer portion (330) of the captured 3D representation (300) based on the identified positions (1-27) of the set of facial landmarks, wherein the inner portion (320) of the captured 3D representation (300) represents the face of the participant (101); extracting (506) the inner portion (320) of the captured 3D representation (300); transmitting (524) the extracted interior portion (320) of the captured 3D representation (300) to the receiving computing device (130); An edge computing device (120, 600) that causes the edge computing device to be operable to perform the steps described above.
2. training (510) an ML model using at least the outer portion (330) of the captured 3D representation (300) and the determined pose of the head; sending (525) the updated ML model to the receiving computing device (130); The edge computing device (120, 600) of claim 1, further operable to:
3. 3. The edge computing device (120, 600) of claim 2, operable to train (510) the ML model further based on the inner portion (320) of the captured 3D representation (300).
4. 4. The edge computing device (120, 600) of claim 1, wherein the captured 3D representation (300) and the inner portion of the captured 3D representation (300) are point cloud, mesh, or depth map images.
5. 1. A receiving computing device (130, 600) for generating a three-dimensional (3D) representation of a head (102) of a participant (101) in a video communication session, the receiving computing device comprising a processing circuit (602) configured to: receiving (523) the head pose from an edge computing device (120); receiving (524) an inner portion (320) of a captured 3D representation (300) of the head (102) from the edge computing device (120), wherein the inner portion (320) of the captured 3D representation (300) represents the face of the participant (101); generating (507) an avatar representation corresponding to an outer portion (330) of the captured 3D representation (300) using a machine learning (ML) model trained on human heads with the received pose of the head as input; Merging (508) the received interior portion (320) of the captured 3D representation (300) and the generated avatar representation into a merged 3D representation of the head (102); a receiving computing device (130, 600) that causes the receiving computing device (130, 600) to be operable to perform the steps:
6. 6. The receiving computing device (130, 600) of claim 5, further operable to display (509) the merged 3D representation of the head (102) using a display device (131).
7. 7. The receiving computing device (130, 600) of claim 6, wherein the display device (131) is any one of a computer display, a television, an augmented reality (AR) device, a virtual reality (VR) device, a mixed reality (MR) device, an extended reality (XR) device, and a head-mounted display (HMD) device.
8. 8. The receiving computing device (130, 600) of any one of claims 5 to 7, wherein the ML model is trained on the head (102) of the participant (101).
9. 9. The receiving computing device (130, 600) of claim 5, wherein the captured 3D representation (300), the inner portion of the captured 3D representation (300), the merged 3D representation, and the avatar representation are point cloud, mesh, or depth map images.
10. 10. A system of computing devices (120, 130, 600) for generating a three-dimensional (3D) representation of a head (102) of a participant (101) in a video communication session, the system comprising an edge computing device (120) according to any one of claims 1 to 4 and a receiving computing device (130) according to any one of claims 5 to 9.
11. 1. A method (700) for generating a three-dimensional (3D) representation of a head (102) of a participant (101) in a video communication session, the method being performed by an edge computing device (120, 600), comprising: receiving (701) a captured 3D representation (300) of the head (102) from a sending computing device (110); identifying (702) locations (1-27) of a set of facial landmarks in the captured 3D representation (300), the set of facial landmarks including facial landmarks that delineate boundaries of a human face; determining (703) a pose of the head (102); transmitting the determined pose of the head (102) to a receiving computing device (130) and determining (704) a boundary (310) between an inner portion (320) and an outer portion (330) of the captured 3D representation (300) based on the identified positions (1-17) of the set of facial landmarks, wherein the inner portion (320) of the captured 3D representation (300) represents the face of the participant (101); extracting (707) the inner portion (320) of the captured 3D representation (300); transmitting the extracted interior portion (320) of the captured 3D representation (300) to the receiving computing device (130); A method (700) comprising:
12. training (706) an ML model using at least the outer portion (330) of the captured 3D representation (300) and the determined pose of the head (102); sending the updated ML model to the receiving computing device (130); 12. The method (700) of claim 11, further comprising:
13. The method (700) of claim 12, wherein the ML model is trained (706) further based on the interior portion (320) of the captured 3D representation (300).
14. 14. The method (700) of any one of claims 11 to 13, wherein the captured 3D representation (300) and the interior portion of the captured 3D representation (300) are point cloud, mesh, or depth map images.
15. 1. A method (700) for generating a three-dimensional (3D) representation of a head (102) of a participant (101) in a video communication session, the method being performed by a receiving computing device (130, 600), comprising: receiving the head pose from an edge computing device (120); receiving an inner portion (320) of a captured 3D representation (300) of the head (102) from the edge computing device (120), the inner portion (320) of the captured 3D representation (300) representing the face of the participant (101); generating an avatar representation corresponding to an outer portion (330) of the captured 3D representation (300) using a machine learning (ML) model trained on a human head, with the received pose of the head as input; merging the received interior portion (320) of the captured 3D representation (300) and the generated avatar representation into a merged 3D representation of the head (102); A method (700) comprising:
16. 16. The method (700) of claim 15, further comprising displaying (709) the merged 3D representation of the head (102) using a display device (131).
17. 17. The method of claim 16, wherein the display device is one of a computer display, a television, an augmented reality (AR) device, a virtual reality (VR) device, a mixed reality (MR) device, an extended reality (XR) device, and a head-mounted display (HMD) device.
18. 18. The method (700) of any one of claims 15 to 17, wherein obtaining (701) a captured 3D representation (300) of the head (102) comprises capturing the 3D representation of the head using a 3D sensor (111).
19. 20. The method (700) of claim 18, wherein the 3D sensor (111) comprises one or more of a 3D camera, a LiDAR, and an optical 3D sensor.
20. 20. The method (700) of any one of claims 15 to 19, wherein the ML model is trained on the head (102) of the participant (101).
21. 21. The method (700) of any one of claims 15 to 20, wherein the captured 3D representation (300), the inner portion of the captured 3D representation (300), the merged 3D representation, and the avatar representation are point cloud, mesh, or depth map images.
22. 22. A computer program (605) comprising instructions that, when executed by a computing device (600), cause the computing device (600) to perform a method according to any one of claims 15 to 21.
23. A computer readable memory having stored thereon a computer program (605) according to claim 22.
Citation Information
Patent Citations
Avatar creation user interface
JP2019207670A
Methods and systems for constructing an animated 3D facial model from a 2d facial image
US20200020173A1
Virtual 3D communications with actual to virtual cameras optical axes compensation
US20210360195A1