Method and apparatus for immersive video conferencing

The method uses neural networks to encode and decode 3D head models, including separate face and hair models, addressing facial posture and hair rendering issues in telepresence systems, achieving accurate and efficient image reconstruction.

JP2026509906APending Publication Date: 2026-03-25INTERDIGITALCE PATENT HLDG SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-08
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing immersive telepresence systems face challenges in achieving proper facial posture and eye contact, and faithful rendering of the hair region, particularly in telepresence video conferencing, due to limitations in head pose adjustment and hair modeling.

Method used

A method involving neural networks for encoding and decoding semantic description data of 3D geometric and photometric head models, including separate 3D models of the face and hair, to synthesize a composite image that ensures accurate hair rendering and proper eye contact, using autoencoders and Generative Adversarial Networks (GANs) for image reconstruction.

Benefits of technology

The method achieves faithful reconstruction of both the face and hair, ensuring accurate eye contact and reducing data transmission requirements, while maintaining photorealism and independence from input image resolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509906000001_ABST
    Figure 2026509906000001_ABST
Patent Text Reader

Abstract

A method and apparatus for encoding / decoding semantic description data representing a 3D head model including facial and hair modeling. Such a method and apparatus implements a neural network. In one embodiment, an image including a user's head is encoded by extracting semantic description data representing the 3D geometric and photometric models of the user's face and semantic description data representing the 3D model of the hair region of the user's head. For example, the semantic description data representing the 3D model of the user's hair includes a set of guide curves representing the shape of hair wisps in the input image and the main hair color. In another embodiment, an image of the user's head in a virtual environment is generated from a composite image of the user's face and a composite image of the user's hair using the received semantic description data representing the 3D head model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This embodiment generally relates to a method and apparatus for encoding / decoding semantic description data representing a 3D hair model for immersive telepresence. This embodiment also generally relates to a method and apparatus for encoding or decoding based on a neural network.

Background Art

[0002] Cross-reference of related applications This application claims the benefit of European Patent Application No. 23305338.8, filed on March 13, 2023, the entire disclosure of which is incorporated herein by reference.

[0003] background Telepresence refers to, for example, the use of virtual reality technology to apparently participate in a distant event. A common application can be found in telepresence video conferencing systems that immerse participants in a single common environment. Specifically, in a typical use case, such a system intends for a user sitting at a conference room table to have the impression that other participants are sitting at the same table in the same conference room and looking directly at him or her while talking.

[0004] Immersive telepresence systems require computer vision processing for the capture of remote participants in order to achieve their goals. Usually, these captures are 2D videos acquired by commercially available cameras. The first problem is the head pose. That is, the position and orientation of the head of the remote participant in the received image need to be changed on the receiving side to establish eye contact with the user within the user's viewing device. The second problem is the rendering of the hair region, for example, to conform to the head pose or to be faithful to the actual appearance of the remote participant.

[0005] Therefore, an efficient immersive telepresence system is desirable to address two known problems with telepresence: (a) achieving proper facial posture and eye positioning to support proper eye contact, and (b) faithful rendering of the hair area. [Overview of the Initiative]

[0006] According to various embodiments, methods and apparatus are provided for encoding and decoding semantic description data representing 3D geometric and photometric head models for immersive telepresence. The head model includes at least a 3D model of the user's face and a 3D model of the user's hair.

[0007] According to one embodiment, a method is provided which includes receiving semantic description data representing a 3D model of the face region of the user's head in an image. A method is provided which includes receiving semantic description data representing a 3D model of the hair region of the user's head in an image, compositing an image of the face region of the user's head, compositing an image of the hair region of the user's head, and generating an image of the user's head from the composite image of the face region of the user's head and the composite image of the hair region of the user's head.

[0008] According to another embodiment, the method includes receiving an input image including the user's head; determining semantic description data representing a 3D model of the face region of the user's head in the input image by applying a neural network to the image; determining semantic description data representing a 3D model of the hair region of the user's head in the input image by applying a neural network to the input image; and providing semantic description data representing a 3D model of the face region of the user's head and semantic description data representing a 3D model of the hair region of the user's head for synthesizing an image of the user's head. According to one embodiment, up to two networks are used for face detection and face encoding, respectively, to determine semantic description data representing a 3D model of the face region. According to one embodiment, up to three networks are used for hair segmentation, guide curve estimation, and guide curve encoding, respectively, to determine semantic description data representing a 3D model of the hair region.

[0009] One or more embodiments also provide a device comprising one or more processors configured to perform any embodiment of the method described above.

[0010] One or more embodiments also provide a computer program that, when executed by one or more processors, includes instructions causing one or more processors to perform any of the methods of the embodiments described above. One or more embodiments also provide a computer-readable storage medium storing instructions for editing a video shot, instructions for encoding at least one image or video, or instructions for decoding at least one image or video, according to any of the embodiments described above.

[0011] One or more embodiments also provide a bitstream containing image or video data encoded according to any embodiment of the encoding method described above. One or more embodiments also provide a computer-readable storage medium storing the bitstream described above.

[0012] One or more embodiments also provide a method for transmitting a bitstream containing image or video data encoded according to any embodiment of the encoding method described herein. One or more embodiments also provide an apparatus for receiving a bitstream containing image or video data encoded according to any embodiment of the encoding method described herein. [Brief explanation of the drawing]

[0013] [Figure 1] A block diagram of a system capable of carrying out an embodiment of this model is shown. [Figure 2] A block diagram of a system that can implement an embodiment of this model according to another embodiment is shown. [Figure 3] A schematic diagram of a telepresence system that can implement an embodiment of this model is shown below. [Figure 4] A schematic diagram shows a telepresence system equipped with an encoder and a decoder that can carry out an embodiment of this model. [Figure 5] A schematic diagram of a telepresence system comprising an encoder and decoder that implements the method according to the present invention is shown. [Figure 6] Examples of components of a 3D environment model according to the embodiment are shown. [Figure 7] A method for decoding at least one image according to an embodiment is shown. [Figure 8] A method for encoding at least one image according to an embodiment is shown. [Figure 9]This example shows two remote devices communicating via a communication network related to this principle. [Figure 10] The syntax of a signal relating to one example of this principle is shown below. [Modes for carrying out the invention]

[0014] We will immediately explain this principle in the specific case of immersive video conferencing. However, this principle is not limited to video conferencing and can be directly and clearly derived to any telepresence system in which user representation is driven by remote capture of the user's head by a camera. Such systems include, but are not limited to, gaming frameworks in which users are represented by avatars and their movements and expressions are driven by remote live video capture, or, more generally, frameworks that fall within the realm of the metaverse in which participants interact in a virtual environment through an embodiment as an avatar, and the appearance, movements, and expressions of the avatar are driven by remote live video capture of the participant's head. Some examples of such frameworks can host commercial applications, such as e-learning, e-tourism, and e-commerce.

[0015] Figure 3 schematically shows a telepresence system in which an embodiment of this embodiment may be implemented, according to an embodiment.

[0016] The telepresence system in Figure 3 includes three communication devices D1, D2, and D3 connected via a communication network. Communication device D1 includes a camera that captures a scene as a series of images forming video data. The captured scene here consists of at least the head of a first user P1, for example, standing in front of a table T. Communication device D1 further includes a display for rendering the video, and remote users P2 and P3 are displayed in an immersive environment that makes them appear to be looking at user P1, for example, in front of the display at the same table T. As shown on the right side of Figure 3, user P1 is also displayed by the display of communication device D2 or D3 in an immersive environment that makes them appear to be looking at user P2 or P3, for example, in front of the display at the same table T.

[0017] To enable the rendering of a common immersive environment in the telepresence system, communication device D1 includes a transmitter / encoder used to process captured video data and provide descriptive data for synthesizing the head of user P1 on remote communication device D2 or D3, as described later. Communication device D1 also includes a receiver / decoder for receiving and processing the descriptive data provided by remote communication devices D2 and D3 and rendering the heads of users P2 and P3 in the immersive video displayed by communication device D1.

[0018] Similarly, the communication device D2 can play an immersive video such that a user P2 watching the immersive video played on the device D2 gets the impression that the user P1 is sitting at the same table T and directly looking at him / her while he / she is talking. This is made possible, for example, by a receiver / decoder of the device D2 implementing a semantic compression scheme that extracts, encodes, transmits, and decodes a 3D model of the face of the participant, as shown in, for example, FIG. 4. This is in contrast to conventional compression systems where an image is encoded as an array of pixels regardless of its semantic content. Thus, according to the embodiment of FIG. 4 or FIG. 5, instead of transmitting / encoding 2D video data representing the head of the first user, the transmitter / encoder of the device D1 processes the 2D video data by applying an encoder to the video data to obtain semantic description data representing a 3D model of the head of the first user in the first video data, e.g., a 3D geometry and photometric model. And it provides semantic description data for rendering the head of the first user in the immersive video. According to different alternative embodiments shown in FIG. 4 or FIG. 5, the head can include the face of the user, or the head can include the face and the hair of the user. Advantageously, the embodiment of FIG. 4 or FIG. 5 enables a significant reduction in the amount of data transmitted over the communication network. In addition to the low bitrate, those skilled in the art will understand that the semantic data does not depend on the resolution of the video data on which it is displayed. Thus, the compression efficiency is more important as the resolution of the displayed video is higher.

[0019] Advantageously, the above method could also be initiated on the user's smartphone, the user's laptop, or deployed in the cloud of a social network.

[0020] Figure 4 schematically shows a telepresence system with an encoder and a decoder in which aspects of the present embodiment may be implemented, according to an embodiment. One possible semantic compression scheme for video conferencing is described in European Patent Application No. 22306339.7 filed in September. It was filed in December 2022 by the same applicant and is shown in Figure 4. In each image 410 captured on the transmitting side, a face region is detected 420. The detected face sub-image is mapped to a semantic 3D model using a pre-trained autoencoder neural network composed of a face model encoder 430 and a face model decoder 450. This autoencoder is person-general because it is trained on a large collection of photos including a wide variety of identities, facial expressions, lighting, and head poses.

[0021] The face model encoder, also referred to as the transmitter, includes an encoding module 430 that performs tasks corresponding to the 3D model extraction module. Semantic description data representing the parametric 3D face model 440 of the sender's face is extracted from the captured two-dimensional video of the sender's face. This model corresponds to the inner part of the face including the eyes, nose, and mouth, but not the hair at this stage. However, according to a variant embodiment, facial hair such as beards, eyebrows, and mustaches is managed not by a hair model but by a face region model (as a specific texture). The sender semantic description data includes at least ● an indication of the sender's head pose, i.e., the 3D rotation and translation of the face within the image with respect to the front parallel view point, ● an indication of the identity representing the sender's physiognomy with a neutral expression, e.g., a 3D mesh representing the 3D geometry of the sender's face with a neutral expression, ● an indication of the expression representing an emotional expression with respect to the sender's neutral expression, e.g., a plurality of displacements of the vertices of the 3D mesh resulting from the expression, which usually results from showing emotions and / or speaking, and ● Appearance indications that represent the texture of the sender's face Includes.

[0022] The face model is transmitted over the network, then decoded at the receiving end to reconstruct the face subimages captured at the sending end. Importantly, the semantic face model deals only with the inner regions of the face, i.e., the areas surrounding the eyes, nose, and mouth, but does not encode the hair and bust regions.

[0023] In each receiving device, the face model decoder 450 retrieves semantic description data of the sender's 3D face model and generates an image of the inside 460 of the sender's face. According to certain features of this embodiment, the head pose and appearance representing the texture of the sender's face in the 3D model of the sender's face are acquired to further drive the computation of the sender's face image, which can be overlaid on top of the rendering of the environment. The generator module 470, located in the backend of the semantic compression pipeline in Figure 4, reconstructs the image 480 that will be displayed to the receiving participant. Its purpose is dual. First, it adds realism to the synthetic reconstruction of the inside region 460 of the face, driven by the semantic 3D face model. Second, it reconstructs a complete image, essentially the hair and bust region and the background, by "hallucinating" the missing elements. The generator 470 is individual-specific. The generator 470 is a neural network trained offline on images or videos of the face of the person in question. Therefore, the missing elements it generates are determined by the content of the training images and video, not by the content of the live capture in the video conference session.

[0024] In addition to allowing the receiving end to edit components of the 3D model, such as head pose, another advantage of semantic face compression lies in the compactness of the transmitting model, which enables transmission at very low bitrates. Importantly, the model content, and therefore the bitrate, is independent of the input image resolution.

[0025] However, this semantic compression scheme still presents several problems. The first problem with the semantic compression scheme in Figure 4 is that, generally, only the inner region of the face in the reconstructed image on the receiving end is faithful to the appearance of the person on the transmitting end. In fact, other elements of the head, particularly the hair area, are hallucinated by the generator based on the content of the video images of the person trained offline. The hairstyle of the person recorded in the offline training video and images is hallucinated by the generator and displayed to participants in the video conference session, even if the person has changed their hairstyle in the meantime. This is a major drawback of the compression scheme and can result in a discrepancy between the ground truth appearance of the person captured by the transmitting device and the appearance of the person displayed to other participants in the video conference session.

[0026] The second problem with the semantic compression method in Figure 4 is that, unlike the inner facial region, the hair region should not be edited, for example, to adjust head posture to achieve eye contact. In fact, the hair region in the reconstructed image is hallucinated by the generator module and is not based on a 3D model of the person's hair. While the generator can learn to some extent to adapt the rendering of the hair region to the head posture in the input inner facial image, it is unlikely to be faithful to ground truth as if a 3D model of hair were sent and its geometric elements were edited to match the target head posture.

[0027] At least one embodiment avoids these two problems by separately estimating, encoding, and transmitting a semantic 3D model of the person's hair, in addition to the semantic 3D model of the face. As a result, the reconstruction of the hair region within the head image is combined with the reconstruction of the inner region of the face, making it possible to provide the generator with a nearly complete composite image of the head upon input. This ensures that the appearance of the hair in the composite image matches the hair features in the transmitted image. Furthermore, it frees the generator from the task of distorting the hair region in addition to the background and other facial regions of the image, thereby potentially improving the quality of the reconstructed image at the receiving end. Moreover, the transmitted 3D semantic model of the hair can be edited at the receiving end, and in particular, its position, scale, and orientation can be adjusted to ensure that the hair region in the reconstructed image matches the desired rendering viewpoint in the receiving device of the video conference participant.

[0028] Figure 5 schematically illustrates a telepresence system with an encoder and decoder that implements a method according to at least one embodiment. For clarity, but without loss of generality, the telepresence scheme focuses on a scenario with only two participants: a sender and a receiver.

[0029] In the configuration of Figure 5, the upper path of the figure describes the encoding and decoding of the inner region of the face in the detected head region in the input image, as shown in Figure 4 and described in EP Patent Application No. 22306339.7 filed in September 12, 2022. In the preliminary stage, upon receiving an input image 510 containing the user's head, it is input to the system. The neural network face autoencoder, consisting of a neural network face encoder and a subsequent neural network face decoder, is trained on a large collection of faces with various facial features, head poses, expressions, and luminescence. The face encoder 530 outputs a handcrafted semantic 3D face model 540 consisting of several vectors representing the following components: ● Translation, rotation, and scaling that define the viewpoint of the face based on a frontal parallel viewpoint where the face is viewed from the front. ● An identity that represents the sender's facial features with a neutral expression, for example, an indication of a 3D mesh that represents the 3D geometry of the sender's face with a neutral expression, ● Typically, as a result of showing emotion and / or uttering, an indication of an emotional expression relative to the sender's neutral expression, such as the multiple displacements of vertices of a 3D mesh caused by a facial expression. ● An appearance indicator representing the texture of the user's face on the surface of a 3D mesh.

[0030] According to this embodiment, the head pose of the 3D model of the sender's face is not transmitted over the network, but is determined during decoding to enable eye contact in an immersive environment. Therefore, as previously shown in Figure 4, in steps 520 and 530, semantic description data representing the 3D model of the user's head and face region in the input image is determined by applying a neural network to the image.

[0031] The face decoder implements a differentiable image formation model that reconstructs the face region within a head image.

[0032] A notable feature of this embodiment is that the lower path corresponds to the upper path for the hair region. Hair consists of numerous fine fibers (strands) of the order of 100,000 rooted in the scalp. A simplified model of this complex 3D geometry can be obtained by considering that the strands are grouped into wisps, and each wisp can be thought of as a cylinder having a similar 3D shape to the strands. The central strand of each wisp is called a guide curve. A set of guide curves for all the wisps in a haircut forms a set of guide curves, providing a simplified 3D hair model from which a hair image can be reconstructed. Those skilled in the art will note that the wisps can be made smaller as needed to model isolated strands or small groups of strands.

[0033] According to at least one embodiment, hair regions in an input image of a video conferencing system are encoded into a set of guide curves, compressed, and transmitted over the network as a compact hair model. At the receiving end, the set of guide curves is decompressed and mapped to the rendering of the hair regions by a neural network. Thus, referring to the lower path of the block diagram in Figure 5, the semantic description data representing a 3D model of the hair regions of the user's head in the input image is determined by applying a neural network to the input image. According to at least one embodiment, the semantic description data representing a 3D model of the user's hair includes a set of guide curves representing the shape of the hair wisps in the input image. More specifically, the processing of the hair portion in the input image includes a first segmentation step 522 of the hair regions in the input image. For example, the hair regions in the input image are segmented using a dedicated neural network trained in pairs, where the pair includes a face image as the first element of the pair and a corresponding hair segmentation mask as the second element of the pair. For example, such a network can be implemented as described in "Two-stage human hair segmentation in the wild using deepshape prior" by Y. Yan, S. Duffner, X. Naturel, A. Berthelier, C. Garcia, C. Blanc, and T. Chateau in PatternRecognitionLetters, vol. 136, pp. 293-300, 2020. The segmented hair regions are fed into a guide curve estimator NN, which generates a set of guide curves in step 524. Information representing the hair color in the segmented hair regions, such as the dominant hair color, is extracted from the segmented hair regions and sent via the network to a decoder to render the hair regions. In yet another variant embodiment, the segmentation mask of the hair regions is sent via the network to the decoder as side information to assist in the rendering step.A neural network trained on a dataset is used for this purpose, and the dataset contains pairs of image hair regions with corresponding sets of guide curves. For example, the NN guide curve estimator 524 can be implemented as described in "HairNet: Single-view hair reconstruction using a convolutional neural network" by Y. Zhou, L. Hu, W. Chen, H. Kung, X. Tong and H. Li at the European Conference on Computer Vision, 2018. In this dataset, the sets of guide curves are typically handcrafted by the artist after photographs of haircuts taken from different viewpoints. Some of these photographs provide image hair region items in pairs. Advantageously, the sets of guide curves for all haircuts in the dataset are resampled to the same number of guide curves at fixed, predetermined root positions on the scalp surface. This normalization of the guide curve estimation network output data facilitates its convergence.

[0034] According to at least one modified embodiment, a guide curve set autoencoder network is used to reduce the amount of data transmitted over the network. In fact, a guide curve set typically consists of hundreds of curves, each represented by a fairly large number of parameters. However, in practice, adjacent guide curves have similar shapes. This redundancy can be used to reduce the dimensionality of information within a guide curve set. This autoencoder network is trained on a dataset of guide curves, such as the guide curve set output by step 524. This consists of a guide curve set encoder 535 and a guide curve set decoder 552. The encoder network 535 encodes the guide curve set at the output of step 524 into a latent code of a predetermined dimension substantially smaller than the dimensions of the input guide curve set, thereby providing a compact representation of the set. This latent code is transmitted over the network and then decoded at the receiving end by the guide curve set decoder 552 to return an approximation of the encoded guide curve set. Therefore, in step 552, the compressed representation of the guide curve set is decoded, and a 3D model of the user's hair is reconstructed, including a guide curve set representing the shape of the hair whip and the main hair color.

[0035] On the decoder side, a neural network is used to render an image of hair from the – presumably – denser set of guide curves in the output of step 552, performing the inverse function of step 524.

[0036] This network 556 is trained on the same dataset as in step 524, but with the inputs and outputs reversed. It takes a decoded set of guide curves that provides a 3D model of the target hair represented in the input image 510, as well as information representing the hair color transmitted through the network, as input to synthesize an image of the hair region. If necessary, to assist the reconstruction process, a segmentation mask of the hair region image within its bounding box may be transmitted in parallel to the geometry of the set of guide curves and used as an additional input to the neural network to ensure that the hair pixels are synthesized in the correct regions of the output image. In any variation, if the density of the guide curves in the output of step 552 is too low to represent the 3D hair model, a guide curve interpolation block 554 is used to generate a denser hair model at the receiver. The synthesized image of the user's face and the synthesized image of the user's hair are synthesized into a synthesized image of the user's head 560. In other words, the inner face image region output by the face decoder 550 in the upper path and the hair region image output by the hair model renderer 556 in the lower path are combined to form a more complete, but still incomplete, reconstruction of the entire head image. What is still missing after this combination is the face region between the inner face and the hair, as well as the jaw and neck region and the image background. This problem is solved by applying a Generative Adversarial Network 570 to the combined image of the user's head to generate an image of the user's head. In a modified embodiment, the output of the combination is fed into a generator network, which is usually complemented by a discriminator network to form a Generative Adversarial Network. The generator and discriminator are trained together. The generator is fed the combined inner face and hair image, which acts as a conditional input and drives the estimation of the complete image in its output. The generator creates a plausible, realistic head image that roughly preserves the appearance of the inner face and hair areas in its input by hallucinating the missing regions.

[0037] According to this principle, semantic description data representing a 3D model of the face region of the user's head and semantic description data representing a 3D model of the hair region of the user's head are transmitted from the encoder to the decoder in order to synthesize an image of the user's head. According to a particular embodiment, the generated image 580 of the user's head is faithful to the input image of the user's head used to obtain the semantic description data representing the 3D model of the user's face and the semantic description data representing the 3D model of the user's hair. According to a particular embodiment, the image is part of a video, and the steps of this method are repeated for each input image.

[0038] Figure 6 schematically shows an example of the components of a 3D environment model according to an embodiment. The immersive environment in which video conference participants are represented is obtained from a predetermined 3D model of the scene. For example, this scene could represent a room with a floor, walls, and windows, further including a table 610 and chairs 620 around this table 610. In this example, if there are multiple participants in the video conference system, each user is assigned a predetermined chair 620 which is represented as sitting in an immersive video displayed on the receiver devices of other participants. On each receiver device, the image of the virtual environment is calculated by rendering a projection of the aforementioned 3D environment model onto the image plane of a predetermined virtual camera 630, in particular with respect to position, orientation, and optical parameters, including focal length.

[0039] Figure 7 shows a general method 700 for decoding semantic description data representing a user's head and generating an image of the user's head in an immersive environment, according to an embodiment. Semantic description data representing a user's head, such as P1 in Figure 3, includes two parts: semantic description data representing a 3D model of the user's face and semantic description data representing a 3D model of the user's hair. For example, both semantic description data form part of a bitstream. In step 710, semantic description data representing a 3D model of the user's face (i.e., inside the face) is received. For example, the 3D model of the user's face includes an identity indication representing the facial features of sender P1 in Figure 3 with a neutral expression; an expression indication representing an emotional expression with respect to sender P1's neutral expression; and an appearance indication representing the color of sender P1's face, such as an indication of the surface texture of a 3D mesh. In step 730, an image of the sender's face is synthesized from the received semantic description data representing the 3D geometric and photometric model of the remote user P1's face and the rigid head pose of the remote user's face. In the example scene of Figure 6 described above, the rigid head pose 650 is calculated so that the image obtained from the overlay represents the sender sitting in a predetermined seat assigned to him or her, with the face looking at the virtual camera. For example, the 3D head pose model consists of scale, translation, and rotation components defined in the 3D coordinate system 640 of a given 3D scene model, as shown in Figure 6. In step 720, semantic description data representing the 3D model of the user's hair is received. For example, the semantic description data representing the 3D model of the user's hair includes a set of guide curves representing the shape of the hair wisps and the main hair color. Advantageously, a compressed representation of the set of guide curves is decoded in a subsequent step (not shown in Figure 7). In a variation, the set of guide curves can be interpolated to obtain a denser set of guide curves (not shown in Figure 7). Next, in step 740, an image of the user's hair is synthesized from the received semantic description data representing a 3D hair model.In a modified version, the synthesis of the user's hair image can further take a rigid head pose of the remote user's face as input to render the user's head in a desired head pose. Finally, in step 750, a photorealistic image of the user's head is generated from the synthesized face and hair regions. According to this principle, the generated image of the user's head is faithful to the appearance of the user's face and hair in the input image used to obtain semantic description data representing the 3D model of the user's face and semantic description data representing the 3D model of the user's hair. In a modified version, generation step 750 is performed using a Generative Adversarial Network (GAN). The GAN performs two functions: firstly, to hallucinate parts of the input image that are not encoded and not transmitted over the network, such as the background; and secondly, to make the rendering of the face and head image model synthesized in steps 730 and 740 more photorealistic. This GAN is a person-specific neural network trained offline on images or videos of the target person's face. Therefore, the missing elements it generates are determined by the content of the training images and video, not by the content of the live capture in the video conference session. According to another variation, the foreground region, excluding the background and representing the user in the image output by the GAN, is segmented from the image output by the GAN and overlaid on a predetermined rendering of the virtual environment to produce the final rendered image displayed to the receiving user. Advantageously, the training images and video for the GAN are captured against a uniform background. Since the frames output by the GAN replicate this uniform background, the foreground region representing the sender in the rendered image can be effectively and efficiently extracted using color keying techniques known from the prior art. While this principle is not limited to GANs for the generation step, those skilled in the art will understand that GANs are currently the most effective implementation for generating such photorealistic images from a synthetic view.

[0040] Advantageously, Method 700 reduces the amount of data transmitted to the video conferencing system by decoding semantic description data of a 3D head model instead of two-dimensional video data. Furthermore, the method advantageously achieves proper posing of the rendered face to support proper eye contact and provides a faithful reconstruction of both the user's face and hair captured within the input image. In yet another variation, the generated image is part of a video, and decoding is repeated for each frame / image of the video.

[0041] Figure 8 shows a general method 800 for encoding semantic description data representing a user's head, according to an embodiment. In the first step 810, an input image containing the user's face is received. For example, the input image is a frame from a video. In the encoding step 820, a task corresponding to the 3D face model NN encoder in Figure 4 or Figure 5 is applied to the input image to obtain semantic description data representing a 3D geometric and photometric model of the user's face. This model covers only the interior of the head, i.e., the region including the eyes, nose, and mouth, and does not cover the hair. In a modified version, preliminary steps of face detection and cropping in the input image are performed before the encoding step. In a modified version, the semantic description data in the output of the encoding step 820 is ● An identity representing the facial features of a user with a neutral expression, for example, an indication of a 3D mesh representing the 3D geometry of the sender's face with a neutral expression. ● Indication of multiple displacements of vertices of a 3D mesh resulting from facial expressions, such as facial expressions, regarding the user's neutral expression and expressions that represent emotional expressions. ●Appearance indicators representing the texture of the user's face on the surface of the 3D mesh, and ● Rigid head pose indication representing 3D rotation and translation of the face in the input video image relative to a frontal parallel viewpoint. Includes.

[0042] In principle, the autoencoder in the modified version of Figure 4 requires a complete face model to function, so these components are extracted. However, since some of them are replaced by components in the immersive environment, it is not necessary to transmit all of them. Advantageously, this saves even more bitrate.

[0043] In step 830, a task corresponding to the NN-based encoding of the semantic 3D hair model in Figure 5 is applied to the input image to obtain semantic description data representing the 3D geometric and photometric models of the user's hair. An NN guide curve estimator is applied to the input image to generate semantic description data representing the 3D model of the hair region of the user's head in the input image. According to a modified embodiment, the encoding step 830 may also include segmenting the hair region in the input image and providing the segmented hair region to the NN guide curve estimator to generate a set of guide curves representing the shape of the hair whip and the main hair color in the input image. According to yet another modified embodiment, the encoding step 830 may further include applying a guide curve NN encoder to the set of guide curves in order to reduce the amount of semantic description data being transmitted. In step 840, the semantic description data representing the 3D model of the face region of the user's head and the semantic description data representing the 3D model of the hair region of the user's head are provided to a remote decoder to synthesize and render the user's head in the input image in a virtual environment.

[0044] Figure 1 shows a block diagram of a system in which an embodiment of this embodiment may be implemented. Figure 1 schematically shows a communication device, such as the video conferencing device of Figure 3, according to the embodiment.

[0045] According to one embodiment, the above-described method is implemented as an instruction that causes one or more processors to execute the method steps.

[0046] In one embodiment, Figure 1 illustrates a block diagram of an example of a system in which the various embodiments and configurations described above can be implemented. System 100 can be implemented as a device comprising various components described below and configured to perform one or more of the configurations described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected consumer electronics, and servers. The elements of System 100 can be implemented individually or in combination as a single integrated circuit, multiple ICs, and / or separate components. For example, in at least one embodiment, the processing and encoder / decoder elements of System 100 are distributed across multiple ICs and / or separate components. In various embodiments, System 100 is communicably coupled to other systems or other electronic devices, for example, via a communication bus or via dedicated input and / or output ports. In various embodiments, System 100 is configured to perform one or more configurations described in this application.

[0047] System 100 includes, for example, at least one processor 110 configured to execute instructions loaded therein in order to implement various embodiments described in this application. The processor 110 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140 which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 140 may, in non-limiting examples, include an internal storage device, an attached storage device, and / or a network-accessible storage device.

[0048] According to the embodiment, the system 100 includes, for example, an encoder / decoder module 130 configured to process data and provide encoded or decoded video, the encoder / decoder module 130 of which may include its own processor and memory. The encoder / decoder module 130 represents a module that may be included in the device to perform encoding and / or decoding functions. As is well known, the device may include one or both of the encoding and decoding modules. Furthermore, the encoder / decoder module 130 may be implemented as a separate element of the system 100 or may be incorporated into the processor 110 as a combination of hardware and software, as is known to those skilled in the art.

[0049] Program code loaded onto the processor 110 to perform various embodiments described in this application may be stored in a storage device 140 and subsequently loaded onto memory 120 for execution by the processor 110. Depending on the various embodiments, one or more of the processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video shots, mosaic images, warps, 3D models, color conversion information, visibility maps, matrices, variables, and one of the intermediate or final results from the processing of equations, formulas, operations, and arithmetic logic.

[0050] In some embodiments, memory within the processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for the preprocessing steps and / or processing required during video editing of the methods described herein. However, in other embodiments, memory outside the processing device (for example, the processing device may be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be memory 120 and / or storage device 140, for example, dynamic volatile memory and / or non-volatile flash memory.

[0051] Inputs to the elements of system 100 may be provided via various input devices, as shown in block 105. Such input devices include, but are not limited to, (i) an RF section for receiving RF signals wirelessly transmitted by a broadcasting station, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0052] In various embodiments, the input devices of block 105 are associated with their respective input processing elements, as is known in the art. For example, the RF portion may be associated with an element suitable for (i) selecting a desired frequency (also called signal selection or band-limiting of a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting it to a narrower frequency band to select a signal frequency band called a channel in a particular embodiment, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF portion of various embodiments includes one or more elements that perform these functions, e.g., frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may also include a tuner that performs these various functions, e.g., down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements perform frequency selection by receiving an RF signal transmitted over a wired (e.g., cable) medium, filtering it, down-converting it, and filtering it again to a desired frequency band. Various embodiments rearrange the order of the elements described above (and others), remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.

[0053] Furthermore, the USB and / or HDMI terminals may include their respective interface processors for connecting the system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, can be implemented, for example, in a separate input processing IC or within the processor 110, as needed. Similarly, aspects of USB or HDMI interface processing can be implemented, for example, in a separate interface IC or within the processor 110, as needed. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 110 and an encoder / decoder 130 operating in conjunction with memory and storage elements, which process the data stream as needed for presentation on an output device.

[0054] Various elements of system 100 can be provided within an integrated housing. Within the integrated housing, the various elements are interconnected using an internal bus known in the art, such as an IC bus, wiring, and printed circuit board, and data can be transmitted between them.

[0055] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card. The communication channel 190 may be implemented, for example, in a wired and / or wireless medium.

[0056] In various embodiments, the data is streamed to system 100 using a Wi-Fi network such as IEEE 802.11. In these embodiments, the Wi-Fi signal is received via a communication channel 190 and a communication interface 150 adapted for Wi-Fi communication. In these embodiments, the communication channel 190 is typically connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, a set-top box that distributes data via an HDMI connection of input block 105 is used to provide the streamed data to system 100. In yet another embodiment, an RF connection of input block 105 is used to provide the streamed data to system 100.

[0057] System 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripheral devices 185. In various embodiments, the other peripheral devices 185 include one or more standalone DVRs, disc players, stereo systems, lighting systems, and other devices that provide functions based on the output of System 100. In various embodiments, control signals are communicated between System 100 and the display 165, speaker 175, or other peripheral devices 185 using signaling such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output devices may be communicably coupled to System 100 via dedicated connections through their respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to System 100 using a communication channel 190 via a communication interface 150. The display 165 and speaker 175 may be integrated into a single unit with other components of System 100, such as an electronic device, for example, a television. In various embodiments, the display interface 160 includes a display driver, such as a timing controller (TCon) chip.

[0058] Alternatively, the display 165 and speaker 175 may be separated from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments where the display 165 and speaker 175 are external components, the output signal may be provided via a dedicated output connection, for example, an HDMI port, a USB port, or a COMP output.

[0059] Figure 2 shows a block diagram of a system in which an aspect of this embodiment may be implemented according to another embodiment. Figure 2 schematically shows a communication device, for example, the communication device of Figure 3 according to an embodiment. Figure 2 shows one embodiment of a device using the method described above. The device includes a processor 210 which can be interconnected to a memory 220 via at least one port. Both the processor 210 and the memory 220 can further have one or more additional interconnections to external connections. The processor 210 is also configured to receive or output images, encode at least one image, or decode at least one image using the method described above.

[0060] According to an example of this principle shown in Figure 9, in a transmission context between two remote devices A and B over a communication network NET, device A includes a processor with memory RAM and ROM configured to implement one of the embodiments of a method for encoding at least one image, as described in relation to the figure. Figures 4, 5, or 8, and device B includes a processor with memory RAM and ROM configured to implement one of the embodiments of a method for decoding at least one image, as described in relation to Figures 4, 5, or 7. For example, the network is a broadcast network adapted to broadcast / transmit an encoded image from device A to a decoding device including device B. The signal intended to be transmitted by device A carries at least one bitstream containing encoded data representing at least one image.

[0061] Figure 10 shows an example of the syntax of such a signal when at least one encoded image is transmitted via a packet-based transmission protocol. Each transmitted packet P includes a header H and a payload PAYLOAD.

[0062] Various methods are described herein, each of which includes one or more steps or actions to achieve the described method. Unless a particular order of steps or actions is required for the proper operation of the method, the order and / or use of any particular steps and / or actions may be modified or combined. Furthermore, terms such as “first,” “second,” etc., may be used in various embodiments to modify elements, components, steps, operations, etc., for example, “first decode” and “second decode.” The use of such terms does not imply any ordering of the modified operations unless specifically required. Thus, in this example, the first decode does not need to be performed before the second decode, and may occur, for example, before, during, or overlapping with the second decode.

[0063] Unless otherwise indicated or technically excluded, the embodiments described in this application may be used individually or in combination.

[0064] Various numerical values ​​are used in this application. Certain values ​​are for illustrative purposes only, and the embodiments described are not limited to these specific values.

[0065] The implementations and embodiments described herein may be carried out, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if described only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the described features may also be carried out in other forms (e.g., apparatus or programs). Apparatus may be carried out, for example, in appropriate hardware, software, and firmware. Methods may be carried out, for example, in apparatus, for example, a processor. A processor refers to a general processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. A processor also includes communication devices, for example, computers, mobile phones, personal digital assistants ("PDAs"), and other devices that facilitate the communication of information between end users.

[0066] References to “one embodiment,” “embodiment,” “one implementation,” or “implementation,” as well as other variations thereof, mean that certain features, structures, characteristics, etc., described in relation to the embodiment are included in at least one embodiment. Therefore, the appearances of the phrases “in one embodiment,” “in one embodiment,” or “in one implementation,” or “in implementation,” as well as other variations, appearing in various places throughout this application do not necessarily all refer to the same embodiment.

[0067] Furthermore, this application may also refer to "determining" various types of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory. Furthermore, this application may also mean "accessing" various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0068] Furthermore, this application can mean "receiving" various types of information. Receiving, like "accessing," is intended to be a broad term. Receiving information can include, for example, accessing information or retrieving information (for example, from memory), one or more of these. Moreover, "receiving" is typically involved in some way during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0069] The use of any of the following " / ", "and / or", and "at least one of" should be understood to include, for example, in the cases of "A / B", "A and / or B", and "at least one of A and B", the selection of only the first listed option (A), or only the second listed option (B), or the selection of both options (A and B). As further examples, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to include the selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or the selection of all three options (A, B, and C). This can be extended to many items listed, as will be obvious to those skilled in the art.

[0070] Furthermore, as used herein, the term “signal” refers, in particular, to indicating something to a corresponding decoder. Thus, in embodiments, the same parameters are used on both the encoder and decoder sides. For example, an encoder can transmit (explicitly signal) certain parameters to a decoder so that the decoder can use the same particular parameters. Conversely, if the decoder already has certain parameters, like any other, signaling can be used without transmission (implicit signaling) simply to allow the decoder to know and select the particular parameters. Bit saving is achieved in various embodiments by avoiding the transmission of actual functions. It should be understood that signaling can be achieved in various ways. For example, in various embodiments, one or more syntax elements, flags, etc., are used to signal information to a corresponding decoder. The above concerns the verb form of the word “signal,” but the word “signal” can also be used as a noun in this specification.

[0071] As will be apparent to those skilled in the art, implementations can generate a variety of signals formatted to carry information that can be stored or transmitted. The information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is well known. The signal may be stored in a processor-readable medium.

Claims

1. Receiving semantic description data representing a 3D model of the user's head and face region in the image, The system receives semantic description data representing a 3D model of the hair region of the user's head in the aforementioned image, The image of the face region of the user's head is synthesized, The image of the hair region of the user's head is synthesized, The process involves generating an image of the user's head from the composite image of the face region of the user's head and the composite image of the hair region of the user's head. Includes, method.

2. The method according to claim 1, wherein the semantic description data representing the 3D model of the hair region includes a set of guide curves representing the shape of the hair wisp and the main hair color.

3. The method according to claim 2, further comprising decoding a compressed representation of the set of guide curves.

4. The method according to claim 2, further comprising interpolating the set of guide curves to obtain a set of guide curves with higher density.

5. The above image is generated by The composite image of the face region and the composite image of the hair region are combined with the composite image of the user's head. Applying a generative adversarial network to the synthesized image of the user's head to generate an image of the user's head. The method according to claim 1, further comprising:

6. The process involves receiving a segmentation mask of the hair region of the user's head in the aforementioned image, The synthesis of the image of the hair region of the user's head is performed using the segmentation mask. The method according to claim 5, further comprising:

7. The method according to any one of claims 1 to 6, wherein the generated image is part of a video.

8. The method according to claim 1, wherein the generated image of the user's head is faithful to an input image of the user's head used to obtain the semantic description data representing the 3D model of the user's face and the semantic description data representing the 3D model of the user's hair.

9. Receiving an input image that includes the user's head, By applying an NN encoder to the aforementioned image, semantic description data representing a 3D model of the user's head and face region in the input image is determined. By applying an NN guide curve estimator to the input image, semantic description data representing a 3D model of the hair region of the user's head in the input image is determined. The semantic description data representing the 3D model of the face region of the user's head and the semantic description data representing the 3D model of the hair region of the user's head are provided to synthesize an image of the user's head. Methods that include...

10. The method according to claim 9, wherein the semantic description data representing the 3D model of the user's hair includes a set of guide curves representing the shape of the hair wisps in the input image and the main hair color.

11. Determining the semantic description data representing the 3D model of the user's hair in the input image is: Segmenting the hair region within the input image, The segmented hair region is supplied to the NN guide curve estimator to generate the set of guide curves. The method according to claim 10, further comprising:

12. Providing semantic description data to synthesize the image of the user's head is, Applying a guide curve NN encoder to the set of guide curves reduces the amount of semantic description data transmitted. The method according to claim 11, further comprising:

13. The method according to claim 11, further comprising providing a segmentation mask for the hair region.

14. The method according to any one of claims 9 to 13, wherein the input image is part of a video.

15. The method according to claim 9, wherein the input image is part of a video, and the semantic description data representing a 3D model of the user's hair is determined for each input image in the video.

16. The system receives semantic description data representing a 3D model of the user's head and face region within the image. Semantic description data representing a 3D model of the hair region of the user's head in the aforementioned image is received. Images of the face region of the user's head are combined, Images of the hair region of the user's head are combined, An image of the user's head is generated from the composite image of the face region of the user's head and the composite image of the hair region of the user's head. A device including one or more processors configured in such a manner.

17. The system receives an input image that includes the user's head. By applying an NN encoder to the aforementioned image, semantic description data representing a 3D model of the user's head and face region in the input image is determined. By applying an NN guide curve estimator to the input image, semantic description data representing a 3D model of the hair region of the user's head in the input image is determined. The semantic description data representing the 3D model of the face region of the user's head and the semantic description data representing the 3D model of the hair region of the user's head are provided to synthesize an image of the user's head. A device including one or more processors configured in such a manner.

18. A device in a video conferencing system, The apparatus according to claim 17, At least one device according to claim 16, A device that includes this.

19. A device in a video conferencing system, The apparatus according to claim 17, At least one device according to claim 16, An antenna configured to receive a bitstream, wherein the bitstream includes semantic description data representing a 3D model of a remote user's head. A display configured to display the generated image, including the synthesized head of the remote user, in an immersive video, A device that includes this.

20. A bitstream containing semantic description data for rendering a user's head from an input image, wherein the bitstream includes semantic description data representing a 3D model of the hair region of the user's head in the input image, and the semantic description data representing the 3D model of the hair region includes a set of guide curves representing the shape of the hair wisps in the input image and the main hair color.

21. The bitstream further includes semantic description data representing a 3D model of the face region of the user's head in the input image, and the semantic description data representing the 3D model of the face region is The aforementioned neutral facial expression represents the user's identity, Regarding the user's neutral facial expression, an indication of facial expression that represents the deformation of the face caused by expressing an emotional expression or speaking, and An appearance indicator representing the texture of the user's face, The bitstream of claim 20, including the bitstream of claim 20.

22. A computer-readable medium storing the bitstream according to claim 20 or 21.

23. A computer-readable medium storing instructions for causing one or more processors to perform the method described in any one of claims 1 to 15.