Generative face video coding with disentangled background

Generative face video coding with disentangled and dynamic backgrounds addresses the limitations of conventional codecs by separating foreground and background, achieving improved bandwidth efficiency and reconstruction fidelity through segmentation and inpainting techniques.

WO2025217233A1PCT designated stage Publication Date: 2025-10-16DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/023770
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-01
Filing Date
2025-04-09
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Conventional video codecs face limitations in disentangling foreground and background in video sequences, leading to entangled backgrounds and static backgrounds that lack dynamic features, which affect bandwidth efficiency and reconstruction fidelity.

Method used

Generative face video coding with disentangled and dynamic backgrounds, utilizing segmentation and inpainting models to separate foreground and background, and transmitting the inpainted background with additional motion information for dynamic replication at the decoder side.

Benefits of technology

Achieves significant bitrate savings and improved background fidelity, with up to 60% bitrate savings and enhanced subjective quality in reconstructed videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025023770_16102025_PF_FP_ABST
    Figure US2025023770_16102025_PF_FP_ABST
Patent Text Reader

Abstract

Methods and apparatus employing Generative Face Video Coding (GFVC) techniques wherein images are segmented into foreground and background parts and wherein the foreground and background parts are subjected to different respective sets of processing operations at the encoder and / or the decoder.
Need to check novelty before this filing date? Find Prior Art

Description

GENERATIVE FACE VIDEO CODING WITH DISENTANGLED BACKGROUNDCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of Indian Provisional Patent Application Ser. No. 202411028978, filed on 9 April 2024, and Indian Provisional Patent Application Ser. No.202411050189, filed on 1 July 2024, each of which is incorporated by reference herein in its entirety.FIELD OF THE DISCLOSURE

[0002] Various example embodiments relate generally to video coding and, more specifically but not exclusively, to Generative Face Video Coding (GFVC).BACKGROUND

[0003] Generative Face Video Coding (GFVC) techniques use deep learning models for compact facial feature representation and face video generation that enables high quality communication with ultra-low bandwidth. GFVC typically uses an analysis model at the encoder side for representing the motion parameters and a synthesis model at the decoder side for face video reconstruction. In at least some cases, GFVC has a potential to surpass conventional video codecs in terms of the coding efficiency and performance.BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS

[0004] Disclosed herein are various embodiments of GFVC techniques wherein images are segmented into foreground and background parts and wherein the foreground and background parts are subjected to different respective sets of processing operations at the encoder and / or the decoder.

[0005] According to one example, a generative face- video-coding method comprises: obtaining a background picture and a sequence of foreground pictures by applying image segmentation to images corresponding to a face video that includes a first picture and a sequence of second pictures; generating a compressed video bitstream by applying video compression to the first picture and the background picture; generating a face feature representation by processing the sequence of foreground pictures with a neural network; and transmitting thecompressed video bitstream and the face feature representation through a communication channel.

[0006] According to another example, a generative face- video-coding method comprises: decoding a compressed video bitstream to obtain a first picture and a background picture encoded in the compressed video bitstream; obtaining a first foreground picture via image segmentation applied to the first picture; processing the first foreground picture and a face feature representation with a neural network to obtain a sequence of second foreground pictures; and fusing each of the second foreground pictures with a respective background picture to obtain a sequence of video frames representing a face video, the respective background picture being based on the background picture obtained via the decoding.

[0007] According to yet another example, a generative face-video-coding method comprises: obtaining a base picture by applying image segmentation to a first picture of a face video and replacing a background portion of the segmented first picture by a corresponding chroma key background; generating a compressed video bitstream by applying video compression to the base picture; obtaining a sequence of driving pictures by applying image segmentation to a sequence of second pictures of the face video and replacing background portions of the segmented second picture by respective chroma key backgrounds; generating a face feature representation by processing the sequence of driving pictures with a neural network; and transmitting the compressed video bitstream and the face feature representation through a communication channel.

[0008] According to yet another example, a generative face-video-coding method comprises: decoding a compressed video bitstream to obtain a base picture encoded in the compressed video bitstream, the base picture having a corresponding chroma key background; processing the base picture and a face feature representation with a neural network to obtain a sequence of driving pictures, each of the driving pictures having a respective chroma key background; and fusing the base picture and each of the driving pictures with a background picture to obtain a sequence of video frames representing a face video encoded in the compressed video bitstream and in the face feature representation.

[0009] According to yet another example, provided is a non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the above methods.

[0010] According to yet another example, provided is an apparatus for generative face- video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: obtain a background picture and a sequence of foreground pictures by applying image segmentation to images corresponding to a face video that includes a first picture and a sequence of second pictures; generate a compressed video bitstream by applying video compression to the first picture and the background picture; generate a face feature representation by processing the sequence of foreground pictures with a first neural network; and transmit the compressed video bitstream and the face feature representation through a communication channel.

[0011] According to yet another example, provided is an apparatus for generative face- video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: decode a compressed video bitstream to obtain a first picture and a background picture encoded in the compressed video bitstream; obtain a first foreground picture via image segmentation applied to the first picture; process the first foreground picture and a face feature representation with a neural network to obtain a sequence of second foreground pictures; and fuse each of the second foreground pictures with a respective background picture to obtain a sequence of video frames representing a face video, the respective background picture being based on the background picture obtained via the decoding.

[0012] According to yet another example, provided is an apparatus for generative face- video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: decode a compressed video bitstream to obtain a base picture encoded in the compressed video bitstream, the base picture having a corresponding chroma key background; process the base picture and a face feature representation with a neural network to obtain a sequence of driving pictures, each of the driving pictures having a respective chroma key background; and fuse the base picture and each of the driving pictures with a background picture to obtain a sequence of video frames representing a face video encoded in the compressed video bitstream and in the face feature representation.

[0013] According to yet another example, provided is an apparatus for generative face- video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with theat least one processor, cause the apparatus at least to: obtain a base picture by applying image segmentation to a first picture of a face video and replacing a background portion of the segmented first picture by a corresponding chroma key background; generate a compressed video bitstream by applying video compression to the base picture; obtain a sequence of driving pictures by applying image segmentation to a sequence of second pictures of the face video and replacing background portions of the segmented second picture by respective chroma key backgrounds; generate a face feature representation by processing the sequence of driving pictures with a neural network; and transmit the compressed video bitstream and the face feature representation through a communication channel.BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Other aspects, features, and benefits of various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings, in which:

[0015] FIG. 1 is a block diagram illustrating a generic framework of GFVC algorithms according to some examples.

[0016] FIGS. 2A-2B are block diagrams illustrating extensions of Versatile Supplemental Enhancement Information (VSEI) for supporting generative face video coding according to some examples.

[0017] FIG. 3 is a block diagram illustrating a Compact Feature learning for Temporal Evolution (CFTE) workflow according to one example.

[0018] FIG. 4 is a block diagram illustrating a (GFVC) architecture according to some examples.

[0019] FIGS. 5A-5D pictorially illustrate certain operations of a segmentation module used in the GFVC architecture of FIG. 4 according to some examples.

[0020] FIGS. 6A-6B pictorially illustrate segmentation masks produced by the segmentation module used in the GFVC architecture of FIG. 4 according to one example.

[0021] FIGS. 7A-7D pictorially illustrate certain operations of an inpainting module used in the GFVC architecture of FIG. 4 according to one example.

[0022] FIG. 8 is a diagram graphically illustrating a bitstream transmitted in the GFVC architecture of FIG. 4 according to some examples.

[0023] FIGS. 9-10 are block diagrams illustrating modifications to the GFVC architecture of FIG. 4 according to some examples.

[0024] FIG. 11 is a diagram graphically illustrating an encoding process in the modified GFVC architecture of FIG. 9 according to some examples.

[0025] FIG. 12 is a diagram graphically illustrating an encoding process in the modified GFVC architecture of FIG. 10 according to some examples.

[0026] FIG. 13 is a block diagram illustrating possible sources of dynamic backgrounds according to some examples.

[0027] FIGS. 14A-14B pictorially illustrate a dynamic background caused by camera motion according to one example.

[0028] FIG. 15 is a diagram illustrating a process of updating a dynamic background based on camera motion according to one example.

[0029] FIG. 16 is a block diagram illustrating a process for deriving the background displacement due to camera / background motion according to some examples.

[0030] FIGS. 17A-17B pictorially illustrate a dynamic background caused by foreground motion according to one example.

[0031] FIG. 18 is a diagram illustrating implementations of the inpainting approach according to one example.

[0032] FIGS. 19A and 19B pictorially compare inpainting results obtained based on a single frame and based on multiple frames according to one example.

[0033] FIG. 20 is a block diagram illustrating a workflow of inpainting operations based on multiple frames according to one example.

[0034] FIG. 21 graphically illustrates two different information aggregation methods that can be implemented in the information aggregation module of the workflow of FIG. 20 according to some examples.

[0035] FIG. 22 a block diagram illustrating a GFVC architecture according to additional examples.

[0036] FIG. 23 is a diagram graphically illustrating an encoding process in the GFVC architecture of FIG. 22 according to some examples.

[0037] FIG. 24 graphically illustrates example improvements that can be achieved with the GFVC architecture of FIG. 4 according to some examples.

[0038] FIG. 25 is a block diagram of an example computing device, one or more instances of which can be used to implement various GFVC architectures, according to some examples.DETAILED DESCRIPTION

[0039] Face video communications is currently a significant area of growth due to various applications thereof, such as video conferencing, online education, online-consultation, live broadcasting, etc. For such applications, visual face data occupy most of the transmission bandwidth. Accordingly, an important component of face video communications is directed to compact characterization of facial information for low-bandwidth, high-quality, and / or high- fidelity transmission.

[0040] Conventional video codecs, such as H.264 / Advanced Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), and H.266 / Versatile Video Coding (VVC), implement tradeoffs between the applicable bit-rate constraints and reconstruction quality. While these video codecs enable popular applications of video conferencing and live broadcasting, there is still a significant room for improvements. For example, statistical characteristics of face data remain underutilized. Leveraging such statistical characteristics can potentially provide an avenue for improvements in the bandwidth efficiency of face video communications. In some examples, Model-Based Coding (MBC) can provide enhanced video compression by utilizing strong face priors, but the quality of face synthesis still needs to be improved. We have realized that application of deep generative (machine learning) models can improve the MBC via high quality face reconstruction.

[0041] FIG. 1 is a block diagram illustrating a generic framework (100) of GFVC algorithms according to some examples. At an encoder side (102) of the framework (100), an encoder (110) configured to use one the traditional codecs (e.g., AVC, HEVC, VVC) compresses a base picture of a face video to generate a first coded bitstream (112). Additionally, a plurality of driving pictures (118) is fed to a deep learning analysis model (120) for compact features extraction.The extracted features are then coded, with a parameter encoding module (124), into a second coded bitstream (128). At a decoder side (104) of the framework (100), a conventional decoder (114) operates to decode the received first coded bitstream (112) to generate a decoded base picture (116). A parameter decoding module (132) operates to decode the received second coded bitstream (128) to recover the compact representation generated with the parameter encoding module (124). A deep learning synthesis model (136) receives the decoded base picture (116) along with recovered compact representation and generates a plurality of driving pictures (138). Hence, instead of compressing the driving pictures (118) directly, the corresponding compact representation is transmitted from the encoder side (102) to the decoder side (104) of the framework (100), which leads to a significant reduction in the transmission volume (in terms of the actually transmitted bits).

[0042] Based on the above-indicated approach, a generic GFVC framework, such as the framework (100) (sometimes referred to as Compact Feature learning for Temporal Evolution inference, or CFTE), can be configured, e.g., to utilize a U-Net for inter-frames’ compact feature extraction, and VVC for the base picture compression. At the decoder side (104), dense motion and occlusion maps are generated utilizing the decoded base picture (116) and compact features, which are then input to a generative adversarial network (GAN) for generating the final frames.

[0043] One limitation of the above-described framework (100) and some other similar approaches, such as First Order Motion Model (FOMM), Face Video two Vide (FV2V), etc., is the entanglement between the foreground and background such that, in the decoded video sequence, the background surrounding the foreground may move with the later. Another limitation is in terms of the fidelity of the background, and that the background motion in the original video sequence may not be replicated in the decoded sequence. Hence, it typically contains only a static background, which lacks any dynamic features.

[0044] To address at least some of the above-indicated limitations, various embodiments disclosed herein use generative face video coding with disentangled and dynamic background. Example features of such embodiments may include some or all of the following:• A novel scheme for disentangled background in the GFVC, wherein: o A modified architecture includes additional segmentation and inpainting models, o The segmentation model separates the foreground and background information from the base and driving pictures. o The inpainting model completes the segmented background to fill in missing details.o The generated background is transmitted along with the base picture which is utilized for disentanglement at the decoder side.• A novel scheme for dynamic background in the GFVC, wherein: o The inpainted background is transmitted periodically or as needed, o Additionally, the information based on the relative offsets between the pictures is also transmitted. o The above information is utilized to model dynamic motion at the decoder side.

[0045] FIGS. 2A-2B are block diagrams illustrating extensions of Versatile Supplemental Enhancement Information (VSEI) for supporting generative face video coding according to some examples. In Joint Video Experts Team (JVET), a generative face video (GFV) SEI message is proposed and adopted in the document “JVET-AG2032: Technologies under consideration for future extensions of VSEI (version 3),” (Ref. [1]), which is incorporated herein by reference in its entirety. Two cases are supported based on the syntax gfv_drive_pic_fusion flag, as illustrated in FIGS, 2A- and 2B, respectively. Case (A), illustrated in FIG. 2A, supports an original Compact Feature learning for Temporal Evolution (CFTE) framework indicated by the setting gfv_drive_pic_fusion_flag=0. Case (B), illustrated in FIG. 2B, aims at supporting the fusion technology, various examples of which are described in more detail herein below, by setting the gfv_drive_pic_fusion_flag=l. In some examples, the disclosed fusion technology can beneficially provide about 60% bitrate savings based on the Deep Image Structure and Texture Similarity (DISTS) or Learned Perceptual Image Patch Similarity (LPIPS) image quality metrics. Subjectively, the disclosed fusion technology also shows a significant benefit in at least some examples. To improve on certain limitations of the above GFV SEI messages in Case (B), additional new syntax is provided herein below.

[0046] FIG. 3 is a block diagram illustrating a CFTE workflow (300) according to one example. The CFTE workflow (300) includes an encoder (310) and a decoder (330). Therein, the encoder (310) includes the following modules:1. A conventional video encoder (312) for compressing a base picture (302).2. A compact feature representation module (316), comprising: a) A deep learning-based architecture (feature extractor) for the extraction of compact face features from driving pictures (306). b) A feature coding module that is utilized for compression of the interpredicted residuals of the compact face features.With these modules, the encoding process can be implemented using the following operations: (i) the base picture (302) is compressed by the VVC encoder (312); (ii) the feature extractor isutilized for representing the driving frames (306) with a compact feature matrix of size 4x4; and (iii) the extracted compact features are inter-predicted and quantized, and the residuals are entropy-coded to generate a corresponding output bitstream (318).

[0047] In the example shown, the decoder (330) includes the following modules:1. A VVC decoder (332) for reconstructing the base picture (302) to generate a decoded base picture (334).2. A motion estimation and frame generation module (336), comprising: a) A component for Entropy decoding and compensation configured to reconstruct the compact features. b) A component for generation of the final video through reconstructed features and the decoded base picture (334).An output video sequence (340) is obtained through the following operations: (i) calculating the sparse motion field using the compact features of the base picture (334) and the reconstructed driving pictures; (ii) generating a pixel- wise dense motion map and an occlusion map based on the sparse motion field; and (iii) inputting the decoded base picture, pixel-wise dense motion map and the occlusion map to a deep generative model that generates the output video (340).

[0048] FIG. 4 is a block diagram illustrating a GFVC architecture (400) according to some examples. The GFVC architecture (400) includes an encoder (410) and a decoder (430). In operation, the GFVC architecture (400) enables a disentangled background in the reconstructed video sequence at the decoder end. The GFVC architecture (400) also contains segmentation and inpainting modules (404, 408) at the encoder end. Additional instances (e.g., copies) of the segmentation module (408) are also utilized at the decoder (430). In the example shown, the decoder 430 also includes a VVC decoder (432), a motion estimation and frame generation module (436), an artifact reduction module (440), and a fusion module (450). In some examples, the VVC decoder (432) and the motion estimation and frame generation module (436) can be similar to the VVC decoder (332) and the motion estimation and frame generation module (336), respectively, of the decoder (330).

[0049] In operation, an inpainted background base picture (414) and a compact feature representation (416) along with a base picture (402) are transmitted from the encoder (410) to the decoder (430) as indicated in FIG. 4. The decoder (430) then operates to fuse a decoded background base picture (444) and reconstructed segmented driving pictures (448) using the various above-mentioned modules of the decoder (430), as indicated in FIG. 4, to generate a sequence of background stabilized driving pictures (460). Also generated during the illustrateddecoding process are a decoded base picture (434), a segmented base picture (442), and recovered segmented driving pictures (446).

[0050] In at least some examples, the GFVC architecture (400) provides an improvement to the CFTE workflow (300) illustrated in FIG. 3 via modifications and additions of modules that enable disentanglement between the background and the foreground in the reconstructed video. Note that the encoder (410) in the GFVC architecture (400) illustrated in FIG. 4 is shown in the out-of-loop configuration. However, various embodiments are not so limited. In some embodiments, the encoder (410) can be configured to operate in the in-loop configuration in which a segmentation tool operates on the encoded and then decoded images to obtain substantially similar processing latency at the encoder (410) and the decoder (430). In some examples, the in-loop configuration of the encoder (410) is configured to use the artifact reduction module (440) prior to the segmentation tool (404).

[0051] In general, conventional video codecs can be used to encode the base pictures and background. In the GFVC architecture (400), the VVC codec (412, 432) is utilized for two different frames: 1) the base picture (402, 434); and 2) the background base picture (414, 444). The latter is obtained from the former through segmentation and inpainting (404, 408). These two frames are encoded and transmitted at the beginning of a corresponding bitstream (422) and / or at other time instances, e.g., when updates are needed. In the example shown, the VVC codec (412, 432) is used. However, other video codecs (e.g., AVC, HEVC, and machine- leaming-based codecs) can similarly be used in other examples of the GFVC architecture (400).

[0052] The segmentation module (404) is utilized for foreground and background separation in the base picture (402) and the driving pictures (406). In one example, the output from an instance of the segmentation module (404) includes: (1) a segmentation mask; and (2) the separated foreground and background frames. Therein:• For the base picture (402), the segmentation mask is given to the inpainting module (408) to obtain the full background.• For a driving picture (406), the extracted foreground is merged with a new background, which is generated from the mean values of the RGB channels of the extracted foreground.This approach helps to smooth the fusion process at the decoder (430).

[0053] FIGS. 5A-5D pictorially illustrate certain operations of the segmentation module (404) according to some examples. Comparison of the base picture (FIG. 5A or 5C) and thecorresponding driving picture (FIG. 5B or 5D) reveals that, in the driving picture, the background is replaced with channel-wise mean values after segmentation.

[0054] FIGS. 6A-6B pictorially illustrate segmentation masks (602, 604) produced by the segmentation module (404) according to some examples. In some examples, the segmentation module (404) may benefit from the use of a segmentation tool referred to as Semantic Guided Human Matting. This segmentation tool is described, e.g., in Xiangguang Chen, et al., “Robust Human Matting via Semantic Guidance,” Proceedings of the Asian Conference on Computer Vision (ACCV), 2020, (Ref. [2]), which is incorporated herein by reference in its entirety.

[0055] Herein below, Mnrepresents the segmentation mask of the nthframe. Similarly, Mn' represents the inverted segmentation mask (1 — Mn) of the nthframe. Both of these masks are binary masks. Accordingly, in Mn, the foreground is represented with ones, while the background is represented with zeros. In Mn' , the foreground and background are represented with zeros and ones, respectively. The segmentation mask (602) pictorially shown in FIG. 6A is an example of the mask Mn. The segmentation mask (604) pictorially shown in FIG. 6B is an example of the mask Mn' .

[0056] FIGS. 7A-7D pictorially illustrate certain operations of the inpainting module (408) according to some examples. In operation, the inpainting module (408) takes the segmented background picture (which is obtained by applying the segmentation mask to the original full picture) and fills in the missing pixel information in the foreground pixel locations. In some examples, the segmentation module may benefit from the use of the inpainting tool described in Chaohao Xie, et al., “Image Inpainting with Learnable Bidirectional Attention Maps” (ICCV 2019), Ref.[3]), which is incorporated herein by reference in its entirety.

[0057] FIG. 7A pictorially shows a base picture (702) that may be applied to the segmentation module (404) according to one example. FIG. 7B pictorially shows a segmentation mask Mn(704) produced by the segmentation module (404) based on the base picture (702).FIG. 7C pictorially shows a background frame (706) produced by the segmentation module (404) using the segmentation mask (704) and the base picture (702). FIG. 7D pictorially shows an inpainted background frame (708) produced by the inpainting module (408) based on the background frame (706).

[0058] In some examples, the compact feature representation module (416) is implemented using a deep learning-based architecture (such as the U-Net) for extracting compact face features from the segmented driving pictures (418). As already indicated above, each of the segmenteddriving pictures (418) includes the respective segmented foreground fused with a constant background. Hence, the compact features are extracted by the compact feature representation module (416) from the segmented foreground picture (not from the full image). As a result, the spatial dimensions of the extracted features are significantly smaller than the input size. For example, for a frame size of 256 X 256, the compact features may be of the size 4 x 4. The resulting extracted features are quantized and entropy-coded for the bitstream transmitted to the decoder (430). In effect, these transmitted features represent the facial characteristics and motion. At the decoder (430), these features are used to generate the dense and occlusion motion maps, which are then used to generate frames for the output video (460). It should be noted that various alternative face features, such as key points, face landmarks, etc., can be used in various examples.

[0059] The motion-estimation and frame-generation module (436) is used at the decoder (430) for generating the segmented driving pictures (446). In some examples, the module (436) includes the following stages:• In a first stage of the module (436), the dense motion map and the occlusion map are generated using the compact features of the driving and base pictures.• In a second stage of the module (436), these maps are used as inputs to a deep generative network configured to reconstruct the driving pictures.

[0060] At relatively high values of the Quantization Parameter (QP), such as approximately 52 or larger than a certain threshold QP value (e.g., 45), the decoded base picture (434) may exhibit a degraded reconstructed image, and the segmentation module (404) configured to process the decoded base picture (434) may perform poorly as a result. In some examples, at such relatively high QPs, a key picture may show significant blocky artifacts. To address this problem, the GFVC architecture (400) incorporates the optional compression artifact reduction module (440). In some examples, the compression artifact reduction module (440) is diffusion- model-based. The latter model is applied to the base picture at the higher QPs, which tends to significantly boost the performance of the downstream segmentation module (404) and the overall performance of the decoder (430).

[0061] FIG. 8 is a diagram graphically illustrating an encoding process (800) in the GFVC architecture (400) according to some examples. In the notation used in FIG. 8, “C” represents compact features. The encoding process (800) begins with the transmission of the base picture (402) and the background base picture (414) followed by the transmission of the compactfeatures I of the driving pictures (418). In one example, the encoding process (800) includes the following operations:1. The base picture (402) (as the key frame denoted as K) is segmented with the segmentation module (404) to obtain the corresponding segmented background.2. The inpainting module (408) is used to inpaint the segmented background. The resulting picture serves as the background base picture (414) (also denoted as B).3. The base picture (402), as the first frame (K), e.g., IRAP (instantaneous decoding refresh picture) in a coded sequence, and the background base picture (414), as the second frame (B) following the first frame (K) are encoded through the VVC encoder (412) and transmitted in the output video bitstream (422).For the base and driving pictures (402, 406), the compact features I are extracted from the segmented foreground pictures (418). Then, the compact features are inter-predicted and quantized, and the resulting residuals (802) are entropy encoded (denoted as C in FIG. 8) in a metadata bitstream (424) (e.g., implemented as an SEI message).

[0062] Note that, in the CFTE workflow (300, FIG. 3), the compact feature for the base picture is generated at the decoder (330) using the decoded base picture (334). In contrast, for purposes of the GFV SEI message (424), the compact feature I for the base picture (402) is extracted at the encoder (410), which uses the same process as that used for the driving pictures (418), and then is sent to the decoder (430). In general, better quality can be achieved at the encoder (410) than at the decoder (430) when the quality of the decoded image is low.

[0063] In some examples, the decoding process in GFVC architecture (400) includes the following operations:1. The base picture (402) received with the frame (K) of the bitstream (422) and the background picture (414) received with the frame (B) of the bitstream (422) are decoded using the VVC decoder (432).2. The decoded base picture (434) is repaired by the compression artifact reduction tool (440), when needed.3. The repaired, decoded base picture (434) is segmented using the segmentation tool (404) to obtain the segmentation mask and the corresponding foreground and background. For the foreground image (442), the missing background is filled in by the mean values of the RGB channels of the foreground pixels, which results in a constant background.4. The compact features (802) of the base and driving pictures (402, 418) are decoded. Both (i) the compact features (802) of the base picture and of the driving pictures and(ii) the segmented foreground image (442) are provided to the Motion Estimation and Frame generation module (436) to reconstruct the segmented driving pictures (446), which have the foreground with the constant background.5. Both (i) the decoded background picture (444) and (ii) the segmented driving pictures (448) are given to the fusion module (450) that outputs the sequence of background stabilized driving pictures (460).Note that, in the above processing sequence, the base picture (402) is sent by the encoder (410) for backward compatibility, e.g., in case of the decoder (430) that does not support GFV SEI messaging, so that it can still output the base picture. When backward compatibility is not needed or the background information is not needed, the encoder (410) can send the segmented foreground in lieu of the base picture. This workaround may obviate the need for the corresponding segmentation module (404) at the decoder (430) and possibly also for the artifact reduction module (440), in at least some cases. In some examples, conventional image processing can be used in the fusion module (450). One example is alpha blending shown in the following equation: merged_picture = alpha_mask * foreground_picture + ( 1 - alpha_mask) * background_picture (1) where alpha_mask is derived based on the segmentation mask (see, e.g., 704, FIG. 7B). In some approaches, additional information, such as motion, etc., can be incorporated into alpha_mask to smooth the temporal effects. Alternatively, a neural network model can be used in the fusion module (450) in some implementations thereof.

[0064] Chroma key compositing, or chroma keying, is a visual-effects and post-production technique for compositing (layering) two or more images or video streams together based on color hues (chroma range). This technique is used in many fields to remove a background from the subject of a photo or video. Some use case examples include applications to newscasting, motion picture, and video game industries. A color range in the foreground footage is made transparent, allowing separately filmed background footage or a static image to be inserted into the scene. The chroma keying technique is commonly used in video production and postproduction. This technique is also sometimes referred to as color keying, color- separation overlay (CSO), or by various terms for specific color-related variants, such as the green screen or blue screen. In some examples, the chroma keying can be done with backgrounds of any colorthat is uniform and distinct. However, the green and blue backgrounds are more commonly used because they differ most distinctly in hue from any human skin color.

[0065] In some examples, the “keying” process is used to replace portions of the video that match the pre-selected color by the substitute background video. An important factor for keying is the color separation of the foreground and background. The chroma key can be achieved by a numerical comparison between the video and the pre-selected color. If the color at a particular point on the screen matches (either exactly or within a preselected range), then the contents of that point are replaced by the substitute background contents.

[0066] The following pseudocode provides an example keying algorithm.EXAMPLE PSEUDOCODE FOR KEYING ALGORITHM chromakey algorithm input : fg(i, j) = i, j element of foreground bg(i, j) = i, j element of background tola, tolb returns : out ( i , j ) get Cb and Cr of key color; call them Cb_key and Cr_key for each i , j : get Cb and Cr for pixel value; call them Cb_p and Cr_p let maskl = colorclose (Cb_p, Cr_p, Cb_key, Cr_key, tola, tolb) let mask = 1.0 - maskl out (i, j) = fg(i, j) - mask*key_color + bg(i, j) *mask / / key composition def colorclose (Cb_p, Cr_p, Cb_key, Cr_key, tola, tolb) temp = math . sqrt ( (Cb_key-Cb_p) * *2+ (Cr_key-Cr_p) * *2 ) if temp < tola: return 0.0 else if temp < tolb: return (temp-tola) / (tolb-tola)else : return 1 . 0

[0067] The above algorithm includes converting images into YcbCr space and looking at the CbCr plane therein. The key color defines a point on the plane. The plane is divided into three regions (pseudo code function: colorclose(Cb_p, Cr_p, Cb_key, Cr_key, tola, tolb)): (i) a region close to the key color, (ii) a region far from the key color, and (iii) a middle region. The variable temp provides the distance to the key color. The “close” region has all colors that are less than a first threshold, tola, from the key color. The “far” region has all colors that are farther than a second threshold, tolb, from the key color. The “middle” region has the colors with the distances in-between the thresholds tola and tolb. Then maskl is generated as follows: the “close” region is all background (maskl = 0); the “far” region is all foreground (maskl = 1). The “middle” region is assigned a value based on its distance from the key color. In this example, maskl = temp-tola)! (tolb -tola). For the key composition, mask = 1 - maskl. The above algorithm is aligned with alpha blending (alpha = mask), which can be used in the fusion module (450) or equivalents thereof. The latter feature helps to smooth out the object boundary.

[0068] FIGS. 9-10 are block diagrams illustrating modifications of the GFVC architecture (400) according to some examples. The modified GFVC architecture shown in FIG. 9 is denoted using the reference numeral (900). The modified GFVC architecture shown in FIG. 10 is denoted using the reference numeral (1000). Both of the modified GFVC architectures (900, 1000) use chroma key backgrounds as described in more detail below and reuse some of the components of the GFVC architecture (400). For the description of such reused components, the reader is referred to the above description referring FIGS. 4-8. Note that the use of chroma key backgrounds in the modified GFVC architectures (900, 1000) beneficially enable elimination of instances of the segmentation module 404 at the respective decoders. It also helps in cases when the segmentation at the encoder and at the decoder are different, and visual artifacts may appear at the decoder.

[0069] Referring to FIG. 9, the modified GFVC architecture (900) includes an encoder (910) and a decoder (930). One difference between the encoder (910) and the above-described encoder (410) is that the segmentation module (404) used in the encoder (910) is configured to generate a segmented base picture (902) having the foreground of the base picture (402) and further having a chroma key background. Another difference is that the VVC encoder of the encoder (910) is configured to receive segmented base picture (902) instead of the original base picture (402)(also see FIG. 4). Further, the compact feature representation module (416) operates on segmented driving pictures (918) that differ from the segmented driving pictures (418) (FIG. 4) in that the segmented driving pictures (918) have the chroma key background instead of the above-described constant background.

[0070] At the decoder (930), the VVC decoder (432) generates a decoded base picture (934) by processing the first bitstream (422) carrying the encoded segmented base picture (902). The motion-estimation and frame-generation module (436) uses the decoded base picture (934) to generate segmented driving pictures (946) based on the second bitstream (424). A chroma-key fusion module (950) then outputs a sequence of background stabilized driving pictures (960) based on the segmented driving pictures (946) and the decoded background base picture (444) (also see FIG. 4).

[0071] Referring to FIG. 10, the modified GFVC architecture (1000) includes an encoder (1010) and a decoder (1030). The encoder (1010) differs from the encoder (910) in that the first bitstream (422) generated by the encoder (1010) does not carry a background base picture because such picture is not provided by the segmentation module (404) to the VVC encoder (412). The encoder (910) also lacks the inpainting module (408). The decoder (1030) differs from the decoder (930) in that the chroma-key fusion module (950) receives a virtual background picture (1044) instead of the decoded background base picture (444). The chromakey fusion module (950) then outputs a sequence of background stabilized driving pictures (1060) based on the segmented driving pictures (946) and the virtual background picture (1044).

[0072] FIG. 11 is a diagram graphically illustrating an encoding process (1100) in the modified GFVC architecture (900) according to some examples. The notation used in FIG. 11 is the same as in FIG. 8. The encoding process (1100) includes transmitting of the encoded segmented base picture (902) and the background base picture (414) via the first bitstream (422). The encoding process (1100) also includes transmitting the compact features I of the base picture and (902) and the segmented driving pictures (918) via the second bitstream (424) carrying SEI messages (1102).

[0073] FIG. 12 is a diagram graphically illustrating an encoding process (1200) in the modified GFVC architecture (1000) according to some examples. The notation used in FIG. 12 is the same as in FIG. 8. The encoding process (1200) includes transmitting of the encoded segmented base picture (902) via the first bitstream (422). The encoding process (1200) alsoincludes transmitting the compact features I of the base picture (902) and the segmented driving pictures (918) via the second bitstream (424) carrying SEI messages (1202).

[0074] To support the handling of the chroma-key background at the decoder (930, 1030), a corresponding syntax is added to the conventional generative face video SEI message syntax. The added syntax includes a first set of syntax configured to specify a chroma key value. The added syntax also includes a second set of syntax configured to specify the tolerance threshold of the chroma key value, and / or to specify a range of values. Table 1 includes the conventional syntax and an example the added syntax. Table 1: Generative Face Video SEI Messagegfv_chroma_key_background_present_flag equal to 1 indicates the background of the base picture is chroma key background. Gfv_chroma_key_background_present_flag equal to 0 indicates the background of the base picture is not chroma key background.Note: when gfv_chroma_key_background_present_flag is equal to 1, it indicates to the decoder that the base picture may need fusion with the true background picture (if available at the decoder) or with the virtual background picture. For the drive picture, when gfv_drive_pic_fusion_flag is equal to 1, the drive picture is fused with the decoded background picture; when gfv_drive_pic_fusion_flag is equal to 0, the drive picture may be fused with the virtual background picture. gfv_chroma_key_value[ c ] specifies the chroma key value for the c-th color component, where c equal to 0 refers to the Cb component, and c equal to 1 refers to the Cr component. The length of gfv_chroma_key_value[ c ] is BitDepthc for c set to 0 and 1, respectively. gfv_chroma_key_thr_low specifies the low bound of the chroma key distortion threshold. gfv_chroma_key_thr_high specifies the high bound of the chroma key distortion threshold.

[0075] In some examples, a chroma key value can be specified as the Y / Cb / Cr component value in the same bit depth as the decoded picture (derived using vui_colour_description related information in the Voice User Interface, VUI). In another example, one can use the R / G / B component value or specified color primaries for the R / G / B components. For the tolerance threshold, following the above-described example keying algorithm, one can specify two distance thresholds, tola and tolb, by computing the L2 distance of the Cb and Cr components.One can also specify the color distance using all 3 components, Y / Cb / Cr or R / G / B. In another example, the tolerance value is specified as the percentage of the color difference from each component or some of the components, e.g., in the Y / Cb / Cr or R / G / B space. In yet another example, one can use hue or other suitable color indication. In some example, one can specify the tolerance as the color distance using other suitable color difference metrics, such as the metrics specified in DE2000 or Rec. ITU-R BT.2124 or the AEITP metric.

[0076] In one example, given a sample in the decoded picture, the sample value is specified by decoded_pixel_value[ c ]. The chroma key distance is computed as follows: distance = round ( sqrt( ( decoded_pixel_value

[0000] - gfv_chroma_key_value

[0000] ) 2 + ( decoded_pixel_value

[0001] - gfv_chroma_key_value

[0001] )A2 ) ) (2)When distance is smaller than gfv_chroma_key_thr_low, the sample is considered to have a chroma key value. When distance is larger than or equal to gfv_chroma_key_thr_high, the sample is not considered to have a chroma key value. For distance between gfv_chroma_key_thr_low and gfv_chroma_key_thr_high, the sample is considered to have a value between the chroma key and non-chroma keys.

[0077] In some examples, alternative Chroma Key Syntax representations can be used in H.263 as described in more detail below.

[0078] In H.263 (https: / / www.itu.int / rec / T-REC-H.263-200501-Een), chroma keying information is presented in Annex L: Supplemental enhancement information specification, L.14: Chroma Keying Information. It provides more flexible indication on chroma key value and distance. The below is CKIF from H.263. The following provisions of Annex L can be adopted to support some embodiments disclosed herein.

[0079] The Chroma Keying Information Function (CKIF) indicates that the “chroma keying” technique is used to represent “transparent” and “semi-transparent” pixels in the decoded video pictures. When being presented on the display, “transparent” pixels are not displayed. Instead, a background picture which is either a prior reference picture or is an externally controlled picture is revealed. Semi-transparent pixels are displayed by blending the pixel value in the current picture with the corresponding value in the background picture. One octet is used to indicate the keying color value for each component (F, CB, or CR) which is used for chroma keying. To represent pixels that are to be “semi-transparent”, two threshold values, denoted as Ti and T , areused. Let ex denote the transparency of a pixel; ex = 255 indicates that the pixel is opaque, and ex = 0 indicates that the pixel is transparent. For other values of ex, the resulting value for a pixel should be a weighted combination of the pixel value in the current picture and the pixel value from the background picture (which is specified externally). The values of ex may be used to form an image that is called an “alpha map.” Thus, the resulting value for each component may be:[a - X + (255- a) -Z] / 255 (3) where X is the decoded pixel component value (for Y, CB, or CR), and Z is the corresponding pixel component value from the background picture. The ex value can be calculated as follows. First, the distance of the pixel color from the key color value is calculated:where XY, XB, and XRare the Y, CB, and CRvalues of the decoded pixel color, KY, KB, and KRare the corresponding key color parameters, and AY, AB, and AR. Are keying flag bits which indicate which color components are used as keys. Once the distance d is calculated, the ex value may be computed as specified in the following pseudocode: for each pixel i f ( d<Ti ) then CX = 0 ; el se i f ( d>12 ) then CX = 255 ; el se ex = [ 255- ( d-Ti ) ] / ( T2-T1 )However, the specific method for performing the chroma keying operation in the decoder is not specified herein because normative specification of the method is not needed for interoperability. The corresponding process described herein is provided for illustration purposes in order to generally and illustratively convey an intended interpretation of some data parameters.

[0080] In some examples, the following syntax based on the H.263 Chroma Key Information SEI can be used:Table 2: Generative Face Video SEI Messagegfv_chroma_key_background_present_flag equal to 1 indicates the background of the base picture is chroma key background. gfv_chroma_key_background_present_flag equal to 0 indicates the background of the base picture is not chroma key background. gfv_chroma_key_flag[ c ] equal to 1 indicates the chroma key value for the c-th colour component gfv_chroma_key_value[ c ] is present, where c equal to 0 refers to the Y component, c equal to 1 refers to the Cb component, and c equal to 2 refers to the Cr component. gfv_chroma_key_flag[ c ] equal to 0 indicates that the chroma key value gfv_chroma_key_value[ c ] for the c-th colour component is not present. gfv_chroma_key_value[ c ] specifies the chroma key value for the c-th colour component, where c equal to 0 refers to the Y component, c equal to 1 refers to the Cb component and c equal to 2 refers to the Cr component. The length of gfv_chroma_key_value[ c ] is BitDepthy for c equal to 0, and BitDepthc for c equal to 1 and 2, respectively.When all three gfv_chroma_key_flag[ c ] are equal to 0, the default chroma key value is set as follows: gfv_chroma_key_value

[0000] = ( 50 « ( BitDepthy - 8) ) ) gfv_chroma_key_value

[0001] = ( 220 « ( BitDepthc - 8) ) gfv_chroma_key_value

[0002] = ( 100 « ( BitDepthc - 8) ) gfv_chroma_key_flag[ c ] is reset to 1 for c=0...2. gfv_chroma_key_thr_flag[ i ] equal to 1 indicate gfv_chroma_key_thr_value[ i ] is present, where i = 0,1. gfv_chroma_key_thr_flag[ i ] equal to 0 indicates that gfv_chroma_key_thr_value[ i ] is not present. gfv_chroma_key_thr_value[ i ] specifies the i-th threshold of the distance of the pixel color from the chroma key color value. gfv_chroma_key_thr_value

[0000] specifies the threshold for transparency. gfv_chroma_key_thr_value

[0001] specifies the threshold for opacity.If gfv_chroma_key_thr_flag

[0000] is equal to 0, the default gfv_chroma_key_thr_value

[0000] is set to equal to ( 48 « ( BitDepthy - 8)2) if gfv_chroma_key_thr_flag

[0001] is equal to 0, the default gfv_chroma_key_thr_value

[0001] is set to equal to ( 75 « ( BitDepthy - 8)2)The distance of the pixel color specified as decoded_pixel_value[ c ] from the key color value gfv_chroma_key_value[ c ] (c=0, 1, 2) is computed as:gfv_chroma_key_value[c )2(5) if d < gfv_chroma_key_thr_value

[0000] , the pixel is considered as a transparency pixel; if d > gfv_chroma_key_thr_value

[0001] , the pixel is considered as an opacity pixel.

[0081] Some examples described above deal with various scenarios in which the background is substantially static. Below, we describe modifications to the GFVC architecture (400) of FIG. 4 to provide accommodations for a moving, dynamic background. Based on the provided description, a person of ordinary skill in the pertinent art will be able to make functionally similar modifications to the GFVC architectures (900, 1000) of FIGS. 9-10 without any undue experimentation .

[0082] In some examples, a generative face video (GFV) SEI message is included in JVET- AH2032 technologies (Ref. [4]). JVET-AH0118 (Ref. [5]) describes the use of picture fusion to mitigate an entanglement between the face and background. To segment the face (in the foreground) from the background, a background picture is generated by inpainting and is sent as an additional picture to the receiver (having the decoder). The corresponding features are generated based on the foreground picture and are conveyed via a properly configured GFV SEI message in a bitstream. At the receiver, fusion is used to merge the face generated from the features conveyed in the GFV SEI message with the background picture. As described in JVET- AH0118, segmentation can be performed both at the encoder and at the receiver.

[0083] In some cases, when the respective segmentations are different at the encoder and at the decoder, such differences might cause unwanted artifacts. However, when a chroma key background is used, segmentation at the decoder can be avoided, which enables simplification of the processing implemented at the receiver and avoids possible artifacts caused by the above- mentioned segmentation mismatch between at the encoder and the receiver. Accordingly, insome examples, the following syntax and semantics for the GFV SEI message can be used to support the use of chroma key for picture fusion based on the above considerations. The syntax provided below is generally consistent with the chroma key information function specified in ITU-T H.263.Table 3: Generative Face Video SEI Message} gfv_chroma_key_info_present_flag equal to 1 indicates that the syntax elements gfv_chroma_key_value_present_flag[ c ] and gfv_chroma_key_thr_present_flag[ i ] are present and the syntax elements and gfv_chroma_key_value[ c ] and gfv_chroma_key_thr_value[ i ] might be present. gfv_chroma_key_info_present_flag equal to 0 specifies that the syntax elements gfv_chroma_key_value_present_flag[ c ], gfv_chroma_key_thr_present_flag[ i ], gfv_chroma_key_value[ c ], and gfv_chroma_key_thr_value[ i ] are not present. gfv_chroma_key_value_present_flag[ c ] equal to 1 indicates that the syntax element gfv_chroma_key_value[ c ] is present. gfv_chroma_key_present_flag[ c ] equal to 0 indicates that the syntax element gfv_chroma_key_value[ c ] is not present.The variable ChromaKeyDefaultValueFlag is set equal to !( gfv_chroma_key_value_present_flag

[0000] I I gfv_chroma_key_value_present_flag

[0001] I I gfv_chroma_key_value_present_flag

[0002] ). gfv_chroma_key_value[ c ] specifies the chroma key value corresponding to the c-th colour component as follows:- If ChromaKeyDefaultValueFlag is equal to 1, the variables GfvChromaKey Value [ c ] are specified as follows:- GfvChromaKey Value

[0000] is set equal to 50- GfvChromaKey Value

[0001] is set equal to 220- GfvChromaKey Value

[0002] is set equal to 100- Otherwise, ChromaKeyDefaultValueFlag is equal to 0, the following applies:- if gfv_chroma_key_value_present_flag[ c ] is equal to 1, GfvChromaKey Value [ c ] is set equal to the value of gfv_chroma_key_value[ c ]- otherwise, gfv_chroma_key_value_present_flag[ c ] is equal to 0, GfvChromaKey Value[ c ] is not specified by this Specification. gfv_chroma_key_thr_present_flag[ i ] equal to 1 indicates that the syntax element gfv_chroma_thr_value[ i ] is present. gfv_chroma_key_thr_present_flag[ i ] equal to 0 indicates gfv_chroma_key_thr_value[ i ] is not present.gfv_chroma_key_thr_value[ i ], when present, specifies the i-th chroma key threshold value.When not present, the value of gfv_chroma_key_thr_value[ i ] is inferred as follows:If i is equal to 0, gfv_chroma_key_thr_value

[0000] is set equal to 48Otherwise, i is equal to 1, gfv_chroma_key_thr_value

[0001] is set equal to 75NOTE - The syntax elements gfv_chroma_key_value_present[ c ], gfv_chroma_key_value[ c ], and gfv_chroma_key_thr_value[ i ] could be used to determine a transparency indicator for fusion of the generated face picture and background picture. For example, a transparency indicator, denoted as alpha[ x ] [ y ] , for picture sample value, denoted as I[ c ][ x ][ y ] with bitDepth[ c ], where bitDepth

[0000] is equal to BitDepthy, bitDepth

[0001] and bitDepth

[0002] are equal to BitDepthc, and GfvChromaKey[ c ] values for sample coordinates x, y, and colour components c could be determined as follows: d[ x ][ y ]= 0 for( c = 0; c < 3; C++ ) if( gfv_chroma_key_value_present[ c ] I I ChromaKeyDefaultValueFlag ) d[ x ][ y ] += ( I[ c ][ x ][ y ] / ( 1 « (bitDepth[ c ] - 8) ) -GfvChromaKey Value[ c ] )2(XX) if ( ( d[ x ][ y ] < gfv_chroma_key_thr_value

[0000] ) alpha[ x ] [ y ] = 0 else if ( ( d[ x ][ y ] > gfv_chroma_key_thr_value

[0001] ) alpha[ x ] [ y ] = 1 else alpha[ x ][ y ] = ( d[ x ][ y ] - gfv_chroma_key_thr_value

[0000] ) +( gfv_chroma_key_thr_value

[0001] - gfv_chroma_key_thr_value

[0000] )A value of alpha[ x ][ y ] equal to 0 could indicate transparency. A value of alpha[ x ][ y ] equal to 1 could indicate opacity. Intermediate values of alpha[ x ][ y ] could indicate semitransparency.

[0084] FIG. 13 is a block diagram illustrating possible sources of dynamic backgrounds according to some examples. In the examples shown, a dynamic background (1302) having time-dependent changes therein may result from a camera / background motion (1304) and / or from a foreground motion (1306). Accordingly, two different mechanisms can be used to handle these scenarios. In various examples, we utilize:1. Optical flow to handle the dynamic background (1302) resulting from the camera / background motion. (1304)2. Change in the segmentation masks to handle the dynamic background (1302) resulting from the foreground motion (1306).

[0085] FIGS. 14A-14B pictorially illustrate the dynamic background (1302) caused by the camera motion (1304) according to one example. Comparison of the two shown frames (0, 119) readily reveals that there is an addition of new pixels in the later frame (119).

[0086] For such cases, it may be more efficient to incrementally update the background instead of performing a complete full update. In the incremental update, one can update the existing background by including a set of new added pixels (due to the motion). In some examples, a continuing incremental update includes the following operations (also see FIG. 15):1. Initialize a larger frame with a null background. For example, for a 256 X 256 input frame, initialize a 320 X 320 blank frame.2. Arrange the background of the previous (e.g., the first) frame in the center of the above initial null frame. For example, a 256 X 256 frame will start from pixel (32,32) and end at pixel (287,287) in the initialized larger background.3. For the current n'1' frame, estimate the motion between the previous (n- 1 )-th frame and the current n'1' frame (e.g., as described in more detail below). This estimate will provide the direction and magnitude of the motion represented by displacementsAx, Ay. Therein, the signs of the Ax, Ayprovide information about the direction of the motion. For example, when Ax, = 1, Ay= 1, the corresponding motion is the displacement by one pixel in the top-right direction. When Ax= — 1, Ay= — 1, the corresponding motion is the displacement by one pixel in the bottom-left direction, and so on.4. Partially update the background based on Ax, Ayand further based on the information from the second frame. For example, if Fx.x+255,y;y+255 isthe filled background in the initialized null background, thenAy'sthe updated background. This process is pictorially represented in FIG. 15.5. The above operations (3)-(4) are repeated by moving to the next frame (e.g., by the recursive update n=n+ \ ) until the index n reaches the last frame.

[0087] FIG. 15 is a diagram illustrating a process of updating the background based on the camera motion according to one example. As already indicated above, FIG. 15 provides a visualillustration to the above-described continuous incremental update. More specifically, a first background frame (1502) shown in FIG. 15 has a background (1504) of the first video frame placed in the center of a larger null frame. The first background frame (1502) and the second video frame are subjected to motion-estimation processing (1510) to obtain the direction and magnitude of the motion represented by the above-described displacements Ax, Ay. Displacement vectors (1514) overlaid on the first background frame (1502) schematically illustrate the estimated motion. With the displacement vectors (1514), a partially update (1520) of the background (1504) is performed based on Ax, Ayand further based on the information from the second video frame, thereby producing an updated background (1522) within the background frame (1502).

[0088] FIG. 16 is a block diagram illustrating a process (1600) for deriving the background displacement due to the camera / background motion according to some examples. In the example shown, an optical flow processing (1610) is used for the derivation of the displacement (A: (Ax, Ay)) based on the background / camera motion.

[0089] Let Fnand Mn' represent the nthframe and the corresponding inverted segmentation mask, respectively. Mnrepresents the segmentation mask of the nthframe. In Mn, the foreground has non-zero pixels while the background has zero pixels. In contrast, Mn' has zero and non-zero values for the foreground and background, respectively. In a first set of the operations (1610), we calculate the optical flow between Fnand Fn+ 1, which is represented as O. The obtained optical flow (O) contains the values for the background and foreground pixels’ movements. As we are interested only in the background motion, it is advantageous to mask the foreground pixels. To achieve this masking, we utilize the segmentation masks Mn' and M^+1. These segmentation masks track the foreground pixels’ movements. Hence, Mn' U M^+1provides the pixels belonging to the background in Fnand Fn+ 1. Accordingly, the Hadamard (pixel wise) product operation (1620) applied to O and (M,) U M^+1) manifests the background motion, and a resulting frame (1622) is used to derive the displacement (A).

[0090] In some examples, to calculate A, we accumulate the shift (which may not have an integer value) over several consecutive frames and compare it with a threshold (<z2)- The optical flow for individual frames can be small (for example, 0.01, 0.05). For example, the motion- induced shift can be accumulated until a relatively large total magnitude of the shift is reached so that the contribution of motion-induced changes can be significant. Also, aggregation of the predicted motion is beneficial in terms of better approximating the original motion. When thetotal displacement crosses a2, we map it to the nearest integer to obtain A. This procedure can be implemented using the following algorithm.Displacement Calculation AlgorithmThe symbol “o” denotes the Hadamard (pixel wise) product between 0 and (Mn'U M’n+i).

[0091] FIGS. 17A-17B pictorially illustrate a dynamic background caused by the foreground motion according to one example. Comparison of the two shown frames (0, 39) readily shows that new pixels are revealed in the background of the later frame (39), while some previously visible background pixels are now blocked. With a large amount of foreground motion, larger areas of occluded pixels will be revealed along the timeline. With the extra information of the revealed background pixels, the background can be updated using the collected revealed “ground truth” pixels, instead of the “inpainted” pixels. In some examples, for purposes of tracking the foreground, the use of segmentation masks can be more efficient than the use of the optical flow (1610). An example algorithm that can be used for such tracking is described below.

[0092] This algorithm starts with the segmentation mask of the first frame as a reference, Mref. For the subsequent frames, the algorithm calculates 6 = diff Mn, Mref), where diff -) represents the change between Mnand Mref . The function diff(-) is defined to calculate the change in the foreground pixels. The pixel-wise difference between Mnand Mref can produce the same values for different scenarios. If the foreground pixels completely overlapin Mnand Mref then \Mn— Mref | = 0. Similarly, no overlap between the two frames (foreground pixels have shifted significantly in Mn) will also result in \Mn— Mref \ = 0.

[0093] In some examples, to obtain the unique values for such scenarios, we use the intersection between Mnand Mref . For a complete overlap of the frames, MnA Mref is equal to Mnor Mref. For no overlap between the frames, MnA Mref is null. Hence, MnA Mref decreases with increasing separation between Mnand Mref .

[0094] Next, we defineas the number of selected pixels inSimilarly,IS the number of intersected pixels in MnA Mref. We express the difference between Mnand Mref as the ratio between An-ref and Aref as follows:where || ■ ||Frepresents the Frobenius Norm. Due to the presence of An-ref, the function diff(-) has a lower value for a higher change in the foreground between the frames, and vice versa.When 6 = dif f (Mn, Mref ) is below the selected fixed thresholdthe existing background is updated with the newly revealed background which can be obtained using the methods described below. A higher value of a should be selected for capturing minor displacements, and a lower value of a should be selected for significant displacements.

[0095] FIG. 18 is a diagram illustrating implementations (1800) of the inpainting approach according to one example. In the example shown, two different types of background update algorithms (1801, 1802) are used. The first type (1801) is referred to as the Online Approach. The second type (1802) is referred to as the Look-Back Approach. Both types are described in more detail below.

[0096] Under the Online Approach (1801), we obtain the background of Fnusing the inpainting model (e.g., implemented using the above-described inpainting module (408)), which serves as the new background. As discussed previously, for Fn, 8 crosses the predefined threshold a . Thus, under the online approach (1801), we are substantially ignoring the background information available in the previous frames (Foto Fn-x) and are considering only the current frame for updating the background. It should also be noted that, when the frequency of the updates is high, it may lead to observable flickering in the background of the video sequence, which is undesirable. In one example, the Online Approach (1801) can be implemented using the following algorithm.Online Approach for Background Update Algorithm

[0097] The Look-Back Approach (1802) is designed to address at least some of the aboveindicated limitations of the Online Approach (1801). Under the Look-Back Approach (1802), to avoid information loss, we use an information aggregation module before using the inpainting module (408). The information aggregation module operates to collect available information from the frames Foto Fnbased on masks Moto Mn. Thus, the information aggregation module substantially collects available background information from all previous frames. Afterwards, the remaining background pixels are filled-in using the inpainting module (408). This process is discussed in more detail below. In some examples, we maintain an aggregated buffer to avoid information aggregation from Foto the current frame every time the look-back approach (1802) is invoked. To apply this information at Fn+k, we retrieve the last aggregated frame from the buffer and fill the information based on the frames Fn+1to Fn+k. In one example, the Look- Back Approach (1802) can be implemented using the following algorithm.Look-back Approach for Background Update Algorithm

[0098] FIGS. 19A and 19B pictorially compare inpainting results obtained based on a single frame and based on multiple frames according to one example. Comparison of the shown pictures clearly indicates that more details can be provided when background information is collected from multiple frames. In one example, the corresponding process includes the following operations:1. A blank frame is initialized, and information from the multiple frames is filled to produce the partial aggregated frame.2. The partial aggregated frame is inputted to the inpainting module that outputs the complete background.In one example, this process can be implemented using the following algorithm.Inpainting Based on Multiple Frames Algorithm

[0099] FIG. 20 is a block diagram illustrating a workflow (2000) of inpainting operations based on multiple frames according to one example. The workflow includes an information aggregation module (2010). Different information aggregation methods that can be implemented with the information aggregation module (2010) according to several nonlimiting examples are described in more detail below.

[0100] As an example, according to the information filling process in the initialized blank frame ( ), we consider two different information aggregation methods. For the frame and segmentation mask set (Fn, Mn) followed by the next set (Fn+1, Mn+1), a subset of pixels in the foreground in Fnmay now be in the background in Fn+ 1. Similarly, some pixels may switch from the background to the foreground between the two frames.

[0101] FIG. 21 graphically illustrates two different information aggregation methods (2101, 2102) that can be implemented in the information aggregation module (2010) according to some examples. In FIG. 21, the white areas represent the background, whereas the dark areas represent the foreground. In the first information aggregation method (2101), all of the foreground pixels from the current frame (Fn+ 1) are added to <p. while the pixels that moved from the foreground in Fnto the background in Fn+1are retained from the previous frame (Fn). In the second information aggregation method (2102), all of the foreground pixels from the current frame (fJJ are added to <p. Next, the pixels that switched from the background in Fnto the foreground in Fn+ 1are added from the current frame (Fn+ 1). In this sequence, the information is added in an incremental manner, and only the newly available information cumulatively added.

[0102] In the above description, we discussed the change in the background due to two different factors: the background / camera motion and the foreground motion. Now, we combine both factors and present a unified approach for the background update. In the first step of this approach, we calculate the change in segmentation masks (Moand Mref) and update the background through the online or look-back approach (1801, 1802) as described above. This particular update considers changes in the foreground. In the next step of the unified approach, we utilize the optical flow (1610) to calculate the displacement due to the background or camera motion. Based on this displacement, we incrementally add to the existing background. In one example, the unified approach can be implemented using the following algorithm.Unified Background Update Algorithm

[0103] FIG. 22 a block diagram illustrating a GFVC architecture (2200) according to additional examples. The GFVC architecture (2200) includes an encoder (2210) and a decoder(2230) and represents yet another modification of the GFVC architecture (400) of FIG. 4, wherein a background update module (2202) is added at the encoder (2210) and a transform / cropping module (2232) is added at the decoder (2230). In various examples, the added modules enable the GFVC architecture (2200) to provide accommodations for a moving, dynamic background, e.g., using at least some of the above-described methods and approaches.

[0104] The background update module (2202) is introduced in the encoder (2210) to provide an additional input to the VVC encoder (412). The background update module (2202) tracks changes in the background due to the foreground and / or background motion and makes the corresponding updates to the background, e.g., as described above. The background update module (2202) also sends pertinent metadata via the first bitstream (422) to assist background updates at the decoder (2230).

[0105] The transform / cropping module (2232) is inserted in the decoder (2230) upstream from the fusion module (450). For a dynamic background, the encoder (2210) is configured to transmit a video with a higher spatial resolution to include more background pixels, e.g., as discussed above in reference to FIG. 11. The transform / cropping module (2232) is configured to select the regions and perform a rescale operation to match the base picture resolution. The transform / cropping module (2232) is further configured to pass this processed / rescaled background onto the fusion module (450) for being fused with the foreground.

[0106] FIG. 23 is a diagram graphically illustrating an encoding process (2300) in the GFVC architecture (2200) of FIG. 22 according to some examples. In one example, the encoding process (2300) includes the following operations:1. The first frame (402) is transmitted as the base picture (A) .2. A larger blank background (1502) is created. For example, for a 256 X 256 frame, initialize a 320 X 320 blank frame. Arrange the background of the first frame in the center (1504) of the initialized null frame (1502). A 256 X 256 frame will start from pixel (32,32) and end at pixel (287,287) in the 320 X 320 background frame. This frame is transmitted via the first bitstream (422) as the first background base picture. Along with it, the metadata (AT) containing the starting coordinates and frame dimensions are also transmitted. At the decoder (2230), this information is utilized to crop the larger background frame to the original size. In one example, the metadata is of the form (xstart, y start’ height, width), which is (32,32, 256, 256) for the aboveindicated size example.3. For the base and driving pictures, the compact features (C) (2302) are extracted and transmitted via the second bitstream (424).4. The change in the background (due to the foreground and / or background motion) is continuously tracked through the background update module (2202), e.g., using one or more of the above-described algorithms and methods. a. When there is a significant displacement in the foreground based on the threshold a i. The background B is updated to B', and an update (2304) is transmitted in the first bitstream (422). b. When there is a background change due to the camera motion based on the threshold a2: i. The background B' is incrementally updated to B' + A, and a corresponding update (2306) is added to the first bitstream (422). Derivation of A is already described above. For a nonzero A, the (. start’ y start) will change tDue to the motion, there will be new background information that is not available in the current background. The new information is stitched in the existing background B'. For example, if Ax= 2 and Ay= 1, then JVC will change to (34,33, 256, 256), and the background will extend from (32,32) to (290,289). The decoder (2230) is configured to perform cropping from this larger background based on the metadata JVC. In this example, the cropping area will be from (34,33) to (290,289). c. When there is background change due to a combination of the foreground and background motion (as defined by the thresholds a and <z2) inthe same frame, the above two steps are applied sequentially: first, B' is updated to B'' , which is then incrementally updated to B" + A. A corresponding update (2308) is added to the first bitstream (422), and the metadata JVC are updated accordingly.

[0107] In some examples, to support the dynamic background, the resolutions of the base picture and the background picture can be set to different values. The final generative video (460) will have the same resolution as the base picture. For the compression of the base picture and background picture, different solutions may be applied based on the 2D codec selection. For example, for VVC, because of “reference picture resampling (RPR),” it is allowed for the intra / inter pictures to have different picture resolutions. However, for AVC or HEVC, the RPR is not supported. For such codec selections, when the background picture has a differentresolution, the system can be configured to either code it as an intra picture (in SEI, specifying the background picture as the driving picture) or resample the background picture to the resolution of the base picture, code it as and inter picture, and convert it back to the original resolution at the decoder.

[0108] In one example, the decoding process compatible with the above-described encoding process (2300) includes the following operations:1. The base picture (K) and the background pictures (B) are decoded using the VVC decoder (432).2. The base picture (K) is repaired by the compression artifact reduction tool (440).3. The base picture (K) is segmented using the segmentation tool (404) to obtain the segmentation mask and the corresponding foreground and background. For the foreground image, the decoder (2230) fills in the now missing background by the mean values of the RGB channels in the foreground pixels, which results in the constant background.4. The compact features (C) of the base and driving pictures are decoded. Both (i) the compact features of the base and driving pictures and (ii) the segmented foreground image are provided to the motion estimation and frame generation module (436) to reconstruct the segmented driving pictures (including the foreground with a constant background).5. Transform and crop the decoded background in accordance with the received metadata (2232). The decoded background is of a larger resolution, and the metadata are utilized to crop down the background size from this larger background.6. Both (i) the decoded background picture and (ii) the generated driving picture are given to the fusion module (450) that outputs the corresponding complete frame with a stabilized background for the video (460).

[0109] In some examples, the system is configured to prepare three types of buffers for use with the GFVC architecture (2200) of FIG. 22. Each of these buffers has a larger size than the original video size to accommodate stitched images at different times. In one example, the following buffers are used:1. Ground Truth Buffer. This buffer (denoted as BuffG) is used to accumulate the collected ground truth (registered through the optical flow) pixels along the time domain. For every frame, we keep updating this buffer. In some examples, we can have an auxiliary buffer to indicate whether the pixel BuffG_idx has valid value (e.g., updated from ground truth or still being the same as at the initialization).2. Inpainted Buffer. This buffer is denoted as BuffP. At frame index k, if we decide to send the updated video frame to the client side, we take BuffG to inpaint, thereby getting BuffP(k).3. Video Encoding Buffer. This buffer (denoted as BuffV) is the buffer feeding the video encoder. Parts of this buffer represent different frame indices, e.g., BuffV(k) for the frame index k. When we have an update from BuffP(k), we put it into BuffV(k)=BuffP(k). Otherwise, the buffer repeats itself as BuffV(k)=BuffV(k-l). We let the video codec to handle temporal redundancy. When frame k is identical to frame k- 1, the video codec will signal it as a skip frame. When BuffV (k) ~=BuffV(k-l), the video codec will encode the difference between the two frames using P-frame encoding to reduce the bit rate.With these buffers, the encoding workflow (2300) can be implemented using the following algorithm.Refined GFVC Workflow for Disentangled and Dynamic Background AlgorithmIn this algorithm, dT is a control parameter that can be used to set the update frequency.

[0110] Referring back to FIG. 22, in some examples, the background update module (2202) in the encoder (2210) is enabled only when a change in the background occurs. For a static background (due to the absence of the foreground and / or background motion) the backgroundupdate module (2202) can be disabled, and its output is null. With the background update module (2202) disabled, the GFVC architecture (2200) reduces to the GFVC architecture (400) of FIG. 4, wherein only the background of the first frame is transmitted.

[0111] In the above dynamic algorithms, no update of the background occurs when there is no foreground motion or when it is below the threshold a . Similarly, no update of the background occurs when there is no background / camera motion or when it is below the threshold <z2• For the static background, no update in the background occurs for the same reasons. As such, the static case can be viewed as a special case of the generic dynamic background update. It is to be noted that a static background can be transmitted even for inputs with the dynamic background by setting the thresholds (o^, a2) to relatively large values.

[0112] Based on JVET-AG2032 GFV (generative face video SEI message), additional new syntax can be constructed and used with the above-described GFVC architectures. For the algorithms supporting the dynamic background cases, the resolution of the background picture needs to be known along with the x / y offsets of the starting position in the background picture used for the fusion process. Typically, the resolution of the background picture is known from the coded bitstream. When the picture resolution scaling is not used in the GFVC architecture, only the x / y offset needs to be signaled. In some examples, we can signal the absolute value, or the predicted value from the previously decoded picture. Table 4 provides an SEI message syntax that can be used with the above-described GFVC architecture in some examples. The shown example corresponds to a case in which the cropped background picture is of the same resolution as the base picture.Table 4: Generative Face Video SEI Message.gfv_dec_pic_res_change_flag equal to 1 specifies the decoded picture for driving picture might have different resolution from the base picture and background cropping window offset parameters follow next in the SEI. gfv_dec_pic_res_change_flag equal to 0 specifies the decoded picture for driving picture have the same resolution from the base picture and the background cropping window offset parameters are not present in the SEI. gfv_dec_pic_offset_x and gfv_dec_pic_offset_y specify the top and left position for the background cropping window. When gfv_dec_pic_res_change_flag is equal to 0, the values of gfv_dec_pic_offset_x and gfv_dec_pic_offset_y are inferred to be equal to 0.

[0113] In some examples, a SEI message includes a syntax to specify the camera motion. One embodiment operates to first crop the background window, then apply the affine transform. Another embodiment operates to fist apply the affine transform, then crop the background window. Table 5 shows an example syntax corresponding to the first embodiment.Table 5: Generative Face Video SEI Message.gfv_dec_pic_res_change_flag equal to 1 specifies the decoded picture for driving picture might have different resolution from the base picture and background related parameters follow next in the SEI. gfv_dec_pic_res_change_flag equal to 0 specifies the decoded picture for drivingpicture have the same resolution from the base picture and the background related parameters are not present in the SEI. gfv_dec_pic_left_offset, gfv_dec_pic_right_offset, gfv_dec_pic_top_offset, and gfv_dec_pic_bottome_offset specify the background cropping window. When gfv_dec_pic_res_change_flag is equal to 0, the values of gfv_dec_pic_left_offset, gfv_dec_pic_right_offset, gfv_dec_pic_top_offset, and gfv_dec_pic_bottome_offset are inferred to be equal to 0. gfv_camera_matrix_present_flag equal to 1 specifies the camera matrix is present. gfv_camera_matrix_present_flag equal to 0 specifies the camera matrix is not present. gfv_camera_matrix_element_precision_factor_minusl, gfv_camera_matrix_element_int, gfv_camera_matrix_element_dec have the same semantics as gfv_matrix_element_precision_factor_minusl , gfv_matrix_element_int, and gfv_matrix_element_dec , respectively .Note: in one example, the camera matrix is specified using a 3x3 affine transform.

[0114] FIG. 24 graphically illustrates example improvements that can be achieved with the GFVC architecture (400) of FIG. 4 according to some examples. The bitrate plots (2402, 2404) shown in FIG. 24 are constructed using the Eearned Perceptual Image Patch Similarity (EPIPS) metric. A significant bitrate reduction achieved with the GFVC architecture (400) over the corresponding conventional VVC encoding is evident from a comparison of the bitrate plots (2402, 2404).

[0115] FIG. 25 is a block diagram of an example computing device (2500), one or more instances of which can be used to implement various GFVC architectures, according to some examples. The computing device (2500) of FIG. 25 is illustrated as having a number of components, but any one or more of these components may be omitted or duplicated, as suitable for the application and setting. In some embodiments, some or all of the components included in the computing device (2500) may be attached to one or more motherboards and enclosed in a housing. In some embodiments, some of those components may be fabricated onto a single system-on-a-chip (SoC) (e.g., the SoC may include one or more electronic processing devices (2502) and one or more storage devices (2504)). Additionally, in various embodiments, the computing device (2500) may not include one or more of the components illustrated in FIG. 25,but may include interface circuitry for coupling to the one or more components using any suitable interface (e.g., a Universal Serial Bus (USB) interface, a High-Definition Multimedia Interface (HDMI) interface, a Controller Area Network (CAN) interface, a Serial Peripheral Interface (SPI) interface, an Ethernet interface, a wireless interface, or any other appropriate interface). For example, the computing device (2500) may not include a display device (2510), but may include display device interface circuitry (e.g., a connector and driver circuitry) to which an external display device (2510) may be coupled.

[0116] The computing device (2500) includes a processing device (2502) (e.g., one or more processing devices). As used herein, the terms “electronic processor device” and “processing device” interchangeably refer to any device or portion of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that may be stored in registers and / or memory. In various embodiments, the processing device (2502) may include one or more digital signal processors (DSPs), application-specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GPUs), server processors, or any other suitable processing devices.

[0117] The computing device (2500) also includes a storage device (2504) (e.g., one or more storage devices). In various embodiments, the storage device (2504) may include one or more memory devices, such as random-access memory (RAM) devices (e.g., static RAM (SRAM) devices, magnetic RAM (MRAM) devices, dynamic RAM (DRAM) devices, resistive RAM (RRAM) devices, or conductive-bridging RAM (CBRAM) devices), hard drive-based memory devices, solid-state memory devices, networked drives, cloud drives, or any combination of memory devices. In some embodiments, the storage device (2504) may include memory that shares a die with the processing device (2502). In such an embodiment, the memory may be used as cache memory and include embedded dynamic random-access memory (eDRAM) or spin transfer torque magnetic random-access memory (STT-MRAM), for example. In some embodiments, the storage device (2504) may include non-transitory computer readable media having instructions thereon that, when executed by one or more processing devices (e.g., the processing device (2502)), cause the computing device (2500) to perform any appropriate ones of the methods disclosed herein below or portions of such methods.

[0118] The computing device (2500) further includes an interface device (2506) (e.g., one or more interface devices (2506)). In various embodiments, the interface device (2506) may include one or more communication chips, connectors, and / or other hardware and software to govern communications between the computing device (2500) and other computing devices. Forexample, the interface device (2506) may include circuitry for managing wireless communications for the transfer of data to and from the computing device (2500). The term “wireless” and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data via modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not. Circuitry included in the interface device (2506) for managing wireless communications may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.11 family), IEEE 802.16 standards, Long-Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as “3GPP2”), etc.). In some embodiments, circuitry included in the interface device (2506 for managing wireless communications may operate in accordance with a Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE network. In some embodiments, circuitry included in the interface device (2506 for managing wireless communications may operate in accordance with Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E- UTRAN). In some embodiments, circuitry included in the interface device (2506 for managing wireless communications may operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. In some embodiments, the interface device (2506 may include one or more antennas (e.g., one or more antenna arrays) configured to receive and / or transmit wireless signals.

[0119] In some embodiments, the interface device (2506) may include circuitry for managing wired communications, such as electrical, optical, or any other suitable communication protocols. For example, the interface device (2506) may include circuitry to support communications in accordance with Ethernet technologies. In some embodiments, the interface device (2506) may support both wireless and wired communication, and / or may support multiple wired communication protocols and / or multiple wireless communication protocols. For example, a first set of circuitry of the interface device (2506) may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second set of circuitryof the interface device (2506) may be dedicated to longer-range wireless communications such as global positioning system (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some other embodiments, a first set of circuitry of the interface device (2506) may be dedicated to wireless communications, and a second set of circuitry of the interface device (2506) may be dedicated to wired communications.

[0120] The computing device (2500) also includes battery / power circuitry (2508). In various embodiments, the battery / power circuitry (2508) may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device (2500) to an energy source separate from the computing device (2500) (e.g., to AC line power).

[0121] The computing device (2500) also includes a display device (2510) (e.g., one or multiple individual display devices). In various embodiments, the display device (2510) may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display.

[0122] The computing device (2500) also includes additional input / output (I / O) devices (2512). In various embodiments, the I / O devices (2512) may include one or more data / signal transfer interfaces, audio I / O devices (e.g., microphones or microphone arrays, speakers, headsets, earbuds, alarms, etc.), audio codecs, video codecs, printers, sensors (e.g., thermocouples or other temperature sensors, humidity sensors, pressure sensors, vibration sensors, etc.), image capture devices (e.g., one or more cameras), human interface devices (e.g., keyboards, cursor control devices, such as a mouse, a stylus, a trackball, or a touchpad), etc.

[0123] Depending on the specific embodiment, various components of the interface devices (2506) and / or I / O devices (2512) can be configured to output suitable control signals, receive suitable control / telemetry signals, and receive and transmit data streams. In some examples, the interface devices (2506) and / or I / O devices (2512) include one or more analog-to-digital converters (ADCs) for transforming received analog signals into a digital form suitable for operations performed by the processing device (2502) and / or the storage device (2504). In some additional examples, the interface devices (2506) and / or I / O devices (2512) include one or more digital-to-analog converters (DACs) for transforming digital signals provided by the processing device (2502) and / or the storage device (2504) into an analog form suitable for being transmitted through a communication channel.

[0124] According to an example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-25, provided is a generative face- video-coding method, comprising: obtaining a background picture and a sequence of foreground pictures by applying image segmentation to images corresponding to a face video that includes a first picture and a sequence of second pictures; generating a compressed video bitstream by applying video compression to the first picture and the background picture; generating a face feature representation by processing the sequence of foreground pictures with a neural network; and transmitting the compressed video bitstream and the face feature representation through a communication channel.

[0125] In some embodiments of the above method, applying the image segmentation to the images corresponding to the face video comprises: applying video compression to the first picture to obtain a compressed picture; applying video decompression to the compressed picture to obtain a decompressed picture; and applying image segmentation to the decompressed picture to obtain the background picture.

[0126] In some embodiments of any of the above methods, applying the image segmentation to the images corresponding to the face video further comprises applying artifact reduction to the decompressed picture before the image segmentation.

[0127] In some embodiments of any of the above methods, obtaining the background picture comprises: generating a binary segmentation mask based on the first picture, with a first binary value in the binary segmentation mask representing background pixels of the first picture and a second binary value in the binary segmentation mask representing foreground pixels of the first picture; applying the binary segmentation mask to the first picture to obtain a partial background picture in which the foreground pixels are nulled; and inpainting the foreground pixels of the partial background picture to obtain the background picture.

[0128] In some embodiments of any of the above methods, the method further comprises: updating the background picture to include background changes resulting from one or more of camera motion, background motion, and foreground motion; and including the background picture updates into the compressed video bitstream.

[0129] In some embodiments of any of the above methods, the updating includes one or both of: applying an optical flow algorithm to the sequence of second pictures to determine background displacement due to the camera motion or due to the background motion; and changing the binary segmentation mask to account for the foreground motion.

[0130] In some embodiments of any of the above methods, the updating includes a series of incremental updates, with each of the incremental updates being based on a different respective set of second pictures from the sequence of second pictures.

[0131] In some embodiments of any of the above methods, the method further comprises: applying an optical flow algorithm to a selected pair of second pictures from the sequence of second pictures to obtain a corresponding optical flow frame; generating a union mask by computing a union of two segmentation masks corresponding to the selected pair of second pictures; and determining a background displacement vector using a Hadamard product of the union mask and the corresponding optical-flow frame.

[0132] In some embodiments of any of the above methods, the method further comprises comparing a length of the background displacement vector with a threshold value, wherein an incremental update for the series of incremental updates is performed only when the length exceeds the threshold value.

[0133] In some embodiments of any of the above methods, the method further comprises: computing a difference metric for binary segmentation masks corresponding to a reference frame and a later frame; and when the difference metric is smaller than a threshold value, updating the background picture using inpainting of the latter fame.

[0134] In some embodiments of any of the above methods, the difference metric is computed based on Eq. (1).

[0135] In some embodiments of any of the above methods, the method further comprises transmitting metadata through the communication channel, the metadata including information about the background picture updates configured to assist a decoder with processing the background picture updates.

[0136] In some embodiments of any of the above methods, the metadata include one or more of starting coordinates of a background frame, frame dimensions, and transform parameters.

[0137] In some embodiments of any of the above methods, obtaining the background picture comprises: generating a plurality of binary segmentation masks corresponding to a plurality of second pictures from the sequence of second pictures, with a first binary value in the binary segmentation masks representing background pixels of the second images and a second binary value in the binary segmentation masks representing foreground pixels of the second images; applying the plurality of binary segmentation masks to the plurality of second pictures to obtain acorresponding plurality of partial background pictures in which the foreground pixels are nulled; combining the corresponding plurality of partial background pictures to obtain an aggregated partial-background frame; and inpainting the foreground pixels of the aggregated partialbackground frame to obtain the background picture.

[0138] In some embodiments of any of the above methods, the method further comprises: maintaining a copy of a previous aggregated partial-background frame in a buffer; computing a difference metric for binary segmentation masks corresponding to a reference frame and a later frame; and when the difference metric is smaller than a threshold value, updating the background picture using inpainting of the previous aggregated partial-background frame retrieved from the buffer.

[0139] In some embodiments of any of the above methods, obtaining the sequence of foreground pictures comprises: for each color channel, replacing background pixel values in the second pictures by a respective constant value.

[0140] In some embodiments of any of the above methods, the method further comprises determining the respective constant value for a color channel by computing a mean value of foreground pixel values in said color channel.

[0141] In some embodiments of any of the above methods, the face feature representation is transmitted using a supplemental enhancement information (SEI) message.

[0142] In some embodiments of any of the above methods, the video compression is performed in accordance with a format selected from the group consisting of: an Advanced Video Coding (AVC) format; a High Efficiency Video Coding (HEVC) format; a Versatile Video Coding (VVC) format; and a Model-Based Coding (MBC) format.

[0143] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the above methods.

[0144] According to another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-25, provided is an apparatus for generative face-video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: obtain a background picture and a sequence of foreground pictures by applying imagesegmentation to images corresponding to a face video that includes a first picture and a sequence of second pictures; generate a compressed video bitstream by applying video compression to the first picture and the background picture; generate a face feature representation by processing the sequence of foreground pictures with a first neural network; and transmit the compressed video bitstream and the face feature representation through a communication channel.

[0145] According to yet another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-25, provided is a generative face-video-coding method, comprising: decoding a compressed video bitstream to obtain a first picture and a background picture encoded in the compressed video bitstream; obtaining a first foreground picture via image segmentation applied to the first picture; processing the first foreground picture and a face feature representation with a neural network to obtain a sequence of second foreground pictures; and fusing each of the second foreground pictures with a respective background picture to obtain a sequence of video frames representing a face video, the respective background picture being based on the background picture obtained via the decoding.

[0146] In some embodiments of the above method, the decoding is performed in accordance with a format selected from the group consisting of: an Advanced Video Coding (AVC) format; a High Efficiency Video Coding (HEVC) format; a Versatile Video Coding (VVC) format; and a Model-Based Coding (MBC) format.

[0147] In some embodiments of any of the above methods, the method further comprises receiving the compressed video bitstream and the face feature representation through a communication channel.

[0148] In some embodiments of any of the above methods, the face feature representation is transmitted through the communication channel using a supplemental enhancement information (SEI) message.

[0149] In some embodiments of any of the above methods, obtaining the first foreground picture includes applying artifact reduction to the first picture before the image segmentation.

[0150] In some embodiments of any of the above methods, the background picture has a larger size than each of the video frames representing the face video; and wherein the respective background picture is obtained by at least one of cropping and transforming the backgroundpicture obtained via the decoding, the at least one of cropping and transforming being performed based on corresponding metadata received through the communication channel.

[0151] In some embodiments of any of the above methods, a sequence of different cropping operations is applied to the background picture to produce a dynamic background in the face video.

[0152] In some embodiments of any of the above methods, the dynamic background represents background changes resulting from one or more of camera motion, background motion, and foreground motion.

[0153] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the above methods.

[0154] According to yet another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-25, provided is an apparatus for generative face-video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: decode a compressed video bitstream to obtain a first picture and a background picture encoded in the compressed video bitstream; obtain a first foreground picture via image segmentation applied to the first picture; process the first foreground picture and a face feature representation with a neural network to obtain a sequence of second foreground pictures; and fuse each of the second foreground pictures with a respective background picture to obtain a sequence of video frames representing a face video, the respective background picture being based on the background picture obtained via the decoding.

[0155] According to yet another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-25, provided is a generative face-video-coding method, comprising: obtaining a base picture by applying image segmentation to a first picture of a face video and replacing a background portion of the segmented first picture by a corresponding chroma key background; generating a compressed video bitstream by applying video compression to the base picture; obtaining a sequence of driving pictures by applying image segmentation to a sequence of second pictures of the face video and replacing background portions of the segmented second picture by respective chroma key backgrounds; generating a face feature representation by processing the sequence of drivingpictures with a neural network; and transmitting the compressed video bitstream and the face feature representation through a communication channel.

[0156] In some embodiments of the above method, applying image segmentation to the first picture includes obtaining a background picture; and wherein generating the compressed video bitstream includes applying video compression to the background picture.

[0157] In some embodiments of any of the above methods, obtaining the background picture comprises: generating a binary segmentation mask based on the first picture, with a first binary value in the binary segmentation mask representing background pixels of the first picture and a second binary value in the binary segmentation mask representing foreground pixels of the first picture; applying the binary segmentation mask to the first picture to obtain a partial background picture in which the foreground pixels are nulled; and inpainting the foreground pixels of the partial background picture to obtain the background picture.

[0158] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the above methods.

[0159] According to yet another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-25, provided is an apparatus for generative face-video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: decode a compressed video bitstream to obtain a base picture encoded in the compressed video bitstream, the base picture having a corresponding chroma key background; process the base picture and a face feature representation with a neural network to obtain a sequence of driving pictures, each of the driving pictures having a respective chroma key background; and fuse the base picture and each of the driving pictures with a background picture to obtain a sequence of video frames representing a face video encoded in the compressed video bitstream and in the face feature representation.

[0160] According to yet another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-25, provided is an apparatus for generative face-video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus atleast to: obtain a base picture by applying image segmentation to a first picture of a face video and replacing a background portion of the segmented first picture by a corresponding chroma key background; generate a compressed video bitstream by applying video compression to the base picture; obtain a sequence of driving pictures by applying image segmentation to a sequence of second pictures of the face video and replacing background portions of the segmented second picture by respective chroma key backgrounds; generate a face feature representation by processing the sequence of driving pictures with a neural network; and transmit the compressed video bitstream and the face feature representation through a communication channel.

[0161] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.

[0162] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.

[0163] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.

[0164] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.

[0165] While this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to persons skilled in the art to which the disclosure pertains are deemed to lie within the principle and scope of the disclosure, e.g., as expressed in the following claims.

[0166] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit.

[0167] Some embodiments can be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments can also be embodied in the form of program code recorded in tangible media, such as magnetic recording media, optical recording media, solid state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s). Some embodiments can also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium including being loaded into and / or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus for practicing the patented invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits.

[0168] Unless explicitly stated otherwise, each numerical value and range should be interpreted as being approximate as if the word “about” or “approximately” preceded the value or range.

[0169] The use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not to be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.

[0170] Although the elements in the following method claims, if any, are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence.

[0171] Reference herein to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments necessarily mutually exclusive of other embodiments. The same applies to the term “implementation.”

[0172] Unless otherwise specified herein, the use of the ordinal adjectives “first,” “second,” “third,” etc., to refer to an object of a plurality of like objects merely indicates that different instances of such like objects are being referred to, and is not intended to imply that the like objects so referred-to have to be in a corresponding order or sequence, either temporally, spatially, in ranking, or in any other manner.

[0173] Unless otherwise specified herein, in addition to its plain meaning, the conjunction “if’ may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” which construal may depend on the corresponding specific context. For example, the phrase “if it is determined” or “if [a stated condition] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].”

[0174] Also, for purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developedin which energy is allowed to be transferred between two or more elements, and the interposition of one or more additional elements is contemplated, although not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements.

[0175] As used herein in reference to an element and a standard, the term compatible means that the element communicates with other elements in a manner wholly or partially specified by the standard and would be recognized by other elements as sufficiently capable of communicating with the other elements in the manner specified by the standard. The compatible element does not need to operate internally in a manner specified by the standard.

[0176] The functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and / or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage. Other hardware, conventional and / or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.

[0177] As used in this application, the terms “circuit,” “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not neededfor operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.

[0178] It should be appreciated by those of ordinary skill in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.

[0179] “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” in this specification is intended to introduce some example embodiments, with additional embodiments being described in “DETAILED DESCRIPTION” and / or in reference to one or more drawings. “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

[0180] Various aspects of the present disclosure may be appreciated from the following Enumerated Example Embodiments (EEEs):EEE1. A generative face-video-coding method, comprising: obtaining a background picture and a sequence of foreground pictures by applying image segmentation to images corresponding to a face video that includes a first picture and a sequence of second pictures; generating a compressed video bitstream by applying video compression to the first picture and the background picture; generating a face feature representation by processing the sequence of foreground pictures with a neural network; andtransmitting the compressed video bitstream and the face feature representation through a communication channel.EEE2. The method of EEE1, wherein said applying the image segmentation to the images corresponding to the face video comprises: applying video compression to the first picture to obtain a compressed picture; applying video decompression to the compressed picture to obtain a decompressed picture; and applying image segmentation to the decompressed picture to obtain the background picture.EEE3. The method of EEE2, wherein said applying the image segmentation to the images corresponding to the face video further comprises applying artifact reduction to the decompressed picture before the image segmentation.EEE4. The method of any one of EEE1 to EEE3, wherein obtaining the background picture comprises: generating a binary segmentation mask based on the first picture, with a first binary value in the binary segmentation mask representing background pixels of the first picture and a second binary value in the binary segmentation mask representing foreground pixels of the first picture; applying the binary segmentation mask to the first picture to obtain a partial background picture in which the foreground pixels are nulled; and inpainting the foreground pixels of the partial background picture to obtain the background picture.EEE5. The method of EEE4, further comprising: updating the background picture to include background changes resulting from one or more of camera motion, background motion, and foreground motion; and including the background picture updates into the compressed video bitstream.EEE6. The method of EEE5, wherein the updating includes one or both of: applying an optical flow algorithm to the sequence of second pictures to determine background displacement due to the camera motion or due to the background motion; and changing the binary segmentation mask to account for the foreground motion.EEE7. The method of EEE5, wherein the updating includes a series of incremental updates, with each of the incremental updates being based on a different respective set of second pictures from the sequence of second pictures.EEE8. The method of EEE7, further comprising: applying an optical flow algorithm to a selected pair of second pictures from the sequence of second pictures to obtain a corresponding optical flow frame; generating a union mask by computing a union of two segmentation masks corresponding to the selected pair of second pictures; and determining a background displacement vector using a Hadamard product of the union mask and the corresponding optical-flow frame.EEE9. The method of EEE8, further comprising comparing a length of the background displacement vector with a threshold value, wherein an incremental update for the series of incremental updates is performed only when the length exceeds the threshold value.EEE10. The method of EEE5, further comprising: computing a difference metric for binary segmentation masks corresponding to a reference frame and a later frame; and when the difference metric is smaller than a threshold value, updating the background picture using inpainting of the latter fame.EEE11. The method of EEE10, wherein the difference metric is computed based on Eq. (1).EEE 12. The method of EEE5, further comprising transmitting metadata through the communication channel, the metadata including information about the background picture updates configured to assist a decoder with processing the background picture updates.EEE 13. The method of EEE 12, wherein the metadata include one or more of starting coordinates of a background frame, frame dimensions, and transform parameters.EEE14. The method of EEE1, wherein obtaining the background picture comprises: generating a plurality of binary segmentation masks corresponding to a plurality of second pictures from the sequence of second pictures, with a first binary value in the binarysegmentation masks representing background pixels of the second images and a second binary value in the binary segmentation masks representing foreground pixels of the second images; applying the plurality of binary segmentation masks to the plurality of second pictures to obtain a corresponding plurality of partial background pictures in which the foreground pixels are nulled; combining the corresponding plurality of partial background pictures to obtain an aggregated partial-background frame; and inpainting the foreground pixels of the aggregated partial-background frame to obtain the background picture.EEE15. The method of EEE14, further comprising: maintaining a copy of a previous aggregated partial-background frame in a buffer; computing a difference metric for binary segmentation masks corresponding to a reference frame and a later frame; and when the difference metric is smaller than a threshold value, updating the background picture using inpainting of the previous aggregated partial-background frame retrieved from the buffer.EEE 16. The method of any preceding EEE, wherein obtaining the sequence of foreground pictures comprises: for each color channel, replacing background pixel values in the second pictures by a respective constant value.EEE17. The method of EEE16, further comprising determining the respective constant value for a color channel by computing a mean value of foreground pixel values in said color channel.EEE18. The method of any preceding EEE, wherein the face feature representation is transmitted using a supplemental enhancement information (SEI) message.EEE 19. The method of any preceding EEE, wherein the video compression is performed in accordance with a format selected from the group consisting of: an Advanced Video Coding (AVC) format; a High Efficiency Video Coding (HEVC) format; a Versatile Video Coding (VVC) format; and a Model-Based Coding (MBC) format.EEE20. A generative face-video-coding method, comprising: decoding a compressed video bitstream to obtain a first picture and a background picture encoded in the compressed video bitstream; obtaining a first foreground picture via image segmentation applied to the first picture; processing the first foreground picture and a face feature representation with a neural network to obtain a sequence of second foreground pictures; and fusing each of the second foreground pictures with a respective background picture to obtain a sequence of video frames representing a face video, the respective background picture being based on the background picture obtained via the decoding.EEE21. The method of EEE20, wherein the decoding is performed in accordance with a format selected from the group consisting of: an Advanced Video Coding (AVC) format; a High Efficiency Video Coding (HEVC) format; a Versatile Video Coding (VVC) format; and a Model-Based Coding (MBC) format.EEE22. The method of EEE20 or EEE21, further comprising receiving the compressed video bitstream and the face feature representation through a communication channel.EEE23. The method of EEE22, wherein the face feature representation is transmitted through the communication channel using a supplemental enhancement information (SEI) message.EEE24. The method of any one of EEE20 to EEE23, wherein obtaining the first foreground picture includes applying artifact reduction to the first picture before the image segmentation.EEE25. The method of EEE22, wherein the background picture has a larger size than each of the video frames representing the face video; and wherein the respective background picture is obtained by at least one of cropping and transforming the background picture obtained via the decoding, the at least one of cropping and transforming being performed based on corresponding metadata received through the communication channel.EEE26. The method of EEE25, wherein a sequence of different cropping operations is applied to the background picture to produce a dynamic background in the face video.EEE27. The method of EEE26, wherein the dynamic background represents background changes resulting from one or more of camera motion, background motion, and foreground motion.EEE28. A generative face-video-coding method, comprising: obtaining a base picture by applying image segmentation to a first picture of a face video and replacing a background portion of the segmented first picture by a corresponding chroma key background; generating a compressed video bitstream by applying video compression to the base picture; obtaining a sequence of driving pictures by applying image segmentation to a sequence of second pictures of the face video and replacing background portions of the segmented second picture by respective chroma key backgrounds; generating a face feature representation by processing the sequence of driving pictures with a neural network; and transmitting the compressed video bitstream and the face feature representation through a communication channel.EEE29. The method of EEE28, wherein applying image segmentation to the first picture includes obtaining a background picture; and wherein generating the compressed video bitstream includes applying video compression to the background picture.EEE30. The method of EEE29, wherein obtaining the background picture comprises: generating a binary segmentation mask based on the first picture, with a first binary value in the binary segmentation mask representing background pixels of the first picture and a second binary value in the binary segmentation mask representing foreground pixels of the first picture; applying the binary segmentation mask to the first picture to obtain a partial background picture in which the foreground pixels are nulled; and inpainting the foreground pixels of the partial background picture to obtain the background picture.EEE31. A generative face-video-coding method, comprising: decoding a compressed video bitstream to obtain a base picture encoded in the compressed video bitstream, the base picture having a corresponding chroma key background; processing the base picture and a face feature representation with a neural network to obtain a sequence of driving pictures, each of the driving pictures having a respective chroma key background; and fusing the base picture and each of the driving pictures with a background picture to obtain a sequence of video frames representing a face video encoded in the compressed video bitstream and in the face feature representation.EEE32. The method of EEE31, wherein background picture is a virtual background picture; and wherein the fusing includes replacing the corresponding and respective chroma key backgrounds by respective portions of the virtual background picture.EEE33. The method of EEE31, wherein the decoding includes obtaining the background picture from the compressed video bitstream; and wherein the fusing includes replacing the corresponding and respective chroma key backgrounds by respective portions of the background picture.EEE34. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the methods of EEE1-EEE33.EEE35. An apparatus for generative face- video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: obtain a background picture and a sequence of foreground pictures by applying image segmentation to images corresponding to a face video that includes a first picture and a sequence of second pictures;generate a compressed video bitstream by applying video compression to the first picture and the background picture; generate a face feature representation by processing the sequence of foreground pictures with a first neural network; and transmit the compressed video bitstream and the face feature representation through a communication channel.EEE36. An apparatus for generative face- video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: decode a compressed video bitstream to obtain a first picture and a background picture encoded in the compressed video bitstream; obtain a first foreground picture via image segmentation applied to the first picture; process the first foreground picture and a face feature representation with a neural network to obtain a sequence of second foreground pictures; and fuse each of the second foreground pictures with a respective background picture to obtain a sequence of video frames representing a face video, the respective background picture being based on the background picture obtained via the decoding.EEE37. An apparatus for generative face- video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: decode a compressed video bitstream to obtain a base picture encoded in the compressed video bitstream, the base picture having a corresponding chroma key background; process the base picture and a face feature representation with a neural network to obtain a sequence of driving pictures, each of the driving pictures having a respective chroma key background; and fuse the base picture and each of the driving pictures with a background picture to obtain a sequence of video frames representing a face video encoded in the compressed video bitstream and in the face feature representation.EEE38. An apparatus for generative face-video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: obtain a base picture by applying image segmentation to a first picture of a face video and replacing a background portion of the segmented first picture by a corresponding chroma key background; generate a compressed video bitstream by applying video compression to the base picture; obtain a sequence of driving pictures by applying image segmentation to a sequence of second pictures of the face video and replacing background portions of the segmented second picture by respective chroma key backgrounds; generate a face feature representation by processing the sequence of driving pictures with a neural network; and transmit the compressed video bitstream and the face feature representation through a communication channel.REFERENCESNote: the term JVET refers to the Joint Video Experts Team of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29.1. S. McCarthy, et al., “Technologies under consideration for future extensions of VSEI (version 3),” JVET-AG2032, JVET 33d Meeting, online, output document, uploaded, March 29 2024, ISO / IEC and ITU.2. Xiangguang Chen, et al., “Robust Human Matting via Semantic Guidance,” Proceedings of the Asian Conference on Computer Vision (ACCV), 2020, pp. 2984-2989.3. Chaohao Xie, et al., “Image Inpainting with Learnable Bidirectional Attention Maps,” in Proceedings of the IEEE / CVF international conference on computer vision 2019 (ICCV 2019), pp. 8858-8867.4. S. McCarthy, et al., “Technologies under consideration for future extensions of VSEI (version 4),” JVET-AH2032, JVET 34-th Meeting, Rennes, FR, output document, uploaded June5. 2024, ISO / IEC and ITU.5. S. Gehlot, et al., “AHG9 / AHG16: Showcase for picture fusion for generative face video SEI message,” JVET-AHO188, JVET 34-th Meeting, Rennes, FR, April 2024, ISO / IEC and ITU.

Claims

CLAIMS1. A generative face-video-coding method, comprising: obtaining a background picture and a sequence of foreground pictures by applying image segmentation to images corresponding to a face video that includes a first picture and a sequence of second pictures; generating a compressed video bitstream by applying video compression to the first picture and the background picture; generating a face feature representation by processing the sequence of foreground pictures with a neural network; and transmitting the compressed video bitstream and the face feature representation through a communication channel.

2. The method of claim 1, wherein said applying the image segmentation to the images corresponding to the face video comprises: applying video compression to the first picture to obtain a compressed picture; applying video decompression to the compressed picture to obtain a decompressed picture; and applying image segmentation to the decompressed picture to obtain the background picture.

3. The method of claim 2, wherein said applying the image segmentation to the images corresponding to the face video further comprises applying artifact reduction to the decompressed picture before the image segmentation.

4. The method of any one of claims 1 to 3, wherein obtaining the background picture comprises: generating a binary segmentation mask based on the first picture, with a first binary value in the binary segmentation mask representing background pixels of the first picture and a second binary value in the binary segmentation mask representing foreground pixels of the first picture; applying the binary segmentation mask to the first picture to obtain a partial background picture in which the foreground pixels are nulled; and inpainting the foreground pixels of the partial background picture to obtain the background picture.

5. The method of claim 4, further comprising: updating the background picture to include background changes resulting from one or more of camera motion, background motion, and foreground motion; and including the background picture updates into the compressed video bitstream.

6. The method of claim 5, wherein the updating includes one or both of: applying an optical flow algorithm to the sequence of second pictures to determine background displacement due to the camera motion or due to the background motion; and changing the binary segmentation mask to account for the foreground motion.

7. The method of claim 5, wherein the updating includes a series of incremental updates, with each of the incremental updates being based on a different respective set of second pictures from the sequence of second pictures.

8. The method of claim 7, further comprising: applying an optical flow algorithm to a selected pair of second pictures from the sequence of second pictures to obtain a corresponding optical flow frame; generating a union mask by computing a union of two segmentation masks corresponding to the selected pair of second pictures; and determining a background displacement vector using a Hadamard product of the union mask and the corresponding optical-flow frame.

9. The method of claim 8, further comprising comparing a length of the background displacement vector with a threshold value, wherein an incremental update for the series of incremental updates is performed only when the length exceeds the threshold value.

10. The method of claim 5, further comprising: computing a difference metric for binary segmentation masks corresponding to a reference frame and a later frame; and when the difference metric is smaller than a threshold value, updating the background picture using inpainting of the latter fame.

11. The method of claim 10, wherein the difference metric is computed based on Eq. (1).

12. The method of claim 5, further comprising transmitting metadata through the communication channel, the metadata including information about the background picture updates configured to assist a decoder with processing the background picture updates.

13. The method of claim 12, wherein the metadata include one or more of starting coordinates of a background frame, frame dimensions, and transform parameters.

14. The method of claim 1, wherein obtaining the background picture comprises: generating a plurality of binary segmentation masks corresponding to a plurality of second pictures from the sequence of second pictures, with a first binary value in the binary segmentation masks representing background pixels of the second images and a second binary value in the binary segmentation masks representing foreground pixels of the second images; applying the plurality of binary segmentation masks to the plurality of second pictures to obtain a corresponding plurality of partial background pictures in which the foreground pixels are nulled; combining the corresponding plurality of partial background pictures to obtain an aggregated partial-background frame; and inpainting the foreground pixels of the aggregated partial-background frame to obtain the background picture.

15. The method of claim 14, further comprising: maintaining a copy of a previous aggregated partial-background frame in a buffer; computing a difference metric for binary segmentation masks corresponding to a reference frame and a later frame; and when the difference metric is smaller than a threshold value, updating the background picture using inpainting of the previous aggregated partial-background frame retrieved from the buffer.

16. The method of any preceding claim, wherein obtaining the sequence of foreground pictures comprises: for each color channel, replacing background pixel values in the second pictures by a respective constant value.

17. The method of claim 16, further comprising determining the respective constant value for a color channel by computing a mean value of foreground pixel values in said color channel.

18. The method of any preceding claim, wherein the face feature representation is transmitted using a supplemental enhancement information (SEI) message.

19. The method of any preceding claim, wherein the video compression is performed in accordance with a format selected from the group consisting of: an Advanced Video Coding (AVC) format; a High Efficiency Video Coding (HEVC) format; a Versatile Video Coding (VVC) format; and a Model-Based Coding (MBC) format.

20. A generative face-video-coding method, comprising: decoding a compressed video bitstream to obtain a first picture and a background picture encoded in the compressed video bitstream; obtaining a first foreground picture via image segmentation applied to the first picture; processing the first foreground picture and a face feature representation with a neural network to obtain a sequence of second foreground pictures; and fusing each of the second foreground pictures with a respective background picture to obtain a sequence of video frames representing a face video, the respective background picture being based on the background picture obtained via the decoding.

21. The method of claim 20, wherein the decoding is performed in accordance with a format selected from the group consisting of: an Advanced Video Coding (AVC) format; a High Efficiency Video Coding (HEVC) format; a Versatile Video Coding (VVC) format; and a Model-Based Coding (MBC) format.

22. The method of claim 20 or 21, further comprising receiving the compressed video bitstream and the face feature representation through a communication channel.

23. The method of claim 22, wherein the face feature representation is transmitted through the communication channel using a supplemental enhancement information (SEI) message.

24. The method of any one of claims 20 to 23, wherein obtaining the first foreground picture includes applying artifact reduction to the first picture before the image segmentation.

25. The method of claim 22, wherein the background picture has a larger size than each of the video frames representing the face video; and wherein the respective background picture is obtained by at least one of cropping and transforming the background picture obtained via the decoding, the at least one of cropping and transforming being performed based on corresponding metadata received through the communication channel.

26. The method of claim 25, wherein a sequence of different cropping operations is applied to the background picture to produce a dynamic background in the face video.

27. The method of claim 26, wherein the dynamic background represents background changes resulting from one or more of camera motion, background motion, and foreground motion.

28. A generative face-video-coding method, comprising: obtaining a base picture by applying image segmentation to a first picture of a face video and replacing a background portion of the segmented first picture by a corresponding chroma key background; generating a compressed video bitstream by applying video compression to the base picture; obtaining a sequence of driving pictures by applying image segmentation to a sequence of second pictures of the face video and replacing background portions of the segmented second picture by respective chroma key backgrounds; generating a face feature representation by processing the sequence of driving pictures with a neural network; and transmitting the compressed video bitstream and the face feature representation through a communication channel.

29. The method of claim 28, wherein applying image segmentation to the first picture includes obtaining a background picture; andwherein generating the compressed video bitstream includes applying video compression to the background picture.

30. The method of claim 29, wherein obtaining the background picture comprises: generating a binary segmentation mask based on the first picture, with a first binary value in the binary segmentation mask representing background pixels of the first picture and a second binary value in the binary segmentation mask representing foreground pixels of the first picture; applying the binary segmentation mask to the first picture to obtain a partial background picture in which the foreground pixels are nulled; and inpainting the foreground pixels of the partial background picture to obtain the background picture.

31. A generative face-video-coding method, comprising: decoding a compressed video bitstream to obtain a base picture encoded in the compressed video bitstream, the base picture having a corresponding chroma key background; processing the base picture and a face feature representation with a neural network to obtain a sequence of driving pictures, each of the driving pictures having a respective chroma key background; and fusing the base picture and each of the driving pictures with a background picture to obtain a sequence of video frames representing a face video encoded in the compressed video bitstream and in the face feature representation.

32. The method of claim 31 , wherein background picture is a virtual background picture; and wherein the fusing includes replacing the corresponding and respective chroma key backgrounds by respective portions of the virtual background picture.

33. The method of claim 31 , wherein the decoding includes obtaining the background picture from the compressed video bitstream; and wherein the fusing includes replacing the corresponding and respective chroma key backgrounds by respective portions of the background picture.

34. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the methods of claims 1-33.

35. An apparatus for generative face-video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: obtain a background picture and a sequence of foreground pictures by applying image segmentation to images corresponding to a face video that includes a first picture and a sequence of second pictures; generate a compressed video bitstream by applying video compression to the first picture and the background picture; generate a face feature representation by processing the sequence of foreground pictures with a first neural network; and transmit the compressed video bitstream and the face feature representation through a communication channel.

36. An apparatus for generative face-video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: decode a compressed video bitstream to obtain a first picture and a background picture encoded in the compressed video bitstream; obtain a first foreground picture via image segmentation applied to the first picture; process the first foreground picture and a face feature representation with a neural network to obtain a sequence of second foreground pictures; and fuse each of the second foreground pictures with a respective background picture to obtain a sequence of video frames representing a face video, the respective background picture being based on the background picture obtained via the decoding.

37. An apparatus for generative face-video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: decode a compressed video bitstream to obtain a base picture encoded in the compressed video bitstream, the base picture having a corresponding chroma key background; process the base picture and a face feature representation with a neural network to obtain a sequence of driving pictures, each of the driving pictures having a respective chroma key background; and fuse the base picture and each of the driving pictures with a background picture to obtain a sequence of video frames representing a face video encoded in the compressed video bitstream and in the face feature representation.

38. An apparatus for generative face-video coding, the apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: obtain a base picture by applying image segmentation to a first picture of a face video and replacing a background portion of the segmented first picture by a corresponding chroma key background; generate a compressed video bitstream by applying video compression to the base picture; obtain a sequence of driving pictures by applying image segmentation to a sequence of second pictures of the face video and replacing background portions of the segmented second picture by respective chroma key backgrounds; generate a face feature representation by processing the sequence of driving pictures with a neural network; and transmit the compressed video bitstream and the face feature representation through a communication channel.