Framework for video conferencing based on face restoration

KR102999634B1Active Publication Date: 2026-08-03TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2021-10-01
Publication Date
2026-08-03

Smart Images

  • Figure 112022080235208-PCT00132_ABST
    Figure 112022080235208-PCT00132_ABST
Patent Text Reader

Abstract

A method and apparatus are included, comprising computer code configured to enable a processor or processors to acquire video data, detect at least one face from at least one frame of video data, determine a set of face landmark features of at least one face from at least one frame of video data, and at least partially code the video data by a neural network based on the determined set of face landmark features.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] This application claims priority to U.S. Provisional Application No. 63 / 134,522 filed January 6, 2021 and U.S. Application No. 17 / 490,103 filed September 30, 2021, the entirety of which is expressly incorporated herein by reference.

[0002] The present disclosure relates to video conferencing including face restoration (or face hallucination) capable of restoring realistic details from an actual low-quality (LQ) face to a high-quality (HW) face based on landmark features. Background Technology

[0003] International standardization organizations such as ISO, IEC, and IEEE are actively exploring AI-based video coding technologies, particularly focusing on Deep Neural Network (DNN)-based technologies. Various AhGs have been formed to investigate Neural Network Compression (NNR), Video Coding for Machine (VCM), and Neural Network-based Video Coding (NNVC). China's AITISA and AVS have also established corresponding expert groups to research the standardization of similar technologies.

[0004] Video conferencing has recently become increasingly important, and generally requires low-bandwidth transmission to support collaborative meetings of multiple end users. Compared to general video compression tasks, video in meeting scenarios mostly consists of similar content, namely one or a few speakers who are the main subject of the video and occupy the main part of the overall scene. The unrestricted background can be arbitrarily complex, regardless of whether it is indoors or outdoors, but is less important. Recently, Nvidia’s Maxine video conferencing platform proposed an AI-based framework based on face re-enactment technology. 2D or 3D facial landmarks (one or more of the following and / or their data, such as nose, chin, eyes, proportions, positions, wrinkles, ears, geometries, etc.) ("face landmark(s)" and "face feature(s)" may be considered interchangeable terms here) are extracted from a DNN to capture pose and emotion information of the human face. To capture the shape and texture of the face, these features are transmitted to the decoder along with high-quality features calculated at low frequencies, where a high-quality face is reconstructed by transferring the shape and texture based on pose and expression information from each restored frame. This framework significantly reduces transmission bit consumption because it transmits only pose and expression-related landmark features instead of transmitting original pixels for most frames. However, reconstruction-based frameworks cannot guarantee fidelity to the original facial appearance, and in many cases, dramatic artifacts may occur. For example, they are generally very sensitive to occlusion and large motion, and therefore cannot be strongly utilized in actual video conferencing products.

[0005] As such, there are additional technical flaws, including a lack of compressibility, inaccuracy, and the unnecessary discarding of information related to neural networks.

[0006] According to an exemplary embodiment, there is a method and apparatus comprising a memory configured to store computer program code, and a processor or processors configured to access said computer program code and operate as commanded by said computer program code. The computer program code includes: an acquisition code configured to cause said at least one processor to acquire; a detection code configured to cause said at least one processor to detect at least one face from at least one frame of said video data; a determination code configured to cause said at least one processor to determine a set of face landmarks of at least one face from at least one frame of said video data; and a coding code configured to cause said at least one processor to code said video data at least partially by a neural network based on the determined set of face landmark features.

[0007] According to an exemplary embodiment, the video data comprises an encoded bitstream of the video data, and determining the set of face landmarks comprises upsampling at least one down-sampled sequence obtained by decompressing the encoded bitstream.

[0008] According to an exemplary embodiment, the computer program code further comprises: an additional decision code configured to cause the at least one processor to determine an extended face area (EFA) including an extended boundary area from the detected at least one face area from at least one frame of the video data; and to determine an EFA feature set from the EFA; and an additional coding code configured to cause the at least one processor to code the video data at least partially by the neural network based on the determined face landmark feature set.

[0009] According to an exemplary embodiment, determining the EFA and determining the set of EFA features includes upsampling the at least one downsampled sequence obtained by decompressing the encoded bitstream.

[0010] According to an exemplary embodiment, determining the EFA and determining the set of EFA features further include reconstructing the EFA features for each of the face landmark features of the set of face landmark features by a generative adversarial network.

[0011] According to an exemplary embodiment, coding the video data at least partially by a neural network based on the determined face landmark feature set further comprises aggregating the face landmark set, the reconstructed EFA feature, and the upsampled sequence from upsampling the at least one downsampled sequence to code the video data at least partially by the neural network based on the determined face landmark feature set.

[0012] According to one embodiment, at least one face from at least one frame of the video data is determined to be the largest face among a plurality of faces in at least one frame of the video data.

[0013] According to an exemplary embodiment, the decision code is further configured such that the at least one processor determines a plurality of face landmark feature sets for each of a plurality of faces in at least one frame of the video data, in addition to a face landmark feature set of the at least one face from at least one frame of the video data, and the coding code is further configured such that the at least one processor codes the video data at least partially by the neural network based on the determined face landmark set and the determined plurality of face landmark feature sets. Brief explanation of the drawing

[0014] Additional features, characteristics, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and attached drawings. FIG. 1 is a simplified example of a schematic diagram according to embodiments. FIG. 2 is a simplified example of a schematic diagram according to embodiments. FIG. 3 is a simplified example of a schematic diagram according to embodiments. FIG. 4 is a simplified example of a schematic diagram according to embodiments. FIG. 5 is a simplified example of a diagram according to embodiments. FIG. 6 is a simplified example of a diagram according to embodiments. FIG. 7 is a simplified example of a diagram according to embodiments. FIG. 8 is a simplified example of a diagram according to embodiments. FIG. 9a is a simplified example of a diagram according to embodiments. FIG. 9b is a simplified example of a diagram according to embodiments. FIG. 10 is a simplified example of a flowchart according to embodiments. FIG. 11 is a simplified example of a flowchart according to embodiments. FIG. 12 is a simplified example of a block diagram according to embodiments. FIG. 13 is a simplified example of a block diagram according to embodiments. FIG. 14 is a simplified example of a schematic diagram according to embodiments. Specific details for implementing the invention

[0015] The proposed features discussed below may be used separately or combined in any order. Additionally, the embodiments may be implemented by processing circuits (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a computer-readable non-transient medium.

[0016] FIG. 1 illustrates a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The communication system (100) may include at least two terminals (102, 103) interconnected via a network (105). For unidirectional transmission of data, the first terminal (103) may encode video data at a local location for transmission to another terminal (102) via the network (105). The second terminal (102) may receive the coded video data from the other terminal from the network (105), decode the coded data, and display the restored video data. Unidirectional data transmission may be common in media serving applications, etc.

[0017] FIG. 1 illustrates a second pair of terminals (101 and 104) provided to support bidirectional transmission of coded video that may occur, for example, during a video conference. For bidirectional transmission of data, each terminal (101, 104) can encode video data captured at a local location for transmission to another terminal via a network (105). Each terminal (101, 104) can also receive coded video data transmitted by another terminal, decode the coded data, and display the restored video data on a local display device.

[0018] In FIG. 1, terminals (101, 102, 103, 104) may be exemplified as servers, personal computers, and smartphones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure find applications using laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (105) represents any number of networks that transmit coded video data between terminals (101, 102, 103, 104), including, for example, wired and / or wireless communication networks. The communication network (105) may exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network (105) may not be important to the operation of the present disclosure unless described below.

[0019] FIG. 2 illustrates the placement of a video encoder and a decoder in a streaming environment as an example of an application for the disclosed subject. The disclosed subject may be similarly applied to other video-enabled applications, such as video conferencing, digital TV, and storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0020] A streaming system may include a capture subsystem (203) that may include a video source (201), such as a digital camera, that generates, for example, an uncompressed video sample stream (213). The sample stream (213) may be highlighted with a high data volume compared to an encoded video bitstream and may be processed by an encoder (202) coupled to the camera (201). The encoder (202) may enable or implement aspects of the disclosed subject as described in more detail below, including hardware, software, or a combination thereof. An encoded video bitstream (204) that may be highlighted with a lower data volume compared to the sample stream may be stored in a streaming server (205) for future use. One or more streaming clients (212 and 207) may access the streaming server (205) to retrieve copies (208 and 206) of the encoded video bitstream (204). The client (212) may include a video decoder (211) that decodes an incoming copy of an encoded video bitstream (208) and generates an outgoing video sample stream (210) that can be rendered on a display (209) or another rendering device (not shown). In some streaming systems, the video bitstreams (204, 206, 208) may be encoded according to specific video coding / compression standards. Examples of such standards are mentioned above and are further described herein.

[0021] FIG. 3 may be a functional block diagram of a video decoder (300) according to an embodiment of the present invention.

[0022] The receiver (302) may receive one or more codec video sequences to be decoded by the decoder (300); in the same or different embodiments, one coded video sequence at a time, and the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequences may be received from a channel (301) which may be a hardware / software link to a storage device that stores the encoded video data. The receiver (302) may receive a stream of encoded video data, e.g., encoded audio data and / or ancillary data, along with other data, which may be forwarded to each of them using an entity (not shown). The receiver (302) may separate the coded video sequences from the other data. To prevent network jitter, a buffer memory (303) may be coupled between the receiver (302) and the entropy decoder / parser (304) (hereinafter "parser"). When the receiver (302) receives data from a storage / forwarding device with sufficient bandwidth and controllability or from an isosychronous network, the buffer (303) may not be needed or may be small. For use in a best-effort packet network such as the Internet, the buffer (303) may be needed, may be relatively large, and advantageously may be of adaptive size.

[0023] The video decoder (300) may include a parser (304) to reconstruct symbols (313) from an entropy-coded video sequence. Categories of these symbols include information used to manage the operation of the decoder (300), and potential information for controlling a rendering device, such as a display (312), which is not an essential part of the decoder but can be coupled thereto. The control information for the rendering device(s) may be in the form of a Supplementary Enhancement Information (SEI) message or a fragment of a Video Usability Information parameter set (not shown). The parser (304) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow video coding techniques or standards and may follow principles known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (304) may extract a set of subgroup parameters for at least one of the subgroups of pixels of the video decoder from the coded video sequence based on at least one parameter corresponding to a group. The subgroups may include Groups of Picture (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. The entropy decoder / parser may also extract from the coded video sequence information, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0024] The parser (304) can generate symbols (313) by performing an entropy decoding / parsing operation on a video sequence received from the buffer (303). The parser (304) can receive encoded data and selectively decode specific symbols (313). Additionally, the parser (304) can determine whether specific symbols (313) should be provided to a motion compensation prediction unit (306), a scaler / inverse transform unit (305), an intra prediction unit (307), or a loop filter (311).

[0025] The reconstruction of the symbol (313) may include a number of different units depending on the type of the coded video picture or part thereof (e.g., inter- and intra-pictures, inter- and intra-blocks) and other factors. The units and methods associated with the subgroup control information parsed from the coded video sequence by the parser (304) may be controlled. The flow of this subgroup control information between the parser (304) and the number of units is not illustrated below for clarity.

[0026] Beyond the functional blocks already mentioned, the decoder (300) can be conceptually subdivided into a number of functional units as described below. In actual implementations operating under commercial constraints, many of these units may interact closely with one another and be at least partially integrated with one another. However, the conceptual subdivision into the functional units below is appropriate for illustrating the subject matter disclosed.

[0027] The first unit is a scaler / inverse transform unit (305). The scaler / inverse transform unit (305) receives quantized transform coefficients as well as control information, including the transform to be used, block size, quantization factor, quantization scaling matrix, etc., as symbol(s) (313) from the parser (304). It can output a block containing sample values ​​that can be input to an aggregator (310).

[0028] In some cases, the output samples of the scaler / inverse transform (305) may belong to an intra-coded block; that is, a block that does not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed part of the current picture. This prediction information may be provided by an intra-picture prediction unit (307). In some cases, the intra-picture prediction unit (307) uses surrounding already reconstructed information fetched from the current (partially reconstructed) picture (309) to generate a block of the same size and shape as the block being reconstructed. In some cases, the aggregator (310) adds the prediction information generated by the intra-prediction unit (307) on a sample basis to the output sample information provided by the scaler / inverse transform unit (305).

[0029] In other cases, the output samples of the scaler / inverse conversion unit (305) may be intercoded and potentially belong to a motion-compensated block. In such cases, the motion compensation prediction unit (306) may access the reference picture memory (308) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (313) belonging to the block, these samples may be added to the output of the scaler / inverse conversion unit (referred to as residual samples or residual signals in this case) by the aggregator (310) to generate output sample information. The address within the form of the reference picture memory from which the motion compensation unit fetches prediction samples may be controlled by a motion vector available to the motion compensation unit in the form of a symbol (313) that may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory when a sub-sample exact motion vector is in use, a motion vector prediction mechanism, etc.

[0030] The output sample of the aggregator (310) may be subject to various loop filtering techniques in the loop filter unit (311). The video compression technique may include an in-loop filter technique that is controlled by parameters included in the coded video bitstream and is also made available to the loop filter unit (311) as a symbol (313) from the parser (304), but may respond not only to previously reconstructed and loop-filtered sample values ​​but also to meta-information obtained during the decoding of a previous (in decoding order) part of the coded picture or coded video sequence.

[0031] The output of the loop filter unit (311) may be a sample stream that can be output to the render device (312) as well as stored in the reference picture memory (557) for use in predicting future pictures.

[0032] When fully reconstructed, a specific coded picture can be used as a reference picture for future prediction. When a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (304)), the current reference picture (309) can become part of the reference picture buffer (308), and a fresh current picture memory can be reallocated before starting the reconstruction of the next coded picture.

[0033] The video decoder (300) may perform decoding operations according to a predetermined video compression technology that may be documented in a standard such as ITU-T Rec. H.265. The coded video sequence may follow the syntax specified by the video compression technology or standard in use, in the sense that it complies with the syntax of the video compression technology or standard as specifically stated in the video compression technology document or standard or in the profile document therein. Additionally, for compliance, the complexity of the coded video sequence must be within the range defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0034] In one embodiment, the receiver (302) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence(s). The additional data may be used by the video decoder (300) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form, for example, a time, spatial, or signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.

[0035] FIG. 4 may be a functional block diagram of a video encoder (400) according to one embodiment of the present disclosure.

[0036] The encoder (400) can receive video samples from a video source (401) (not part of the encoder) capable of capturing video image(s) to be coded by the encoder (400).

[0037] A video source (401) may provide a source video sequence to be coded by an encoder (303) in the form of a digital sample stream that may have an arbitrary appropriate bit depth (e.g., 8-bit, 10-bit, 12-bit, …), an arbitrary color space (e.g., BT.601 Y CrCB, RGB, …), and an appropriate sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (401) may be a storage device that stores a pre-prepared video. In a video conferencing system, the video source (401) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that impart motion when viewed in sequence. The picture itself may consist of a spatial array of pixels, where each pixel may contain one or more samples depending on the sampling structure, color space, etc. being used. Those skilled in the art can easily understand the relationship between pixels and samples. The description below focuses on samples.

[0038] According to one embodiment, the encoder (400) can code and compress pictures of a source video sequence into a coded video sequence (410) in real time or under any other time constraint required by the application. Applying an appropriate coding rate is one of the functions of the controller (402). The controller controls other function units as described below and is functionally coupled to these units. The coupling is not shown for clarity. Parameters set by the controller may include rate control-related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify the other functions of the controller (402) because they may belong to a video encoder (400) optimized for a specific system design.

[0039] Some video encoders operate in a manner that a person skilled in the art would easily recognize as a "coding loop." As an oversimplified description, the coding loop may consist of the encoding part of the encoder (402) (hereinafter "source coder") (responsible for generating symbols based on the input picture to be coded and reference picture(s)) and a (local) decoder (406) embedded in the encoder (400) that reconstructs the symbols to generate sample data to be generated by the (remote) decoder (since compression between the symbols and the coded video bitstream in the video compression technique considered in the disclosed subject is lossless). The reconstructed sample stream is input into the reference picture memory (405). Since the decoding of the symbol stream yields an accurate bit result regardless of the decoder location (local or remote), the reference picture buffer content is also bit accurate between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the reference picture samples exactly the same sample values ​​that the decoder "sees" when using the prediction during decoding. The basic principle of reference picture synchronicity (for example, drift occurs if synchronization cannot be maintained due to channel errors) is well known to those skilled in the art.

[0040] The operation of the “local” decoder (406) may be the same as the operation of the “remote” decoder (300) already described in detail above in relation to FIG. 3. Referring briefly to FIG. 4, however, since the symbols are available and the encoding / decoding of symbols for the coded video sequence by the entropy coder (408) and the parser (304) may be lossless, the entropy decoding part of the decoder (300), including the channel (301), receiver (302), buffer (303), and parser (304), may not be fully implemented in the local decoder (406).

[0041] At this point, it can be observed that any decoder technique, excluding parsing / entropy decoding present in the decoder, must necessarily exist in the corresponding encoder in substantially the same functional form. The description of the encoder technique can be abbreviated as it is the inverse of the comprehensively described decoder technique. Further details are required only in specific areas, which are provided below.

[0042] As part of the operation, the source coder (403) may perform motion-compensated predictive coding, which predictively codes the input frame by referencing one or more previously coded frames from a video sequence designated as a "reference frame." In this way, the coding engine (407) codes the difference between a pixel block of the input frame and a pixel block of the reference frame(s) that can be selected as predictive reference(s) for the input frame.

[0043] The local video decoder (406) can decode the coded video data of a frame that can be designated as a reference frame based on the symbols generated by the source coder (403). The operation of the coding engine (407) may advantageously be a lossy process. When the coded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may generally be a replica of the source video sequence with some errors. The local video decoder (406) can replicate the decoding process that can be performed by the video decoder on the reference frame and allow the reconstructed reference frame to be stored in the reference picture cache (405). In this way, the encoder (400) can locally store a copy of the reconstructed reference frame having common content as the reconstructed reference frame to be acquired by the far-end video decoder (without transmission errors).

[0044] The predictor (404) can perform a predictive search for the coding engine (407). That is, for a new frame to be coded, the predictor (404) can search the reference picture memory (405) for sample data (as a candidate reference pixel block) or specific metadata such as a reference picture motion vector, block shape, etc., which can serve as a suitable predictive reference for the new picture. The predictor (404) can operate on a sample block basis to find a suitable predictive reference. In some cases, as determined by the search results obtained by the predictor (404), the input picture may have a predictive reference taken from multiple reference pictures stored in the reference picture memory (405).

[0045] The controller (402) may, for example, manage the coding operation of the video coder (403), including the settings of parameters and subgroup parameters used to encode video data.

[0046] The outputs of all the aforementioned functional units can be entropy-coded in an entropy coder (408). The entropy coder converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable-length coding, arithmetic coding, etc.

[0047] The transmitter (409) may buffer the coded video sequence(s) generated by the entropy coder (408) to prepare them for transmission over a communication channel (411), which may be a hardware / software link to a storage device for storing the encoded video data. The transmitter (409) may merge the coded video data from the video coder (403) with other data to be transmitted, e.g., coded audio data and / or an auxiliary data stream (source not shown).

[0048] The controller (402) can manage the operation of the encoder (400). During coding, the controller (405) can assign a specific coded picture type to each coded picture, which can affect the coding technique that can be applied to each picture. For example, a picture can often be assigned to one of the following frame types.

[0049] An Intra Picture (I-Picture) may be one that can be coded and decoded without using any other frame of the sequence as a prediction source. Some video codecs allow various types of Intra Pictures, including, for example, Independent Decoder Refresh Pictures. Those skilled in the art are aware of these variations of I-Pictures and their respective applications and characteristics.

[0050] The predictive picture (P picture) may be encoded and decoded using intra-prediction or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0051] A bidirectionally predictive picture (B picture) may be coded and decoded using intra-prediction or inter-prediction, which uses up to two motion vectors and reference indices to predict sample values ​​of each block. Similarly, a multiple-predictive picture may use two or more reference pictures and associated metadata for the reconstruction of a single block.

[0052] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples, respectively) and can be coded block by block. A block may be predictively coded by referencing other (already coded) blocks as determined by the coding assignment applied to each picture in the block. For example, a block of picture I may be non-predictively coded or predictively coded by referencing an already coded block of the same picture (spatial prediction or intra prediction). A pixel block of picture P may be non-predictively coded through spatial prediction or temporal prediction by referencing one previously coded reference picture. A block of picture B may be non-predictively coded through spatial prediction or temporal prediction by referencing one or two previously coded reference pictures.

[0053] The video coder (400) may perform coding operations according to a specified video coding technology or standard, such as ITU-T Rec. H.265. In such operations, the video coder (400) may perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Accordingly, the coded video data may follow the syntax specified by the video coding technology or standard used.

[0054] In one embodiment, the transmitter (409) may transmit additional data along with the encoded video. The source coder (403) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as time / space / SNR enhancement layers, redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.

[0055] Figure 5 illustrates the intra prediction modes used in HEVC and JEM. To capture any edge direction present in natural video, the number of directional intra modes is expanded from 33 used in HEVC to 65. The additional directional modes of JEM over HEVC are indicated by dashed arrows in Figure 1(b), while the planar and DC modes remain the same. These denser directional intra prediction modes are applied to all block sizes and both luminal and chroma intra prediction. As shown in Figure 5, the directional intra prediction mode identified by the dashed arrow associated with the odd intra prediction mode index is called the odd intra prediction mode. The directional intra prediction mode identified by the solid arrow associated with the even intra prediction mode index is called the even intra prediction mode. In this document, the directional intra-prediction mode indicated by the solid or dotted arrow in Fig. 5 is also referred to as the angular mode.

[0056] In JEM, a total of 67 intra prediction modes are used for luminal intra prediction. To code the intra modes, a list of most probable modes (MPMs) of size 6 is constructed based on the intra modes of neighboring blocks. If an intra mode is not in the MPM list, a flag is signaled indicating whether the intra mode belongs to the selected modes. JEM-3.0 has 16 selected modes, and these modes are selected uniformly for every fourth angle mode. In JVET-D0114 and JVET-G0060, 16 secondary MPMs are derived to replace the uniformly selected modes.

[0057] FIG. 6 illustrates N reference tiers used in intra-directional mode. There are block unit (611), segment A (601), segment B (602), segment C (603), segment D (604), segment E (605), segment F (606), first reference tier (610), second reference tier (609), third reference tier (608) and fourth reference tier (607).

[0058] In HEVC and JEM, as well as some other standards such as H.264 / AVC, the reference samples used to predict the current block are limited to the nearest reference line (row or column). In multiple reference line intra-prediction methods, the number of candidate reference lines (rows or columns) increases from 1 (i.e., the nearest) to N for the intra-direction mode, where N is an integer greater than or equal to 1. Figure 2 illustrates the concept of a multiple line intra-direction prediction method using a 4×4 prediction unit (PU) as an example. The intra-direction mode can generate a predictor by randomly selecting one of N reference tiers. In other words, a predictor p(x,y) is generated from one of the reference samples S1, S2, …, SN. A flag is signaled to indicate which reference tier was selected for the intra-direction mode. When N is set to 1, the intra-direction prediction method is identical to the existing method of JEM 2.0. In FIG. 6, the reference line (610, 609, 608, 607) consists of six segments (601, 602, 603, 604, 605, 606) along with the top-left reference sample. In this document, the reference tier is also referred to as the reference line. The coordinates of the top-left pixel within the current block unit are (0,0), and the coordinates of the top-left pixel in the first reference line are (-1,-1).

[0059] In JEM, for the luminance component, neighbor samples used for intra-prediction sample generation are filtered before the generation process. Filtering is controlled by a given intra-prediction mode and transform block size. If the intra-prediction mode is DC or the transform block size is equal to 4×4, neighbor samples are not filtered. If the distance between a given intra-prediction mode and a vertical mode (or horizontal mode) is greater than a predefined threshold, the filtering process is activated. For neighbor sample filtering, a [1, 2, 1] filter and a bi-linear filter are used.

[0060] The PDPC (position dependent intra prediction combination) method is an intra prediction method that calls a combination of HEVC-style intra predictions with unfiltered boundary reference samples and filtered boundary reference samples. x , y Each prediction sample located at ) before [ x ][ y ] is calculated as follows.

[0061]

[0062] Here, R x,-1 ,R -1,y and represent the unfiltered reference samples located above and to the left of the current sample (x, y), respectively, and R -1,-1 represents the unfiltered reference sample located at the top-left corner of the current block. The weight is calculated as follows.

[0063] (Formula. 2-2) (Formula. 2-3) (Formula. 2-4) (Formula. 2-5)

[0064] FIG. 7 illustrates a diagram (700) in which DC mode PDPC weights (wL, wT, wTL) for (0, 0) and (1, 0) are placed within a single 4×4 block. When PDPC is applied to DC, planar, horizontal, and vertical intra modes, additional boundary filters, such as HEVC DC mode boundary filters or horizontal / vertical mode edge filters, are not required. FIG. 7 illustrates the definition of reference samples Rx,-1, R-1,y, and R-1,-1 for PDPC applied to the top right diagonal mode. The prediction sample pred(x', y') is located at (x', y') within the prediction block. The coordinate x of the reference sample Rx,-1 is given as x = x' + y' + 1, and the coordinate y of the reference sample R-1,y is similarly given as y = x' + y' + 1.

[0065] FIG. 8 illustrates a Local Illumination Compensation (LIC) diagram (800) based on a linear model for illumination change using a scaling factor a and an offset b, and adaptively enabled or disabled for each intermode coded coding unit (CU).

[0066] When LIC is applied to CU, the least squares error method is used to derive parameters a and b using neighbor samples of the current CU and corresponding reference samples. More specifically, as exemplified in FIG. 8, subsampled (2:1 subsampling) neighbor samples of the CU and corresponding samples from the reference picture (identified as motion information of the current CU or sub-CU) are used. IC parameters are derived and applied separately for each prediction direction.

[0067] When the CU is coded in merge mode, the LIC flag is copied from neighboring blocks in a manner similar to copying motion information in merge mode; otherwise, the LIC flag is signaled to the CU to indicate whether the LIC is applied.

[0068] FIG. 9a illustrates an intra prediction mode (900) used in HEVC. HEVC has a total of 35 intra prediction modes, of which mode 10 is a horizontal mode, mode 26 is a vertical mode, and modes 2, 18, and 34 are diagonal modes. The intra prediction modes are signaled by 3 most probable modes (MPMs) and 32 remaining modes.

[0069] FIG. 9b illustrates an embodiment of VVC in which there are a total of 87 intra prediction modes, where mode 18 is a horizontal mode, mode 50 is a vertical mode, and modes 2, 34, and 66 are diagonal modes. Modes -1 to -10 and modes 67 to 76 are called Wide-Angle Intra Prediction (WAIP) modes.

[0070] The predicted sample pred(x,y) located at position (x, y) is predicted using a linear combination of intra-prediction modes (DC, plane, angle) and reference samples according to the PDPC representation.

[0071] ppred(x,y) = ( wL × R-1,y + wT × Rx,-1 - wTL × R-1,-1 + (64 - wL - wT + wTL) × pred(x,y) + 32 ) >> 6

[0072] Here, Rx,-1 and R-1,y represent reference samples located at the top and left of the current sample (x, y), respectively, and R-1,-1 represents a reference sample located at the top-left corner of the current block.

[0073] In DC mode, the weight for a block with width and height dimensions is calculated as follows.

[0074] wT = 32 >> ( ( y<<1 ) >> nScale ), wL = 32 >> ( ( x<<1 ) >> nScale ), wTL = ( wL>>4 ) + ( wT>>4 )

[0075] Here, wT indicates the weighting factor for the reference sample located on the upper reference line with the same horizontal coordinates, wL indicates the weighting factor for the reference sample located on the left reference line with the same vertical coordinates, wTL indicates the weighting factor for the top-left reference sample of the current block, nScale specifies how quickly the weighting factor decreases along the axis (wL decreases from left to right or wT decreases from top to bottom), i.e., the rate of weighting factor decrease, and is the same along the x-axis (left to right) and y-axis (top to bottom) in the current design. And 32 indicates the initial weighting factor for neighbor samples, which is also the top (left or top-left) weight assigned to the top-left sample in the current CB, and in the PDPC process, the weighting factor of neighbor samples must be less than or equal to the initial weighting factor.

[0076] For planar mode, wTL = 0, for horizontal mode, wTL = wT, and for vertical mode, wTL = wL. PDPC weights can be calculated only by addition and shifting. The value of pred(x,y) can be computed in a single step using Equation 1.

[0077] The methods proposed herein may be used separately or combined in any order. Additionally, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a computer-readable non-transient medium. According to the embodiment, the term block may be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU.

[0078] FIG. 10 illustrates an exemplary flowchart (1000) and will be described with further reference to FIG. 12, which illustrates a workflow (1200) of an exemplary framework according to an exemplary embodiment. The workflow (1200) includes modules such as a face detection and face landmark extraction module (122), a spatial-temporarily (ST) downsampling module (123), a landmark feature compression and transmission module (126), an extended face area (EFA) feature compression and transmission module (127), a face detail reconstruction module (130), an EFA reconstruction module (131), a video compression and transmission module (135), an ST upsampling module (137), and a fusion module (139), and the workflow (1200) also includes various data (121, 124, 125, 128, 129, 132, 133, 134, 136, 138, 140).

[0079] Input video sequence such as data (121), like S101 Given, the Face Detection & Facial Landmark Extraction module (122) first, in S102, each video frame One or more valid faces are determined from. In one embodiment, only the most prominent (e.g., the largest) face is detected, and in another embodiment, all faces within the frame that satisfy a condition (e.g., having a sufficiently large size exceeding a threshold) are detected. In S103, For the j-th face, a set of face landmarks is determined, and face landmark features The set is calculated accordingly, and this is in the decoder It will be used to reconstruct the j-th face in. In S103, all face landmark features of all faces It is combined into data (124), which is encoded and transmitted by the Landmark Feature Compression & Transmission module (126). At the same time in S105, For the j-th face, the Extended Face Area (EFA) can be calculated by extending the boundary area of ​​the original detected face (boundaries such as rectangles, eclipses, or fine-grained segmentation boundaries) to include additional hair, body parts, or background. In S106 and S107, EFA features The sets can be calculated correspondingly, and this In S107, all EFA features of all faces will be used by the decoder to reconstruct the EFA of the j-th face. , combined into data (125), which is encoded and transmitted by the EFA Compression & Transmission module (127).

[0080] According to exemplary embodiments, the face detection and face landmark extraction module (122) is an arbitrary object detection DNN that treats human faces as a special object category, or another DNN architecture specifically designed to find human faces, for each video frame x i An arbitrary face detector may be used to locate face regions. The face detection and face landmark extraction module (122) may also use an arbitrary face landmark detector to locate a predetermined set of face landmarks (e.g., surrounding landmarks such as left / right eyes, nose, mouse, etc.) for each detected face. In some embodiments, a single multi-task DNN may be used to locate the positions of faces and associated landmarks simultaneously. Face landmark features can be an intermediate latent representation computed by a face landmark detector directly used to find the landmarks of the j-th face. The intermediate latent representation is further processed, and face landmark features To calculate this, additional DNNs may be applied. For example, information from feature maps corresponding to individual landmarks around facial parts, such as the right eye, can be aggregated into joint features for the corresponding facial parts. Similarity, EFA features can be an intermediate latent representation calculated by a face detector corresponding to the j-th face. An additional DNN, based on the intermediate latent representation, for example by emphasizing background regions rather than actual face regions, It may also be used to calculate. Various exemplary embodiments may not be limited to face detectors, face landmark detectors, face landmark feature extractors, or EFA feature extractor methods or DNN architectures.

[0081] According to an exemplary embodiment, the landmark feature compression and transmission module (126) may use various methods to efficiently compress face landmark features. In a preferred embodiment, a codebook-based mechanism is used in which a codebook can be generated for each face part (e.g., right eye). Then, for a specific face part of a specific face (e.g., the right eye of the current face in the current frame), the face landmark feature may be represented as a combination of weighted codewords in this codebook. In such a case, the codebook is stored on the decoder side, and it may only be necessary to pass weight coefficients for the codewords to the decoder side to restore the face landmark feature. Similarly, the EFA compression and transmission module (127) may use various methods to compress EFA features. In a preferred embodiment, an EFA codebook in which a specific EFA feature is represented by a weighted combination of EFA codewords is also used, and it may only be necessary to pass weight coefficients for the codewords to restore the EFA feature.

[0082] Meanwhile, the input video sequence , data (121) is obtained by the ST (spatial-temporarily) downsampling module (123) ST is downsampled. Compared to, It can be downsampled spatially, temporally, or both spatially and temporally. When is spatially downsampled, each and has the same timestamp, and Is It is calculated with reduced resolution, for example, by conventional or DNN-based interpolation. When is downsampled in time, each at different timestamps Corresponding to, herek is the downsampling frequency (one frame is To create X All of k (Sampled from several frames). X When is downsampled both spatially and temporally, each is reduced resolution from different timestamps, for example, by conventional or DNN-based interpolation It is calculated from. Then, the downsampled sequence , data (134) is the original HQ input It can be processed as the LQ version of. It can be encoded and transmitted by a video compression and transmission module (135). Any video compression framework may be used by the video compression and transmission module (135), such as HEVC, VVC, NNVC, or end-to-end video coding, according to exemplary embodiments.

[0083] At the decoder side, for example, as described in relation to the flowchart (1100) of FIG. 11 and the various modules of FIG. 12, at S111, the received encoded bitstream is first decompressed at S112 to obtain a decoded downsampled sequence , data (136), decoded EFA features , data (129), and decoded face landmark features , obtain data (128). Each decoded frame downsampled Corresponds to. Each decoded EFA characteristic EFA characteristics Corresponds to. Each decoded feature of each landmark It is a landmark feature Corresponds to. In S113, the decoded downsampled sequence is an upsampled sequence , is transmitted through the ST Up Sample module (137) to generate data (138). Corresponding to the encoder size, this ST Up Sample module performs both spatial and temporal upsampling, or spatial and temporal upsampling, as an inverse operation of the downsampling process in the ST Down Sample module (123). When spatial downsampling is used on the encoder side, spatial upsampling is used, where each at the same timestamp, for example, by conventional interpolation or DNN-based superresolution methods It is up-sampled, Is It has the same resolution as . When temporal downsampling is used on the encoder side, temporal upsampling is used, and here each go And, and Addition between ( k -1) The frames are, for example, and It can be calculated using existing motion interpolation or DNN-based frame synthesis methods based on. When both spatial and temporal downsampling are used on the encoder side, spatial and temporal upsampling are used, where each using existing interpolation or DNN-based superresolution methods By spatially upsampling It is calculated from, and The additional frames in between are, and It is additionally generated using existing motion interpolation or DNN-based frame synthesis methods based on.

[0084] In S114, decoded EFA features is the sequence of the reconstructed EFA Data (133) is transmitted through the EFA reconstruction module (131) to calculate each = is a frame at j EFA set, which is the EFA of the th face Includes decoded landmark features , data (128) is transmitted through the face detail reconstruction module (130), and the restored face details , calculates the sequence of data (132). Each = is a frame at j Face detail expression set corresponding to the th face Includes. In a preferred embodiment, the EFA reconstruction module (131) is a DNN composed of a stack of residual blocks and convolution layers. The face detail reconstruction module (130) is a conditional generative adversarial network (GAN) that conditions landmark features corresponding to different face parts. Timestamp i for To calculate, the EFA reconstruction module (131) uses the decoded EFA features of this time stamp Using only or an EFA of a few neighboring timestamps ( n , m (which is any positive integer) can be used. Similarly, timestamps i for To calculate, the face detail reconstruction module (130) uses the decoded landmark features of this timestamp Using only or an EFA of a few neighbor timestamps It can be used. Afterwards, in S115, restored face details , reconstructed EFA , and upsampled sequence The final reconstructed video sequence that is aggregated together by the fusion module (139) , generates data (140). The fusion module can be a small DNN, where timestamp i at To generate, the fusion module from the same timestamp , , and or use only , You can use, and from several neighbor timestamps The following may be used. Exemplary embodiments may not include any limitations on the DNN architecture of the face detail reconstruction module (130), EFA reconstruction module (131), and / or fusion module (139).

[0085] The purpose of using the EFA function is to improve the reconstruction quality of extended facial regions (e.g., hair, body parts, etc.). In some embodiments, the process associated with EFA may be optional depending on the trade-off between the reconstruction quality and the computation and transmission costs. Accordingly, in FIG. 12, this optional process is indicated by a dashed line such as between the elements (125, 127, 129, 131, and 133).

[0086] Additionally, according to an exemplary embodiment, the proposed framework has several components to be trained, and such training will be described in relation to FIG. 13, which illustrates the workflow (1300) of an exemplary training process according to an exemplary embodiment. The workflow (1300) includes modules such as a face detection and face landmark extraction module (223), an ST downsampling module (222), a landmark feature noise modeling module (226), an EFA feature noise modeling module (227), a face detail reconstruction module (230), an EFA reconstruction module (231), a video noise modeling module (235), an ST upsampling module (237), a fusion module, an adversarial loss calculation module (241), a reconstruction loss calculation module (242), and a compute perceptual loss module (243), and the workflow (1300) also includes various data (221, 224, 225, 229, 228, 232, 233, 236, 238, 240).

[0087] According to an exemplary embodiment, the proposed framework includes several components that need to be trained before deployment, including a face detector, a face landmark detector, a face landmark feature extractor, and an EFA feature extractor of a face detection and face landmark module (122), an EFA reconstruction module (131), and a face detail reconstruction module (130). Optionally, if a learning-based downsampling or upsampling method is used, an ST downsampling (123) module and an ST upsampling (137) module must also be pre-trained. In a preferred embodiment, all these components use a DNN-based method, and the weight parameters of this DNN need to be trained. In another embodiment, some of these components may use a conventional learning-based method, such as a conventional face landmark detector, and the corresponding model parameters also need to be trained. Each DNN-based or conventional learning-based component is first pre-trained individually and then jointly tuned through the training process described in this disclosure.

[0088] For example, FIG. 13 provides the entire workflow (1300) of a preferred embodiment of the training process. For training, the actual video compression and transmission module (135) relies on the video noise modeling module (235). This is because actual video compression involves non-differentiable processes such as quantization. The video noise modeling module (235) recycles random noise into a downsampled sequence In addition to, by miming the true data distribution of the decoded downsampled sequence in the final test phase, the decoded downsampled sequence It is generated in the training process. Therefore, the noise model used by the video noise modeling module (235) generally relies on the actual video compression method used in practice. Similarly, the EFA feature compression and transmission module (127) is replaced with the EFA feature noise modeling module (227), which is the noise In addition to, by mimicking the data distribution of the actual decoded EFA features, the EFA features decoded in the final training step It generates. In addition, the landmark feature compression and transmission module (126) is replaced with the landmark feature noise modeling module (226), which is noise In addition to, by mimicking the true distribution of actually decoded landmark features, the EFA features decoded in the final training stage Generates. An exemplary example calculates the following loss function for training.

[0089] Several types of loss are calculated during the training process to learn learnable components. Distortion loss The difference between the original training sequence and the reconstructed training sequence can be calculated in the reconstruction loss calculation module (242) to measure the difference, for example, and, here, Is and It can be MAE or SSIM between. Importance weight maps can also be used to highlight distortions in the reconstructed face region or different parts of the face region. In addition, perceptual loss It can be calculated in the perceptual loss calculation module, for example, And, here, the feature extraction DNN (e.g., VGG backbone network) is, respectively and Calculate the feature representation based on. and The difference in feature representations calculated based on (e.g., MSE) is used as the perception loss. Allelic loss is the reconstructed input To measure how natural it looks, it can be calculated by the antagonistic loss calculation module (241), for example. is. This is the actual x or reconstructed It is fed to a discriminator (typically a classification DNN like ResNet) to classify whether it is natural or reconstructed, and the classification error (e.g., cross-entropy loss) is It can be used as. Distortion loss , perceptual loss , and antagonistic loss is joint loss It is weighted and combined, and here, a gradient can be calculated to update model parameters through backpropagation.

[0090] (Formula. 1)

[0091] Here, α and β are hyperparameters that balance the importance of different loss conditions.

[0092] Different components may be updated at different times with different update frequencies. In some cases, only some components may be updated periodically after deployment or frequently when new training data becomes available. In some cases, only some model parameters may be updated after deployment. This deployment does not impose restrictions on the optimization method, model update frequency, or the percentage of model parameters to be updated.

[0093] As such, an exemplary embodiment of any one of the workflows (1200 and 1300) represents a new framework for video compression and transmission in video conferencing based on face reconstruction with improved coding efficiency by transmitting LQ frames and face features, a flexible and general framework for spatially, temporally, or spatially-temporally downsampled frames, a flexible and general framework for different DNN architectures, and a flexible and general framework capable of accommodating multiple faces with any background.

[0094] The embodiments also describe a video conferencing framework based on face restoration (or face illusion) that restores realistic details from actual low-quality (LQ) faces to high-quality (HQ) faces. Instead of relying on error-prone shape and texture transmission as in face reproduction methods, it restores details of HQ faces based on LQ face and face landmark features. The exemplary framework disclosed herein can ensure robust quality of the restored faces, which are the core of the actual product. For example, transmission costs can be reduced by transmitting only downsampled frames and face features, and HQ frames can be restored at the decoder side based on downsampled frames and face features.

[0095] The technology described above may be implemented as computer software that uses computer-readable instructions and is physically stored on one or more computer-readable media, or by one or more specifically configured hardware processors. For example, FIG. 14 illustrates a computer system (1400) suitable for implementing a specific embodiment of the disclosed subject matter.

[0096] Computer software may be coded using any suitable machine code or computer language, which generates code containing instructions that can be interpreted, microcode executed, and executed through one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or directly, under assembly, compilation, linking, or similar mechanisms.

[0097] The command can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, etc.

[0098] The components illustrated in FIG. 14 for the computer system (1400) are essentially exemplary and are not intended to suggest any limitations on the scope of use or function of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be interpreted as having any dependencies or requirements related to any one or any combination of the components illustrated in the exemplary embodiments of the computer system (1400).

[0099] The computer system (1400) may include a specific human interface input device. This human interface input device may respond to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, applause), visual input (e.g., gestures), and olfactory input (not shown). The human interface device may also be used to capture specific media that are not directly related to conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., 2D video, 3D video including stereoscopic video).

[0100] The input human interface device may include one or more of a keyboard (1401), a mouse (1402), a trackpad (1403), a touch screen (1410), a joystick (1405), a microphone (1406), a scanner (1408), and a camera (1407) (each only one is shown).

[0101] The computer system (1400) may also include specific human interface output devices. These human interface output devices may stimulate the senses of one or more human users, for example, through tactile output, sound, light and smell / taste. These human interface output devices may include a tactile output device (e.g., a tactile feedback device that includes tactile feedback via a touch screen (1410) or a joystick (1405), but does not function as an input device), an audio output device (e.g., a speaker (1409), headphones (not shown)), a visual output device (e.g., a screen (1410) including a CRT screen, an LCD screen, a plasma screen, or an OLED screen, each of which may or may not have a touch screen input function, each of which may or may not have a tactile feedback capability, some of which may or may not have a two-dimensional visual output or a three-dimensional output via stereographic output means such as virtual reality glasses (not shown), a holographic display and a smoke tank (not shown)), and a printer (not shown).

[0102] The computer system (1400) also includes human-accessible storage devices and associated media, such as optical media including CD / DVD ROM / RW (1420) or similar media with CD / DVD (1411), thumb drives (1722), removable hard drives or solid-state drives (1423), legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0103] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include transmission media, carrier waves, or other transient signals.

[0104] The computer system (1400) may also include an interface (1499) to one or more communication networks (1498). The network (1498) may be, for example, wireless, wired, or optical. The network (1498) may also be local, wide area, metropolitan, automotive and industrial, real-time, latency-tolerant, etc. Examples of networks (1498) include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, automotive and industrial networks including CANBus, etc. A specific network generally requires an external network interface adapter attached to a specific general-purpose data port or peripheral bus (1450 and 1451) (for example, a USB port of the computer system); Others are integrated into the core of the computer system (1400) by means of what is typically attached to the system bus (e.g., an Ethernet interface for a PC computer system or a cellular network interface for a smartphone computer system), as described below. Using any of these networks (1498), the computer system (1400) can communicate with other entities. Such communication may be unidirectional, receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., from a CANbus to a specific CANbus device), or bidirectional, to other computer systems using local or wide-area digital networks, for example. Specific protocols and protocol stacks may be available for each network and network interface as described above.

[0105] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core (1440) of the computer system (1400).

[0106] The core (1440) may include one or more central processing units (CPUs) (1441), graphics processing units (GPUs) (1442), graphics adapters (1417), specialized programmable processing units in the form of Field Programmable Gate Areas (FPGAs) (1443), hardware accelerators (1444) for specific tasks, etc. Along with read-only memory (ROM) (1445), random access memory (1746), and internal mass storage (1447) such as internal non-user accessible hard drives, SSDs, etc., these devices may be connected via a system bus (1448). In some computer systems, the system bus (1448) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached to the core's system bus (1448) directly or via a peripheral bus (1451). Architectures of peripheral buses include PCI, USB, etc.

[0107] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) can be combined to execute specific instructions that can construct the aforementioned computer code. This computer code may be stored in ROM (1445) or RAM (1446). Transitional data may be stored in RAM (1446), but permanent data may be stored, for example, in internal mass storage (1447). Fast storage and retrieval for all memory devices may be enabled by using a cache memory that may be closely associated with one or more CPUs (1441), GPUs (1442), mass storage (1447), ROM (1445), RAM (1446), etc.

[0108] A computer-readable medium may have computer code for performing various computer-implemented operations. The medium and the computer code may be specifically designed and configured for the purposes of this disclosure, or they may be of a type well known and available to those skilled in the field of computer software.

[0109] For example, without limitation, a computer system having an architecture (1400), in particular a core (1440), may provide function as a result of one or more types of processor(s) (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software implemented on a computer-readable medium. Such computer-readable medium may be a medium associated with the user-accessible mass storage introduced above, or a specific storage of the core (1440) having non-transient characteristics, such as a core internal mass storage (1447) or ROM (1445). Software implementing various embodiments of the present disclosure may be stored on such a device and executed by the core (1440). The computer-readable medium may include one or more memory devices or chips depending on specific requirements. Software may enable the core (1440) and, in particular, the processor within it (including a CPU, GPU, FPGA, etc.) to execute the specific process or specific process described herein, including defining data structures stored in RAM (1446) and modifying such data structures according to a process defined in the software. Additionally or alternatively, the computer system may provide functionality as a result of logic that is hardwired or implemented in a circuit (e.g., accelerator (1444)) that can work in place of or with the software to execute the specific process or specific part of the specific process described herein. References to software may include logic, and where appropriate, vice versa. References to computer-readable media may include circuits that store software for execution (e.g., integrated circuits (ICs)), circuits that implement logic for execution, or, where appropriate, both. The present disclosure covers any suitable combination of hardware and software.

[0110] Although the present disclosure describes some exemplary embodiments, there are modifications, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, those skilled in the art will understand that numerous systems and methods can be devised that embody the principles of the present disclosure and are therefore within the spirit and scope thereof, even though they are not expressly illustrated or described in this specification.

Claims

Claim 1 A method performed by at least one processor of a video decoder, comprising: acquiring a bitstream comprising a set of face landmark features for at least one face included in at least one frame of video data and a sequence of encoded video data; decoding the sequence of encoded video data; and reconstructing the video data at least partially by a neural network based on the set of face landmark features, wherein the bitstream further comprises a set of extended face region (EFA) features, and the set of EFA features is determined from an EFA including a boundary region extended from a region of one face included in at least one frame of video data, and the step of reconstructing the video data comprises: reconstructing an EFA feature for each of the face landmark features of the set of face landmark features by a generative adversarial network; and further reconstructing the video data at least partially by the neural network based on the reconstructed set of EFA features. Claim 2 A method according to claim 1, further comprising the step of upsampling a sequence of decoded video data. Claim 3 In paragraph 2, the step of reconstructing the video data comprises the step of reconstructing the video data at least partially by the neural network based additionally on a sequence of upsampled video data. Claim 4 A method according to claim 1, further comprising the step of upsampling a sequence of decoded video data, and the step of reconstructing the video data further comprising the step of reconstructing the video data at least partially by the neural network based on the face landmark feature set, the reconstructed EFA feature set, and the sequence of upsampled video data. Claim 5 A method according to claim 1, wherein at least one face from at least one frame of the video data is determined as the largest face among a plurality of faces in at least one frame of the video data. Claim 6 A method according to claim 1, wherein the neural network comprises a deep neural network (DNN). Claim 7 A video coding device comprising: at least one memory configured to store computer program code; and at least one processor configured to access said computer program code and operate as commanded by said computer program code, wherein said computer program code causes said processor to perform a video coding method according to any one of claims 1 to 4. Claim 8 A computer-readable non-transient medium storing a program that causes a computer to execute a process, wherein the process comprises steps of a method according to any one of claims 1 to 4. Claim 9 delete Claim 10 delete Claim 11 delete Claim 12 delete Claim 13 delete Claim 14 delete Claim 15 delete Claim 16 delete Claim 17 delete Claim 18 delete Claim 19 delete Claim 20 delete