Virtual reference frame for image encoding and decoding

By generating virtual reference frames and encoding image frames based on these frames, the problems of poor video quality and bandwidth resource consumption caused by reference frame limitations in the existing technology are solved, and more efficient video encoding and decoding are achieved.

CN120642334APending Publication Date: 2025-09-12QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480011191.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-14
Filing Date
2024-02-07
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In the prior art, due to the limitation of reference frames in video decoding, the prediction quality of image frames is poor, which in turn affects the video reproduction quality and increases bandwidth resource consumption.

Method used

Virtual reference frames are selectively generated by generating synthetic support data associated with the image frames, and encoded versions of the image frames are encoded based on the virtual reference frames.

Benefits of technology

It improves video quality, reduces bandwidth consumption, and enables more efficient image encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120642334A_ABST
    Figure CN120642334A_ABST
Patent Text Reader

Abstract

An apparatus includes one or more processors configured to obtain a bitstream corresponding to an encoded version of an image frame. The one or more processors are further configured to generate a virtual reference frame based on synthesis support data included in the bitstream based on determining that the bitstream includes a virtual reference frame usage indicator. The one or more processors are further configured to generate a decoded version of the image frame based on the virtual reference frame.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] I. Cross-reference to Related Applications

[0002] This application claims the benefit of priority to commonly owned U.S. non-provisional patent application No. 18 / 168,891, filed on February 14, 2023, the contents of which are expressly incorporated herein by reference in their entirety. Technical Field

[0003] The present disclosure generally relates to image encoding and decoding.

[0004] III. Description of Related Technology

[0005] Technological advances have led to smaller and more powerful computing devices. For example, there are currently a variety of portable personal computing devices, including small, lightweight, and user-friendly wireless phones (such as mobile and smart phones, tablet computers, and laptop computers). These devices can communicate voice and data packets over wireless networks. In addition, many such devices incorporate additional functionality, such as digital still cameras, digital video cameras, digital recorders, and audio file players. In addition, such devices can process executable instructions, including software applications, such as web browser applications, which can be used to access the Internet. As such, these devices can include significant computing power.

[0006] Such computing devices typically incorporate functionality for receiving encoded video data corresponding to compressed image frames from another device. Typically, previously decoded image frames are used as reference frames for predicting decoded image frames. The more suitable such reference frames are for predicting image frames, the more accurately the image frames can be decoded, resulting in higher-quality reproduction of the video data. However, because the reference frames available to conventional decoders are limited to previously decoded image frames, in some cases the available reference frames can only provide suboptimal predictions of the image frames, potentially resulting in reduced-quality video reproduction. While decoding quality can be enhanced by sending additional data to the decoder to generate a higher-quality reproduction of the image frames, transmitting such additional data consumes additional bandwidth resources, which may not be available to devices operating with limited transmission channel capacity. Summary of the Invention

[0007] According to one embodiment of the present disclosure, a device includes one or more processors configured to obtain composition support data associated with an image frame in a sequence of image frames. The one or more processors are further configured to selectively generate a virtual reference frame based on the composition support data. The one or more processors are further configured to generate a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.

[0008] According to another embodiment of the present disclosure, a method includes obtaining, at a device, synthesis support data associated with an image frame in a sequence of image frames. The method also includes selectively generating a virtual reference frame based on the synthesis support data. The method also includes generating, at the device, a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.

[0009] According to another specific implementation of the present disclosure, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to obtain composition support data associated with an image frame in a sequence of image frames. The instructions, when executed by the one or more processors, further cause the one or more processors to selectively generate a virtual reference frame based on the composition support data. The instructions, when executed by the one or more processors, further cause the one or more processors to generate a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.

[0010] According to another embodiment of the present disclosure, an apparatus includes means for obtaining composition support data associated with an image frame in a sequence of image frames. The apparatus also includes means for selectively generating a virtual reference frame based on the composition support data. The apparatus also includes means for generating a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.

[0011] According to another embodiment of the present disclosure, a device includes one or more processors configured to obtain a bitstream corresponding to an encoded version of an image frame. The one or more processors are further configured to generate a virtual reference frame based on synthesis support data included in the bitstream, based on determining that the bitstream includes a virtual reference frame usage indicator. The one or more processors are further configured to generate a decoded version of the image frame based on the virtual reference frame.

[0012] According to another embodiment of the present disclosure, a method includes obtaining, at a device, a bitstream corresponding to an encoded version of an image frame. The method also includes generating, based on determining that the bitstream includes a virtual reference frame usage indicator, a virtual reference frame based on synthesis support data included in the bitstream. The method also includes generating, at the device, a decoded version of the image frame based on the virtual reference frame.

[0013] According to another specific implementation of the present disclosure, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to obtain a bitstream corresponding to an encoded version of an image frame. The instructions, when executed by the one or more processors, further cause the one or more processors to generate a virtual reference frame based on synthesis support data included in the bitstream, based on determining that the bitstream includes a virtual reference frame usage indicator. The instructions, when executed by the one or more processors, further cause the one or more processors to generate a decoded version of the image frame based on the virtual reference frame.

[0014] According to another embodiment of the present disclosure, an apparatus includes means for obtaining a bitstream corresponding to an encoded version of an image frame. The apparatus also includes means for generating a virtual reference frame based on synthesis support data included in the bitstream, the virtual reference frame generated based on determining that the bitstream includes a virtual reference frame usage indicator. The apparatus also includes means for generating a decoded version of the image frame based on the virtual reference frame.

[0015] Other aspects, advantages, and features of the present disclosure will become apparent upon review of the entire application, including the Brief Description of the Drawings, Detailed Description, and Claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a block diagram of certain illustrative aspects of a system operable to generate a virtual reference frame for image encoding according to some examples of the present disclosure.

[0017] Figure 2 is operable to generate a virtual reference frame for image decoding according to some examples of the present disclosure Figure 1 Schematic diagram of the system.

[0018] Figure 3 According to some examples of the present disclosure Figure 1 FIG2 is a diagram of illustrative aspects of operations associated with a frame analyzer and a virtual reference frame generator.

[0019] Figure 4 According to some examples of the present disclosure Figure 1 A diagram of illustrative aspects of operations associated with a synthesis support analyzer of a frame.

[0020] Figure 5 According to some examples of the present disclosure Figure 1 A diagram of illustrative aspects of operations associated with a virtual reference frame generator.

[0021] Figure 6 According to some examples of the present disclosure Figure 1Schematic diagram of illustrative aspects of the operation of a facial virtual reference frame generator and associated operations of a video encoder.

[0022] Figure 7 According to some examples of the present disclosure Figure 1 FIG. 1 is a diagram of illustrative aspects of the motion of a virtual reference frame generator and associated operations of a video encoder.

[0023] Figure 8 According to some examples of the present disclosure Figure 2 A diagram of illustrative aspects of operations associated with a virtual reference frame generator.

[0024] Figure 9 According to some examples of the present disclosure Figure 2 Schematic diagram of illustrative aspects of the operations of a facial virtual reference frame generator and associated video decoder.

[0025] Figure 10 According to some examples of the present disclosure Figure 2 FIG. 1 is a diagram of illustrative aspects of the motion of a virtual reference frame generator and associated operations of a video decoder.

[0026] Figure 11 According to some examples of the present disclosure Figure 1 Schematic diagram of illustrative aspects of the operation of a frame analyzer, a virtual reference frame generator, and a video encoder.

[0027] Figure 12 According to some examples of the present disclosure Figure 2 Schematic diagram of illustrative aspects of the operation of a virtual reference frame generator and a video decoder.

[0028] Figure 13 Examples of an integrated circuit operable to generate a virtual reference frame for image encoding, image decoding, or both, according to some examples of the present disclosure are illustrated.

[0029] Figure 14 is a diagram of a mobile device operable to generate a virtual reference frame for image encoding, image decoding, or both, according to some examples of the present disclosure.

[0030] Figure 15 is a diagram of a wearable electronic device operable to generate a virtual reference frame for image encoding, image decoding, or both, according to some examples of the present disclosure.

[0031] Figure 16 is a diagram of a camera operable to generate a virtual reference frame for image encoding, image decoding, or both, according to some examples of the present disclosure.

[0032] Figure 17 is a diagram of a head-mounted device (such as a virtual reality, mixed reality, or augmented reality head-mounted device) that is capable of generating a virtual reference frame for image encoding, image decoding, or both according to some examples of the present disclosure.

[0033] Figure 18 is a diagram of a first example of a vehicle operable to generate a virtual reference frame for image encoding, image decoding, or both, according to some examples of the present disclosure.

[0034] Figure 19 is a diagram of a second example of a vehicle operable to generate a virtual reference frame for image encoding, image decoding, or both, according to some examples of the present disclosure.

[0035] Figure 20 According to some examples of the present disclosure, Figure 1 FIG2 is a diagram illustrating a specific implementation of a method for generating a virtual reference frame for image encoding performed by a device of FIG.

[0036] Figure 21 According to some examples of the present disclosure, Figure 2 FIG2 is a diagram illustrating a specific implementation of a method for generating a virtual reference frame for image decoding performed by a device.

[0037] Figure 22 is a block diagram of a specific illustrative example of a device operable to generate a virtual reference frame for image encoding, image decoding, or both, according to some examples of the present disclosure. DETAILED DESCRIPTION

[0038] Typically, video decoding involves using a previously decoded image frame as a reference frame for predicting the decoded image frame. In an example, the sequence of image frames includes a first image frame and a second image frame. An encoder encodes the first image frame to generate first coded bits. For example, the encoder uses intra-frame compression to generate the first coded bits.

[0039] The encoder encodes the second image frame to generate second coded bits. For example, the encoder uses a local decoder to decode the first coded bits to generate a first decoded image frame, and encodes the second image frame using the first decoded image frame as a reference frame. For illustration, the encoder determines first residual data based on a difference between the first decoded image frame and the second image frame. The encoder generates second coded bits based on the first residual data. The first coded bits and the second coded bits are transmitted from a first device including the encoder to a second device including the decoder.

[0040] The decoder decodes the first coded bits to generate a first decoded image frame. For example, the decoder performs intra-frame prediction on the first coded bits to generate the first decoded image frame. The decoder decodes the second coded bits to generate residual data for a second decoded image frame. In response to determining that the first decoded image frame is a reference frame for the second decoded image frame, the decoder generates a second decoded image frame based on a combination of the residual data and the first decoded image frame.

[0041] At low bit rate settings (e.g., used during video conferencing), the presence of compression artifacts can degrade video quality. For example, a first compression artifact associated with intra-frame compression may be present in a first decoded image frame. As another example, a second compression artifact associated with decoded residual bits may be present in a second decoded image frame.

[0042] The present invention discloses a system and method for generating a virtual reference frame for image encoding and decoding. In one example, an encoder determines synthetic support data for a second image frame and generates a virtual reference frame for the second image frame based on the synthetic support data. In some implementations, the synthetic support data may include facial landmark data indicating the positions of facial features in the second image frame. In some implementations, the synthetic support data may include motion-based data indicating global motion (e.g., camera movement) detected in the second image frame relative to the first image frame (or the first decoded image frame generated by a local decoder).

[0043] The encoder generates a virtual reference frame based on applying the synthesis support data to the first image frame (or the first decoded image frame). The encoder generates second residual data based on a difference between the virtual reference frame and the second image frame. The encoder generates second coded bits based on the second residual data. The first coded bits, the second coded bits, the synthesis support data, and a virtual reference frame usage indicator are transmitted from the first device to the second device. The virtual reference frame usage indicator indicates virtual reference frame usage.

[0044] The decoder decodes the first coded bits to generate a first decoded image frame. For example, the decoder performs intra-frame prediction on the first coded bits to generate the first decoded image frame. The decoder decodes the second coded bits to generate second residual data. In response to determining that the virtual reference frame usage indicator indicates virtual reference frame usage, the decoder applies synthesis support data to the first decoded image frame to generate the virtual reference frame. In one example, the synthesis support data includes facial landmark data indicating positions of facial features in the second image frame. Applying the facial landmark data to the first decoded image frame includes adjusting the positions of the facial features to more closely match the positions of the facial features indicated in the second image frame. In another example, the synthesis support data includes motion-based data indicating global motion detected in the second image frame relative to the first image frame. Applying the motion-based data to the first decoded image frame includes applying the global motion to the first decoded image frame to generate the virtual reference frame. The decoder applies the second residual data to the virtual reference frame to generate the second decoded image frame.

[0045] Using a virtual reference frame can improve video quality by preserving perceptually important features (e.g., facial landmarks) in the second decoded image frame. In some examples, the encoded version of the synthesis support data and the second residual data (e.g., corresponding to the difference between the virtual reference frame and the second image frame) uses fewer bits than the encoded version of the first residual data (e.g., corresponding to the difference between the first decoded image frame and the second image frame). To illustrate, the second residual data can have smaller magnitudes and smaller overall variance than the first residual data, and thus can be encoded more efficiently (e.g., using fewer bits). In these examples, the virtual reference frame approach can reduce bandwidth usage, improve video quality, or both.

[0046] The following describes certain aspects of the present disclosure with reference to the accompanying drawings. In this description, common features are designated by common reference numerals. As used herein, various terms are used only for the purpose of describing a particular implementation and are not intended to limit the implementation. For example, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. In addition, some features described herein are singular in some implementations and plural in other implementations. For the purpose of illustration, Figure 1 Depicts a system comprising one or more processors ( Figure 1 190), indicating that in some implementations, the device 102 includes a single processor 190, while in other implementations, the device 102 includes multiple processors 190. For ease of reference herein, such features are generally introduced as "one or more" features and are subsequently referred to in the singular unless aspects are being described that relate to multiple ones of the features.

[0047] In some of the figures, multiple instances of a particular type of feature are used. Although the features are physically and / or logically different, the same reference numeral is used for each feature, and the different instances are distinguished by adding a letter to the reference numeral. When features as a group or type are referred to herein (e.g., when no reference is made to a particular one of the features), the reference numeral is used without a distinguishing letter. However, when a particular feature among multiple features of the same type is referred to herein, the reference numeral is used with a distinguishing letter. For example, with reference to Figure 1 , multiple image frames are illustrated and are associated with reference numerals 116A and 116N. When referring to a specific one of these images, such as image frame 116A, the distinguishing letter "A" is used. However, when referring to any one of these image frames or referring to these image frames as a group, the reference numeral 116 is used without a distinguishing letter.

[0048] As used herein, the term "comprise" may be used interchangeably with "include". Additionally, the term "wherein" may be used interchangeably with "wherein". As used herein, "exemplary" indicates an example, specific implementation, and / or aspect and should not be interpreted as limiting or indicating a preference or preferred specific implementation. As used herein, ordinal terms (e.g., "first," "second," "third," etc.) used to modify an element (such as a structure, component, operation, etc.) do not themselves indicate any priority or order of the element relative to another element, but simply distinguish the element from another element with the same name (but using an ordinal term). As used herein, the term "set" refers to one or more specific elements of a particular element, while the term "number" refers to a plurality (e.g., two or more) of specific elements.

[0049] As used herein, "coupling" may include "communicatively coupled," "electrically coupled," or "physically coupled," and may also (or alternatively) include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., a wired network, a wireless network, or a combination thereof), etc. As illustrative, non-limiting examples, two electrically coupled devices (or components) may be included in the same device or in different devices, and may be connected via electronics, one or more connectors, or inductive coupling. In some implementations, two devices (or components) that are communicatively coupled (such as electrically connected) may transmit and receive signals (e.g., digital signals or analog signals) directly or indirectly via one or more wires, buses, networks, etc. As used herein, "direct coupling" may include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without an intermediate component.

[0050] In this disclosure, terms such as "determine," "calculate," "estimate," "shift," "adjust," etc. may be used to describe how to perform one or more operations. It should be noted that such terms should not be interpreted as limiting, and other techniques may be utilized to perform similar operations. Additionally, as referred to herein, "generate," "calculate," "estimate," "use," "select," "access," and "determine" may be used interchangeably. For example, "generating," "calculating," "estimating," or "determining" a parameter (or signal) may refer to actively generating, estimating, calculating, or determining the parameter (or signal), or may refer to using, selecting, or accessing a parameter (or signal) that has already been generated (such as by another component or device).

[0051] refer to Figure 1 , shows certain illustrative aspects of a system 100 configured to generate a virtual reference frame for image encoding and decoding. The system 100 includes a device 102 configured to be coupled to a camera 110, a device 160, or both.

[0052] Device 102 includes input interface 114, one or more processors 190, and modem 170. Input interface 114 is coupled to one or more processors 190 and is configured to couple to camera 110. Input interface 114 is configured to receive camera output 112 from camera 110 and provide camera output 112 to one or more processors 190 as image frames 116.

[0053] One or more processors 190 are coupled to the modem 170 and include a video analyzer 140. The video analyzer 140 includes a frame analyzer 142 coupled to a video encoder 146 via a virtual reference frame (VRF) generator 144. The video encoder 146 is coupled to the modem 170.

[0054] The video analyzer 140 is configured to obtain a sequence of image frames 116, such as image frame 116A, image frame 116N, one or more additional image frames, or a combination thereof. In some implementations, the sequence of image frames 116 may include one or more image frames before image frame 116A, one or more image frames between image frame 116A and image frame 116N, one or more image frames after image frame 116N, or a combination thereof.

[0055] Each of image frames 116 is associated with a frame identifier (ID) 126. For example, image frame 116A has frame identifier 126A, image frame 116N has frame identifier 126N, and so on. In some implementations, frame identifiers 126 indicate the order of image frames 116 in the sequence. In an example, frame identifier 126A having a first value that is less than a second value of frame identifier 126N indicates that image frame 116A precedes image frame 116N in the sequence.

[0056] Video analyzer 140 is configured to selectively generate one or more virtual reference frames (VRFs) for particular ones of image frames 116. Frame analyzer 142 is configured to generate synthetic support data 150N for image frame 116N in response to determining that at least one VRF 156 associated with image frame 116N is to be generated. Synthetic support data 150N may include facial landmark data, motion-based data, or both. For example, frame analyzer 142 is configured to generate facial landmark data as synthetic support data 150N in response to detecting a face in image frame 116N. The facial landmark data indicates the locations of facial features detected in image frame 116N. As another example, frame analyzer 142 is configured to include motion-based data in synthetic support data 150N in response to determining that the motion-based data indicates that global motion in image frame 116N relative to image frame 116A (e.g., a previous image frame in the sequence) is greater than a global motion threshold.

[0057] In an example, frame analyzer 142 is configured to generate a virtual reference frame (VRF) usage indicator 186N having a first value (e.g., 0) in response to determining that a VRF is not to be generated for image frame 116N. For example, frame analyzer 142 is configured to determine that a VRF is not to be generated for image frame 116N in response to determining that no face is detected in image frame 116N and that global motion less than or equal to a global motion threshold is detected in image frame 116N. Alternatively, frame analyzer 142 is configured to generate a VRF usage indicator 186N having a second value (e.g., 1), a third value (e.g., 2), or a fourth value (e.g., 3) in response to determining that at least one VRF 156N is to be generated for image frame 116N. For example, the VRF usage indicator 186N has a second value (e.g., 1) to indicate that the synthesized support data 150N includes facial landmark data, a third value (e.g., 2) to indicate that the synthesized support data 150N includes motion-based data, or a fourth value (e.g., 3) to indicate that the synthesized support data 150N includes both facial landmark data and motion-based data.

[0058] VRF generator 144 is configured to generate one or more VRFs 156N based on synthesis support data 150N in response to determining that VRF usage indicator 186N has a value (e.g., 1, 2, or 3) indicating VRF usage for image frame 116N. Reference list 176 associated with image frame 116 indicates reference frame candidates for image frame 116. In an example, VRF generator 144 is configured to generate reference list 176N associated with image frame 116N, the reference list indicating one or more VRFs 156N. Video encoder 146 is configured to encode image frame 116N based on the reference frame candidates indicated by reference list 176N to generate coded bits 166N.

[0059] Modem 170 is coupled to one or more processors 190 and configured to enable communications with device 160, such as transmitting bitstream 135 to device 160 via wireless transmission. For example, bitstream 135 includes reference list 176N, coded bits 166N, synthesis support data 150N, VRF usage indicator 186N, or a combination thereof.

[0060] In some implementations, the device 102 corresponds to or is included in one of various types of devices. In the illustrative example, the one or more processors 190 are integrated into at least one of the following: Figure 14 The mobile phone or tablet computer device described, as referenced Figure 15 The wearable electronic device described in reference Figure 16 The camera device described or as referenced Figure 17In another illustrative example, one or more processors 190 are integrated into a vehicle, such as a vehicle in which a user is to use a virtual reality, mixed reality, or augmented reality head mounted device. Figure 18 and Figure 19 as further described.

[0061] During operation, the video analyzer 140 obtains a sequence of image frames 116. In a particular example, the input interface 114 receives the camera output 112 from the camera 110 and provides the camera output 112 as the image frames 116 to the video analyzer 140. In another example, the video analyzer 140 obtains the image frames 116 from a storage device, a network device, another component of the device 102, or a combination thereof.

[0062] Video analyzer 140 selectively generates a VRF for image frame 116. In an example, frame analyzer 142 generates synthesis support data 150N, VRF usage indicator 186N, or both based on determining whether at least one VRF is to be generated for image frame 116N, as described with reference to FIG. Figure 3 and Figure 4 Further described. For example, in response to determining that a VRF is not to be generated for image frame 116N, frame analyzer 142 generates a VRF usage indicator 186N having a first value (e.g., 0) indicating no VRF usage. Alternatively, in response to determining that at least the face of person 180 is detected in image frame 116N, frame analyzer 142 adds facial landmark data to synthesis support data 150N and generates a VRF usage indicator 186N having a second value (e.g., 1) indicating facial VRF usage. The facial landmark data indicates the location of facial features of person 180 detected in image frame 116N. According to some aspects, the facial features include at least one of the eyes, eyelids, eyebrows, nose, lips, or facial contour of person 180.

[0063] In yet another example, the frame analyzer 142 generates motion-based data based on a comparison of the image frame 116N and the image frame 116A (e.g., a previous image frame in the sequence). In some implementations, the motion-based data includes motion sensor data indicating motion of an image capture device (e.g., camera 110) associated with the image frame 116N. In some implementations, the motion-based data indicates global motion detected in the image frame 116N relative to the previous image frame (e.g., image frame 116A).

[0064] In response to determining that the motion-based data indicates global motion greater than a global motion threshold, frame analyzer 142 adds the motion-based data to synthesized support data 150N and generates a VRF usage indicator 186N having a third value (e.g., 2) indicating motion VRF usage. In some examples, in response to determining that the motion-based data and facial landmark data are to be used to generate at least one VRF, frame analyzer 142 generates synthesized support data 150N that includes the facial landmark data and the motion-based data, and generates a VRF usage indicator 186N having a fourth value (e.g., 3) indicating both facial VRF usage and motion VRF usage. Frame analyzer 142 provides VRF usage indicator 186N to VRF generator 144. In examples where VRF usage indicator 186N has a value (e.g., 1, 2, or 3) indicating VRF usage, frame analyzer 142 provides synthesized support data 150N to VRF generator 144. In particular aspects, the composition support data 150N, the VRF usage indicator 186N, or both include a frame identifier 126N to indicate an association with the image frame 116N.

[0065] In response to determining that the VRF usage indicator 186N has a first value (e.g., 0) indicating no VRF usage, the VRF generator 144 provides the VRF usage indicator 186N to the video encoder 146 and refrains from passing the reference list 176N to the video encoder 146. Alternatively, in some implementations, in response to determining that the VRF usage indicator 186N has a first value (e.g., 0) indicating no VRF usage, the VRF generator 144 passes an empty list as the reference list 176N to the video encoder 146.

[0066] Alternatively, in response to determining that the VRF usage indicator 186N has a value indicating VRF usage (e.g., 1, 2, or 3), the VRF generator 144 generates one or more VRFs 156N as one or more VRF reference candidates associated with the image frame 116N. For example, in response to determining that the VRF usage indicator 186N has a value indicating facial VRF usage (e.g., 1 or 3), the VRF generator 144 generates at least VRF 156NA based on facial landmark data included in the synthetic support data 150N, as referenced by Figure 5 and Figure 6 In response to determining that the VRF usage indicator 186N has a value indicating motion VRF usage (e.g., 2 or 3), the VRF generator 144 generates at least VRF 156NB based on the motion-based data included in the synthetic support data 150N, as described in reference to Figure 5 and Figure 7 Further described.

[0067] VRF generator 144 generates a reference list 176N to indicate that one or more VRFs 156N are designated as a first set of reference candidates (e.g., VRF reference candidates) for image frame 116N. In an example, reference list 176N includes frame identifier 126N to indicate association with image frame 116N. Reference list 176N includes one or more VRF reference candidate identifiers 172 for the first set of reference candidates. For example, one or more VRF reference candidate identifiers 172 include one or more VRF identifiers 196N for one or more VRFs 156N. For example, one or more VRF reference candidate identifiers 172 include VRF identifier 196NA for VRF 156NA, VRF identifier 196NB for VRF 156NB, one or more additional VRF identifiers for one or more additional VRFs, or a combination thereof. VRF generator 144 provides one or more VRFs 156N, reference list 176N, VRF usage indicator 186N, or a combination thereof to video encoder 146.

[0068] Video encoder 146 is configured to encode image frame 116N to generate coded bits 166N. In certain aspects, video encoder 146 generates a subset of coded bits 166N based at least in part on a second set of reference candidates (e.g., encoder reference candidates) that is distinct from VRF 156. The second set of reference candidates includes one or more previous image frames or one or more previously decoded image frames. In certain implementations, video encoder 146 uses image frame 116A (or a locally decoded image frame corresponding to image frame 116A) as an intra-coded frame (i-frame). In this implementation, the subset of coded bits 166N is based on a residual corresponding to the difference between image frame 116A (or a locally decoded image frame) and image frame 116N. Video encoder 146 adds frame identifier 126A of image frame 116A (or a locally decoded image frame) to one or more encoder reference candidate identifiers 174 of the second set of reference candidates in reference list 176N.

[0069] The video encoder 146 selectively generates one or more subsets of the coded bits 166N based on one or more VRFs 156N. For example, in response to determining that the VRF usage indicator 186N has a particular value indicating VRF usage (e.g., 1, 2, or 3) and the encoder reference candidate count is less than a threshold reference count, the video encoder 146 generates one or more subsets of the coded bits 166N based on the one or more VRFs 156N. Alternatively, in response to determining that the VRF usage indicator 186N has a particular value indicating no VRF usage (e.g., 0), the encoder reference candidate count is greater than or equal to the threshold reference count, or both, the video encoder 146 refrains from generating any of the coded bits 166N based on the VRF 156.

[0070] In a particular aspect, the video encoder 146 determines an encoder reference candidate count based on a count of one or more encoder reference candidate identifiers 174 included in the reference list 176N. In some aspects, the encoder reference candidate count is based on default data, a configuration setting, a user input, a coding configuration of the video encoder 146, or a combination thereof. In some implementations, a threshold reference count is based on default data, a configuration setting, a user input, a coding configuration of the video encoder 146, or a combination thereof.

[0071] Optionally, in some implementations, the VRF generator 144 selectively generates one or more VRFs 156N based on determining that the encoder reference candidate count is less than a threshold reference count. In certain aspects, the VRF generator 144 determines the encoder reference candidate count based on default data, configuration settings, user input, a coding configuration of the video encoder 146, or a combination thereof. In certain aspects, the VRF generator 144 receives the encoder reference candidate count from the video encoder 146.

[0072] In some implementations, the VRF generator 144 determines the threshold VRF count based on a comparison of the threshold reference count and the encoder reference candidate count (e.g., the difference between the two). In these implementations, the VRF generator 144 generates one or more VRFs 156N such that the count of the one or more VRFs 156N is less than or equal to the threshold VRF count.

[0073] In certain aspects, based at least in part on determining that the VRF usage indicator 186N has a particular value (e.g., 1 or 3) indicating facial VRF usage, the video encoder 146 generates a first subset of coded bits 166N based on the VRF 156NA, as described with reference to FIG. Figure 6 Based at least in part on determining that VRF usage indicator 186N has a particular value (e.g., 2 or 3) indicating motion VRF usage, video encoder 146 generates a second subset of coded bits 166N based on VRF 156NB, as described with reference to FIG. Figure 7 Further described.

[0074] Video encoder 146 provides reference list 176N, coded bits 166N, or both to modem 170. Additionally, frame analyzer 142 provides VRF usage indicator 186N, synthesis support data 150N, or both to modem 170. Modem 170 transmits bitstream 135 to device 160. Bitstream 135 includes coded bits 166N, reference list 176N, VRF usage indicator 186N, synthesis support data 150N, or a combination thereof. For example, VRF usage indicator 186N indicates whether any virtual reference frames are to be used to generate the decoded version of image frame 116N.

[0075] In some aspects, the bitstream 135 includes a supplemental enhancement information (SEI) message indicating synthesis support data 150N. In some aspects, the bitstream 135 includes the SEI message including a VRF usage indicator 186N. In certain aspects, the bitstream 135 corresponds to an encoded version of the image frame 116N based at least in part on one or more VRFs 156N, one or more encoder reference candidates associated with one or more encoder reference candidate identifiers 174, or a combination thereof.

[0076] In some implementations, bitstream 135 includes coded bits 166 associated with number of image frames 116, reference lists 176, VRF usage indicators 186, composition support data 150, or a combination thereof. In a particular implementation, bitstream 135 includes reference lists 176 that include a first reference list associated with image frame 116A, a reference list 176N associated with image frame 116N, one or more additional reference lists associated with one or more additional image frames of the sequence, or a combination thereof. For example, reference list 176 may include a first reference list associated with image frame 116A, a reference list 176N associated with image frame 116N, one or more additional reference lists associated with one or more additional image frames of the sequence, or a combination thereof.

[0077] 176 includes one or more VRF identifiers 196 associated with image frame 116A, one or more VRF identifiers 197 associated with image frame 116N, and one or more VRF identifiers 198 associated with image frame 116A.

[0078] VRF identifier 196N, one or more VRF identifiers 196 associated with one or more additional image frames 116, or a combination thereof. As another example, reference list 176 includes one or more frame identifiers 126 as one or more encoder reference candidate identifiers 174 associated with image frame 116A, one or more frame identifiers 126 as one or more encoder reference candidate identifiers 174 associated with image frame 116N, one or more additional frame identifiers 126 as one or more encoder reference candidate identifiers 174 associated with one or more additional image frames 116, or a combination thereof.

[0079] Thus, the system 100 enables generation of VRFs 156 that preserve perceptually important features (e.g., facial landmarks). Technical advantages of using synthetic supporting data 150N (e.g., facial landmark data, motion-based data, or both) to generate one or more VRFs 156N may include one or more VRFs 156N being a closer approximation of an image frame 116N, thereby improving the video quality of the decoded image frame.

[0080] Although the camera 110 is illustrated as being external to the device 102, in other implementations, the camera 110 can be integrated into the device 102. Although the video analyzer 140 is illustrated as obtaining the image frames 116 from the camera 110, in other implementations, the video analyzer 140 can obtain the image frames 116 from another component of the device 102 (e.g., a graphics processor), another device (e.g., a storage device, a network device, etc.), or a combination thereof. The camera 110 is illustrated as an example of an image capture device, and in some implementations, the video analyzer 140 can obtain the image frames 116 from various types of image capture devices, such as an extended reality (XR) device, a vehicle, the camera 110, a graphics processor, or a combination thereof.

[0081] Although the frame analyzer 142, VRF generator 144, video encoder 146, and modem 170 are illustrated as separate components, in other implementations, two or more of the frame analyzer 142, VRF generator 144, video encoder 146, or modem 170 may be combined into a single component. Although the frame analyzer 142, VRF generator 144, and video encoder 146 are illustrated as being included in a single device (e.g., device 102), in other implementations, one or more operations described herein with reference to the frame analyzer 142, VRF generator 144, or video encoder 146 may be performed at another device. Optionally, in some implementations, the video analyzer 140 may receive the image frames 116, the synthesis support data 150, or both from another device.

[0082] refer to Figure 2 , showing certain illustrative aspects of system 100. System 100 is operable to generate a virtual reference frame for image decoding. Device 160 is configured to be coupled to display device 210, device 102, or both.

[0083] Device 102 includes an output interface 214, one or more processors 290, and a modem 270. Output interface 214 is coupled to one or more processors 290 and is configured to couple to display device 210.

[0084] Modem 270 is coupled to one or more processors 290 and configured to enable communications with device 102, such as receiving bitstream 135 from device 102 via wireless transmission. For example, bitstream 135 includes reference list 176N, coded bits 166N, synthesis support data 150N, VRF usage indicator 186N, or a combination thereof.

[0085] One or more processors 290 are coupled to the modem 270 and include a video generator 240. The video generator 240 includes a bitstream analyzer 242 coupled to a VRF generator 244 and a video decoder 246. The VRF generator 244 is coupled to the video decoder 246. The bitstream analyzer 242 is also coupled to the modem 270.

[0086] The bitstream analyzer 242 is configured to obtain the bitstream 135 corresponding to the bitstream from the modem 270. Figure 1 For example, the bitstream 135 includes coded bits 166N, a VRF usage indicator 186N, a reference list 176N, or a combination thereof. If the bitstream 135 includes the VRF usage indicator 186N having a specific value (e.g., 1, 2, or 3) indicating VRF usage, the bitstream 135 also includes compositing support data 150N.

[0087] The bitstream analyzer 242 is configured to, in response to determining that the bitstream 135 includes a VRF usage indicator 186N having a particular value (e.g., 1, 2, or 3) indicating VRF usage, extract the synthesis support data 150N from the bitstream 135 and provide the synthesis support data 150N to the VRF generator 244. In some implementations, the bitstream analyzer 242 is configured to provide the VRF usage indicator 186N, the reference list 176N, or both to the VRF generator 244. The bitstream analyzer 242 is configured to provide the coded bits 166N, the reference list 176N, or both to the video decoder 246.

[0088] VRF generator 244 is configured to selectively generate one or more VRFs 256N for use in generating a decoded version of image frame 116N. For example, VRF generator 244 is configured to determine whether to use at least one VRF to generate a decoded version of image frame 116N based on synthesis support data 150N, reference list 176N, VRF usage indicator 186N, or a combination thereof associated with image frame 116N. VRF generator 244 is configured to generate one or more VRFs 256N based on synthesis support data 150N in response to determining that at least one VRF is to be used. For example, VRF generator 244 is configured to generate one or more VRFs 256N based on facial landmark data, motion-based data, or both indicated by synthesis support data 150N.

[0089] Video decoder 246 is configured to generate a sequence of image frames 216 corresponding to decoded versions of the sequence of image frames 116. In an example, image frames 216 include image frame 216A, image frame 216N, one or more additional image frames, or a combination thereof. Each of image frames 216 is associated with a frame identifier 126. For example, image frame 216A, which corresponds to a decoded version of image frame 116A, includes frame identifier 126A for image frame 116A. As another example, image frame 216N, which corresponds to a decoded version of image frame 116N, includes frame identifier 126N for image frame 116N.

[0090] The video decoder 246 is configured to selectively generate an image frame 216 based on the corresponding one or more VRFs 256. For example, the video decoder 246 is configured to generate an image frame 216N based on the coded bits 166N, one or more VRFs 256N, the reference list 176N, or a combination thereof. In some implementations, the video generator 240 is configured to provide the image frames 216 to the display device 210 via the output interface 214. In a particular implementation, the video generator 240 is configured to provide the image frames 216 to the display device 210 in the playback order indicated by the frame identifiers 126. For example, during forward playback and based on determining that the frame identifier 126A is less than the frame identifier 126N, the video generator 240 provides the image frame 216A to the display device 210 for playback earlier than the image frame 216N. In this particular example, the person 280 can view the image frame 216 displayed by the display device 210.

[0091] In some implementations, the device 160 corresponds to or is included in one of various types of devices. In the illustrative example, the one or more processors 290 are integrated into at least one of the following: Figure 14 The mobile phone or tablet computer device described, as referenced Figure 15 The wearable electronic device described in reference Figure 16 The camera device described or as referenced Figure 17 In another illustrative example, one or more processors 290 are integrated into a vehicle, such as a vehicle in which a user is to use a virtual reality, mixed reality, or augmented reality head mounted device. Figure 18 and Figure 19 as further described.

[0092] During operation, the video generator 240 obtains Figure 116N, coded bits 166N, VRF usage indicator 186N, reference list 176N, or a combination thereof, associated with image frame 116N. In some examples, bitstream 135 also includes compositing support data 150N associated with image frame 116N. In certain aspects, coded bits 166N, VRF usage indicator 186N, reference list 176N, compositing support data 150N, or a combination thereof, indicate a frame identifier 126N for image frame 116N.

[0093] In a particular example, video generator 240 obtains bitstream 135 via modem 270. In another example, video generator 240 obtains bitstream 135 from a storage device, a network device, another component of device 160, or a combination thereof.

[0094] Video generator 240 selectively generates a VRF for determining a decoded version of image frame 116. In an example, in response to determining that bitstream 135 does not include VRF usage indicator 186N or that VRF usage indicator 186N has a first value (e.g., 0) indicating no VRF usage, bitstream analyzer 242 determines that a VRF will not be used to generate image frame 216N corresponding to the decoded version of image frame 116N. Alternatively, in response to determining that bitstream 135 includes VRF usage indicator 186N having a specific value (e.g., 1, 2, or 3) indicating VRF usage, bitstream analyzer 242 determines that at least one VRF will be used to generate image frame 216N.

[0095] In response to determining that at least one VRF is to be used to generate image frame 216N, bitstream analyzer 242 provides synthesis support data 150N, reference list 176N, VRF usage indicator 186N, or a combination thereof to VRF generator 244 to generate at least one VRF. Bitstream analyzer 242 also provides coded bits 166N, reference list 176N, or both to video decoder 246 to generate image frame 216N. In some examples, bitstream analyzer 242, VRF generator 244, or both provide VRF usage indicator 186N to video decoder 246.

[0096] In response to determining that the bitstream 135 includes a VRF usage indicator 186N having a specific value (e.g., 1, 2, or 3) indicating VRF usage, the VRF generator 244 generates one or more VRFs 256N as one or more VRF reference candidates to be used for generating the image frame 216N. For example, in response to determining that the VRF usage indicator 186N has a specific value (e.g., 1 or 3) indicating facial VRF usage, the VRF generator 244 generates at least VRF 256NA based on facial landmark data included in the synthesis support data 150N, as referenced. Figure 8 and Figure 9 In response to determining that the VRF usage indicator 186N has a particular value (e.g., 2 or 3) indicating motion VRF usage, the VRF generator 244 generates at least VRF 256NB based on the motion-based data included in the synthetic support data 150N, as described in reference to FIG. Figure 8 and Figure 10 Further described.

[0097] As reference Figure 1 As depicted, reference list 176N includes one or more VRF reference candidate identifiers 172. For example, one or more VRF reference candidate identifiers 172 include VRF identifier 196NA of VRF 156NA, VRF identifier 196NB of VRF 156NB, one or more additional VRF identifiers of one or more additional VRFs, or a combination thereof.

[0098] The VRF generator 244 assigns one or more VRF identifiers 196N to one or more VRFs 256N. In a specific example, in response to determining that the facial landmark data is associated with the VRF identifier 196NA, the VRF generator 244 assigns the VRF identifier 196NA to a VRF 256NA generated based on the facial landmark data. Thus, the VRF 256NA corresponds to the VRF 196NA generated in the face. Figure 1 In another example, in response to determining that the motion-based data is associated with the VRF identifier 196NB, the VRF generator 244 assigns the VRF identifier 196NB to the VRF 256NB generated based on the motion-based data. Thus, the VRF 256NB corresponds to the VRF 156NA generated at the video analyzer 140. Figure 1 The VRF generator 244 provides one or more VRFs 256N to the video decoder 246.

[0099] The video decoder 246 is configured to generate an image frame 216N based on at least the coded bits 166N (eg, Figure 1 In certain aspects, the video decoder 246 selectively generates the image frame 216N based on one or more VRFs 256N. Figure 1 As described, the reference list 176N includes one or more VRF reference candidate identifiers 172 for a first set of reference candidates (e.g., one or more VRFs 256N), one or more encoder reference candidate identifiers 174 for a second set of reference candidates (e.g., one or more previously decoded image frames 216), or a combination thereof.

[0100] In a particular example, reference list 176N is empty, and video decoder 246 generates image frame 216N by processing (eg, decoding) coded bits 166N independently of any reference candidates. As an illustrative example, image frame 216N may correspond to an i-frame.

[0101] In a specific example, the video decoder 246 selects one or more of the reference candidates indicated in the reference list 176N to generate the image frame 216N based on a selection criterion. The selection criterion may be based on user input, default data, a configuration setting, a threshold reference count, or a combination thereof. In an example, if the reference list 176N does not indicate any of the first set of reference candidates (e.g., one or more VRFs 256N), the video decoder 246 selects one or more of the second set of reference candidates (e.g., encoder reference candidates). Alternatively, if the reference list 176N indicates at least one of the one or more VRFs 256N, the video decoder 246 generates the image frame 216N based on the one or more VRFs 256N and independently of the encoder reference candidates.

[0102] The video decoder 246 applies the coded bits 166N (e.g., residual) to the selected one of the reference candidates to generate a decoded image frame. For example, the video decoder 246 applies a first subset of the coded bits 166N to the VRF 256NA to generate a first decoded image frame, such as the reference candidate. Figure 9 As another example, the video decoder 246 applies a second subset of the coded bits 166N to the VRF 256NB to generate a second decoded image frame, as described in reference to FIG. Figure 10 In yet another example, the video decoder 246 applies the third subset of the coded bits 166N to the image frame 216A to generate a third decoded image frame.

[0103] In a particular embodiment in which the video decoder 246 selects a single reference candidate from the reference candidates (e.g., VRF 256NA, VRF 256NB, or image frame 216A), the corresponding decoded image frame (e.g., the first decoded image frame, the second decoded image frame, or the third decoded image frame) is designated as image frame 216N.

[0104] In certain implementations where the video decoder 246 selects multiple reference candidates (e.g., VRF 256NA, VRF 256NB, and image frame 216A), the video decoder 246 generates the image frame 216N based on a combination of corresponding decoded image frames (e.g., the first decoded image frame, the second decoded image frame, and the third decoded image frame). For example, the video decoder 246 generates the image frame 216N by averaging the decoded image frames (e.g., the first decoded image frame, the second decoded image frame, and the third decoded image frame) on a pixel-by-pixel basis or using information in the bitstream 135 that indicates how to combine the decoded image frames (e.g., weights of a weighted sum of the decoded image frames).

[0105] In the illustrative example, video generator 240 provides image frame 216N to display device 210 via output interface 214. Alternatively, in some implementations, video generator 240 provides image frame 216N to a storage device, a network device, a user device, or a combination thereof.

[0106] Thus, system 200 enables the generation of decoded image frames (e.g., image frame 216N) using VRFs 256 that preserve perceptually important features (e.g., facial landmarks). Technical advantages of using synthetic support data 150N (e.g., facial landmark data, motion-based data, or both) to generate one or more VRFs 256N may include one or more VRFs 256N being a closer approximation of image frame 116N (compared to image frame 216A), thereby improving the video quality of image frame 216N.

[0107] Although display device 210 is illustrated as being external to device 160, in other implementations, display device 210 may be integrated into device 160. Although video generator 240 is illustrated as receiving bitstream 135 from device 160 via modem 270, in other implementations, video generator 240 may obtain bitstream 135 from another component of device 102 (e.g., a graphics processor), another device (e.g., a storage device, a network device, etc.), or a combination thereof. In particular implementations, device 102, device 160, or both may include a replica of video analyzer 140 and a replica of video generator 240. For example, the video analyzer 140 of the device 102 generates the bitstream 135 based on the image frame 116 received from the camera 110, the video analyzer 140 stores the bitstream 135 in a memory, the video generator 240 of the device 102 retrieves the bitstream 135 from the memory, the video generator 240 generates the image frame 216 based on the bitstream 135, and the video generator 240 provides the image frame 216 to the display device.

[0108] Although the bitstream analyzer 242, the VRF generator 244, the video decoder 246, and the modem 270 are illustrated as separate components, in other implementations, two or more of the bitstream analyzer 242, the VRF generator 244, the video decoder 246, or the modem 270 may be combined into a single component. Although the bitstream analyzer 242, the VRF generator 244, and the video decoder 246 are illustrated as being included in a single device (e.g., the device 160), in other implementations, one or more operations described herein with reference to the bitstream analyzer 242, the VRF generator 244, or the video decoder 246 may be performed at another device.

[0109] refer to Figure 3 , a diagram 300 showing illustrative aspects of operations associated with the frame analyzer 142 and the VRF generator 144 according to some examples of the present disclosure. The frame analyzer 142 includes a visual analysis engine 312 coupled to a composition support analyzer 314.

[0110] Visual analysis engine 312 includes face detector 302, facial landmark detector 304, and global motion detector 306. Face detector 302 uses facial recognition technology to generate a face detection indicator 318N that indicates whether at least one face is detected in image frame 116N. For example, face detection indicator 318N has a first value (e.g., 0) to indicate that no face is detected in image frame 116N or a second value (e.g., 1) to indicate that at least one face is detected in image frame 116N.

[0111] In response to determining that the face detection indicator 318N indicates that at least one face is detected in the image frame 116N, the facial landmark detector 304 uses facial analysis techniques to generate facial landmark data 320N indicating the locations of facial features detected in the image frame 116N and includes the facial landmark data 320N in the synthesized support data 150N, as described with reference to FIG. Figure 6 Further described.

[0112] Global motion detector 306 uses global motion detection techniques to generate motion detection indicator 316N that indicates whether at least a threshold global motion is detected in image frame 116N relative to image frame 116 A. For example, motion detection indicator 316N has a first value (e.g., 0) indicating that at least a threshold global motion is not detected in image frame 116N or a second value (e.g., 1) indicating that at least a threshold global motion is detected in image frame 116N.

[0113] The global motion detector 306 uses motion analysis techniques to generate motion-based data 322N indicating global motion detected in the image frame 116N and, in response to determining that the motion detection indicator 316N indicates that at least a threshold global motion is detected in the image frame 116N, includes the motion-based data 322N in the synthesized support data 150N, as described with reference to FIG. Figure 7 As further described. In certain implementations, global motion detector 306 generates motion-based data 322N (e.g., a global motion vector) based on a comparison of image frame 116A and image frame 116N. In some implementations, global motion detector 306 also or alternatively receives sensor data indicating a first position of camera 110 at a first capture time of image frame 116A and a second position of camera 110 at a second capture time of image frame 116N. Global motion detector 306 determines global motion based on a comparison of the first position and the second position (e.g., a difference between the first position and the second position). In response to determining that the global motion is greater than a threshold global motion, global motion detector 306 generates motion-based data 322N indicating the difference between the second position and the first position. Visual analysis engine 312 provides motion detection indicator 316N and face detection indicator 318N to composition support analyzer 314.

[0114] The composition support analyzer 314 generates the VRF usage indicator 186N based on the motion detection indicator 316N, the face detection indicator 318N, or both. For example, the VRF usage indicator 186N has a first value (e.g., 0) indicating no VRF usage, which corresponds to the first value (e.g., 0) of the motion detection indicator 316N and the first value (e.g., 0) of the face detection indicator 318N. In another example, the VRF usage indicator 186N has a second value (e.g., 1) indicating facial VRF usage and no motion VRF usage, which corresponds to the first value (e.g., 0) of the motion detection indicator 316N and the second value (e.g., 1) of the face detection indicator 318N. The VRF usage indicator 186N has a third value (e.g., 2) indicating motion VRF usage and no facial VRF usage, which corresponds to the second value (e.g., 1) of the motion detection indicator 316N and the first value (e.g., 0) of the face detection indicator 318N. The VRF usage indicator 186N has a fourth value (e.g., 3) indicating motion VRF usage and face VRF usage, which corresponds to the second value (e.g., 1) of the motion detection indicator 316N and the second value (e.g., 1) of the face detection indicator 318N. In a specific implementation, each of the motion detection indicator 316N and the face detection indicator 318N is a single-bit value, and the VRF usage indicator 186N is a two-bit value corresponding to the concatenation of the motion detection indicator 316N and the face detection indicator 318N.

[0115] The frame analyzer 142 provides the VRF usage indicator 186N to the VRF generator 144. When the VRF usage indicator 186N has a specific value (e.g., 1, 2, or 3) indicating VRF usage, the frame analyzer 142 also provides the synthesized support data 150N to the VRF generator 144. In response to determining that the VRF usage indicator 186N has a specific value (e.g., 1 or 3) indicating that the synthesized support data 150N includes facial landmark data 320N, the VRF generator 144 generates a VRF 156NA based on the facial landmark data 320N, as described with reference to FIG. Figure 6 As further described. VRF generator 144 generates a VRF identifier 196NA for VRF 156NA and adds VRF identifier 196NA to one or more VRF reference candidate identifiers 172 of reference list 176N, as referenced Figure 1 described.

[0116] In response to determining that the VRF usage indicator 186N has a particular value (e.g., 2 or 3) indicating that the synthetic support data 150N includes motion-based data 322N, the VRF generator 144 generates a VRF 156NB based on the motion-based data 322N, as shown in FIG. Figure 7 As further described. VRF generator 144 generates a VRF identifier 196NB for VRF 156NB and adds VRF identifier 196NB to one or more VRF reference candidate identifiers 172 of reference list 176N, as referenced Figure 1 described.

[0117] The visual analysis engine 312 including both the facial landmark detector 304 and the global motion detector 306 is provided as an exemplary implementation. Alternatively, in some implementations, the visual analysis engine 312 may include a single one of the facial landmark detector 304 or the global motion detector 306, and the synthesized support data 150N may include a corresponding one of the facial landmark data 320N or the motion-based data 322N. Technical advantages of the visual analysis engine 312 including a single one of the facial landmark detector 304 or the global motion detector 306 may include less hardware used by the visual analysis engine 312, lower memory usage, fewer computation cycles, or a combination thereof. Technical advantages of the visual analysis engine 312 including both the facial landmark detector 304 and the global motion detector 306 as compared to including a single one of the facial landmark detector 304 or the global motion detector 306 may include enhanced image frame reproduction quality, reduced transmission resource usage, or both. Another technical advantage of the visual analytics engine 312 including both the facial landmark detector 304 and the global motion detector 306 may include compatibility with decoders that include support for face VRF, motion VRF, or both.

[0118] refer to Figure 4 , showing a method for generating a composite support analyzer 314 associated with the composite support analyzer 314 according to some examples of the present disclosure Figure 1 4. In a particular aspect, the composition support analyzer 314 initializes the VRF usage indicator 186N to a first value (eg, 0) indicating no VRF usage.

[0119] At 402, the synthesis support analyzer 314 determines Figure 1 The one or more encoder reference candidate identifiers 174 indicate whether the encoder reference candidate count is less than a threshold reference count.

[0120] In response to determining at 402 that the encoder reference candidate count is not less than (i.e., greater than or equal to) the threshold reference count, the synthesis support analyzer 314 outputs at 404 a VRF having a first value (e.g., 0) indicating no VRF usage. Figure 1 Alternatively, in response to determining at 402 that the count of encoder reference candidates is less than the threshold reference count, the synthesis support analyzer 314 determines at 406 that Figure 3 Whether the face detection indicator 318N indicates that at least one face is detected in the image frame 116N.

[0121] In response to determining that face detection indicator 318N indicates that at least one face is detected in image frame 116N, composition support analyzer 314 updates VRF usage indicator 186N to a second value (e.g., 1) to indicate facial VRF usage at 408. At 410, composition support analyzer 314 determines whether the sum of the encoder reference candidate count and one is less than a threshold reference count.

[0122] In response to determining at 406 that the face detection indicator 318N indicates that no face was detected in the image frame 116N, or determining at 410 that the sum of the encoder reference candidate count and one is less than the threshold reference count, the composition support analyzer 314 determines at 412 Figure 3 Whether the motion detection indicator 316N indicates that global motion greater than a threshold is detected in the image frame 116N.

[0123] In response to determining at 412 that motion detection indicator 316N indicates that global motion greater than a threshold is detected in image frame 116N, composition support analyzer 314 updates VRF usage indicator 186N to indicate motion VRF usage. For example, in response to determining that VRF usage indicator 186N has a first value (e.g., 0) indicating no facial VRF usage, composition support analyzer 314 sets VRF usage indicator 186N to a third value (e.g., 2) indicating both motion VRF usage and no facial VRF usage. As another example, in response to determining that VRF usage indicator 186N indicates a second value (e.g., 1) indicating facial VRF usage, composition support analyzer 314 sets VRF usage indicator 186N to a fourth value (e.g., 3) indicating motion VRF usage in addition to facial VRF usage.

[0124] Alternatively, in response to determining at 410 that the sum of the encoder reference candidate count and one is greater than or equal to the threshold reference count, or determining at 412 that the motion detection indicator 316N indicates that no global motion greater than the threshold is detected in the image frame 116N, the synthesis support analyzer 314 outputs the VRF usage indicator 186N indicating no motion VRF usage. For example, the synthesis support analyzer 314 refrains from updating the VRF usage indicator 186N having a first value (e.g., 0) indicating no VRF usage or having a second value (e.g., 1) indicating both face VRF usage and no motion VRF usage.

[0125] Diagram 400 is an illustrative example of operations performed by the synthesis support analyzer 314. Optionally, in some implementations, the synthesis support analyzer 314 may generate the VRF usage indicator 186N based on a single one of the motion detection indicator 316N or the face detection indicator 318N. Alternatively, in some implementations where the VRF usage indicator 186N is based on the face detection indicator 318N and not on the motion detection indicator 316N, the synthesis support analyzer 314 performs operations 402, 404, 406, and 408 and does not perform operations 410, 412, 414, and 416. For illustration, in response to determining at 402 that the encoder reference candidate count is less than the threshold reference count and determining at 406 that the face detection indicator 318N indicates that at least one face is detected in the image frame 116N, the synthesis support analyzer 314 outputs at 408 the VRF usage indicator 186N having a second value (e.g., 1) indicating face VRF usage. Alternatively, in response to determining at 402 that the encoder reference candidate count is greater than or equal to the threshold reference count, or determining at 406 that the face detection indicator 318N indicates that no face is detected in the image frame 116N, the synthesis support analyzer 314 proceeds to 404 and outputs a VRF usage indicator 186N having a first value (e.g., 0) indicating no VRF usage.

[0126] Alternatively, in some implementations in which the VRF usage indicator 186N is based on the motion detection indicator 316N and not on the face detection indicator 318N, the composition support analyzer 314 performs operations 402, 404, 412, and 414 and does not perform operations 406, 408, 410, and 416. To illustrate, in response to determining at 402 that the encoder reference candidate count is less than the threshold reference count and determining at 412 that the motion detection indicator 316N indicates that at least a threshold global motion is detected in the image frame 116N, the composition support analyzer 314 outputs at 414 the VRF usage indicator 186N having a third value (e.g., 2) indicating motion VRF usage. Alternatively, in response to determining at 402 that the encoder reference candidate count is greater than or equal to the threshold reference count, or determining at 412 that the motion detection indicator 316N indicates that no global motion greater than the threshold is detected in the image frame 116N, the synthesis support analyzer 314 proceeds to 404 and outputs a VRF usage indicator 186N having a first value (e.g., 0) indicating no VRF usage.

[0127] refer to Figure 5 , a diagram 500 showing illustrative aspects of operations associated with the VRF generator 144 according to some examples of the present disclosure. The VRF generator 144 includes a facial VRF generator 504 and a motion VRF generator 506.

[0128] In response to determining that the VRF usage indicator 186N has a particular value (e.g., 1 or 3) indicating facial VRF usage, the facial VRF generator 504 processes the image frame 116A (or a locally decoded version of the image frame 116A) based on the facial landmark data 320N to generate the VRF 156NA, as described with reference to FIG. Figure 6 As further described, face VRF generator 504 assigns VRF identifier 196NA to VRF 156NA and adds VRF identifier 196NA to one or more VRF reference candidate identifiers 172 in reference list 176N.

[0129] In response to determining that the VRF usage indicator 186N has a particular value (e.g., 2 or 3) indicating motion VRF usage, the motion VRF generator 506 processes the image frame 116A (or a locally decoded version of the image frame 116A) based on the motion-based data 322N to generate the VRF 156NB, as described with reference to FIG. Figure 7 As further described, motion VRF generator 506 assigns VRF identifier 196NB to VRF 156NB and adds VRF identifier 196NB to one or more VRF reference candidate identifiers 172 in reference list 176N.

[0130] The VRF generator 144 including both the face VRF generator 504 and the motion VRF generator 506 is provided as an illustrative example. Alternatively, in some implementations, the VRF generator 144 may include a single one of the face VRF generator 504 or the motion VRF generator 506. Technical advantages of including a single one of the face VRF generator 504 or the motion VRF generator 506 may include less hardware used by the VRF generator 144, lower memory usage, fewer computation cycles, or a combination thereof. Technical advantages of the VRF generator 144 including both the face VRF generator 504 and the motion VRF generator 506 compared to including a single one of the facial landmark detector 304 or the global motion detector 306 may include enhanced image frame reproduction quality, reduced transmission resource usage, or both. Another technical advantage of the visual analysis engine 312 including both the facial landmark detector 304 and the global motion detector 306 may include compatibility with decoders that include support for the face VRF, the motion VRF, or both.

[0131] refer to Figure 6 , diagram 600 showing illustrative aspects of operations associated with the facial VRF generator 504 and the video encoder 146, according to some examples of the present disclosure.

[0132] In response to determining that the VRF usage indicator 186N has a particular value (e.g., 1 or 3) indicating facial VRF usage, the facial VRF generator 504 applies the facial landmark data 320N to the image frame 116A (or a locally decoded version of the image frame 116A). For example, the facial landmark data 320N indicates the location of facial features in the image frame 116N. Figure 6 A graphical representation of facial landmark data 320N illustrating the positioning of facial features detected in image frame 116N is shown in FIG. For illustration, a person's eyes may be depicted as being more open in image frame 116N relative to the depiction of eyes in image frame 116A.

[0133] Applying facial landmark data 320N to image frame 116A (or a locally decoded version of image frame 116A) adjusts the positioning of facial features in image frame 116A (or a locally decoded version of image frame 116A) to generate VRF 156NA as an estimate of image frame 116N. For illustration, the adjusted positioning of facial features in VRF 156NA may more closely match the positioning (or relative positioning) of facial features in image frame 116N. In a specific implementation, facial VRF generator 504 generates a facial model corresponding to the positioning of facial features detected in image frame 116A. Facial VRF generator 504 updates the facial model based on the updated positioning of facial features indicated in facial landmark data 320N. Facial VRF generator 504 generates VRF 156NA corresponding to the updated facial model.

[0134] Facial landmark data 320N indicating the locations of facial features detected in image frame 116N is provided as an illustrative example. Alternatively, in some implementations, facial landmark data 320N indicates locations of facial features detected in image frame 116N that are different from (e.g., updated from) the locations of facial features detected in image frame 116A.

[0135] In certain implementations, the facial VRF generator 504 includes a trained model (eg, a neural network). The facial VRF generator 504 uses the trained model to process the image frame 116A (or a locally decoded version of the image frame 116A) and the facial landmark data 320N to generate the VRF 156NA.

[0136] The face VRF generator 504 provides the VRF 156NA to the video encoder 146. The video encoder 146 determines residual data 604 based on a comparison of the image frame 116N and the VRF 156NA (e.g., the difference between the two). The video encoder 146 generates coded bits 606N corresponding to the residual data 604. For example, the video encoder 146 encodes the residual data 604 to generate the coded bits 606N. The coded bits 606N are included as a coded bit associated with the face VRF usage. Figure 1 6N corresponds to fewer bits than the encoded version of the first residual data based on the difference between image frame 116A (or a locally decoded version of image frame 116A) and image frame 116N. In an example, residual data 604 has a smaller magnitude and a smaller overall variance than the first residual data, and thus, residual data 604 can be encoded more efficiently (e.g., using fewer bits). Technical advantages of providing facial landmark data 320N and residual data 604 (rather than the first residual data) in bitstream 135 can include using fewer resources (e.g., bandwidth, time, or both).

[0137] refer to Figure 7 , diagram 700 showing illustrative aspects of operations associated with the motion VRF generator 506 and the video encoder 146 according to some examples of the present disclosure.

[0138] In response to determining that VRF usage indicator 186N has a particular value (e.g., 2 or 3) indicating motion VRF usage, motion VRF generator 506 applies motion-based data 322N to image frame 116A (or a locally decoded version of image frame 116A). For example, motion-based data 322N indicates global motion (e.g., rotation, translation, or both) detected in image frame 116N relative to image frame 116A (or a locally decoded version of image frame 116A). In another example, motion-based data 322N indicates global motion of the camera moving to the left between a first capture time of image frame 116A and a second capture time of image frame 116N.

[0139] Applying motion-based data 322N to image frame 116A (or a locally decoded version of image frame 116A) applies global motion to image frame 116A (or a locally decoded version of image frame 116A) to generate VRF 156NB as an estimate of image frame 116N. For example, motion VRF generator 506 uses motion-based data 322N to warp image frame 116A (or a locally decoded version of image frame 116A) to generate VRF 156NB. In certain implementations, motion VRF generator 506 includes a trained model (e.g., a neural network). Motion VRF generator 506 uses the trained model to process image frame 116A (or a locally decoded version of image frame 116A) and motion-based data 322N to generate VRF 156NB. For example, image frame 116A (or a locally decoded version of image frame 116A) and motion-based data 322N are provided as inputs to the trained model, and the output of the trained model indicates VRF 156NB.

[0140] The motion VRF generator 506 provides the VRF 156NB to the video encoder 146. The video encoder 146 determines residual data 704 based on a comparison of the image frame 116N and the VRF 156NB (e.g., the difference between the two). The video encoder 146 generates coded bits 706N corresponding to the residual data 704. For example, the video encoder 146 encodes the residual data 704 to generate the coded bits 706N. The coded bits 706N are included as a bit in association with the use of the motion VRF. Figure 1 16A (or a locally decoded version of image frame 116A) and image frame 116N. In certain aspects, motion-based data 322N and coded bits 706N correspond to fewer bits than an encoded version of first residual data based on the difference between image frame 116A (or a locally decoded version of image frame 116A) and image frame 116N. In an example, residual data 704 has a smaller magnitude and a smaller overall variance than the first residual data, and thus, residual data 704 can be encoded more efficiently (e.g., using fewer bits). Technical advantages of providing motion-based data 322N and residual data 704 (rather than the first residual data) in bitstream 135 can include using fewer resources (e.g., bandwidth, time, or both).

[0141] refer to Figure 8 , a diagram 800 showing illustrative aspects of operations associated with the VRF generator 244 according to some examples of the present disclosure. The VRF generator 244 includes a face VRF generator 804 and a motion VRF generator 806.

[0142] In response to determining that the VRF usage indicator 186N has a particular value (e.g., 1 or 3) indicating facial VRF usage, the facial VRF generator 804 processes the image frame 216A based on the facial landmark data 320N to generate the VRF 256NA, as shown in FIG. Figure 9 In response to determining that reference list 176N includes VRF identifier 196NA associated with face VRF usage, facial landmark data 320N is associated with VRF identifier 196NA, or both, face VRF generator 804 assigns VRF identifier 196NA to VRF 256NA.

[0143] In response to determining that the VRF usage indicator 186N has a particular value (e.g., 2 or 3) indicating motion VRF usage, the motion VRF generator 806 processes the image frame 216A based on the motion-based data 322N to generate the VRF 256NB, as shown in FIG. Figure 10 In response to determining that reference list 176N includes VRF identifier 196NB associated with motion VRF usage, motion-based data 322N is associated with VRF identifier 196NB, or both, motion VRF generator 806 assigns VRF identifier 196NB to VRF 256NB.

[0144] The VRF generator 244 including both the face VRF generator 804 and the motion VRF generator 806 is provided as an illustrative example. Alternatively, in some implementations, the VRF generator 244 may include a single one of the face VRF generator 804 or the motion VRF generator 806. Technical advantages of including a single one of the face VRF generator 804 or the motion VRF generator 806 may include less hardware used by the VRF generator 244, lower memory usage, fewer computation cycles, or a combination thereof. Technical advantages of the VRF generator 244 including both the face VRF generator 804 and the motion VRF generator 806, as compared to including a single one of the face VRF generator 804 or the motion VRF generator 806, may include enhanced image frame reproduction quality, reduced transmission resource usage, or both. Another technical advantage of the VRF generator 244 including both the face VRF generator 804 and the motion VRF generator 806 may include compatibility with encoders that include support for the face VRF, the motion VRF, or both.

[0145] refer to Figure 9 , diagram 900 showing illustrative aspects of operations associated with the facial VRF generator 804 and the video decoder 246, according to some examples of the present disclosure.

[0146] In response to determining that the VRF usage indicator 186N has a particular value (eg, 1 or 3) indicating facial VRF usage, the facial VRF generator 804 applies the facial landmark data 320N to the image frame 216A.

[0147] Applying facial landmark data 320N to image frame 216A adjusts the positioning of facial landmarks in image frame 216A to more closely match the positioning (or relative positioning) of facial landmarks in image frame 116N, thereby generating VRF 256NA. In certain aspects, facial VRF generator 804 generates a facial model corresponding to the positioning of facial landmarks detected in image frame 216A. Facial VRF generator 804 updates the facial model based on the updated positioning of facial landmarks indicated in facial landmark data 320N. Facial VRF generator 804 generates VRF 256NA corresponding to the updated facial model.

[0148] In certain implementations, the facial VRF generator 804 includes a trained model (eg, a neural network). The facial VRF generator 804 uses the trained model to process the image frame 216A and the facial landmark data 320N to generate the VRF 256NA.

[0149] The face VRF generator 804 provides the VRF 256NA to the video decoder 246. The video decoder 246 decodes the coded bits 606N (e.g., a first subset of the coded bits 166N associated with the face VRF usage) to generate residual data 604. The face VRF generator 804 generates the image frame 216N based on the combination of the VRF 256NA and the residual data 604. In particular aspects, the facial landmark data 320N and the coded bits 606N correspond to fewer bits than the encoded version of the first residual data based on the difference between the image frame 216A and the image frame 116N. Technical advantages of using the facial landmark data 320N and the residual data 604 to generate the image frame 216N may include using limited bits of the bitstream 135 to generate the image frame 216N as a better approximation of the image frame 116N.

[0150] refer to Figure 10 , diagram 1000 showing illustrative aspects of operations associated with the motion VRF generator 806 and the video decoder 246 according to some examples of the present disclosure.

[0151] In response to determining that the VRF usage indicator 186N has a particular value (eg, 2 or 3) indicating motion VRF usage, the motion VRF generator 806 applies the motion-based data 322N to the image frame 216A.

[0152] Applying motion-based data 322N to image frame 216A applies global motion to image frame 216A to generate VRF 256NB. For example, motion VRF generator 806 warps image frame 216A based on motion-based data 322N to generate VRF 256NB. In certain implementations, motion VRF generator 806 includes a trained model (e.g., a neural network). Motion VRF generator 806 uses the trained model to process image frame 216A and motion-based data 322N to generate VRF 256NB. For example, motion VRF generator 806 provides image frame 216A and motion-based data 322N as inputs to the trained model, and the output of the trained model indicates VRF 256NB.

[0153] Motion VRF generator 806 provides VRF 256NB to video decoder 246. Video decoder 246 decodes coded bits 706N (e.g., a second subset of coded bits 166N associated with motion VRF usage) to generate residual data 704. Motion VRF generator 806 generates image frame 216N based on the combination of VRF 256NB and residual data 704. In particular aspects, motion-based data 322N and coded bits 706N correspond to fewer bits than an encoded version of first residual data based on the difference between image frame 216A and image frame 116N. Technical advantages of using motion-based data 322N and residual data 704 to generate image frame 216N may include using limited bits of bitstream 135 to generate image frame 216N that is a better approximation of image frame 116N.

[0154] Provided based on reference Figure 9 The VRF 256NA corresponding to the facial landmark data 320N described or as referenced Figure 10 The video decoder 246 generates the image frame 216N based on the VRF 256NB corresponding to the motion-based data 322N as an illustrative example. Alternatively, in some implementations, the video decoder 246 generates the image frame 216N based on both the facial landmark data 320N and the motion-based data 322N. As an illustrative example, the video decoder 246 applies the facial landmark data 320N to the image frame 216A to generate the VRF 256NA, as shown in FIG. Figure 9 As described above, the video encoder 146 applies the motion-based data 322N to the VRF 256NA to generate the VRF 256NB. The video decoder 246 applies the residual data 704 to the VRF 156NB to generate the image frame 216N. In this example, the video encoder 146 applies the facial landmark data 320N to the image frame 116A to generate the VRF 156NA, as shown in FIG. Figure 6As depicted, motion-based data 322N is determined based on a comparison of VRF 156NA and image frame 116N, motion-based data 322N is applied to VRF 156NA to generate VRF 156NB, and residual data 704 is determined based on a comparison of VRF 156NB and image frame 116N.

[0155] refer to Figure 11 , diagram 1100 showing illustrative aspects of the operation of the frame analyzer 142, VRF generator 144, and video encoder 146 according to some examples of the present disclosure.

[0156] Each of the frame analyzer 142 and the video encoder 146 is configured to receive a sequence of image frames 116 (such as a sequence of consecutively captured frames of image data), exemplified by a first image frame (F1) 116A, a second image frame (F2) 116B, and one or more additional image frames including an Nth image frame (FN) 116N (where N is an integer greater than two). The frame analyzer 142 is configured to output a sequence of VRF usage indicators, including a first VRF usage indicator (V1) 186A, a second VRF usage indicator (V2) 186B, and one or more additional VRF usage indicators including an Nth VRF usage indicator (VN) 186N. The frame analyzer 142 is also configured to output a corresponding set of synthetic support data 150, which is illustrated as a second synthetic support data (S2) 150B and one or more additional sets of synthetic support data including an Nth synthetic support data (SN) 150N, when the VRF usage indicator 186 has a particular value indicating VRF usage (e.g., 1, 2, or 3).

[0157] The VRF generator 144 is configured to receive a sequence of VRF usage indicators and a corresponding set of synthesized support data. The VRF generator 144 is configured to selectively generate one or more VRFs 156 based on the synthesized support data, which are exemplified as one or more second VRFs (R2) 156B and one or more additional sets of VRFs including one or more Nth VRFs (RN) 156N.

[0158] The video encoder 146 is configured to generate a sequence of coded bits 166 and a sequence of reference lists 176 corresponding to the sequence of image frames 116. The sequence of coded bits 166 is exemplified as a first coded bit (E1) 166A, a second coded bit (E2) 166B, and one or more additional sets of coded bits including an Nth coded bit (EN) 166N. The sequence of reference lists 176 is exemplified as a first reference list (L1) 176A, a second reference list (L2) 176B, and one or more additional reference lists including an Nth reference list (LN) 176N. The video encoder 146 is configured to selectively generate one or more sets of coded bits 166 based on a corresponding VRF 156 and output corresponding synthesis support data.

[0159] During operation, frame analyzer 142 processes a first image frame (F1) 116A to generate a first VRF usage indicator (V1) 186A. In response to determining that first VRF usage indicator (V1) 186A has a specific value (e.g., 0) indicating no VRF usage, frame analyzer 142 refrains from generating corresponding synthesis support data. In response to determining that first VRF usage indicator (V1) 186A has a specific value (e.g., 0) indicating no VRF usage, VRF generator 144 refrains from generating any VRF associated with first image frame (F1) 116A. In response to determining that first VRF usage indicator (V1) 186A has a specific value (e.g., 0) indicating no VRF usage, video encoder 146 generates first coded bits (E1) 166A independent of any VRF. Video encoder 146 outputs first coded bits (E1) 166A and a first reference list (L1) 176A. In a specific example, the video encoder 146 generates the first coded bit (E1) 166A independently of any reference frame, and the reference list 176A is empty. In another example, the video encoder 146 generates the first coded bit (E1) 166A based on a previous frame in the sequence of image frames 116, and the reference list 176A indicates the previous frame.

[0160] The frame analyzer 142 processes the second image frame (F2) 116B to generate a second VRF usage indicator (V2) 186B. In response to determining that the second VRF usage indicator (V2) 186B has a specific value (e.g., 1, 2, or 3) indicating VRF usage, the frame analyzer 142 generates second synthesis support data (S2) 150B for the second image frame (F2) 116B. In response to determining that the second VRF usage indicator (V2) 186B has a specific value (e.g., 1, 2, or 3) indicating VRF usage, the VRF generator 144 generates one or more second VRFs (R2) 156B associated with the second image frame (F2) 116B. In response to determining that the second VRF usage indicator (V2) 186B has a specific value (e.g., 1, 2, or 3) indicating VRF usage, the video encoder 146 generates second coded bits (E2) 166B based on the one or more second VRFs (R2) 156B. Video encoder 146 outputs second coded bits (E2) 166B, second synthesis support data (S2) 150B, and a second reference list (L2) 176B. Reference list 176B includes one or more VRF identifiers for one or more second VRFs 156B. In some examples, reference list 176B may also include one or more identifiers of one or more previous frames in the sequence of image frames 116 that can be used as reference frames. In some examples, second coded bits (E2) 166B includes one or more subsets of coded bits corresponding to the one or more reference frames indicated in reference list 176B.

[0161] Similarly, the frame analyzer 142 processes the Nth image frame (FN) 116N to generate an Nth VRF usage indicator (VN) 186N. In response to determining that the Nth VRF usage indicator (VN) 186N has a specific value (e.g., 1, 2, or 3) indicating VRF usage, the frame analyzer 142 generates Nth synthetic support data (SN) 150N for the Nth image frame (FN) 116N. In response to determining that the Nth VRF usage indicator (VN) 186N has a specific value (e.g., 1, 2, or 3) indicating VRF usage, the VRF generator 144 generates one or more Nth VRFs (RN) 156N associated with the Nth image frame (FN) 116N.

[0162] In response to determining that the Nth VRF usage indicator (VN) 186N has a particular value (e.g., 1, 2, or 3) indicating VRF usage, the video encoder 146 generates Nth coded bits (EN) 166N based on one or more Nth VRFs (RN) 156N. The video encoder 146 outputs Nth coded bits (EN) 166N, Nth synthetic support data (SN) 150N, and Nth reference list (LN) 176N. Reference list 176N includes one or more VRF identifiers for one or more Nth VRFs (RN) 156N. In some examples, reference list 176B may also include one or more identifiers of one or more previous frames in the sequence of image frames 116 that can be used as reference frames. In some examples, Nth coded bits (EN) 166N include one or more subsets of coded bits corresponding to the one or more reference frames indicated in reference list 176N.

[0163] By dynamically generating coding bits based on virtual reference frames, decoding accuracy may be improved for image frames for which synthetic support data (eg, facial data, motion-based data, or both) may be generated.

[0164] refer to Figure 12 , diagram 1200 showing illustrative aspects of the operation of the VRF generator 244 and the video decoder 246 according to some examples of the present disclosure.

[0165] The VRF generator 244 is configured to receive a set of synthetic support data and generate a corresponding set of VRFs. The set of synthetic support data is exemplified as the second synthetic support data (S2) 150B and one or more additional sets of synthetic support data including the Nth synthetic support data (SN) 150N. The set of VRFs is exemplified as one or more second VRFs (R2) 256B and one or more additional sets of VRFs including one or more Nth VRFs (RN) 256N.

[0166] The video decoder 246 is configured to receive a sequence of coded bits 166 and a sequence of reference lists 176. The sequence of coded bits 166 is illustrated as a first coded bit (E1) 166A, a second coded bit (E2) 166B, and one or more additional sets of coded bits including an Nth coded bit (EN) 166N. The sequence of reference lists 176 is illustrated as a first reference list (L1) 176A, a second reference list (L2) 176B, and one or more additional reference lists including an Nth reference list (LN) 176N.

[0167] The video decoder 246 is configured to generate a sequence of decoded image frames 216 based on the sequence of coded bits 166 and the sequence of reference list 176. The sequence of decoded image frames 216 is illustrated as a first image frame (D1) 216A, a second image frame (D2) 216B, and one or more additional image frames including an Nth image frame (DN) 216N. The video decoder 246 is configured to selectively generate a decoded image frame based on a corresponding VRF 256.

[0168] During operation, the video decoder 246 processes the first coded bits (E1) 166A based on the first reference list (L1) 176A to generate a first image frame (D1) 216A. In response to determining that the first reference list (L1) 176A indicates that there is no VRF associated with the first coded bits (E1) 166A, the video decoder 246 generates the first image frame (D1) 216A independently of any VRF. In a particular implementation, the video decoder 246 receives a sequence of VRF usage indicators 186. In this implementation, in response to determining that the first VRF usage indicator (V1) 186A has a particular value (e.g., 0) indicating no VRF usage, the video decoder 246 generates the first image frame (D1) 216A independently of any VRF.

[0169] The VRF generator 244 processes the second synthetic support data (S2) 150B to generate one or more second VRFs (R2) 256B. The video decoder 246 processes the second coded bits (E2) 166B based on the second reference list (L2) 176B to generate a second image frame (D2) 216B. In response to determining that the second reference list (L2) 176B indicates identifiers of one or more second VRFs (R2) 256B associated with the second coded bits (E2) 166B, the video decoder 246 generates a second image frame (D2) 216B based on the one or more second VRFs (R2) 256B.

[0170] Similarly, the VRF generator 244 processes the Nth synthetic support data (SN) 150N to generate one or more Nth VRFs (RN) 256N. The video decoder 246 processes the Nth coded bits (EN) 166N based on the Nth reference list (LN) 176N to generate the Nth image frame (DN) 216N. In response to determining that the Nth reference list (LN) 176N indicates the identifier of one or more Nth VRFs (RN) 256N associated with the Nth coded bits (EN) 166N, the video decoder 246 generates the Nth image frame (DN) 216N based on the one or more Nth VRFs (RN) 256N.

[0171] By dynamically generating decoded image frames based on virtual reference frames, decoding accuracy can be improved for available image frames (e.g., the second image frame (D2) 216B and the Nth image frame (DN) 216N) for which synthetic support data (e.g., facial data, motion-based data, or both) is used.

[0172] Figure 13 Implementation 1300 of device 102 is depicted as an integrated circuit 1302 including one or more processors 1390. In certain aspects, one or more processors 1390 include one or more processors 190, one or more processors 290, or a combination thereof. Integrated circuit 1302 also includes a signal input 1304 (such as one or more bus interfaces) to enable receipt of input data 1328 for processing. Integrated circuit 1302 includes video analyzer 140, video generator 240, or both. Integrated circuit 1302 also includes a signal output 1306 (such as a bus interface) to enable transmission of output data 1330. In certain examples, input data 1328 includes image frames 116, and output data 1330 includes reference list 176, coded bits 166, VRF usage indicator 186, synthesis support data 150, bitstream 135, or a combination thereof. In another example, the input data 1328 includes the reference list 176 , the coded bits 166 , the VRF usage indicator 186 , the composition support data 150 , the bitstream 135 , or a combination thereof, and the output data 1330 includes the image frame 216 .

[0173] Integrated circuit 1302 enables virtual reference frame based encoding and decoding to be implemented as a component in a system such as Figure 14 The depicted mobile phone or tablet, such as Figure 15 The wearable electronic devices depicted, such as Figure 16 The camera depicted, such as Figure 17 Virtual reality, mixed reality, or augmented reality headsets as depicted or Figure 18 or Figure 19 The means of transport depicted.

[0174] Figure 14 An implementation 1400 is depicted in which device 102, device 160, or both include a mobile device 1402 (such as a phone or tablet) as an illustrative, non-limiting example. Mobile device 1402 includes camera 110 and display 1404. In certain aspects, display 1404 corresponds to Figure 21404 . Components of processor(s) 190 and processor(s) 290, including video analyzer 140 and video generator 240, are integrated into mobile device 1402 and are illustrated using dashed lines to indicate internal components that are generally not visible to a user of mobile device 1402. In a specific example, video analyzer 140 operates to detect image frame 116 or bitstream 135 and then processes it to perform one or more operations at mobile device 1402, such as launching a graphical user interface or otherwise displaying other information at display screen 1404 (e.g., via an integrated “smart assistant” application). For example, display screen 1404 indicates that image frame 116 is being processed to generate bitstream 135, or that bitstream 135 is being processed to generate image frame 216.

[0175] Figure 15 An implementation 1500 is depicted in which device 102, device 160, or both include a wearable electronic device 1502 (illustrated as a "smart watch"). Video analyzer 140, video generator 240, camera 110, or a combination thereof are integrated into wearable electronic device 1502.

[0176] In certain examples, video analyzer 140 or video generator 240 operates to detect image frame 116 or bitstream 135, respectively, and then processes them to perform one or more operations at wearable electronic device 1502, such as launching a graphical user interface or otherwise displaying other information at display screen 1504. For example, display screen 1504 indicates that image frame 116 is being processed to generate bitstream 135, that bitstream 135 is being processed to generate image frame 216, or for playout of the generated image frame 216, such as in a streaming video example.

[0177] In a particular example, the wearable electronic device 1502 includes a haptic device that provides a haptic notification (e.g., a vibration) in response to detecting the image frame 116 or the bitstream 135. For example, the haptic notification can cause a user to look toward the wearable electronic device 1502 to see a displayed notification indicating that the image frame 116 is being processed to generate a bitstream 135 that is available for transmission to another, or a displayed notification indicating that the bitstream 135 is being processed to generate an image frame 216 that is available for viewing. Thus, the wearable electronic device 1502 can alert a hearing-impaired user or a user wearing a head-mounted device that the bitstream 135 is available for transmission or that the image frame 216 is available for viewing.

[0178] Figure 16An implementation 1600 is depicted in which device 102, device 160, or both include a portable electronic device corresponding to camera device 1602. Video analyzer 140, video generator 240, or both are included in camera device 1602. In certain aspects, camera device 1602 corresponds to or includes Figure 1 During operation, in response to receiving a spoken command identified as user voice, camera device 1602 may perform operations in response to the spoken user command, such as adjusting image or video capture settings, image or video playback settings, image or video capture instructions, generating bitstream 135 based on image frames 116, or processing bitstream 135 to display image frames 216 at a display screen, as illustrative examples.

[0179] Figure 17 An implementation 1700 is depicted in which device 102, device 160, or both include a portable electronic device corresponding to a virtual reality, mixed reality, or augmented reality head-mounted device 1702. The video analyzer 140, the video generator 240, the camera 110, or a combination thereof is integrated into the head-mounted device 1702. User voice activity detection can be performed based on audio signals received from a microphone of the head-mounted device 1702. A visual interface device is positioned in front of the user's eyes to enable display of augmented reality, mixed reality, or virtual reality images or scenes to the user while wearing the head-mounted device 1702. In certain examples, the visual interface device is configured to display a notification indicating that the image frame 116 is processed to generate the bitstream 135, to display a notification indicating that the bitstream 135 is processed to generate the image frame 216, or to play out the generated image frame 216, such as in a streaming video example.

[0180] Figure 18 An implementation 1800 is depicted in which device 102, device 160, or both corresponds to or is integrated within a vehicle 1802, illustrated as a manned or unmanned aerial vehicle (e.g., a package delivery drone). Video analyzer 140, video generator 240, camera 110, or a combination thereof is integrated into vehicle 1802. User voice activity detection can be performed based on audio signals received from a microphone of vehicle 1802, such as delivery instructions for an authorized user of vehicle 1802. In a specific example, vehicle 1802 includes a visual interface device configured to display a notification indicating processing of image frame 116 to generate bitstream 135 or processing of bitstream 135 to generate image frame 216. In specific aspects, image frame 116 corresponds to an image of the recipient of the package, an image of the assembly or installation of the delivery product, or a combination thereof. In specific aspects, image frame 216 corresponds to assembly or installation instructions.

[0181] Figure 19 Another embodiment 1900 is depicted in which the device 102, the device 160, or both corresponds to or is integrated within a vehicle 1902 (illustrated as a car). The vehicle 1902 includes one or more processors 1390, which include the video analyzer 140, the video generator 240, or both. The vehicle 1902 also includes a camera 110. User voice activity detection can be performed based on audio signals received from a microphone of the vehicle 1902. In some implementations, user voice activity detection can be performed based on audio signals received from an internal microphone, such as for voice commands from an authorized passenger. In some implementations, user voice activity detection can be performed based on audio signals received from an external microphone, such as an authorized user of the vehicle. In a particular implementation, in response to receiving a spoken command identified as user voice, the voice activation system initiates one or more operations of the vehicle 1902 based on one or more keywords (e.g., "unlock," "start engine," "play music," "show weather forecast," "play video," "transfer video," or another voice command), such as by providing feedback or information via the display 1920 or one or more speakers. For example, the display 1920 can provide information indicating that the image frame 116 has been processed to generate the bitstream 135 ready for transmission, the bitstream 135 has been processed to generate the image frame 216 ready for display, or for playing out the generated image frame 216, such as in a streaming video example.

[0182] refer to Figure 20 , shows a specific implementation of method 2000 for image encoding using virtual reference frames. In certain aspects, one or more operations of method 2000 are performed by frame analyzer 142, VRF generator 144, video encoder 146, video analyzer 140, one or more processors 190, device 102, Figure 1 It is executed by at least one of the systems 100 or a combination thereof.

[0183] Method 2000 includes obtaining, at 2002, composition support data associated with an image frame in a sequence of image frames. For example, Figure 1 The frame analyzer 142 obtains the synthesis support data 150N associated with the image frame 116N in the sequence of image frames 116, as shown in FIG. Figure 1 and Figure 3 described.

[0184] The method 2000 also includes selectively generating a virtual reference frame based on the synthetic support data at 2004. For example, Figure 1The VRF generator 144 selectively generates one or more VRFs 156N based on the synthesized support data 150N, as shown in FIG. Figure 1 and Figures 3 to 7 described.

[0185] The method 2000 also includes generating a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame at 2006. For example, Figure 1 The video encoder 146 generates a bitstream 135 corresponding to an encoded version of the image frame 116N based at least in part on one or more VRFs 156N, as shown in FIG. Figure 1 、 Figure 6 and Figure 7 described.

[0186] Thus, the method 2000 enables generation of VRFs 156 that preserve perceptually important features (e.g., facial landmarks). Technical advantages of using synthetic supporting data 150N (e.g., facial landmark data, motion-based data, or both) to generate one or more VRFs 156N may include generating one or more VRFs 156N that are closer approximations of image frames 116N, thereby improving the video quality of decoded image frames.

[0187] Figure 20 The method 2000 may be implemented by a field programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a digital signal processor (DSP), a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 20 The method 2000 may be performed by a processor executing instructions, such as reference Figure 22 described.

[0188] refer to Figure 21 , shows a specific implementation of method 2100 for image decoding using a virtual reference frame. In certain aspects, one or more operations of method 2100 are performed by Figure 1 Device 160, system 100, Figure 2 The video code is executed by at least one of the bitstream analyzer 242, the VRF generator 244, the video decoder 246, the video generator 240, the one or more processors 290, or a combination thereof.

[0189] The method 2100 includes obtaining a bitstream corresponding to an encoded version of an image frame at 2102. For example, Figure 2 The bitstream analyzer 242 obtains the bitstream 135 corresponding to the encoded version of the image frame 116N, as shown in FIG. Figure 2 described.

[0190] The method 2100 also includes generating a virtual reference frame based on the synthesis support data included in the bitstream based on determining that the bitstream includes a virtual reference frame usage indicator at 2104. For example, in response to determining that the bitstream 135 includes a VRF usage indicator 186N having a particular value (e.g., 1, 2, or 3) indicating VRF usage, Figure 2 The VRF generator 244 generates one or more VRFs 256N based on the synthetic support data 150N included in the bitstream 135, as shown in FIG. Figure 2 described.

[0191] The method 2100 also includes generating a decoded version of the image frame based on the virtual reference frame at 2106. For example, Figure 2 The video decoder 246 generates an image frame 216N (eg, a decoded version of the image frame 116N) based on one or more VRFs 256N, as shown in FIG. Figure 2 described.

[0192] Thus, the method 2100 enables the generation of decoded image frames (e.g., image frame 216N) using VRFs 256 that preserve perceptually important features (e.g., facial landmarks). Technical advantages of using synthetic support data 150N (e.g., facial landmark data, motion-based data, or both) to generate one or more VRFs 256N may include using one or more VRFs 256N that are closer approximations of image frame 216N, thereby improving the video quality of image frame 116N.

[0193] Figure 21 The method 2100 may be implemented by an FPGA device, an ASIC, a processing unit (such as a CPU), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 21 The method 2100 may be performed by a processor executing instructions, such as reference Figure 22 described.

[0194] refer to Figure 22 , a block diagram of a particular exemplary implementation of a device is depicted and generally designated 2200. In various implementations, the device 2200 may have Figure 22 More or fewer components may be used than those illustrated. In an exemplary implementation, the device 2200 may correspond to Figure 1 In an exemplary embodiment, the device 2200 may execute the reference Figures 1 to 21 One or more operations described.

[0195] In certain implementations, the device 2200 includes a processor 2206 (e.g., a CPU). The device 2200 may include one or more additional processors 2210 (e.g., one or more DSPs). In certain aspects, Figure 1 The one or more processors 190 may correspond to the processor 2206, the processor 2210, or a combination thereof. In certain aspects, Figure 2 The one or more processors 290 may correspond to processor 2206, processor 2210, or a combination thereof. Processor 2210 may include a speech and music coder-decoder (CODEC) 2208, which may include a speech coder ("vocoder") encoder 2236, a vocoder decoder 2238, or both. Processor 2210 may include video analyzer 140, video generator 240, or both.

[0196] The device 2200 may include a memory 2286 and a CODEC 2234. The memory 2286 may include instructions 2256 that are executable by one or more additional processors 2210 (or processor 2206) to implement the functionality described with reference to the video analyzer 140, the video generator 240, or both. The device 2200 may include a modem 2270 coupled to an antenna 2252 via a transceiver 2250. In certain aspects, the modem 2270 includes Figure 1 Modem 170, Figure 2 modem 270 or both.

[0197] Device 2200 may include a display 2228 coupled to a display controller 2226. In certain aspects, display 2228 includes Figure 2 The display device 210 may be coupled to the speaker 2292, the microphone 2212, the camera 110, or a combination thereof. The CODEC 2234 may include a digital-to-analog converter (DAC) 2202, an analog-to-digital converter (ADC) 2204, or both. In certain implementations, the CODEC 2234 may receive an analog signal from the microphone 2212, convert the analog signal to a digital signal using the analog-to-digital converter 2204, and provide the digital signal to the voice and music codec 2208. The voice and music codec 2208 may process the digital signal. In certain implementations, the voice and music codec 2208 may provide the digital signal to the CODEC 2234. The CODEC 2234 may convert the digital signal to an analog signal using the digital-to-analog converter 2202 and may provide the analog signal to the speaker 2292.

[0198] In a particular implementation, the device 2200 can be included in a system-in-package or system-on-chip device 2222. In a particular implementation, the memory 2286, the processor 2206, the processor 2210, the display controller 2226, the CODEC 2234, and the modem 2270 are included in the system-in-package or system-on-chip device 2222. In a particular implementation, the input device 2230 and the power source 2244 are coupled to the system-in-package or system-on-chip device 2222. Additionally, in a particular implementation, as in Figure 22 , the display 2228, camera 110, input device 2230, speaker 2292, microphone 2212, antenna 2252, and power source 2244 are external to the system-in-package or system-on-chip device 2222. In a particular implementation, each of the display 2228, camera 110, input device 2230, speaker 2292, microphone 2212, antenna 2252, and power source 2244 can be coupled to a component of the system-in-package or system-on-chip device 2222, such as an interface or controller.

[0199] Device 2200 may include a smart speaker, a speaker bar, a mobile communication device, a smart phone, a cellular phone, a laptop, a computer, a tablet, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a head-mounted device, an augmented reality head-mounted device, a mixed reality head-mounted device, a virtual reality head-mounted device, an aircraft, a home automation system, a voice-activated device, a wireless speaker and voice-activated device, a portable electronic device, an automobile, a computing device, a communication device, an Internet of Things (IoT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.

[0200] In conjunction with the described implementations, an apparatus includes means for obtaining composition support data associated with an image frame in a sequence of image frames. For example, the means for obtaining composition support data may correspond to Figure 1 The frame analyzer 142, the video analyzer 140, the modem 170, the one or more processors 190, the device 102, the system 100, Figure 3 The face detector 302, the facial landmark detector 304, the global motion detector 306, the visual analysis engine 312, the modem 2270, the transceiver 2250, the antenna 2252, the processor 2206, the processor 2210, the device 2200, one or more other circuits or components configured to obtain synthetic support data, or any combination thereof.

[0201] The apparatus further includes means for selectively generating a virtual reference frame based on the synthetic support data. For example, the means for selectively generating a virtual reference frame may correspond to Figure 1 VRF generator 144, video analyzer 140, one or more processors 190, device 102, system 100, Figure 5 The facial VRF generator 504, the motion VRF generator 506, the processor 2206, the processor 2210, the device 2200, one or more other circuits or components configured to selectively generate a virtual reference frame, or any combination thereof.

[0202] The apparatus further comprises means for generating a bitstream corresponding to an encoded version of an image frame based at least in part on the virtual reference frame. For example, the means for generating the bitstream may correspond to Figure 1 The video encoder 146, the video analyzer 140, the modem 170, one or more processors 190, the device 102, the system 100, the modem 2270, the transceiver 2250, the antenna 2252, the processor 2206, the processor 2210, the device 2200, one or more other circuits or components configured to generate a bitstream, or any combination thereof.

[0203] Also in conjunction with the described implementation, the apparatus includes means for obtaining a bitstream corresponding to an encoded version of an image frame. For example, the means for obtaining a bitstream may correspond to Figure 2 device 160, system 100, modem 270, bitstream analyzer 242, video generator 240, one or more processors 290, modem 2270, transceiver 2250, antenna 2252, processor 2206, processor 2210, device 2200, one or more other circuits or components configured to obtain a bitstream, or any combination thereof.

[0204] The apparatus further includes means for generating a virtual reference frame based on the synthesis support data included in the bitstream, the virtual reference frame being generated based on determining that the bitstream includes a virtual reference frame usage indicator. For example, the means for generating the virtual reference frame may correspond to Figure 1 Device 160, system 100, Figure 2 The VRF generator 244, the video generator 240, one or more processors 290, the processor 2206, the processor 2210, the device 2200, one or more other circuits or components configured to generate a virtual reference frame, or any combination thereof.

[0205] The apparatus further includes means for generating a decoded version of the image frame based on the virtual reference frame. For example, the means for generating the virtual reference frame may correspond to Figure 1 Device 160, system 100, Figure 2 The VRF generator 244, the video generator 240, one or more processors 290, the processor 2206, the processor 2210, the device 2200, one or more other circuits or components configured to generate a virtual reference frame, or any combination thereof.

[0206] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as memory 2286) includes instructions (e.g., instructions 2256) that, when executed by one or more processors (e.g., processor(s) 190, processor(s) 2210, or processor 2206), cause the one or more processors to obtain composition support data (e.g., composition support data 150N) associated with an image frame (e.g., image frame 116N) of a sequence of image frames (e.g., image frame 116). The instructions, when executed by the one or more processors, further cause the one or more processors to selectively generate a virtual reference frame (e.g., one or more VRFs 156N) based on the composition support data. The instructions, when executed by the one or more processors, further cause the one or more processors to generate a bitstream (e.g., bitstream 135) corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.

[0207] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as memory 2286) includes instructions (e.g., instructions 2256) that, when executed by one or more processors (e.g., processor(s) 290, processor(s) 2210, or processor 2206), cause the one or more processors to obtain a bitstream (e.g., bitstream 135) corresponding to an encoded version of an image frame (e.g., image frame 116N). The instructions, when executed by the one or more processors, further cause the one or more processors to generate a virtual reference frame (e.g., one or more VRFs 256N) based on synthesis support data (e.g., synthesis support data 150N) included in the bitstream, based on determining that the bitstream includes a virtual reference frame usage indicator (e.g., VRF usage indicator 186N). The instructions, when executed by the one or more processors, further cause the one or more processors to generate a decoded version of the image frame based on the virtual reference frame.

[0208] Certain aspects of the present disclosure are described below in a related set of embodiments:

[0209] According to embodiment 1, a device includes: one or more processors, wherein the one or more processors are configured to: obtain a bitstream corresponding to an encoded version of an image frame; based on determining that the bitstream includes a virtual reference frame usage indicator, generate a virtual reference frame based on synthetic support data included in the bitstream; and generate a decoded version of the image frame based on the virtual reference frame.

[0210] Embodiment 2 includes the apparatus of embodiment 1, wherein the synthetic support data comprises facial landmark data, motion-based data, or a combination thereof.

[0211] Embodiment 3 includes the apparatus of embodiment 1 or embodiment 2, wherein the bitstream indication includes a first set of reference candidates for the virtual reference frame.

[0212] Embodiment 4 includes the apparatus of embodiment 3, wherein the bitstream indicates one or more additional first sets of reference candidates including one or more additional virtual reference frames associated with one or more additional image frames in the sequence of image frames.

[0213] Embodiment 5 includes the apparatus of any one of Embodiments 1 to 4, wherein the bitstream further indicates a second set of reference candidates comprising one or more previously decoded image frames.

[0214] Embodiment 6 includes the apparatus of any one of embodiments 1 to 5, wherein the bitstream includes a supplemental enhancement information (SEI) message indicating the synthesis support data.

[0215] Embodiment 7 includes a device according to any one of embodiments 1 to embodiment 6, wherein the synthetic support data includes facial landmark data indicating the location of facial features, and wherein the one or more processors are configured to generate the virtual reference frame based at least in part on a previously decoded image frame and the location of the facial features.

[0216] Embodiment 8 includes an apparatus according to any one of embodiments 1 to 7, wherein the synthetic support data includes motion-based data indicating global motion, and wherein the one or more processors are configured to generate the virtual reference frame based at least in part on a previously decoded image frame and the global motion.

[0217] Embodiment 9 includes an apparatus according to any one of embodiments 1 to 8, wherein the one or more processors are configured to use motion-based data to warp a previously decoded image frame to generate the virtual reference frame, wherein the synthetic support data includes the motion-based data.

[0218] Embodiment 10 includes the apparatus of any one of Embodiments 1 to 9, wherein the one or more processors are configured to generate the virtual reference frame using a trained model.

[0219] Embodiment 11 includes the apparatus of embodiment 10, wherein the trained model comprises a neural network.

[0220] Embodiment 12 includes an apparatus according to embodiment 10 or embodiment 11, wherein the input to the trained model includes the synthetic support data and at least one previously decoded image frame.

[0221] Embodiment 13 includes the apparatus of any one of Embodiments 1 to 12, further comprising a modem configured to receive the bitstream from a second device.

[0222] Embodiment 14 includes the apparatus of any one of Embodiments 1 to 13, further comprising a display device configured to display the decoded version of the image frame.

[0223] According to embodiment 15, a method includes: obtaining a bitstream corresponding to an encoded version of an image frame at a device; generating a virtual reference frame based on synthetic support data included in the bitstream based on determining that the bitstream includes a virtual reference frame usage indicator; and generating a decoded version of the image frame based on the virtual reference frame at the device.

[0224] Embodiment 16 includes the method of embodiment 15, wherein the synthetic support data includes facial landmark data, motion-based data, or a combination thereof.

[0225] Embodiment 17 includes the method of embodiment 15 or embodiment 16, wherein the bitstream indicates a first set of reference candidates including the virtual reference frame.

[0226] Embodiment 18 includes the method of embodiment 17, wherein the bitstream indication includes one or more additional first sets of reference candidates including one or more additional virtual reference frames associated with one or more additional image frames in the sequence of image frames.

[0227] Embodiment 19 includes the method of any one of Embodiments 15 to 18, wherein the bitstream further indicates a second set of reference candidates comprising one or more previously decoded image frames.

[0228] Embodiment 20 includes the method of any one of Embodiments 15 to 19, wherein the bitstream includes a supplemental enhancement information (SEI) message indicating the synthesis support data.

[0229] Embodiment 21 includes a method according to any one of embodiments 15 to embodiment 20, further comprising generating the virtual reference frame based at least in part on a previously decoded image frame and the location of facial features, wherein the synthetic support data includes facial landmark data indicating the location of the facial features.

[0230] Embodiment 22 includes a method according to any one of embodiments 15 to embodiment 21, further comprising generating the virtual reference frame based at least in part on a previously decoded image frame and a global motion, wherein the synthetic support data includes motion-based data indicating the global motion.

[0231] Embodiment 23 includes a method according to any one of embodiments 15 to embodiment 22, further comprising using motion-based data to warp a previously decoded image frame to generate the virtual reference frame, wherein the synthetic support data includes the motion-based data.

[0232] Embodiment 24 includes the method of any one of Embodiments 15 to 23, further comprising generating the virtual reference frame using a trained model.

[0233] Embodiment 25 comprises the method of embodiment 24, wherein the trained model comprises a neural network.

[0234] Embodiment 26 includes a method according to embodiment 24 or embodiment 25, wherein the input to the trained model includes the synthetic support data and at least one previously decoded image frame.

[0235] Embodiment 27 includes the method of any one of Embodiments 15 to 26, further comprising receiving the bitstream from a second device via a modem.

[0236] Embodiment 28 includes the method of any one of Embodiments 15 to 27, further comprising displaying the decoded version of the image frame at a display device.

[0237] According to embodiment 29, a device includes: a memory configured to store instructions; and a processor configured to execute the instructions to perform the method according to any one of embodiments 15 to 28.

[0238] According to embodiment 30, a non-transitory computer-readable medium stores instructions, which, when executed by a processor, cause the processor to perform the method according to any one of embodiments 15 to 28.

[0239] According to embodiment 31, an apparatus includes means for performing the method according to any one of embodiments 15 to 28.

[0240] According to embodiment 32, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to: obtain a bitstream corresponding to an encoded version of an image frame; based on determining that the bitstream includes a virtual reference frame usage indicator, generate a virtual reference frame based on synthetic support data included in the bitstream; and generate a decoded version of the image frame based on the virtual reference frame.

[0241] According to embodiment 33, an apparatus includes: a component for obtaining a bitstream corresponding to an encoded version of an image frame; a component for generating a virtual reference frame based on synthetic support data included in the bitstream, the virtual reference frame being generated based on determining that the bitstream includes a virtual reference frame usage indicator; and a component for generating a decoded version of the image frame based on the virtual reference frame.

[0242] According to embodiment 34, a device includes: one or more processors, wherein the one or more processors are configured to: obtain synthesis support data associated with an image frame in a sequence of image frames; selectively generate a virtual reference frame based on the synthesis support data; and generate a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.

[0243] Embodiment 35 includes the apparatus of embodiment 34, wherein the synthetic support data comprises facial landmark data, motion-based data, or a combination thereof.

[0244] Embodiment 36 includes the apparatus of embodiment 34 or embodiment 35, wherein the bitstream includes the synthesis support data.

[0245] Embodiment 37 includes the apparatus of any one of Embodiments 34 to 36, wherein the one or more processors are configured to generate a first set of reference candidates including the virtual reference frame.

[0246] Embodiment 38 includes the apparatus of embodiment 37, wherein the bitstream indicates the first set of reference candidates.

[0247] Embodiment 39 includes a device according to embodiment 37 or embodiment 38, wherein the one or more processors are configured to generate one or more additional first sets of reference candidates including one or more additional virtual reference frames associated with one or more additional image frames in the sequence of image frames.

[0248] Embodiment 40 includes the apparatus of any one of Embodiments 34 to 39, wherein the bitstream further indicates a second set of reference candidates comprising one or more previously decoded image frames.

[0249] Embodiment 41 includes an apparatus according to embodiment 40, wherein the one or more processors are configured to generate the virtual reference frame based at least in part on determining that the count of reference frames in the second set of reference candidates is less than a threshold reference count of a decoding configuration.

[0250] Embodiment 42 includes an apparatus according to any one of Embodiments 34 to 41, wherein the one or more processors are configured to generate the virtual reference frame based at least in part on detecting a face in the image frame.

[0251] Embodiment 43 includes an apparatus according to any one of Embodiments 34 to 42, wherein the one or more processors are configured to: obtain motion-based data associated with the image frame; and generate the virtual reference frame based at least in part on determining that the motion-based data indicates global motion greater than a global motion threshold.

[0252] Embodiment 44 includes the apparatus of any one of Embodiments 34 to 43, wherein the bitstream includes a supplemental enhancement information (SEI) message indicating the synthesis support data.

[0253] Embodiment 45 includes an apparatus according to any one of Embodiments 34 to 44, wherein the synthesis support data includes facial landmark data indicating the location of facial features in the image frame.

[0254] Embodiment 46 includes the apparatus of embodiment 45, wherein the facial feature comprises at least one of an eye, an eyelid, an eyebrow, a nose, a lip, or a facial contour.

[0255] Embodiment 47 includes the apparatus of any one of Embodiments 34 to 46, wherein the composition support data includes motion sensor data indicative of motion of an image capture device associated with the image frame.

[0256] Embodiment 48 includes the device of embodiment 47, wherein the image capture device comprises at least one of an extended reality (XR) device, a vehicle, or a camera.

[0257] Embodiment 49 includes a device according to any one of Embodiments 34 to 48, wherein the one or more processors are configured to use motion-based data to warp a previously decoded image frame to generate the virtual reference frame, wherein the synthetic support data includes the motion-based data.

[0258] Embodiment 50 includes the apparatus of any one of Embodiments 34 to 49, wherein the bitstream includes a supplemental enhancement information (SEI) message indicating use of a virtual reference frame for generating a decoded version of the image frame.

[0259] Embodiment 51 includes the apparatus of any one of Embodiments 34 to 50, wherein the one or more processors are configured to generate the virtual reference frame using a trained model.

[0260] Embodiment 52 includes an apparatus according to embodiment 51, wherein the trained model comprises a neural network.

[0261] Embodiment 53 includes an apparatus according to embodiment 51 or embodiment 52, wherein the input to the trained model includes the synthetic support data and at least one previously decoded image frame.

[0262] Embodiment 54 includes the apparatus of any one of Embodiments 34 to 53, further comprising a modem configured to transmit the bitstream to a second device.

[0263] Embodiment 55 includes a device according to any one of Embodiments 34 to 54, further comprising a camera configured to capture the image frame.

[0264] According to embodiment 56, a method includes: obtaining, at a device, synthesis support data associated with an image frame in a sequence of image frames; selectively generating a virtual reference frame based on the synthesis support data; and generating, at the device, a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.

[0265] Embodiment 57 includes the method of embodiment 56, wherein the synthetic support data includes facial landmark data, motion-based data, or a combination thereof.

[0266] Embodiment 58 includes the method of embodiment 56 or embodiment 57, wherein the bitstream includes the synthesis support data.

[0267] Embodiment 59 includes the method of any one of Embodiments 56 to 58, further comprising generating a first set of reference candidates including the virtual reference frame.

[0268] Embodiment 60 includes the method of embodiment 59, wherein the bitstream indicates the first set of reference candidates.

[0269] Embodiment 61 includes a method according to embodiment 59 or embodiment 60, which further includes generating one or more additional first sets of reference candidates including one or more additional virtual reference frames associated with one or more additional image frames in the sequence of image frames.

[0270] Embodiment 62 includes the method of any one of Embodiments 56 to 61, wherein the bitstream further indicates a second set of reference candidates comprising one or more previously decoded image frames.

[0271] Embodiment 63 includes the method of embodiment 62, further comprising generating the virtual reference frame based at least in part on determining that a count of reference frames in the second set of reference candidates is less than a threshold reference count of a coding configuration.

[0272] Embodiment 64 includes the method of any one of Embodiments 56 to 63, further comprising generating the virtual reference frame based at least in part on detecting a face in the image frame.

[0273] Example 65 includes a method according to any one of Examples 56 to 64, further comprising: obtaining motion-based data associated with the image frame; and generating the virtual reference frame based at least in part on determining that the motion-based data indicates global motion greater than a global motion threshold.

[0274] Embodiment 66 includes the method of any one of Embodiments 56 to 65, wherein the bitstream includes a supplemental enhancement information (SEI) message indicating the synthesis support data.

[0275] Embodiment 67 includes a method according to any one of Embodiments 56 to 66, wherein the synthesis support data includes facial landmark data indicating the location of facial features in the image frame.

[0276] Embodiment 68 includes a method according to embodiment 67, wherein the facial feature includes at least one of an eye, an eyelid, an eyebrow, a nose, a lip, or a facial contour.

[0277] Embodiment 69 includes the method of any one of Embodiments 56 to 68, wherein the synthesis support data includes motion sensor data indicative of motion of an image capture device associated with the image frame.

[0278] Embodiment 70 includes the method of embodiment 69, wherein the image capture device comprises at least one of an extended reality (XR) device, a vehicle, or a camera.

[0279] Embodiment 71 includes a method according to any one of Embodiments 56 to 70, further comprising using motion-based data to warp a previously decoded image frame to generate the virtual reference frame, wherein the synthetic support data includes the motion-based data.

[0280] Embodiment 72 includes the method of any one of Embodiments 56 to 71, wherein the bitstream includes a supplemental enhancement information (SEI) message indicating the use of a virtual reference frame for generating a decoded version of the image frame.

[0281] Example 73 includes the method of any one of Examples 56 to 72, further comprising generating the virtual reference frame using a trained model.

[0282] Embodiment 74 comprises the method of embodiment 73, wherein the trained model comprises a neural network.

[0283] Embodiment 75 includes a method according to embodiment 73 or embodiment 74, wherein the input to the trained model includes the synthetic support data and at least one previously decoded image frame.

[0284] Embodiment 76 includes the method of any one of Embodiments 56 to 75, further comprising transmitting the bitstream to a second device via a modem.

[0285] Embodiment 77 includes the method of any one of Embodiments 56 to 76, further comprising receiving the image frame from a camera.

[0286] According to embodiment 78, a device includes: a memory configured to store instructions; and a processor configured to execute the instructions to perform the method according to any one of embodiments 56 to 77.

[0287] According to embodiment 79, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform a method according to any one of embodiments 56 to 77.

[0288] According to embodiment 80, an apparatus includes means for performing the method according to any one of embodiments 56 to 77.

[0289] According to embodiment 81, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to: obtain synthesis support data associated with an image frame in a sequence of image frames; selectively generate a virtual reference frame based on the synthesis support data; and generate a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.

[0290] According to embodiment 82, an apparatus comprises: a component for obtaining synthetic support data associated with an image frame in a sequence of image frames; a component for selectively generating a virtual reference frame based on the synthetic support data; and a component for generating a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.

[0291] It will also be apparent to those skilled in the art that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in conjunction with the specific implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of the two. Various illustrative components, blocks, configurations, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or processor-executable instructions depends on the specific application and the design constraints imposed on the overall system. A skilled artisan may implement the described functionality in different ways for each specific application, and such specific implementation decisions should not be interpreted as causing a departure from the scope of this disclosure.

[0292] The steps of the method or algorithm described in conjunction with the specific implementation disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. In an alternative, the storage medium may be integral with the processor. The processor and the storage medium may reside in an application specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. In an alternative, the processor and the storage medium may reside in a computing device or a user terminal as discrete components.

[0293] The preceding description of the disclosed aspects is provided to enable those skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but should be accorded the broadest scope possible consistent with the principles and novel features as defined by the following claims.

Claims

1. A device, comprising: One or more processors configured to: obtaining a bitstream corresponding to an encoded version of an image frame; generating a virtual reference frame based on synthesis support data included in the bitstream based on determining that the bitstream includes a virtual reference frame usage indicator; as well as A decoded version of the image frame is generated based on the virtual reference frame. 2 . The apparatus of claim 1 , wherein the synthetic support data comprises facial landmark data, motion-based data, or a combination thereof. 3 . The apparatus of claim 1 , wherein the bitstream indication includes a first set of reference candidates for the virtual reference frame. 4 . The apparatus of claim 3 , wherein the bitstream indication comprises one or more additional first sets of reference candidates for one or more additional virtual reference frames associated with one or more additional image frames in the sequence of image frames. 5 . The apparatus of claim 1 , wherein the bitstream further indicates a second set of reference candidates comprising one or more previously decoded image frames.

6. The apparatus of claim 1, wherein the bitstream includes a supplemental enhancement information (SEI) message indicating the synthesis support data.

7. The apparatus of claim 1 , wherein the synthesis support data comprises facial landmark data indicating locations of facial features, and wherein the one or more processors are configured to generate the virtual reference frame based at least in part on a previously decoded image frame and the locations of the facial features.

8. The apparatus of claim 1 , wherein the synthesis support data comprises motion-based data indicative of global motion, and wherein the one or more processors are configured to generate the virtual reference frame based at least in part on a previously decoded image frame and the global motion.

9. The apparatus of claim 1, wherein the one or more processors are configured to warp a previously decoded image frame using motion-based data to generate the virtual reference frame, wherein the synthesis support data includes the motion-based data.

10. The apparatus of claim 1, wherein the one or more processors are configured to generate the virtual reference frame using a trained model, and wherein input to the trained model includes the synthetic support data and at least one previously decoded image frame.

11. The device of claim 1, further comprising a modem configured to receive the bit stream from a second device.

12. The device of claim 1, further comprising a display device configured to display the decoded version of the image frame.

13. A method comprising: obtaining, at the device, a bitstream corresponding to an encoded version of the image frame; generating a virtual reference frame based on synthesis support data included in the bitstream based on determining that the bitstream includes a virtual reference frame usage indicator; as well as A decoded version of the image frame is generated at the device based on the virtual reference frame.

14. A device comprising: One or more processors configured to: obtaining composition support data associated with an image frame in the sequence of image frames; selectively generating a virtual reference frame based on the synthesis support data; as well as A bitstream is generated corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.

15. The apparatus of claim 14, wherein the synthetic support data comprises facial landmark data, motion-based data, or a combination thereof.

16. The apparatus of claim 14, wherein the bitstream comprises the synthesis support data. 17 . The apparatus of claim 14 , wherein the one or more processors are configured to generate a first set of reference candidates comprising the virtual reference frame.

18. The apparatus of claim 17, wherein the bitstream indicates the first set of reference candidates.

19. The apparatus of claim 17, wherein the one or more processors are configured to generate one or more additional first sets of reference candidates comprising one or more additional virtual reference frames associated with one or more additional image frames in the sequence of image frames.

20. The apparatus of claim 14, wherein the bitstream further indicates a second set of reference candidates comprising one or more previously decoded image frames, and wherein the one or more processors are configured to generate the virtual reference frame based at least in part on determining that a count of reference frames in the second set of reference candidates is less than a threshold reference count of a decoding configuration.

21. The device of claim 14, wherein the one or more processors are configured to generate the virtual reference frame based at least in part on detecting a face in the image frame.

22. The apparatus of claim 14, wherein the one or more processors are configured to: obtaining motion-based data associated with the image frame; and The virtual reference frame is generated based at least in part on a determination that the motion-based data indicates global motion greater than a global motion threshold.

23. The apparatus of claim 14, wherein the composition support data comprises facial landmark data indicating locations of facial features in the image frame.

24. The device of claim 14, wherein the composition support data comprises motion sensor data indicative of motion of an image capture device associated with the image frame.

25. The device of claim 24, wherein the image capture device comprises at least one of an extended reality (XR) device, a vehicle, or a camera.

26. The device of claim 14, wherein the one or more processors are configured to warp a previously decoded image frame using motion-based data to generate the virtual reference frame, wherein the synthesis support data includes the motion-based data.

27. The apparatus of claim 14, wherein the bitstream includes a supplemental enhancement information (SEI) message indicating use of a virtual reference frame for generating a decoded version of the image frame.

28. The apparatus of claim 14, wherein the one or more processors are configured to generate the virtual reference frame using a trained model, and wherein input to the trained model includes the synthetic support data and at least one previously decoded image frame.

29. The device of claim 14, further comprising a modem configured to transmit the bit stream to a second device.

30. A method comprising: obtaining, at a device, composition support data associated with an image frame in a sequence of image frames; selectively generating a virtual reference frame based on the synthesis support data; as well as A bitstream is generated at the device corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.