Virtual reference frames for image encoding and decoding
Generating virtual reference frames using synthetic support data addresses the limitations of conventional decoders by enhancing prediction accuracy and reducing bandwidth consumption.
Patent Information
- Application Number
- JP2025545255
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-14
- Filing Date
- 2024-02-07
- Publication Date
- 2026-02-13
AI Technical Summary
Conventional image decoders rely on limited previously decoded frames as reference frames, leading to suboptimal predictions and degraded video quality, especially at low bitrate settings, while transmitting additional data for improved quality consumes bandwidth resources.
Generate a virtual reference frame based on synthetic support data, such as facial landmark and motion-based data, to enhance prediction accuracy and reduce bandwidth usage.
Improves video quality by preserving important features and reduces bandwidth requirements through efficient encoding and decoding processes.
Smart Images

Figure 2026505343000001_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of priority to commonly owned U.S. Non-Provisional Patent Application No. 18 / 168,891, filed February 14, 2023, the entire contents of which are expressly incorporated herein by reference.
[0002] TECHNICAL FIELD This disclosure relates generally to image encoding and decoding. [Background technology]
[0003] Advances in technology have resulted in smaller and more powerful computing devices. For example, there are now a variety of portable personal computing devices that are small, lightweight, and easily carried by users, including wireless telephones such as mobile phones and smartphones, tablet computers, and laptop computers. These devices can communicate voice and data packets over wireless networks. Furthermore, many such devices incorporate additional functionality, such as digital still cameras, digital video cameras, digital recorders, and audio file players. Such devices can also process executable instructions, including software applications, such as web browser applications that can be used to access the Internet. Thus, these devices can contain significant computing power.
[0004] Such computing devices often incorporate functionality for receiving encoded video data corresponding to compressed image frames from another device. Typically, previously decoded image frames are used as reference frames for predicting the decoded image frame. The more suitable such reference frames are for predicting an image frame, the more accurately the image frame can be decoded, resulting in higher-quality reproduction of the video data. However, because the reference frames available to conventional decoders are limited to previously decoded image frames, in some situations the available reference frames may only provide suboptimal predictions of the image frame, thus resulting in degraded-quality video reproduction. While decoding quality can be improved by transmitting additional data to the decoder to generate a higher-quality reproduction of the image frame, sending such additional data consumes more bandwidth resources that may be unavailable to devices operating with limited transmission channel capacity. Summary of the Invention [Means for solving the problem]
[0005] According to one implementation of the present disclosure, a device includes one or more processors configured to obtain synthetic support data associated with an image frame of a sequence of image frames. The one or more processors are also configured to selectively generate a virtual reference frame based on the synthetic support data. The one or more processors are further configured to generate a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.
[0006] According to another implementation aspect of the present disclosure, a method includes, at a device, obtaining synthetic support data associated with an image frame of a sequence of image frames. The method also includes selectively generating a virtual reference frame based on the synthetic support data. The method further includes generating, at the device, a bitstream corresponding to an encoded version of the image frame that is based at least in part on the virtual reference frame.
[0007] According to another implementation of the present disclosure, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to obtain synthetic support data associated with an image frame of a sequence of image frames. The instructions, when executed by the one or more processors, also cause the one or more processors to selectively generate a virtual reference frame based on the synthetic support data. The instructions, when executed by the one or more processors, further cause the one or more processors to generate a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.
[0008] According to another implementation of the present disclosure, an apparatus includes means for obtaining synthetic support data associated with an image frame of a sequence of image frames. The apparatus also includes means for selectively generating a virtual reference frame based on the synthetic support data. The apparatus further includes means for generating a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.
[0009] According to another implementation of the present disclosure, a device includes one or more processors configured to obtain a bitstream corresponding to an encoded version of an image frame. The one or more processors are also configured to generate a virtual reference frame based on synthetic support data included in the bitstream based on a determination that the bitstream includes a virtual reference frame usage indicator. The one or more processors are further configured to generate a decoded version of the image frame based on the virtual reference frame.
[0010] According to another implementation of the present disclosure, a method includes, at a device, obtaining a bitstream corresponding to an encoded version of an image frame. The method also includes, based on a determination that the bitstream includes a virtual reference frame usage indicator, generating a virtual reference frame based on synthetic support data included in the bitstream. The method further includes, at the device, generating a decoded version of the image frame based on the virtual reference frame.
[0011] According to another implementation of the present disclosure, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to obtain a bitstream corresponding to an encoded version of an image frame. The instructions, when executed by the one or more processors, also cause the one or more processors to generate a virtual reference frame based on synthetic support data included in the bitstream based on a determination that the bitstream includes a virtual reference frame usage indicator. The instructions, when executed by the one or more processors, further cause the one or more processors to generate a decoded version of the image frame based on the virtual reference frame.
[0012] According to another implementation of the present disclosure, an apparatus includes means for obtaining a bitstream corresponding to an encoded version of an image frame. The apparatus also includes means for generating a virtual reference frame based on synthetic support data included in the bitstream, the virtual reference frame being generated based on a determination that the bitstream includes a virtual reference frame usage indicator. The apparatus further includes means for generating a decoded version of the image frame based on the virtual reference frame.
[0013] Other aspects, advantages, and features of the present disclosure will become apparent after review of the entire application, including the following sections: Brief Description of the Drawings, Detailed Description, and the Claims. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a block diagram of a particular illustrative aspect of a system operable to generate a virtual reference frame for image encoding, in accordance with some examples of the present disclosure. [Figure 2] 2 is a diagram of the system of FIG. 1 operable to generate a virtual reference frame for image decoding, in accordance with some examples of the present disclosure. [Figure 3] 2A-2C illustrate exemplary aspects of operations associated with the frame analyzer and virtual reference frame generator of FIG. 1, in accordance with some examples of the present disclosure. [Figure 4] 2 is a diagram of an example aspect of operations associated with a synthesis support analyzer of the frame analyzer of FIG. 1 in accordance with some examples of the present disclosure. [Figure 5] 2A-2C illustrate example aspects of operations associated with the virtual reference frame generator of FIG. 1, in accordance with some examples of the present disclosure. [Figure 6] 2 is a diagram of an example aspect of operations associated with a facial virtual reference frame generator and video encoder of the virtual reference frame generator of FIG. 1, in accordance with some examples of the present disclosure. [Figure 7]2 is a diagram of an example aspect of the operation associated with the virtual reference frame generator and video encoder of FIG. 1 in accordance with some examples of the present disclosure. [Figure 8] 3A-3C illustrate example aspects of operations associated with the virtual reference frame generator of FIG. 2, in accordance with some examples of the present disclosure. [Figure 9] 3 is a diagram of an example aspect of operations associated with a facial virtual reference frame generator and video decoder of the virtual reference frame generator of FIG. 2, in accordance with some examples of the present disclosure. [Figure 10] 3 is a diagram of an example aspect of the operation associated with the motion virtual reference frame generator and video decoder of the virtual reference frame generator of FIG. 2 in accordance with some examples of the present disclosure. [Figure 11] 2A-2C are diagrams of example aspects of the operation of the frame analyzer, virtual reference frame generator, and video encoder of FIG. 1 in accordance with some examples of the present disclosure. [Figure 12] 3A-3C are diagrams of example aspects of the operation of the virtual reference frame generator and video decoder of FIG. 2 in accordance with some examples of the present disclosure. [Figure 13] 1 illustrates an example integrated circuit operable to generate a virtual reference frame for image encoding, image decoding, or both, according to some examples of the present disclosure. [Figure 14] FIG. 1 is a diagram of a mobile device operable to generate a virtual reference frame for image encoding, image decoding, or both, in accordance with some examples of the present disclosure. [Figure 15] FIG. 1 is a diagram of a wearable electronic device operable to generate a virtual reference frame for image encoding, image decoding, or both, according to some examples of the present disclosure. [Figure 16] 1 is a diagram of a camera operable to generate a virtual reference frame for image encoding, image decoding, or both, in accordance with some examples of the present disclosure. [Figure 17]1 is a diagram of a headset, such as a virtual reality headset, a mixed reality headset, or an augmented reality headset, operable to generate a virtual reference frame for image encoding, image decoding, or both, in accordance with some examples of the present disclosure. [Figure 18] 1 is a diagram of a first example vehicle operable to generate a virtual reference frame for image encoding, image decoding, or both, in accordance with some examples of the present disclosure. [Figure 19] FIG. 10 is a diagram of a second example vehicle operable to generate a virtual reference frame for image encoding, image decoding, or both, in accordance with some examples of the present disclosure. [Figure 20] 2 is a diagram of a specific implementation of a method for generating a virtual reference frame for image encoding that may be performed by the device of FIG. 1 in accordance with some examples of the present disclosure. [Figure 21] 3 is a diagram of a specific implementation of a method for generating a virtual reference frame for image decoding that may be performed by the device of FIG. 2 in accordance with some examples of the present disclosure. [Figure 22] 1 is a block diagram of a particular illustrative example of a device operable to generate a virtual reference frame for image encoding, image decoding, or both, in accordance with some examples of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0015] Typically, video decoding involves using a previously decoded image frame as a reference frame to predict a decoded image frame. In one example, a sequence of image frames includes a first image frame and a second image frame. An encoder encodes the first image frame to generate first coded bits. For example, the encoder uses intraframe compression to generate the first coded bits.
[0016] The encoder encodes the second image frame to generate second coded bits. For example, the encoder uses a local decoder to decode the first coded bits to generate a first decoded image frame and encodes the second image frame using the first decoded image frame as a reference frame. For example, the encoder determines first residual data based on a difference between the first decoded image frame and the second image frame. The encoder generates second coded bits based on the first residual data. The first coded bits and the second coded bits are transmitted from a first device including the encoder to a second device including a decoder.
[0017] The decoder decodes the first coded bits to generate a first decoded image frame. For example, the decoder performs intra-frame prediction on the first coded bits to generate the first decoded image frame. The decoder decodes the second coded bits to generate residual data for the second decoded image frame. In response to determining that the first decoded image frame is a reference frame for the second decoded image frame, the decoder generates the second decoded image frame based on a combination of the residual data and the first decoded image frame.
[0018] At low bitrate settings (e.g., used during video conferencing), the presence of compression artifacts can degrade video quality. For example, a first compression artifact associated with intraframe compression may be present in a first decoded image frame. As another example, a second compression artifact associated with decoded residual bits may be present in a second decoded image frame.
[0019] Systems and methods for generating a virtual reference frame for image encoding and decoding are disclosed. In one example, an encoder determines synthesis support data for a second image frame and generates a virtual reference frame for the second image frame based on the synthesis support data. In some implementations, the synthesis support data can include facial landmark data indicating locations of facial features in the second image frame. In some implementations, the synthesis support data can include motion-based data indicating global motion (e.g., camera motion) detected in the second image frame relative to the first image frame (or a first decoded image frame generated by a local decoder).
[0020] The encoder generates a virtual reference frame based on application of the synthetic support data to the first image frame (or the first decoded image frame). The encoder generates second residual data based on a difference between the virtual reference frame and the second image frame. The encoder generates second coded bits based on the second residual data. The first coded bits, the second coded bits, the synthetic support data, and the virtual reference frame usage indicator are transmitted from the first device to the second device. The virtual reference frame usage indicator indicates the use of the virtual reference frame.
[0021] The decoder decodes the first coded bits to generate a first decoded image frame. For example, the decoder performs intraframe prediction on the first coded bits to generate the first decoded image frame. The decoder decodes the second coded bits to generate second residual data. In response to determining that the virtual reference frame usage indicator indicates virtual reference frame usage, the decoder applies synthesis support data to the first decoded image frame to generate the virtual reference frame. In one example, the synthesis support data includes facial landmark data indicating a position of a facial feature in the second image frame. Applying the facial landmark data to the first decoded image frame includes adjusting a position of the facial feature to more closely match a position of the facial feature indicated in the second image frame. In another example, the synthesis support data includes motion-based data indicating global motion detected in the second image frame relative to the first image frame. Applying the motion-based data to the first decoded image frame includes applying the global motion to the first decoded image frame to generate the virtual reference frame. The decoder applies the second residual data to the virtual reference frame to generate a second decoded image frame.
[0022] Using a virtual reference frame can improve video quality by preserving perceptually important features (e.g., facial landmarks) in the second decoded image frame. In some examples, the encoded version of the synthetic support data and the second residual data (e.g., corresponding to the differences between the virtual reference frame and the second image frame) uses fewer bits than the encoded version of the first residual data (e.g., corresponding to the differences between the first decoded image frame and the second image frame). Illustratively, the second residual data may have smaller numerical values and an overall smaller variance compared to the first residual data, so the second residual data may be encoded more efficiently (e.g., using fewer bits). In these examples, the virtual reference frame approach can reduce bandwidth usage, improve video quality, or both.
[0023] Certain aspects of the present disclosure are described below with reference to the drawings. In this description, common features are indicated by common reference numerals. As used herein, various terms are used only for the purpose of describing particular implementations and are not intended to limit the implementations. For example, the singular forms "a," "an," and "the" are intended to include the plural unless the context clearly dictates otherwise. Furthermore, some features described herein are singular in some implementations and plural in other implementations. To illustrate, FIG. 1 shows a device 102 including one or more processors ("processor(s)" 190 in FIG. 1), indicating that in some implementations, the device 102 includes a single processor 190 and in other implementations, the device 102 includes multiple processors 190. For ease of reference herein, such features are generally introduced as "one or more" features and subsequently referred to in the singular unless an aspect relating to multiple features is described.
[0024] In some figures, multiple instances of a particular type of feature are used. Although these features are physically and / or logically different, the same reference number is used for each, and the different instances are distinguished by the addition of a letter to the reference number. When features as a group or type are referred to herein (e.g., when no specific one of the features is referenced), the reference number is used without the distinguishing letter. However, when one specific feature of multiple features of the same type is referred to herein, the reference number is used with the distinguishing letter. For example, with reference to FIG. 1, multiple image frames are shown and associated with reference numbers 116A and 116N. When referring to a specific one of these images, such as image frame 116A, the distinguishing letter "A" is used. However, when referring to these image frames as any one or group of these image frames, the reference number 116 is used without the distinguishing letter.
[0025] As used herein, the terms “comprise,” “comprises,” and “comprising” may be used interchangeably with “include,” “includes,” or “including.” Additionally, the term “wherein” may be used interchangeably with “where.” As used herein, “exemplary” denotes an example, implementation, and / or aspect and should not be construed as limiting or as indicating a preferred or preferred implementation. As used herein, ordinal terms (e.g., “first,” “second,” “third,” etc.) used to modify an element, such as a structure, component, operation, etc., do not in themselves indicate a priority or order of the element with respect to other elements, but merely distinguish the element from other elements having the same name (apart from the use of ordinal terms). As used herein, the term “set” refers to one or more of a particular element, and the term “plurality” refers to multiple (e.g., two or more) of a particular element.
[0026] As used herein, "coupled" may include "communicatively coupled," "electrically coupled," or "physically coupled," as well as (or alternatively) any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof), etc. Two devices (or components) that are electrically coupled may be included in the same device or in different devices and may be connected via electronic components, one or more connectors, or inductive coupling, as illustrative, non-limiting examples. In some implementations, two devices (or components) that are communicatively coupled, such as in electrical communication, may send and receive signals (e.g., digital or analog signals) directly or indirectly via one or more wires, buses, networks, etc. As used herein, "directly coupled" may include two devices coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) with no intervening components.
[0027] In this disclosure, terms such as "determining," "calculating," "estimating," "shifting," "adjusting," and the like may be used to describe how one or more operations are performed. It should be noted that such terms should not be construed as limiting, and other techniques may be utilized to perform similar operations. Additionally, as referred to herein, "generating," "calculating," "estimating," "using," "selecting," "accessing," and "determining" may be used interchangeably. For example, "generating," "calculating," "estimating," or "determining" a parameter (or signal) may refer to actively generating, estimating, calculating, or determining a parameter (or signal), or may refer to using, selecting, or accessing a parameter (or signal) that has already been generated by another component or device.
[0028] 1, a particular exemplary embodiment of a system 100 configured to generate a virtual reference frame for image encoding and decoding is shown. System 100 includes a device 102 configured to be coupled to a camera 110, a device 160, or both.
[0029] Device 102 includes an input interface 114, one or more processors 190, and a modem 170. Input interface 114 is coupled to one or more processors 190 and configured to be coupled to camera 110. Input interface 114 is configured to receive camera output 112 from camera 110 and to provide camera output 112 to one or more processors 190 as image frames 116.
[0030] The one or more processors 190 are coupled to the modem 170 and include a video analyzer 140. The video analyzer 140 includes a frame analyzer 142 coupled to a video encoder 146 via a virtual reference frame (VRF) generator 144. The video encoder 146 is coupled to the modem 170.
[0031] The video analyzer 140 is configured to obtain a sequence of image frames 116, such as image frame 116A, image frame 116N, one or more additional image frames, or a combination thereof. In some implementations, the sequence of image frames 116 may include one or more image frames before image frame 116A, one or more image frames between image frame 116A and image frame 116N, one or more image frames after image frame 116N, or a combination thereof.
[0032] Each of the image frames 116 is associated with a frame identifier (ID) 126. For example, image frame 116A has frame identifier 126A, image frame 116N has frame identifier 126N, etc. In some implementations, the frame identifier 126 indicates the order of the image frames 116 within the sequence. In one example, a frame identifier 126A having a first value that is less than a second value of the frame identifier 126N indicates that the image frame 116A precedes the image frame 116N in the sequence.
[0033] The video analyzer 140 is configured to selectively generate one or more virtual reference frames (VRFs) for a particular one of the image frames 116. The frame analyzer 142 is configured to generate synthesis support data 150N for the image frame 116N in response to determining that at least one VRF 156 associated with the image frame 116N should be generated. The synthesis support data 150N may include facial landmark data, motion-based data, or both. For example, the frame analyzer 142 is configured to generate facial landmark data as the synthesis support data 150N in response to detecting a face in the image frame 116N. The facial landmark data indicates the location of facial features detected in the image frame 116N. As another example, the frame analyzer 142 is configured to include motion-based data in the synthesis support data 150N in response to determining that the motion-based data indicates that global motion in the image frame 116N relative to the image frame 116A (e.g., the previous image frame in the sequence) is greater than a global motion threshold.
[0034] In one example, the frame analyzer 142 is configured to generate a virtual reference frame (VRF) usage indicator 186N having a first value (e.g., 0) in response to determining that a VRF should not be generated for the image frame 116N. For example, the frame analyzer 142 is configured to determine that a VRF should not be generated for the image frame 116N in response to determining that no face is detected in the image frame 116N and global motion less than or equal to a global motion threshold is detected in the image frame 116N. Alternatively, the frame analyzer 142 is configured to generate a VRF usage indicator 186N having a second value (e.g., 1), a third value (e.g., 2), or a fourth value (e.g., 3) in response to determining that at least one VRF 156N should be generated for the image frame 116N. For example, the VRF usage indicator 186N has a second value (e.g., 1) indicating that the synthesis support data 150N includes facial landmark data, a third value (e.g., 2) indicating that the synthesis support data 150N includes motion-based data, or a fourth value (e.g., 3) indicating that the synthesis support data 150N includes both facial landmark data and motion-based data.
[0035] The VRF generator 144 is configured to generate one or more VRFs 156N based on the synthesis support data 150N in response to determining that the VRF usage indicator 186N has a value (e.g., 1, 2, or 3) that indicates VRF usage for the image frame 116N. The reference list 176 associated with the image frame 116 indicates reference frame candidates for the image frame 116. In one example, the VRF generator 144 is configured to generate a reference list 176N associated with the image frame 116N that indicates one or more VRFs 156N. The video encoder 146 is configured to encode the image frame 116N based on the reference frame candidates indicated by the reference list 176N to generate coded bits 166N.
[0036] The modem 170 is coupled to one or more processors 190 and configured to enable communication with the device 160, such as for sending the bitstream 135 to the device 160 via wireless transmission. For example, the bitstream 135 includes the reference list 176N, the coded bits 166N, the synthesis support data 150N, the VRF usage indicator 186N, or a combination thereof.
[0037] In some implementations, device 102 corresponds to or is included in one of various types of devices. In the illustrated example, one or more processors 190 are incorporated into at least one of a mobile phone or tablet computing device described with reference to FIG. 14, a wearable electronic device described with reference to FIG. 15, a camera device described with reference to FIG. 16, or a virtual reality, mixed reality, or augmented reality headset described with reference to FIG. 17. In another illustrative example, one or more processors 190 are incorporated into a vehicle, as further described with reference to FIGS. 18 and 19.
[0038] During operation, the video analyzer 140 obtains a sequence of image frames 116. In a particular example, the input interface 114 is configured to receive the camera output 112 from the camera 110 and provide the camera output 112 as image frames 116 to the video analyzer 140. In another example, the video analyzer 140 obtains the image frames 116 from a storage device, a network device, another component of the device 102, or a combination thereof.
[0039] The video analyzer 140 selectively generates VRFs for the image frames 116. In one example, the frame analyzer 142 generates the synthesis support data 150N, the VRF usage indicator 186N, or both, based on a determination of whether at least one VRF should be generated for the image frame 116N, as further described with reference to FIGS. 3 and 4. For example, in response to determining that no VRF should be generated for the image frame 116N, the frame analyzer 142 generates the VRF usage indicator 186N having a first value (e.g., 0) indicating no VRF usage. Alternatively, in response to determining that at least the face of the person 180 is detected in the image frame 116N, the frame analyzer 142 adds facial landmark data to the synthesis support data 150N and generates the VRF usage indicator 186N having a second value (e.g., 1) indicating facial VRF usage. The facial landmark data indicates the locations of facial features of the person 180 detected in the image frame 116N. According to some embodiments, the facial features include at least one of the person's 180 eyes, eyelids, eyebrows, nose, lips, or facial contours.
[0040] In yet another example, the frame analyzer 142 generates motion-based data based on a comparison of the image frame 116N with the image frame 116A (e.g., a previous image frame in the sequence). In some implementations, the motion-based data includes motion sensor data indicative of movement of the image capture device (e.g., the camera 110) associated with the image frame 116N. In some implementations, the motion-based data indicates global motion detected in the image frame 116N relative to the previous image frame (e.g., the image frame 116A).
[0041] In response to determining that the motion-based data indicates global motion greater than the global motion threshold, the frame analyzer 142 adds the motion-based data to the synthesis support data 150N and generates a VRF usage indicator 186N having a third value (e.g., 2) indicating motion VRF usage. In some examples, in response to determining that the motion-based data and the facial landmark data should be used to generate at least one VRF, the frame analyzer 142 generates synthesis support data 150N including the facial landmark data and the motion-based data and generates a VRF usage indicator 186N having a fourth value (e.g., 3) indicating both facial VRF usage and motion VRF usage. The frame analyzer 142 provides the VRF usage indicator 186N to the VRF generator 144. In examples in which the VRF usage indicator 186N has a value (e.g., 1, 2, or 3) indicating VRF usage, the frame analyzer 142 provides the synthesis support data 150N to the VRF generator 144. In certain aspects, the compositing support data 150N, the VRF usage indicator 186N, or both, includes a frame identifier 126N to indicate an association with the image frame 116N.
[0042] In response to determining that the VRF usage indicator 186N has a first value (e.g., 0) indicating no VRF usage, the VRF generator 144 provides the VRF usage indicator 186N to the video encoder 146 and refrains from passing the reference list 176N to the video encoder 146. Optionally, in some implementations, in response to determining that the VRF usage indicator 186N has a first value (e.g., 0) indicating no VRF usage, the VRF generator 144 passes an empty list as the reference list 176N to the video encoder 146.
[0043] Alternatively, in response to determining that the VRF usage indicator 186N has a value indicating VRF usage (e.g., 1, 2, or 3), the VRF generator 144 generates one or more VRFs 156N as one or more candidate VRF references associated with the image frame 116N. For example, in response to determining that the VRF usage indicator 186N has a value indicating facial VRF usage (e.g., 1 or 3), the VRF generator 144 generates at least a VRF 156N based on facial landmark data included in the synthesis support data 150N, as further described with reference to Figures 5 and 6. In response to determining that the VRF usage indicator 186N has a value indicating motion VRF usage (e.g., 2 or 3), the VRF generator 144 generates at least a VRF 156N based on motion-based data included in the synthesis support data 150N, as further described with reference to Figures 5 and 7.
[0044] The VRF generator 144 generates a reference list 176N to indicate that one or more VRFs 156N are designated as a first set of reference candidates (e.g., VRF reference candidates) for the image frame 116N. In one example, the reference list 176N includes a frame identifier 126N to indicate an association with the image frame 116N. The reference list 176N includes one or more VRF reference candidate identifiers 172 of the first set of reference candidates. For example, the one or more VRF reference candidate identifiers 172 include one or more VRF identifiers 196N of the one or more VRFs 156N. Illustratively, the one or more VRF reference candidate identifiers 172 include a VRF identifier 196NA of the VRF 156NA, a VRF identifier 196NB of the VRF 156NB, one or more additional VRF identifiers of the one or more additional VRFs, or a combination thereof. The VRF generator 144 provides one or more VRFs 156N, reference lists 176N, VRF usage indicators 186N, or combinations thereof to the video encoder 146.
[0045] Video encoder 146 is configured to encode image frame 116N to generate coded bits 166N. In a particular aspect, video encoder 146 generates the subset of coded bits 166N based at least in part on a second set of reference candidates (e.g., encoder reference candidates) separate from VRF 156. The second set of reference candidates includes one or more previous image frames or one or more previously decoded image frames. In a particular implementation, video encoder 146 uses image frame 116A (or a locally decoded image frame corresponding to image frame 116A) as an i-frame (intra-coded frame). In this implementation, subset 166N of coded bits is based on a residual corresponding to the difference between image frame 116A (or a locally decoded image frame) and image frame 116N. The video encoder 146 adds the frame identifier 126A of the image frame 116A (or a locally decoded image frame) to one or more encoder reference candidate identifiers 174 of a second set of reference candidates in the reference list 176N.
[0046] The video encoder 146 selectively generates one or more subsets of the coded bits 166N based on one or more VRFs 156N. For example, the video encoder 146 generates one or more subsets of the coded bits 166N based on one or more VRFs 156N in response to determining that the VRF usage indicator 186N has a particular value (e.g., 1, 2, or 3) indicating VRF usage and that the encoder reference candidate count is less than a threshold reference count. Alternatively, the video encoder 146 refrains from generating any of the coded bits 166N based on a VRF 156 in response to determining that the VRF usage indicator 186N has a particular value (e.g., 0) indicating no VRF usage, determining that the encoder reference candidate count is equal to or greater than the threshold reference count, or both.
[0047] In certain aspects, the video encoder 146 determines the encoder reference candidate count based on a count of one or more encoder reference candidate identifiers 174 included in the reference list 176N. In some aspects, the encoder reference candidate count is based on default data, configuration settings, user input, a coding configuration of the video encoder 146, or a combination thereof. In some implementations, the threshold reference count is based on default data, configuration settings, user input, a coding configuration of the video encoder 146, or a combination thereof.
[0048] Optionally, in some implementations, the VRF generator 144 selectively generates one or more VRFs 156N based on a determination that the encoder reference candidate count is less than a threshold reference count. In particular aspects, the VRF generator 144 determines the encoder reference candidate count based on default data, configuration settings, user input, the coding configuration of the video encoder 146, or a combination thereof. In particular aspects, the VRF generator 144 receives the encoder reference candidate count from the video encoder 146.
[0049] In some implementations, the VRF generator 144 determines the threshold VRF count based on a comparison of (e.g., the difference between) the threshold reference count and the encoder reference candidate count. In these implementations, the VRF generator 144 generates one or more VRFs 156N such that the count of the one or more VRFs 156N is less than or equal to the threshold VRF count.
[0050] In particular aspects, video encoder 146 generates a first subset of coded bits 166N based on VRF 156NA based at least in part on a determination that VRF usage indicator 186N has a particular value (e.g., 1 or 3) that indicates facial VRF usage, as further described with reference to Figure 6. Video encoder 146 generates a second subset of coded bits 166N based on VRF 156NB based at least in part on a determination that VRF usage indicator 186N has a particular value (e.g., 2 or 3) that indicates motion VRF usage, as further described with reference to Figure 7.
[0051] Video encoder 146 provides reference list 176N, coded bits 166N, or both to modem 170. In addition, frame analyzer 142 provides VRF usage indicator 186N, synthesis support data 150N, or both to modem 170. Modem 170 transmits bitstream 135 to device 160. Bitstream 135 includes coded bits 166N, reference list 176N, VRF usage indicator 186N, synthesis support data 150N, or a combination thereof. For example, VRF usage indicator 186N indicates whether any virtual reference frames should be used to generate a decoded version of image frame 116N.
[0052] In some aspects, the bitstream 135 includes a supplemental enhancement information (SEI) message that indicates the synthesis support data 150N. In some aspects, the bitstream 135 includes an SEI message that includes a VRF usage indicator 186N. In particular aspects, the bitstream 135 corresponds to an encoded version of the image frame 116N that is based at least in part on one or more VRFs 156N, one or more encoder reference candidates associated with one or more encoder reference candidate identifiers 174, or a combination thereof.
[0053] In some implementations, the bitstream 135 includes coded bits 166, reference lists 176, VRF usage indicators 186, synthesis support data 150, or combinations thereof, associated with multiple image frames 116. In particular implementations, the bitstream 135 includes a reference list 176 that includes a first reference list associated with image frame 116A, a reference list 176N associated with image frame 116N, one or more additional reference lists associated with one or more additional image frames of the sequence, or a combination thereof. For example, the reference list 176 includes one or more VRF identifiers 196 associated with image frame 116A, one or more VRF identifiers 196N associated with image frame 116N, one or more VRF identifiers 196 associated with one or more additional image frames 116, or a combination thereof. As another example, the reference list 176 includes one or more frame identifiers 126 as one or more encoder reference candidate identifiers 174 associated with the image frame 116A, one or more frame identifiers 126 as one or more encoder reference candidate identifiers 174 associated with the image frame 116N, one or more additional frame identifiers 126 as one or more encoder reference candidate identifiers 174 associated with one or more additional image frames 116, or a combination thereof.
[0054] Thus, system 100 enables generating VRFs 156 that preserve perceptually important features (e.g., facial landmarks). A technical advantage of using synthesis support data 150N (e.g., facial landmark data, motion-based data, or both) to generate one or more VRFs 156N may include improved video quality of the decoded image frames because the one or more VRFs 156N are closer approximations of image frame 116N.
[0055] Although the camera 110 is shown as being external to the device 102, in other implementations the camera 110 may be integrated into the device 102. Although the video analyzer 140 is shown as obtaining the image frames 116 from the camera 110, in other implementations the video analyzer 140 can obtain the image frames 116 from another component of the device 102 (e.g., a graphics processor), another device (e.g., a storage device, a network device, etc.), or a combination thereof. The camera 110 is shown as an example of an image capture device, and in some implementations the video analyzer 140 can obtain the image frames 116 from various types of image capture devices, such as an augmented reality (XR) device, a vehicle, the camera 110, a graphics processor, or a combination thereof.
[0056] Although frame analyzer 142, VRF generator 144, video encoder 146, and modem 170 are shown as separate components, in other implementations, two or more of frame analyzer 142, VRF generator 144, video encoder 146, or modem 170 may be combined into a single component. Although frame analyzer 142, VRF generator 144, and video encoder 146 are shown as being included in a single device (e.g., device 102), in other implementations, one or more operations described herein with reference to frame analyzer 142, VRF generator 144, or video encoder 146 may be performed in another device. Optionally, in some implementations, video analyzer 140 can receive image frames 116, synthesis support data 150, or both from another device.
[0057] 2, a particular exemplary embodiment of system 100 is shown. System 100 is operable to generate a virtual reference frame for image decoding. Device 160 is configured to be coupled to display device 210, device 102, or both.
[0058] Device 102 includes an output interface 214, one or more processors 290, and a modem 270. Output interface 214 is coupled to one or more processors 290 and is configured to be coupled to a display device 210.
[0059] The modem 270 is coupled to the one or more processors 290 and is configured to enable communication with the device 102, such as to receive the bitstream 135 via wireless transmission from the device 102. For example, the bitstream 135 includes the reference list 176N, the coded bits 166N, the synthesis support data 150N, the VRF usage indicator 186N, or a combination thereof.
[0060] The one or more processors 290 are coupled to the modem 270 and include a video generator 240. The video generator 240 includes a bitstream analyzer 242 coupled to a VRF generator 244 and a video decoder 246. The VRF generator 244 is coupled to the video decoder 246. The bitstream analyzer 242 is also coupled to the modem 270.
[0061] The bitstream analyzer 242 is configured to obtain, from the modem 270, data from the bitstream 135 corresponding to an encoded version of the image frame 116N of Figure 1. Illustratively, the bitstream 135 includes the encoded bits 166N, the VRF usage indicator 186N, the reference list 176N, or a combination thereof. If the bitstream 135 includes the VRF usage indicator 186N having a particular value (e.g., 1, 2, or 3) indicating VRF usage, the bitstream 135 also includes the synthesis support data 150N.
[0062] In response to determining that the bitstream 135 includes a VRF usage indicator 186N having a particular value (e.g., 1, 2, or 3) indicating VRF usage, the bitstream analyzer 242 is configured to extract synthesized support data 150N from the bitstream 135 and provide the synthesized support data 150N to the VRF generator 244. In some implementations, the bitstream analyzer 242 is configured to provide the VRF usage indicator 186N, the reference list 176N, or both to the VRF generator 244. The bitstream analyzer 242 is configured to provide the coded bits 166N, the reference list 176N, or both to the video decoder 246.
[0063] The VRF generator 244 is configured to selectively generate one or more VRFs 256N for generating a decoded version of the image frame 116N. For example, the VRF generator 244 is configured to determine whether at least one VRF should be used to generate a decoded version of the image frame 116N based on the synthesis support data 150N, the reference list 176N, the VRF usage indicator 186N, or a combination thereof associated with the image frame 116N. In response to determining that at least one VRF should be used, the VRF generator 244 is configured to generate the one or more VRFs 256N based on the synthesis support data 150N. For example, the VRF generator 244 is configured to generate the one or more VRFs 256N based on facial landmark data, motion-based data, or both indicated by the synthesis support data 150N.
[0064] The video decoder 246 is configured to generate a sequence of image frames 216 corresponding to a decoded version of the sequence of image frames 116. In one example, the image frames 216 include image frame 216A, image frame 216N, one or more additional image frames, or a combination thereof. Each of the image frames 216 is associated with a frame identifier 126. For example, the image frame 216A corresponding to the decoded version of image frame 116A includes the frame identifier 126A of image frame 116A. As another example, the image frame 216N corresponding to the decoded version of image frame 116N includes the frame identifier 126N of image frame 116N.
[0065] The video decoder 246 is configured to selectively generate the image frame 216 based on the corresponding one or more VRFs 256. For example, the video decoder 246 is configured to generate the image frame 216N based on the coded bits 166N, the one or more VRFs 256N, the reference list 176N, or a combination thereof. In some implementations, the video generator 240 is configured to provide the image frame 216 to the display device 210 via the output interface 214. In a particular implementation, the video generator 240 is configured to provide the image frame 216 to the display device 210 in the playback order indicated by the frame identifier 126. For example, during forward playback, the video generator 240 provides the image frame 216A to the display device 210 for playback earlier than the image frame 216N based on a determination that the frame identifier 126A is less than the frame identifier 126N. In a particular example, a person 280 can view the image frame 216 displayed by the display device 210.
[0066] In some implementations, device 160 corresponds to or is included in one of various types of devices. In the illustrated example, one or more processors 290 are incorporated into at least one of a mobile phone or tablet computing device described with reference to FIG. 14, a wearable electronic device described with reference to FIG. 15, a camera device described with reference to FIG. 16, or a virtual reality, mixed reality, or augmented reality headset described with reference to FIG. 17. In another illustrative example, one or more processors 290 are integrated into a vehicle, as further described with reference to FIGS. 18 and 19.
[0067] During operation, the video generator 240 obtains a bitstream 135 corresponding to an encoded version of the image frame 116N of FIG. 1. For example, the bitstream 135 includes the encoded bits 166N, the VRF usage indicator 186N, the reference list 176N, or a combination thereof associated with the image frame 116N. In some examples, the bitstream 135 also includes synthesis support data 150N associated with the image frame 116N. In certain aspects, the encoded bits 166N, the VRF usage indicator 186N, the reference list 176N, the synthesis support data 150N, or a combination thereof, indicate the frame identifier 126N of the image frame 116N.
[0068] In a particular example, video generator 240 obtains bitstream 135 via modem 270. In another example, video generator 240 obtains bitstream 135 from a storage device, a network device, another component of device 160, or a combination thereof.
[0069] The video generator 240 selectively generates a VRF for determining a decoded version of the image frame 116. In one example, the bitstream analyzer 242 determines that no VRF should be used to generate the image frame 216N corresponding to the decoded version of the image frame 116N in response to determining that the bitstream 135 does not include the VRF usage indicator 186N or that the VRF usage indicator 186N has a first value (e.g., 0) indicating no VRF usage. Alternatively, the bitstream analyzer 242 determines that at least one VRF should be used to generate the image frame 216N in response to determining that the bitstream 135 includes the VRF usage indicator 186N having a particular value (e.g., 1, 2, or 3) indicating VRF usage.
[0070] In response to determining that at least one VRF should be used to generate the image frame 216N, the bitstream analyzer 242 provides the synthesis support data 150N, the reference list 176N, the VRF usage indicator 186N, or a combination thereof, to the VRF generator 244 to generate the at least one VRF. The bitstream analyzer 242 also provides the coded bits 166N, the reference list 176N, or both to the video decoder 246 to generate the image frame 216N. In some examples, the bitstream analyzer 242, the VRF generator 244, or both provide the VRF usage indicator 186N to the video decoder 246.
[0071] In response to determining that the bitstream 135 includes a VRF usage indicator 186N having a particular value (e.g., 1, 2, or 3) indicating VRF usage, the VRF generator 244 generates one or more VRFs 256N as one or more candidate VRF references to be used to generate the image frame 216N. For example, in response to determining that the VRF usage indicator 186N has a particular value (e.g., 1 or 3) indicating facial VRF usage, the VRF generator 244 generates at least a VRF 256N based on facial landmark data included in the synthesis support data 150N, as will be further described with reference to Figures 8 and 9. In response to determining that the VRF usage indicator 186N has a particular value (e.g., 2 or 3) indicating motion VRF usage, the VRF generator 244 generates at least a VRF 256N based on motion-based data included in the synthesis support data 150N, as will be further described with reference to Figures 8 and 10.
[0072] 1, the reference list 176N includes one or more VRF reference candidate identifiers 172. For example, the one or more VRF reference candidate identifiers 172 include a VRF identifier 196NA for VRF 156NA, a VRF identifier 196NB for VRF 156NB, one or more additional VRF identifiers of one or more additional VRFs, or a combination thereof.
[0073] The VRF generator 244 assigns one or more VRF identifiers 196N to one or more VRFs 256N. In a particular example, in response to determining that the facial landmark data is associated with the VRF identifier 196NA, the VRF generator 244 assigns the VRF identifier 196NA to the VRF 256NA generated based on the facial landmark data. Thus, the VRF 256NA corresponds to the VRF 156NA generated by the video analyzer 140 of FIG. 1 . In another example, in response to determining that the motion-based data is associated with the VRF identifier 196NB, the VRF generator 244 assigns the VRF identifier 196NB to the VRF 256NB generated based on the motion-based data. Thus, the VRF 256NB corresponds to the VRF 156NB generated by the video analyzer 140 of FIG. 1 . The VRF generator 244 provides the one or more VRFs 256N to the video decoder 246.
[0074] The video decoder 246 is configured to generate an image frame 216N (e.g., a decoded version of the image frame 116N of FIG. 1 ) based on at least the coded bits 166N. In certain aspects, the video decoder 246 selectively generates the image frame 216N based on one or more VRFs 256N. As described with reference to FIG. 1 , the reference list 176N includes one or more VRF reference candidate identifiers 172 of a first set of reference candidates (e.g., one or more VRFs 256N), one or more encoder reference candidate identifiers 174 of a second set of reference candidates (e.g., one or more previously decoded image frames 216), or a combination thereof.
[0075] In a particular example, reference list 176N is empty, and video decoder 246 generates image frame 216N by processing (e.g., decoding) coded bits 166N without regard to any reference candidates. As an illustrative example, image frame 216N may correspond to an i-frame.
[0076] In a particular example, the video decoder 246 selects one or more of the reference candidates indicated in the reference list 176N to generate the image frame 216N based on a selection criterion. The selection criterion may be based on user input, default data, configuration settings, a threshold reference count, or a combination thereof. In one example, the video decoder 246 selects one or more of the second set of reference candidates (e.g., encoder reference candidates) if the reference list 176N does not indicate any of the first set of reference candidates (e.g., one or more VRFs 256N). Alternatively, the video decoder 246 generates the image frame 216N based on the one or more VRFs 256N, independently of the encoder reference candidates, if the reference list 176N indicates at least one of the one or more VRFs 256N.
[0077] Video decoder 246 applies coded bits 166N (e.g., residual) to a selected one of the reference candidates to generate a decoded image frame. For example, video decoder 246 applies a first subset of coded bits 166N to VRF 256NA to generate a first decoded image frame, as further described with reference to FIG. 9. As another example, video decoder 246 applies a second subset of coded bits 166N to VRF 256NB to generate a second decoded image frame, as further described with reference to FIG. 10. In yet another example, video decoder 246 applies a third subset of coded bits 166N to image frame 216A to generate a third decoded image frame.
[0078] In certain implementations in which video decoder 246 selects a single reference candidate from among the reference candidates (e.g., VRF256NA, VRF256NB, or image frame 216A), the corresponding decoded image frame (e.g., the first decoded image frame, the second decoded image frame, or the third decoded image frame) is designated as image frame 216N.
[0079] In certain implementations in which the video decoder 246 selects multiple reference candidates (e.g., VRF 256NA, VRF 256NB, and image frame 216A), the video decoder 246 generates image frame 216N based on a combination of corresponding decoded image frames (e.g., the first decoded image frame, the second decoded image frame, and the third decoded image frame). For example, the video decoder 246 generates image frame 216N by averaging the decoded image frames (e.g., the first decoded image frame, the second decoded image frame, and the third decoded image frame) pixel by pixel or by using information in the bitstream 135 that indicates how to combine the decoded image frames (e.g., the weights of a weighted sum of the decoded image frames).
[0080] In the illustrative example, video generator 240 provides image frames 216N to display device 210 via output interface 214. Optionally, in some implementations, video generator 240 provides image frames 216N to a storage device, a network device, a user device, or a combination thereof.
[0081] Thus, system 200 enables generating a decoded image frame (e.g., image frame 216N) using a VRF 256 that preserves perceptually significant features (e.g., facial landmarks). A technical advantage of generating one or more VRFs 256N using synthesis support data 150N (e.g., facial landmark data, motion-based data, or both) may include improved video quality of image frame 216N because one or more VRFs 256N are closer approximations of image frame 116N (compared to image frame 216A).
[0082] Although display device 210 is shown as being external to device 160, in other implementations, display device 210 may be incorporated into device 160. Although video generator 240 is shown as receiving bitstream 135 from device 160 via modem 270, in other implementations, video generator 240 can obtain bitstream 135 from another component of device 102 (e.g., a graphics processor), another device (e.g., a storage device, a network device, etc.), or a combination thereof. In particular implementations, device 102, device 160, or both may include a copy of video analyzer 140 and a copy of video generator 240. For example, a video analyzer 140 of the device 102 generates a bitstream 135 from image frames 116 received from the camera 110, the video analyzer 140 stores the bitstream 135 in memory, a video generator 240 of the device 102 retrieves the bitstream 135 from memory, the video generator 240 generates image frames 216 from the bitstream 135, and the video generator 240 provides the image frames 216 to a display device.
[0083] Although bitstream analyzer 242, VRF generator 244, video decoder 246, and modem 270 are shown as separate components, in other implementations, two or more of bitstream analyzer 242, VRF generator 244, video decoder 246, or modem 270 may be combined into a single component. Although bitstream analyzer 242, VRF generator 244, and video decoder 246 are shown as being included in a single device (e.g., device 160), in other implementations, one or more operations described herein with reference to bitstream analyzer 242, VRF generator 244, or video decoder 246 may be performed in another device.
[0084] 3, a diagram 300 of an example aspect of operations associated with frame analyzer 142 and VRF generator 144 is shown, in accordance with some examples of the present disclosure. Frame analyzer 142 includes a visual analysis engine 312 coupled to a synthesis support analyzer 314.
[0085] The visual analytics engine 312 includes a face detector 302, a facial landmark detector 304, and a global motion detector 306. The face detector 302 uses facial recognition technology to generate a face detection indicator 318N that indicates whether at least one face is detected in the image frame 116N. For example, the face detection indicator 318N has a first value (e.g., 0) that indicates no face is detected in the image frame 116N, or a second value (e.g., 1) that indicates at least one face is detected in the image frame 116N.
[0086] As further described with reference to FIG. 6, in response to determining that the face detection indicator 318N indicates that at least one face has been detected in the image frame 116N, the facial landmark detector 304 uses facial analysis techniques to generate facial landmark data 320N indicating the locations of facial features detected in the image frame 116N and includes the facial landmark data 320N in the synthesis support data 150N.
[0087] The global motion detector 306 uses a global motion detection technique to generate a motion detection indicator 316N that indicates whether at least a threshold global motion is detected in the image frame 116N relative to the image frame 116A. For example, the motion detection indicator 316N has a first value (e.g., 0) that indicates that at least a threshold global motion is not detected in the image frame 116N, or a second value (e.g., 1) that indicates that at least a threshold global motion is detected in the image frame 116N.
[0088] The global motion detector 306 generates motion-based data 322N indicative of global motion detected in the image frame 116N using motion analysis techniques, as further described with reference to FIG. 7 , and includes the motion-based data 322N in the synthesis support data 150N in response to determining that the motion detection indicator 316N indicates that at least a threshold global motion has been detected in the image frame 116N. In particular implementations, the global motion detector 306 generates the motion-based data 322N (e.g., global motion vectors) based on a comparison of the image frame 116A and the image frame 116N. In some implementations, the global motion detector 306 also, or alternatively, receives sensor data indicative of a first position of the camera 110 at a first capture time of the image frame 116A and a second position of the camera 110 at a second capture time of the image frame 116N. The global motion detector 306 determines global motion based on a comparison of (e.g., a difference between) the first position and the second position. The global motion detector 306 generates motion-based data 322N indicative of a difference between the second position and the first position in response to determining that the global motion is greater than a threshold global motion. The vision analytics engine 312 provides a motion detection indicator 316N and a face detection indicator 318N to the synthesis support analyzer 314.
[0089] The synthesis support analyzer 314 generates a VRF usage indicator 186N based on the motion detection indicator 316N, the face detection indicator 318N, or both. For example, the VRF usage indicator 186N has a first value (e.g., 0) indicating no VRF usage, corresponding to the first value (e.g., 0) of the motion detection indicator 316N and the first value (e.g., 0) of the face detection indicator 318N. In another example, the VRF usage indicator 186N has a second value (e.g., 1) indicating no motion VRF usage and face VRF usage, corresponding to the first value (e.g., 0) of the motion detection indicator 316N and the second value (e.g., 1) of the face detection indicator 318N. The VRF usage indicator 186N has a third value (e.g., 2) indicating motion VRF usage and no face VRF usage, corresponding to the second value (e.g., 1) of the motion detection indicator 316N and the first value (e.g., 0) of the face detection indicator 318N. The VRF usage indicator 186N has a fourth value (e.g., 3) indicating motion VRF usage and face VRF usage, corresponding to the second value (e.g., 1) of the motion detection indicator 316N and the second value (e.g., 1) of the face detection indicator 318N. In certain implementations, the motion detection indicator 316N and the face detection indicator 318N are each one-bit values, and the VRF usage indicator 186N is a two-bit value corresponding to the concatenation of the motion detection indicator 316N and the face detection indicator 318N.
[0090] The frame analyzer 142 provides the VRF usage indicator 186N to the VRF generator 144. If the VRF usage indicator 186N has a particular value (e.g., 1, 2, or 3) indicating VRF usage, the frame analyzer 142 also provides the synthesis support data 150N to the VRF generator 144. In response to determining that the VRF usage indicator 186N has a particular value (e.g., 1 or 3) indicating that the synthesis support data 150N includes facial landmark data 320N, the VRF generator 144 generates a VRF 156N based on the facial landmark data 320N, as further described with reference to FIG. 6. The VRF generator 144 generates a VRF identifier 196N for the VRF 156N, as described with reference to FIG. 1, and adds the VRF identifier 196N to one or more VRF reference candidate identifiers 172 of the reference list 176N.
[0091] In response to determining that the VRF usage indicator 186N has a particular value (e.g., 2 or 3) indicating that the synthesis support data 150N includes motion-based data 322N, the VRF generator 144 generates a VRF 156NB based on the motion-based data 322N, as further described with reference to Figure 7. The VRF generator 144 generates a VRF identifier 196NB for the VRF 156NB, as described with reference to Figure 1, and adds the VRF identifier 196NB to one or more VRF reference candidate identifiers 172 in the reference list 176N.
[0092] A visual analytics engine 312 including both the facial landmark detector 304 and the global motion detector 306 is provided as an exemplary implementation. Optionally, in some implementations, the visual analytics engine 312 may include one of the facial landmark detector 304 or the global motion detector 306, and the synthesis support data 150N may include a corresponding one of the facial landmark data 320N or the motion-based data 322N. Technical advantages of a visual analytics engine 312 including one of the facial landmark detector 304 or the global motion detector 306 may include less hardware, less memory usage, fewer computing cycles, or a combination thereof, used by the visual analytics engine 312. Technical advantages of a visual analytics engine 312 including both the facial landmark detector 304 and the global motion detector 306 may include improved image frame reproduction quality, reduced use of transmission resources, or both, compared to one including one of the facial landmark detector 304 or the global motion detector 306. Another technical advantage of the visual analytics engine 312 including both the facial landmark detector 304 and the global motion detector 306 may include compatibility with decoders that include support for face VRF, motion VRF, or both.
[0093] 4, a diagram 400 of an example aspect of operations associated with the synthesis support analyzer 314 for generating the VRF usage indicator 186N of FIG. 1 is shown, in accordance with certain examples of the present disclosure. In a particular aspect, the synthesis support analyzer 314 initializes the VRF usage indicator 186N to a first value (e.g., 0) indicating no VRF usage.
[0094] At 402, the synthesis support analyzer 314 determines whether the encoder reference candidate count indicated by one or more encoder reference candidate identifiers 174 of FIG. 1 is less than a threshold reference count.
[0095] In response to determining 402 that the encoder reference candidate count is not less than (i.e., is greater than or equal to) the threshold reference count, the synthesis support analyzer 314 outputs 404 the VRF usage indicator 186N of Figure 1 having a first value (e.g., 0) indicating no VRF usage. Alternatively, in response to determining 402 that the encoder reference candidate count is less than the threshold reference count, the synthesis support analyzer 314 determines 406 whether the face detection indicator 318N of Figure 3 indicates that at least one face has been detected in the image frame 116N.
[0096] In response to determining that the face detection indicator 318N indicates that at least one face has been detected in the image frame 116N, the synthesis support analyzer 314 updates the VRF usage indicator 186N to a second value (e.g., 1) to indicate VRF usage for the face at 408. At 410, the synthesis support analyzer 314 determines whether the sum of the encoder reference candidate count and 1 is less than the threshold reference count.
[0097] In response to determining at 406 that the face detection indicator 318N indicates that no face is detected in the image frame 116N or determining at 410 that the sum of the encoder reference candidate count and 1 is less than the threshold reference count, the synthesis support analyzer 314 determines at 412 whether the motion detection indicator 316N of FIG. 3 indicates that global motion greater than a threshold has been detected in the image frame 116N.
[0098] In response to determining that the motion detection indicator 316N indicates that global motion greater than a threshold has been detected in the image frame 116N, the synthesis support analyzer 314 updates 412 the VRF use indicator 186N to indicate motion VRF use. For example, in response to determining that the VRF use indicator 186N has a first value (e.g., 0) indicating no facial VRF use, the synthesis support analyzer 314 sets the VRF use indicator 186N to a third value (e.g., 2) indicating motion VRF use and no facial VRF use. As another example, in response to determining that the VRF use indicator 186N indicates a second value (e.g., 1) indicating facial VRF use, the synthesis support analyzer 314 sets the VRF use indicator 186N to a fourth value (e.g., 3) to indicate motion VRF use in addition to facial VRF use.
[0099] Alternatively, the synthesis support analyzer 314 outputs a VRF usage indicator 186N indicating no motion VRF usage in response to determining 410 that the sum of the encoder reference candidate count and 1 is greater than or equal to the threshold reference count or determining 412 that the motion detection indicator 316N indicates that no global motion greater than the threshold is detected in the image frame 116N. For example, the synthesis support analyzer 314 refrains from updating the VRF usage indicator 186N having a first value (e.g., 0) indicating no VRF usage or having a second value (e.g., 1) indicating face VRF usage and no motion VRF usage.
[0100] Diagram 400 is an illustrative example of operations performed by synthesis support analyzer 314. Optionally, in some implementations, synthesis support analyzer 314 may generate VRF usage indicator 186N based on one of motion detection indicator 316N or face detection indicator 318N. Optionally, in some implementations in which VRF usage indicator 186N is based on face detection indicator 318N and not on motion detection indicator 316N, synthesis support analyzer 314 performs operations 402, 404, 406, and 408 and does not perform operations 410, 412, 414, and 416. To illustrate, in response to determining 402 that the encoder reference candidate count is less than the threshold reference count and determining 406 that the face detection indicator 318N indicates that at least one face has been detected in the image frame 116N, the synthesis support analyzer 314 outputs 408 a VRF usage indicator 186N having a second value (e.g., 1) indicating face VRF usage. Alternatively, in response to determining 402 that the encoder reference candidate count is greater than or equal to the threshold reference count or determining 406 that the face detection indicator 318N indicates that no face has been detected in the image frame 116N, the synthesis support analyzer 314 proceeds to 404 and outputs 404 a VRF usage indicator 186N having a first value (e.g., 0) indicating no VRF usage.
[0101] Optionally, in some implementations in which the VRF usage indicator 186N is based on the motion detection indicator 316N and not based on the face detection indicator 318N, the synthesis support analyzer 314 performs operations 402, 404, 412, and 414 and does not perform operations 406, 408, 410, and 416. Illustratively, in response to determining at 402 that the encoder reference candidate count is less than a threshold reference count and determining at 412 that the motion detection indicator 316N indicates that at least a threshold global motion has been detected in the image frame 116N, the synthesis support analyzer 314 outputs at 414 the VRF usage indicator 186N having a third value (e.g., 2) indicating motion VRF usage. Alternatively, in response to determining at 402 that the encoder reference candidate count is greater than or equal to the threshold reference count, or determining at 412 that the motion detection indicator 316N indicates that no global motion greater than the threshold is detected within the image frame 116N, the synthesis support analyzer 314 proceeds to 404 and outputs a VRF usage indicator 186N having a first value (e.g., 0) indicating no VRF usage.
[0102] 5, a diagram 500 of an example aspect of operations associated with VRF generator 144 is shown, in accordance with some examples of the present disclosure. VRF generator 144 includes a face VRF generator 504 and a motion VRF generator 506.
[0103] In response to determining that the VRF usage indicator 186N has a particular value (e.g., 1 or 3) indicating facial VRF usage, the face VRF generator 504 processes the image frame 116A (or a locally decoded version of the image frame 116A) based on the facial landmark data 320N to generate the VRF 156NA, as further described with reference to Figure 6. The face VRF generator 504 assigns a VRF identifier 196NA to the VRF 156NA and adds the VRF identifier 196NA to one or more VRF reference candidate identifiers 172 in the reference list 176N.
[0104] In response to determining that the VRF usage indicator 186N has a particular value (e.g., 2 or 3) indicating motion VRF usage, the motion VRF generator 506 processes the image frame 116A (or a locally decoded version of the image frame 116A) based on the motion-based data 322N to generate a VRF 156NB, as further described with reference to Figure 7. The motion VRF generator 506 assigns a VRF identifier 196NB to the VRF 156NB and adds the VRF identifier 196NB to one or more VRF reference candidate identifiers 172 in the reference list 176N.
[0105] A VRF generator 144 including both a face VRF generator 504 and a motion VRF generator 506 is provided as an illustrative example. Optionally, in some implementations, the VRF generator 144 may include one of the face VRF generator 504 or the motion VRF generator 506. Technical advantages of including one of the face VRF generator 504 or the motion VRF generator 506 may include less hardware, less memory usage, fewer computation cycles, or a combination thereof, used by the VRF generator 144. Technical advantages of a VRF generator 144 including both the face VRF generator 504 and the motion VRF generator 506 may include improved image frame reproduction quality, reduced use of transmission resources, or both, compared to including one of the facial landmark detector 304 or the global motion detector 306. Another technical advantage of the visual analytics engine 312 including both the facial landmark detector 304 and the global motion detector 306 may include compatibility with decoders that include support for face VRF, motion VRF, or both.
[0106] Referring to FIG. 6, a diagram 600 of an example aspect of operations associated with face VRF detector 504 and video encoder 146 is shown, in accordance with some examples of this disclosure.
[0107] In response to determining that the VRF usage indicator 186N has a particular value (e.g., 1 or 3) indicating facial VRF usage, the facial VRF generator 504 applies facial landmark data 320N to the image frame 116A (or a locally decoded version of the image frame 116A). For example, the facial landmark data 320N indicates the locations of facial features within the image frame 116N. A graphical representation of the facial landmark data 320N is shown in FIG. 6, which illustrates the locations of facial features detected in the image frame 116N. To illustrate, a person's eyes may be depicted in the image frame 116N as being more widely open relative to the depiction of the eyes in the image frame 116A.
[0108] Applying the facial landmark data 320N to the image frame 116A (or a locally decoded version of the image frame 116A) adjusts the positions of the facial features in the image frame 116A (or a locally decoded version of the image frame 116A) to generate the VRF 156NA as an estimate of the image frame 116N. Illustratively, the adjusted positions of the facial features in the VRF 156NA may more closely match the positions (or relative positions) of the facial features in the image frame 116N. In a particular implementation, the face VRF generator 504 generates a face model corresponding to the positions of the facial features detected in the image frame 116A. The face VRF generator 504 updates the face model based on the updated positions of the facial features indicated in the facial landmark data 320N. The face VRF generator 504 generates the VRF 156NA corresponding to the updated face model.
[0109] Facial landmark data 320N indicating positions of facial features detected in image frame 116N is provided as an illustrative example. Optionally, in some implementations, facial landmark data 320N indicates positions of facial features detected in image frame 116N that are separate from (e.g., updated to) the positions of facial features detected in image frame 116A.
[0110] In particular implementations, the face VRF generator 504 includes a trained model (e.g., a neural network) that the face VRF generator 504 uses to process the image frame 116A (or a locally decoded version of the image frame 116A) and the facial landmark data 320N to generate the VRF 156N.
[0111] The facial VRF generator 504 provides the VRF 156NA to the video encoder 146. The video encoder 146 determines residual data 604 based on a comparison of (e.g., the difference between) the image frame 116N and the VRF 156NA. The video encoder 146 generates coded bits 606N corresponding to the residual data 604. For example, the video encoder 146 encodes the residual data 604 to generate the coded bits 606N. The coded bits 606N are included as a first subset of the coded bits 166N of FIG. 1 associated with the facial VRF specification. In certain aspects, the facial landmark data 320N and the coded bits 606N correspond to fewer bits compared to an coded version of the first residual data based on the difference between the image frame 116A (or a locally decoded version of the image frame 116A) and the image frame 116N. In one example, the residual data 604 has smaller numerical values and an overall smaller variance compared to the first residual data, and therefore the residual data 604 may be encoded more efficiently (e.g., using fewer bits). A technical advantage of providing the facial landmark data 320N and the residual data 604 (instead of the first residual data) in the bitstream 135 may include using fewer resources (e.g., bandwidth, time, or both).
[0112] Referring to FIG. 7, a diagram 700 of an example aspect of operations associated with the motion VRF generator 506 and the day video encoder 146 is shown, in accordance with some examples of this disclosure.
[0113] In response to determining that the VRF usage indicator 186N has a particular value (e.g., 2 or 3) indicating motion VRF usage, the motion VRF generator 506 applies the motion-based data 322N to the image frame 116A (or a locally decoded version of the image frame 116A). For example, the motion-based data 322N indicates global motion (e.g., rotation, translation, or both) detected in the image frame 116N relative to the image frame 116A (or a locally decoded version of the image frame 116A). In another example, the motion-based data 322N indicates global motion of the camera moving left between the first capture time of the image frame 116A and the second capture time of the image frame 116N.
[0114] Applying the motion-based data 322N to the image frame 116A (or a locally decoded version of the image frame 116A) applies global motion to the image frame 116A (or a locally decoded version of the image frame 116A) to generate the VRF 156NB as an estimate of the image frame 116N. For example, the motion VRF generator 506 uses the motion-based data 322N to warp the image frame 116A (or a locally decoded version of the image frame 116A) to generate the VRF 156NB. In certain implementations, the motion VRF generator 506 includes a trained model (e.g., a neural network). The motion VRF generator 506 uses the trained model to process the image frame 116A (or a locally decoded version of the image frame 116A) and the motion-based data 322N to generate the VRF 156NB. For example, image frame 116A (or a locally decoded version of image frame 116A) and motion-based data 322N are provided as inputs to a trained model, and the output of the trained model represents VRF 156NB.
[0115] The motion VRF generator 506 provides the VRF 156NB to the video encoder 146. The video encoder 146 determines residual data 704 based on a comparison of (e.g., the difference between) the image frame 116N and the VRF 156NB. The video encoder 146 generates coded bits 706N corresponding to the residual data 704. For example, the video encoder 146 encodes the residual data 704 to generate coded bits 706N. The coded bits 706N are included as a second subset of the coded bits 166N of FIG. 1 associated with the motion VRF specification. In certain aspects, the motion-based data 322N and the coded bits 706N correspond to fewer bits compared to a coded version of the first residual data based on the difference between the image frame 116A (or a locally decoded version of the image frame 116A) and the image frame 116N. In one example, the residual data 704 has smaller numerical values and an overall smaller variance compared to the first residual data, and therefore the residual data 704 may be coded more efficiently (e.g., using fewer bits). A technical advantage of providing the motion-based data 322N and the residual data 704 (instead of the first residual data) in the bitstream 135 may include using fewer resources (e.g., bandwidth, time, or both).
[0116] 8, a diagram 800 of an example aspect of operations associated with VRF generator 244 is shown, in accordance with some examples of the present disclosure. VRF generator 244 includes a face VRF generator 804 and a motion VRF generator 806.
[0117] In response to determining that the VRF usage indicator 186N has a particular value indicating facial VRF usage (e.g., 1 or 3), the facial VRF generator 804 processes the image frame 216A based on the facial landmark data 320N to generate a VRF 256NA, as further described with reference to Figure 9. In response to determining that the reference list 176N includes a VRF identifier 196NA associated with a facial VRF usage, determining that the facial landmark data 320N is associated with the VRF identifier 196NA, or both, the facial VRF generator 804 assigns the VRF identifier 196NA to the VRF 256NA.
[0118] In response to determining that the VRF usage indicator 186N has a particular value indicating motion VRF usage (e.g., 2 or 3), the motion VRF generator 806 processes the image frame 216A based on the motion-based data 322N to generate a VRF 256NB, as described further with reference to Figure 10. In response to determining that the reference list 176N includes a VRF identifier 196NB associated with motion VRF usage, determining that the motion-based data 322N is associated with the VRF identifier 196NB, or both, the motion VRF generator 806 assigns the VRF identifier 196NB to the VRF 256NB.
[0119] A VRF generator 244 including both a face VRF generator 804 and a motion VRF generator 806 is provided as an illustrative example. Optionally, in some implementations, the VRF generator 244 may include one of the face VRF generator 804 or the motion VRF generator 806. Technical advantages of including one of the face VRF generator 804 or the motion VRF generator 806 may include less hardware, less memory usage, fewer computation cycles, or a combination thereof, used by the VRF generator 244. Technical advantages of a VRF generator 244 including both the face VRF generator 804 and the motion VRF generator 806 may include improved image frame reproduction quality, reduced use of transmission resources, or both, compared to including one of the face VRF generator 804 or the motion VRF generator 806. Another technical advantage of the VRF generator 244, which includes both a face VRF generator 804 and a motion VRF generator 806, may include compatibility with encoders that include support for face VRFs, motion VRFs, or both.
[0120] Referring to FIG. 9, a diagram 900 of an example aspect of operations associated with the face VRF generator 804 and the video decoder 246 is shown, in accordance with some examples of this disclosure.
[0121] In response to determining that the VRF usage indicator 186N has a particular value (eg, 1 or 3) indicating facial VRF usage, the facial VRF generator 804 applies the facial landmark data 320N to the image frame 216A.
[0122] Applying the facial landmark data 320N to the image frame 216A adjusts the positions of the facial landmarks in the image frame 216A to more closely match the positions (or relative positions) of the facial landmarks in the image frame 116N to generate the VRF 256NA. In certain aspects, the face VRF generator 804 generates a face model corresponding to the positions of the facial landmarks detected in the image frame 216A. The face VRF generator 804 updates the face model based on the updated positions of the facial landmarks indicated in the facial landmark data 320N. The face VRF generator 804 generates the VRF 256NA corresponding to the updated face model.
[0123] In certain implementations, the face VRF generator 804 includes a trained model (e.g., a neural network) that the face VRF generator 804 uses to process the image frame 216A and the facial landmark data 320N to generate the VRF 256N.
[0124] The facial VRF generator 804 provides the VRF 256NA to the video decoder 246. The video decoder 246 decodes the coded bits 606N (e.g., a first subset of the coded bits 166N associated with the facial VRF specification) to generate the residual data 604. The facial VRF generator 804 generates the image frame 216N based on a combination of the VRF 256NA and the residual data 604. In certain aspects, the facial landmark data 320N and the coded bits 606N correspond to fewer bits compared to an encoded version of the first residual data based on the differences between the image frame 216A and the image frame 116N. A technical advantage of generating the image frame 216N using the facial landmark data 320N and the residual data 604 may include generating an image frame 216N that is a better approximation of the image frame 116N using limited bits of the bitstream 135.
[0125] Referring to FIG. 10, a diagram 1000 of an example aspect of operations associated with the motion VRF generator 806 and the video decoder 246 is shown, in accordance with some examples of this disclosure.
[0126] The motion VRF generator 806 applies the motion-based data 322N to the image frame 216A in response to determining that the VRF usage indicator 186N has a particular value (eg, 2 or 3) indicating motion VRF usage.
[0127] Applying the motion-based data 322N to the image frame 216A applies global motion to the image frame 216A to generate VRF 256NB. For example, the motion VRF generator 806 warps the image frame 216A based on the motion-based data 322N to generate VRF 256NB. In particular implementations, the motion VRF generator 806 includes a trained model (e.g., a neural network). The motion VRF generator 806 uses the trained model to process the image frame 216A and the motion-based data 322N to generate VRF 256NB. For example, the motion VRF generator 806 provides the image frame 216A and the motion-based data 322N as inputs to the trained model, and the output of the trained model represents VRF 256NB.
[0128] The motion VRF generator 806 provides the VRF 256NB to the video decoder 246. The video decoder 246 decodes the coded bits 706N (e.g., a second subset of the coded bits 166N associated with the motion VRF specification) to generate the residual data 704. The motion VRF generator 806 generates the image frame 216N based on a combination of the VRF 256NB and the residual data 704. In certain aspects, the motion-based data 322N and the coded bits 706N correspond to fewer bits compared to an encoded version of the first residual data based on the differences between the image frame 216A and the image frame 116N. A technical advantage of generating the image frame 116N using the motion-based data 322N and the residual data 704 may include generating an image frame 216N that is a better approximation of the image frame 216N using limited bits of the bitstream 135.
[0129] Generating the image frame 216N based on either VRF256NA corresponding to the facial landmark data 320N as described with reference to Figure 9 or VRF256NB corresponding to the motion-based data 322N as described with reference to Figure 10 is provided as an illustrative example. Optionally, in some implementations, the video decoder 246 generates the image frame 216N based on both the facial landmark data 320N and the motion-based data 322N. As an illustrative example, the video decoder 246 applies the facial landmark data 320N to the image frame 216A to generate VRF256NA, and applies the motion-based data 322N to VRF256NA to generate VRF256NB, as described with reference to Figure 9. The video decoder 246 applies the residual data 704 to VRF156NB to generate the image frame 216N. In this example, the video encoder 146 applies facial landmark data 320N to image frame 116A to generate VRF156NA, determines motion-based data 322N based on a comparison between VRF156NA and image frame 116N, applies motion-based data 322N to VRF156NA to generate VRF156NB, and determines residual data 704 based on a comparison between VRF156NB and image frame 116N, as described with reference to FIG. 6.
[0130] Referring to FIG. 11, a diagram 1100 of an example aspect of operations associated with the frame analyzer 142, the VRF generator 144, and the video encoder 146 is shown, in accordance with some examples of this disclosure.
[0131] Each of the frame analyzer 142 and the video encoder 146 is configured to receive a sequence of image frames 116, such as a sequence of consecutively captured frames of image data, shown as a first image frame (F1) 116A, a second image frame (F2) 116B, and one or more additional image frames including an Nth image frame (FN) 116N, where N is an integer greater than 2. The frame analyzer 142 is configured to output a sequence of VRF usage indicators, including a first VRF usage indicator (V1) 186A, a second VRF usage indicator (V2) 186B, and one or more additional VRF usage indicators including an Nth VRF usage indicator (VN) 186N. The frame analyzer 142 is also configured to output a corresponding set of synthetic support data 150 denoted as second synthetic support data (S2) 150B and one or more additional sets of synthetic support data including Nth synthetic support data (SN) 150N when the VRF usage indicator 186 has a particular value (e.g., 1, 2, or 3) indicating VRF usage.
[0132] The VRF generator 144 is configured to receive a sequence of VRF usage indicators and a corresponding set of synthesis support data. The VRF generator 144 is configured to selectively generate, based on the synthesis support data, one or more VRFs 156 denoted as one or more second VRFs (R2) 156B and one or more additional sets of VRFs including one or more Nth VRFs (RN) 156N.
[0133] The video encoder 146 is configured to generate a sequence of coded bits 166 and a sequence of reference lists 176 corresponding to the sequence of image frames 116. The sequence of coded bits 166 is shown as one or more additional sets of coded bits including a first coded bit (E1) 166A, a second coded bit (E2) 166B, and an Nth coded bit (EN) 166N. The sequence of reference lists 176 is shown as one or more additional reference lists including a first reference list (L1) 176A, a second reference list (L2) 176B, and an Nth reference list (LN) 176N. The video encoder 146 is configured to selectively generate the one or more sets of coded bits 166 based on the corresponding VRF 156 and output corresponding synthesis support data.
[0134] During operation, the frame analyzer 142 processes the first image frame (F1) 116A to generate a first VRF usage indicator (V1) 186A. In response to determining that the first VRF usage indicator (V1) 186A has a particular value (e.g., 0) indicating no VRF usage, the frame analyzer 142 refrains from generating corresponding synthesis support data. In response to determining that the first VRF usage indicator (V1) 186A has a particular value (e.g., 0) indicating no VRF usage, the VRF generator 144 refrains from generating a VRF associated with the first image frame (F1) 116A. In response to determining that the first VRF usage indicator (V1) 186A has a particular value (e.g., 0) indicating no VRF usage, the video encoder 146 generates a first coded bit (E1) 166A independent of any VRF. The video encoder 146 outputs a first coded bit (E1) 166A and a first reference list (L1) 176A. In a particular example, the video encoder 146 generates the first coded bit (E1) 166A without reference to any reference frame, and the reference list 176A is empty. In another example, the video encoder 146 generates the first coded bit (E1) 166A based on a previous frame in the sequence of image frames 116, and the reference list 176A indicates the previous frame.
[0135] The frame analyzer 142 processes the second image frame (F2) 116B to generate a second VRF usage indicator (V2) 186B. In response to determining that the second VRF usage indicator (V2) 186B has a particular value (e.g., 1, 2, or 3) indicating VRF usage, the frame analyzer 142 generates second synthesis support data (S2) 150B for the second image frame (F2) 116B. In response to determining that the second VRF usage indicator (V2) 186B has a particular value (e.g., 1, 2, or 3) indicating VRF usage, the VRF generator 144 generates one or more second VRFs (R2) 156B associated with the second image frame (F2) 116B. In response to determining that the second VRF usage indicator (V2) 186B has a particular value (e.g., 1, 2, or 3) indicating VRF usage, the video encoder 146 generates second coded bits (E2) 166B based on one or more second VRFs (R2) 156B. The video encoder 146 outputs the second coded bits (E2) 166B, second synthesis support data (S2) 150B, and a second reference list (L2) 176B. The reference list 176B includes one or more VRF identifiers for the one or more second VRFs 156B. In some examples, the reference list 176B also includes one or more identifiers of one or more previous frames in the sequence of image frames 116 that may be used as reference frames. In some examples, the second coded bits (E2) 166B include one or more subsets of coded bits corresponding to one or more reference frames indicated in the reference list 176B.
[0136] Similarly, the frame analyzer 142 processes the Nth image frame (FN) 116N to generate an Nth VRF usage indicator (VN) 186N. In response to determining that the Nth VRF usage indicator (VN) 186N has a particular value indicating VRF usage (e.g., 1, 2, or 3), the frame analyzer 142 generates an Nth synthetic support data (SN) 150N for the Nth image frame (FN) 116N. In response to determining that the Nth VRF usage indicator (VN) 186N has a particular value indicating VRF usage (e.g., 1, 2, or 3), the VRF generator 144 generates one or more Nth VRFs (RN) 156N associated with the Nth image frame (FN) 116N.
[0137] In response to determining that the Nth VRF usage indicator (VN) 186N has a particular value (e.g., 1, 2, or 3) indicating VRF usage, the video encoder 146 generates an Nth coded bit (EN) 166N based on one or more Nth VRFs (RN) 156N. The video encoder 146 outputs an Nth coded bit (EN) 166N, an Nth synthesis support data (SN) 150N, and an Nth reference list (LN) 176N. The reference list 176N includes one or more VRF identifiers for one or more Nth VRFs (RN) 156N. In some examples, the reference list 176B may also include one or more identifiers of one or more previous frames in the sequence of image frames 116 that may be used as reference frames. In some examples, the Nth coded bit (EN) 166N includes one or more subsets of coded bits corresponding to one or more reference frames indicated in the reference list 176N.
[0138] By dynamically generating coded bits based on a virtual reference frame, decoding accuracy may be improved for image frames for which synthetic support data (e.g., facial data, motion-based data, or both) may be generated.
[0139] Referring to FIG. 12, a diagram 1200 of an example aspect of operations associated with VRF generator 244 and video decoder 246 is shown, in accordance with some examples of this disclosure.
[0140] The VRF generator 244 is configured to receive a set of synthesized support data and generate a corresponding set of VRFs. The set of synthesized support data is shown as one or more additional sets of synthesized support data including second synthesized support data (S2) 150B and Nth synthesized support data (SN) 150N. The set of VRFs is shown as one or more additional sets of VRFs including one or more second VRFs (R2) 256B and one or more Nth VRFs (RN) 256N.
[0141] The video decoder 246 is configured to receive a sequence of coded bits 166 and a sequence of reference lists 176. The sequence of coded bits 166 is shown as one or more additional sets of coded bits including a first coded bit (E1) 166A, a second coded bit (E2) 166B, and an Nth coded bit (EN) 166N. The sequence of reference lists 176 is shown as one or more additional reference lists including a first reference list (L1) 176A, a second reference list (L2) 176B, and an Nth reference list (LN) 176N.
[0142] The video decoder 246 is configured to generate a sequence of decoded image frames 216 based on the sequence of coded bits 166 and the sequence of reference list 176. The sequence of decoded image frames 216 is shown as a first image frame (D1) 216A, a second image frame (D2) 216B, and one or more additional image frames including an Nth image frame (DN) 216N. The video decoder 246 is configured to selectively generate the decoded image frames based on corresponding VRFs 256.
[0143] During operation, the video decoder 246 processes the first coded bits (E1) 166A based on the first reference list (L1) 176A to generate a first image frame (D1) 216A. In response to determining that the first reference list (L1) 176A indicates no VRF associated with the first coded bits (E1) 166A, the video decoder 246 generates the first image frame (D1) 216A without any VRF association. In a particular implementation, the video decoder 246 receives a sequence of VRF usage indicators 186. In this implementation, the video decoder 246 generates the first image frame (D1) 216A without any VRF association in response to determining that the first VRF usage indicator (V1) 186A has a particular value (e.g., 0) indicating no VRF association.
[0144] The VRF generator 244 processes the second synthesis support data (S2) 150B to generate one or more second VRFs (R2) 256B. The video decoder 246 processes the second coded bits (E2) 166B based on the second reference list (L2) 176B to generate a second image frame (D2) 216B. In response to determining that the second reference list (L2) 176B indicates identifiers of one or more second VRFs (R2) 256B associated with the second coded bits (E2) 166B, the video decoder 246 generates the second image frame (D2) 216B based on the one or more second VRFs (R2) 256B.
[0145] Similarly, the VRF generator 244 processes the Nth synthesis support data (SN) 150N to generate one or more Nth VRFs (RN) 256N. The video decoder 246 processes the Nth coded bits (EN) 166N based on the Nth reference list (LN) 176N to generate the Nth image frame (DN) 216N. In response to determining that the Nth reference list (LN) 176N indicates identifiers of one or more Nth VRFs (RN) 256N associated with the Nth coded bits (EN) 166N, the video decoder 246 generates the Nth image frame (DN) 216N based on the one or more Nth VRFs (RN) 256N.
[0146] By dynamically generating decoded image frames based on a virtual reference frame, the accuracy of decoding of image frames (e.g., the second image frame (D2) 216B and the Nth image frame (DN) 216N) for which synthetic support data (e.g., face data, motion-based data, or both) is available can be improved.
[0147] 13 illustrates an implementation 1300 of device 102 as an integrated circuit 1302 that includes one or more processors 1390. In particular aspects, one or more processors 1390 include one or more processors 190, one or more processors 290, or a combination thereof. The integrated circuit 1302 also includes a signal input 1304, such as one or more bus interfaces, that allows input data 1328 to be received for processing. The integrated circuit 1302 includes a video analyzer 140, a video generator 240, or both. The integrated circuit 1302 also includes a signal output 1306, such as a bus interface, that allows output data 1330 to be transmitted. In particular examples, the input data 1328 includes image frames 116, and the output data 1330 includes reference list 176, coded bits 166, VRF usage indicator 186, synthesis support data 150, bitstream 135, or a combination thereof. In another example, the input data 1328 includes the reference list 176, the coded bits 166, the VRF usage indicator 186, the synthesis support data 150, the bitstream 135, or a combination thereof, and the output data 1330 includes the image frame 216.
[0148] The integrated circuit 1302 enables the implementation of image encoding and decoding based on a virtual reference frame as a component in a system such as a mobile phone or tablet as depicted in FIG. 14, a wearable electronic device as depicted in FIG. 15, a camera as depicted in FIG. 16, a virtual reality headset, a mixed reality headset, or an augmented reality headset shown in FIG. 17, or a vehicle as depicted in FIG. 18 or FIG. 19.
[0149] 14 shows, as an illustrative, non-limiting example, an implementation 1400 in which device 102, device 160, or both, include a mobile device 1402, such as a phone or tablet. Mobile device 1402 includes camera 110 and display screen 1404. In particular aspects, display screen 1404 may correspond to display device 210 of FIG. 2. Components of one or more processors 190 and one or more processors 290, including video analyzer 140 and video generator 240, are incorporated within mobile device 1402 and are shown using dashed lines to indicate internal components that are not normally visible to a user of mobile device 1402. In particular examples, video analyzer 140 operates to detect image frames 116 or bitstream 135, which are then processed to perform one or more actions on mobile device 1402, such as to launch a graphical user interface or, possibly, display other information on display screen 1404 (e.g., via an integrated “smart assistant” application). For example, display screen 1404 may show image frame 116 being processed to generate bitstream 135 or bitstream 135 being processed to generate image frame 216 .
[0150] 15 shows an implementation 1500 in which the device 102, the device 160, or both, include a wearable electronic device 1502, designated as a "smart watch." The video analyzer 140, the video generator 240, the camera 110, or a combination thereof, is incorporated into the wearable electronic device 1502.
[0151] In particular examples, the video analyzer 140 or the video generator 240 operates to detect image frames 116 or bitstream 135, respectively, which are then processed to perform one or more actions on the wearable electronic device 1502, such as to launch a graphical user interface or possibly display other information on the display screen 1504. For example, the display screen 1504 indicates that the image frames 116 have been processed to generate the bitstream 135, that the bitstream 135 has been processed to generate the image frames 216, or that the generated image frames 216 are used to play out, such as in the example of streaming video.
[0152] In particular examples, the wearable electronic device 1502 includes a haptic device that provides a haptic notification (e.g., vibrates) in response to detecting an image frame 116 or a bitstream 135. For example, the haptic notification may cause the user to look at the wearable electronic device 1502 to see a displayed notification indicating the processing of the image frame 116 to generate a bitstream 135 available for transmission to another user, or the processing of the bitstream 135 to generate an image frame 216 available for viewing. The wearable electronic device 1502 may thus alert a user who is hearing impaired or wearing a headset that a bitstream 135 is available for transmission or that an image frame 216 is available for viewing.
[0153] 16 illustrates an implementation 1600 in which device 102, device 160, or both, include a portable electronic device corresponding to camera device 1602. Video analyzer 140, video generator 240, or both are included in camera device 1602. In certain aspects, camera device 1602 corresponds to or includes camera 110 of FIG. 1. In operation, in response to receiving a verbal command identified as a user utterance, camera device 1602 can perform an action according to the verbal user command, such as to adjust image or video capture settings, image or video playback settings, image or video capture instructions, generate bitstream 135 based on image frames 116, or process bitstream 135 to display image frames 216 on a display screen, as illustrative examples.
[0154] 17 illustrates an implementation 1700 in which device 102, device 160, or both include a portable electronic device corresponding to a virtual reality, mixed reality, or augmented reality headset 1702. A video analyzer 140, a video generator 240, a camera 110, or a combination thereof is incorporated into the headset 1702. User voice activity detection may be performed based on an audio signal received from a microphone of the headset 1702. A visual interface device is positioned in front of the user's eyes to enable augmented reality, mixed reality, or virtual reality images or scenes to be displayed to the user while the headset 1702 is being worn. In a particular example, the visual interface device is configured to display a notification indicating the processing of image frames 116 to generate bitstream 135, or to display a notification indicating the processing of bitstream 135 to generate image frames 216, or is used for playout of generated image frames 216, such as in a streaming video example.
[0155] FIG. 18 illustrates an implementation 1800 in which device 102, device 160, or both correspond to or are incorporated into a vehicle 1802, which may be depicted as a manned or unmanned aerial device (e.g., a package delivery drone). A video analyzer 140, a video generator 240, a camera 110, or a combination thereof, is incorporated into vehicle 1802. User voice activity detection may be performed based on an audio signal received from a microphone of vehicle 1802, such as for delivery instructions from an authorized user of vehicle 1802. In a particular example, vehicle 1802 includes a visual interface device configured to display a notification indicating the processing of image frame 116 to generate bitstream 135 or the processing of bitstream 135 to generate image frame 216. In certain aspects, image frame 116 corresponds to an image of a recipient of the package, an image of assembly or installation of a delivered product, or a combination thereof. In certain aspects, image frame 216 corresponds to assembly or installation instructions.
[0156] 19 shows another implementation 1900 in which device 102, device 160, or both, correspond to or are incorporated into a vehicle 1902, shown as a car. Vehicle 1902 includes one or more processors 1390 that include video analyzer 140, video generator 240, or both. Vehicle 1902 also includes camera 110. User voice activity detection may be performed based on audio signals received from a microphone of vehicle 1902. In some implementations, user voice activity detection may be performed based on audio signals received from an internal microphone, such as for voice commands from an authorized occupant. In some implementations, user voice activity detection may be performed based on audio signals received from an external microphone, such as for an authorized user of the vehicle. In particular implementations, in response to receiving a verbal command identified as a user utterance, the voice-activated system initiates one or more actions of the vehicle 1902 based on one or more keywords (e.g., "unlock," "start engine," "play music," "show weather," "play video," "send video," or another voice command), such as by providing feedback or information via the display 1920 or one or more speakers. Illustratively, the display 1920 may provide information indicating that the image frames 116 have been processed to generate a bitstream 135 ready for transmission, that the bitstream 135 has been processed to generate the image frames 216 ready for display, or that will be used to play out the generated image frames 216, such as in the example of streaming video.
[0157] 20, shown is a particular implementation of a method 2000 of image encoding using a virtual reference frame. In particular aspects, one or more operations of the method 2000 are performed by at least one of the frame analyzer 142, the VRF generator 144, the video encoder 146, the video analyzer 140, the one or more processors 190, the device 102, the system 100, or a combination thereof of FIG.
[0158] The method 2000 includes obtaining synthetic support data associated with an image frame of the sequence of image frames, at 2002. For example, the frame analyzer 142 of Figure 1 obtains synthetic support data 150N associated with image frame 116N of the sequence of image frames 116, as described with reference to Figures 1 and 3.
[0159] The method 2000 also includes selectively generating virtual reference frames based on the synthetic support data, at 2004. For example, the VRF generator 144 of FIG. 1 selectively generates one or more VRFs 156N based on the synthetic support data 150N, as described with reference to FIGS. 1 and 3-7.
[0160] The method 2000 further includes generating a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame, at 2006. For example, the video encoder 146 of FIG. 1 generates the bitstream 135 corresponding to an encoded version of the image frame 116N based at least in part on one or more VRFs 156N, as described with reference to FIGS. 1, 6, and 7.
[0161] Thus, method 2000 enables generating VRFs 156 that preserve perceptually important features (e.g., facial landmarks). A technical advantage of generating one or more VRFs 156N using synthesis support data 150N (e.g., facial landmark data, motion-based data, or both) may include improved video quality of the decoded image frames by generating one or more VRFs 156N that are closer approximations of image frame 116N.
[0162] The method 2000 of Figure 20 may be performed by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a digital signal processor (DSP), a controller, another hardware device, a firmware device, or any combination thereof. As an example, the method 2000 of Figure 20 may be performed by a processor executing instructions such as those described with reference to Figure 22.
[0163] 21, illustrated is a particular implementation of a method 2100 of image decoding using a virtual reference frame. In particular aspects, one or more operations of the method 2100 are performed by at least one of the device 160, the system 100 of FIG. 1, the bitstream analyzer 242, the VRF generator 244, the video decoder 246, the video generator 240, the one or more processors 290 of FIG. 2, or a combination thereof.
[0164] Method 2100 includes obtaining a bitstream corresponding to an encoded version of an image frame, at 2102. For example, bitstream analyzer 242 of Figure 2 obtains bitstream 135 corresponding to an encoded version of image frame 116N, as described with reference to Figure 2.
[0165] Method 2100 also includes, at 2104, generating virtual reference frames based on synthesized support data included in the bitstream based on a determination that the bitstream includes a virtual reference frame usage indicator. For example, VRF generator 244 of FIG. 2 generates one or more VRFs 256N based on synthesized support data 150N included in bitstream 135, as described with reference to FIG. 2, in response to determining that bitstream 135 includes a VRF usage indicator 186N having a particular value (e.g., 1, 2, or 3) indicating VRF usage.
[0166] The method 2100 further includes generating a decoded version of the image frame based on the virtual reference frame, at 2106. For example, the video decoder 246 of FIG. 2 generates the image frame 216N (e.g., the decoded version of the image frame 116N) based on one or more VRFs 256N, as described with reference to FIG.
[0167] Thus, method 2100 enables generating a decoded image frame (e.g., image frame 216N) using a VRF 256 that preserves perceptually significant features (e.g., facial landmarks). A technical advantage of generating one or more VRFs 256N using synthesis support data 150N (e.g., facial landmark data, motion-based data, or both) may include improved video quality of image frame 116N due to the use of one or more VRFs 256N that are closer approximations of image frame 216N.
[0168] The method 2100 of Figure 21 may be performed by an FPGA device, an ASIC, a processing unit such as a CPU, a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, the method 2100 of Figure 21 may be performed by a processor executing instructions such as those described with reference to Figure 22.
[0169] 22, a block diagram of a particular example implementation of a device is shown, generally designated 2200. In various implementations, device 2200 may have more or fewer components than those shown in FIG. 22. In an example implementation, device 2200 may correspond to device 102, receiving device 160 of FIG. 1, or both. In an example implementation, device 2200 may perform one or more of the operations described with reference to FIGS. 1-21.
[0170] In particular implementations, device 2200 includes a processor 2206 (e.g., a CPU). Device 2200 may include one or more additional processors 2210 (e.g., one or more DSPs). In particular aspects, one or more processors 190 of FIG. 1 correspond to processor 2206, processor 2210, or a combination thereof. In particular aspects, one or more processors 290 of FIG. 2 correspond to processor 2206, processor 2210, or a combination thereof. Processor 2210 may include a speech and music coder-decoder (codec) 2208, which includes a voice coder (“vocoder”) encoder 2236, a vocoder decoder 2238, or both. Processor 2210 may include video analyzer 140, video generator 240, or both.
[0171] Device 2200 may include memory 2286 and codec 2234. Memory 2286 may include instructions 2256 executable by one or more additional processors 2210 (or processor 2206) to perform the functionality described with reference to video analyzer 140, video generator 240, or both. Device 2200 may include modem 2270 coupled to antenna 2252 via transceiver 2250. In certain aspects, modem 2270 includes modem 170 of FIG. 1, modem 270 of FIG. 2, or both.
[0172] The device 2200 may include a display 2228 coupled to a display controller 2226. In certain aspects, the display 2228 includes the display device 210 of FIG. 2. The speaker 2292, the microphone 2212, the camera 110, or a combination thereof may be coupled to the codec 2234. The codec 2234 may include a digital-to-analog converter (DAC) 2202, an analog-to-digital converter (ADC) 2204, or both. In certain implementations, the codec 2234 may receive an analog signal from the microphone 2212, convert the analog signal to a digital signal using the analog-to-digital converter 2204, and provide the digital signal to the speech and music codec 2208. The speech and music codec 2208 may process the digital signal. In certain implementations, the speech and music codec 2208 can provide the digital signal to the codec 2234. The codec 2234 can convert the digital signal to an analog signal using a digital-to-analog converter 2202 and can provide the analog signal to a speaker 2292 .
[0173] In certain implementations, the device 2200 may be included in a system-in-package or system-on-chip device 2222. In certain implementations, the memory 2286, the processor 2206, the processor 2210, the display controller 2226, the codec 2234, and the modem 2270 are included in the system-in-package or system-on-chip device 2222. In certain implementations, the input device 2230 and the power supply 2244 are coupled to the system-in-package or system-on-chip device 2222. Furthermore, in certain implementations, the display 2228, the camera 110, the input device 2230, the speaker 2292, the microphone 2212, the antenna 2252, and the power supply 2244 are external to the system-in-package or system-on-chip device 2222, as shown in FIG. In particular implementations, each of the display 2228, camera 110, input device 2230, speaker 2292, microphone 2212, antenna 2252, and power supply 2244 may be coupled to a component of the system-in-package or system-on-chip device 2222, such as an interface or controller.
[0174] Device 2200 may include a smart speaker, a speaker bar, a mobile communication device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aviation vehicle, a home automation system, a voice-activated device, a wireless speaker and a voice-activated device, a portable electronic device, an automobile, a computing device, a communication device, an internet-of-thing (IoT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.
[0175] In accordance with the described implementation, the apparatus includes means for acquiring synthesis support data associated with an image frame of the sequence of image frames. For example, the means for acquiring synthesis support data can correspond to frame analyzer 142, video analyzer 140, modem 170, one or more processors 190, device 102, system 100, face detector 302, facial landmark detector 304, global motion detector 306, vision analysis engine 312, modem 2270, transceiver 2250, antenna 2252, processor 2206, processor 2210, device 2200, one or more other circuits or components configured to acquire synthesis support data, or any combination thereof.
[0176] The apparatus also includes means for selectively generating a virtual reference frame based on the synthetic support data. For example, the means for selectively generating a virtual reference frame may correspond to VRF generator 144, video analyzer 140, one or more processors 190, device 102, system 100, face VRF generator 504, motion VRF generator 506, processor 2206, processor 2210, device 2200, one or more other circuits or components configured to selectively generate a virtual reference frame, or any combination thereof, of FIG.
[0177] The apparatus further includes means for generating a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame. For example, the means for generating the bitstream may correspond to video encoder 146, video analyzer 140, modem 170, one or more processors 190, device 102, system 100, modem 2270, transceiver 2250, antenna 2252, processor 2206, processor 2210, device 2200, one or more other circuits or components configured to generate the bitstream, or any combination thereof, of FIG.
[0178] Also, in accordance with the described implementation, the apparatus includes means for obtaining a bitstream corresponding to the encoded version of the image frame. For example, the means for obtaining the bitstream may correspond to device 160, system 100, modem 270, bitstream analyzer 242, video generator 240, one or more processors 290, modem 2270, transceiver 2250, antenna 2252, processor 2206, processor 2210, device 2200, one or more other circuits or components configured to obtain the bitstream, or any combination thereof, of FIG.
[0179] The apparatus also includes means for generating a virtual reference frame based on synthesis support data included in the bitstream, the virtual reference frame being generated based on a determination that the bitstream includes a virtual reference frame usage indicator. For example, the means for generating a virtual reference frame may correspond to device 160, system 100 of FIG. 1, VRF generator 244, video generator 240, one or more processors 290, processor 2206, processor 2210, device 2200, one or more other circuits or components configured to generate a virtual reference frame, or any combination thereof.
[0180] The apparatus further includes means for generating a decoded version of the image frame based on the virtual reference frame. For example, the means for generating the virtual reference frame may correspond to device 160, system 100 of FIG. 1, VRF generator 244, video generator 240, one or more processors 290 of FIG. 2, processor 2206, processor 2210, device 2200, one or more other circuits or components configured to generate the virtual reference frame, or any combination thereof.
[0181] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 2286) includes instructions (e.g., instructions 2256) that, when executed by one or more processors (e.g., one or more processors 190, one or more processors 2210, or processor 2206), cause the one or more processors to obtain synthetic support data (e.g., synthetic support data 150N) associated with an image frame (e.g., image frame 116N) of a sequence of image frames (e.g., image frame 116). The instructions, when executed by the one or more processors, also cause the one or more processors to selectively generate virtual reference frames (e.g., one or more VRFs 156N) based on the synthetic support data. The instructions, when executed by the one or more processors, further cause the one or more processors to generate a bitstream (e.g., bitstream 135) corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.
[0182] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 2286) includes instructions (e.g., instructions 2256) that, when executed by one or more processors (e.g., one or more processors 290, one or more processors 2210, or processor 2206), cause the one or more processors to obtain a bitstream (e.g., bitstream 135) corresponding to an encoded version of an image frame (e.g., image frame 116N). The instructions, when executed by the one or more processors, also cause the one or more processors to generate a virtual reference frame (e.g., one or more VRFs 256N) based on synthesis support data (e.g., synthesis support data 150N) included in the bitstream based on a determination that the bitstream includes a virtual reference frame usage indicator (e.g., VRF usage indicator 186N). The instructions, when executed by one or more processors, further cause the one or more processors to generate a decoded version of the image frame based on the virtual reference frame.
[0183] Certain aspects of the present disclosure are described below in a set of interrelated examples.
[0184] According to Example 1, a device includes one or more processors configured to obtain a bitstream corresponding to an encoded version of an image frame, and based on a determination that the bitstream includes a virtual reference frame usage indicator, generate a virtual reference frame based on synthetic support data included in the bitstream, and generate a decoded version of the image frame based on the virtual reference frame.
[0185] Example 2 includes the device of example 1, wherein the synthesis support data includes facial landmark data, motion-based data, or a combination thereof.
[0186] Example 3 includes the device of example 1 or example 2, wherein the bitstream indicates a first set of reference candidates that includes a virtual reference frame.
[0187] Example 4 includes the device of example 3, wherein the bitstream indicates one or more additional first sets of reference candidates including one or more additional virtual reference frames associated with one or more additional image frames of the sequence of image frames.
[0188] Example 5 includes the device of any of examples 1 to 4, wherein the bitstream further indicates a second set of reference candidates including one or more previously decoded image frames.
[0189] Example 6 includes the device of any of examples 1 to 5, wherein the bitstream includes a supplemental enhancement information (SEI) message indicating synthesis support data.
[0190] Example 7 includes the device of any of Examples 1 to 6, wherein the synthesis support data includes facial landmark data indicating locations of facial features, and wherein the one or more processors are configured to generate a virtual reference frame based at least in part on previously decoded image frames and the locations of the facial features.
[0191] Example 8 includes the device of any of Examples 1 to 7, wherein the synthesis support data includes motion-based data indicative of global motion, and wherein the one or more processors are configured to generate the virtual reference frame based at least in part on the previously decoded image frame and the global motion.
[0192] Example 9 includes the device of any of Examples 1 to 8, wherein the one or more processors are configured to warp previously decoded image frames using motion-based data to generate a virtual reference frame, and the synthesis support data includes the motion-based data.
[0193] Example 10 includes the device of any of Examples 1 to 9, wherein the one or more processors are configured to use the trained model to generate the virtual reference frame.
[0194] Example 11 includes the device of example 10, wherein the trained model includes a neural network.
[0195] Example 12 includes the device of example 10 or example 11, wherein the input to the trained model includes synthetic support data and at least one previously decoded image frame.
[0196] Example 13 includes the device of any of Examples 1 to 12, further including a modem configured to receive the bitstream from the second device.
[0197] Example 14 includes the device of any of Examples 1 to 13, further including a display device configured to display a decoded version of the image frame.
[0198] According to Example 15, a method includes, at a device, obtaining a bitstream corresponding to an encoded version of an image frame; based on a determination that the bitstream includes a virtual reference frame usage indicator, generating a virtual reference frame based on synthetic support data included in the bitstream; and generating, at the device, a decoded version of the image frame based on the virtual reference frame.
[0199] Example 16 includes the method of example 15, wherein the synthesis support data includes facial landmark data, motion-based data, or a combination thereof.
[0200] Example 17 includes the method of example 15 or example 16, wherein the bitstream indicates a first set of reference candidates that includes a virtual reference frame.
[0201] Example 18 includes the method of Example 17, wherein the bitstream indicates one or more additional first sets of reference candidates including one or more additional virtual reference frames associated with one or more additional image frames of the sequence of image frames.
[0202] Example 19 includes the method of any of examples 15-18, wherein the bitstream further indicates a second set of reference candidates including one or more previously decoded image frames.
[0203] Example 20 includes the method of any of examples 15 to 19, wherein the bitstream includes a supplemental enhancement information (SEI) message indicating the synthesis support data.
[0204] Example 21 includes the method of any of Examples 15 to 20, further including generating a virtual reference frame based at least in part on the previously decoded image frame and the positions of the facial features, wherein the synthesis support data includes facial landmark data indicating the positions of the facial features.
[0205] Example 22 includes the method of any of Examples 15 to 21, further including generating a virtual reference frame based at least in part on the previously decoded image frame and the global motion, wherein the synthesis support data includes motion-based data indicative of the global motion.
[0206] Example 23 includes the method of any of Examples 15 to 22, further including warping the previously decoded image frame using motion-based data to generate the virtual reference frame, and the synthesis support data includes the motion-based data.
[0207] Example 24 includes the method of any of Examples 15 to 23, further including using the trained model to generate the virtual reference frame.
[0208] Example 25 includes the method of example 24, wherein the trained model includes a neural network.
[0209] Example 26 includes the method of example 24 or example 25, wherein the input to the trained model includes synthetic support data and at least one previously decoded image frame.
[0210] Example 27 includes the method of any of examples 15-26, further including receiving the bitstream from the second device via the modem.
[0211] Example 28 includes the method of any of examples 15-27, further including displaying the decoded version of the image frame on a display device.
[0212] According to Example 29, a device includes a memory configured to store instructions and a processor configured to execute the instructions to perform the method of any of Examples 15 to 28.
[0213] According to Example 30, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform a method according to any of Examples 15 to 28.
[0214] According to Example 31, an apparatus includes means for performing the method of any of Examples 15 to 28.
[0215] According to Example 32, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to obtain a bitstream corresponding to an encoded version of an image frame, generate a virtual reference frame based on synthetic support data included in the bitstream based on a determination that the bitstream includes a virtual reference frame usage indicator, and generate a decoded version of the image frame based on the virtual reference frame.
[0216] According to Example 33, an apparatus includes means for obtaining a bitstream corresponding to an encoded version of an image frame; means for generating a virtual reference frame based on synthesis support data included in the bitstream; means for generating the virtual reference frame based on a determination that the bitstream includes a virtual reference frame usage indicator; and means for generating a decoded version of the image frame based on the virtual reference frame.
[0217] According to Example 34, a device includes one or more processors configured to obtain synthetic support data associated with an image frame of a sequence of image frames, selectively generate a virtual reference frame based on the synthetic support data, and generate a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.
[0218] Example 35 includes the device of example 34, wherein the synthesis support data includes facial landmark data, motion-based data, or a combination thereof.
[0219] Example 36 includes the device of example 34 or example 35, wherein the bitstream includes synthesis support data.
[0220] Example 37 includes the device of any of Examples 34 to 36, wherein the one or more processors are configured to generate a first set of reference candidates comprising a virtual reference frame.
[0221] Example 38 includes the device of example 37, wherein the bitstream indicates a first set of reference candidates.
[0222] Example 39 includes the device of Example 37 or Example 38, wherein the one or more processors are configured to generate one or more additional first sets of reference candidates including one or more additional virtual reference frames associated with one or more additional image frames of the sequence of image frames.
[0223] Example 40 includes the device of any of Examples 34 to 39, wherein the bitstream further indicates a second set of reference candidates including one or more previously decoded image frames.
[0224] Example 41 includes the device of Example 40, wherein the one or more processors are configured to generate the virtual reference frame based at least in part on determining that a count of reference frames in the second set of reference candidates is less than a threshold reference count of the coding configuration.
[0225] Example 42 includes the device of any of Examples 34 to 41, wherein the one or more processors are configured to generate a virtual reference frame based at least in part on detecting a face in the image frame.
[0226] Example 43 includes the device of any of Examples 34 to 42, wherein the one or more processors are configured to obtain motion-based data associated with the image frame and generate the virtual reference frame based at least in part on a determination that the motion-based data indicates global motion greater than a global motion threshold.
[0227] Example 44 includes the device of any of examples 34 to 43, wherein the bitstream includes a supplemental enhancement information (SEI) message indicating synthesis support data.
[0228] Example 45 includes the device of any of examples 34 to 44, wherein the synthesis support data includes facial landmark data indicating locations of facial features within the image frames.
[0229] Example 46 includes the device of example 45, wherein the facial features include at least one of eyes, eyelids, eyebrows, nose, lips, or facial contours.
[0230] Example 47 includes the device of any of examples 34 to 46, wherein the synthesis support data includes motion sensor data indicative of motion of the image capture device associated with the image frame.
[0231] Example 48 includes the device of example 47, wherein the image capture device includes at least one of an augmented reality (XR) device, a vehicle, or a camera.
[0232] Example 49 includes the device of any of Examples 34 to 48, wherein the one or more processors are configured to warp previously decoded image frames using motion-based data to generate a virtual reference frame, and the synthesis support data includes the motion-based data.
[0233] Example 50 includes the device of any of Examples 34 to 49, wherein the bitstream includes a supplemental enhancement information (SEI) message indicating virtual reference frame usage for generating the decoded version of the image frame.
[0234] Example 51 includes the device of any of Examples 34 to 50, wherein the one or more processors are configured to use the trained model to generate the virtual reference frame.
[0235] Example 52 includes the device of example 51, wherein the trained model includes a neural network.
[0236] Example 53 includes the device of example 51 or example 52, wherein the input to the trained model includes synthetic support data and at least one previously decoded image frame.
[0237] Example 54 includes the device of any of Examples 34 to 53, further including a modem configured to transmit the bitstream to the second device.
[0238] Example 55 includes the device of any of Examples 34 to 54, further including a camera configured to capture the image frames.
[0239] According to Example 56, a method includes obtaining, at a device, synthetic support data associated with one image frame of a sequence of image frames, selectively generating a virtual reference frame based on the synthetic support data, and generating, at the device, a bitstream corresponding to an encoded version of the image frame that is based at least in part on the virtual reference frame.
[0240] Example 57 includes the method of example 56, wherein the synthesis support data includes facial landmark data, motion-based data, or a combination thereof.
[0241] Example 58 includes the method of example 56 or example 57, wherein the bitstream includes synthesis support data.
[0242] Example 59 includes the method of any of Examples 56 to 58, further including generating a first set of reference candidates that includes the virtual reference frame.
[0243] Example 60 includes the method of example 59, wherein the bitstream indicates a first set of reference candidates.
[0244] Example 61 includes the method of Example 59 or Example 60, further including generating one or more additional first sets of reference candidates including one or more additional virtual reference frames associated with one or more additional image frames of the sequence of image frames.
[0245] Example 62 includes the method of any of Examples 56 to 61, wherein the bitstream further indicates a second set of reference candidates including one or more previously decoded image frames.
[0246] Example 63 includes the method of Example 62, further including generating the virtual reference frame based at least in part on determining that the count of the reference frame in the second set of reference candidates is less than a threshold reference count of the coding configuration.
[0247] Example 64 includes the method of any of Examples 56 to 63, further including generating a virtual reference frame based at least in part on detecting a face in the image frame.
[0248] Example 65 includes the method of any of Examples 56 to 64, further including obtaining motion-based data associated with the image frame; and generating the virtual reference frame based at least in part on determining that the motion-based data indicates global motion greater than a global motion threshold.
[0249] Example 66 includes the method of any of Examples 56 to 65, wherein the bitstream includes a supplemental enhancement information (SEI) message indicating the synthesis support data.
[0250] Example 67 includes the method of any of examples 56 to 66, wherein the synthesis support data includes facial landmark data indicating locations of facial features within the image frames.
[0251] Example 68 includes the method of example 67, wherein the facial features include at least one of eyes, eyelids, eyebrows, nose, lips, or facial contours.
[0252] Example 69 includes the method of any of examples 56-68, wherein the synthesis support data includes motion sensor data indicative of motion of an image capture device associated with the image frame.
[0253] Example 70 includes the method of example 69, wherein the image capture device includes at least one of an augmented reality (XR) device, a vehicle, or a camera.
[0254] Example 71 includes the method of any of Examples 56 to 70, further including warping the previously decoded image frame using motion-based data to generate the virtual reference frame, and the synthesis support data includes the motion-based data.
[0255] Example 72 includes the method of any of Examples 56 to 71, wherein the bitstream includes a supplemental enhancement information (SEI) message indicating use of a virtual reference frame to generate a decoded version of the image frame.
[0256] Example 73 includes the method of any of Examples 56 to 72, further including using the trained model to generate the virtual reference frame.
[0257] Example 74 includes the method of example 73, wherein the trained model includes a neural network.
[0258] Example 75 includes the method of example 73 or example 74, wherein the input to the trained model includes synthetic support data and at least one previously decoded image frame.
[0259] Example 76 includes the method of any of Examples 56 to 75, further including transmitting the bitstream to the second device via a modem.
[0260] Example 77 includes the method of any of Examples 56 to 76, further including receiving the image frame from the camera.
[0261] According to Example 78, a device includes a memory configured to store instructions and a processor configured to execute the instructions to perform the method of any of Examples 56 to 77.
[0262] According to Example 79, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform a method according to any of Examples 56 to 77.
[0263] According to Example 80, an apparatus includes means for performing the method of any of Examples 56 to 77.
[0264] According to Example 81, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to obtain synthetic support data associated with image frames of a sequence of image frames, selectively generate virtual reference frames based on the synthetic support data, and generate a bitstream corresponding to encoded versions of the image frames based at least in part on the virtual reference frames.
[0265] According to Example 82, an apparatus includes means for obtaining synthetic support data associated with an image frame of a sequence of image frames, means for selectively generating a virtual reference frame based on the synthetic support data, and means for generating a bitstream corresponding to an encoded version of the image frame based at least in part on the virtual reference frame.
[0266] Those skilled in the art will further appreciate that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. Various illustrative components, blocks, configurations, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or as processor-executable instructions depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0267] The steps of a method or algorithm described in connection with the implementations disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a computing device or user terminal.
[0268] The previous description of the disclosed aspects is provided to enable any person skilled in the art to make or use the disclosed aspects. Various modifications of these aspects will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features defined by the following claims.
Claims
1. obtaining a bitstream corresponding to an encoded version of the image frame; generating a virtual reference frame based on synthesis support data included in the bitstream based on determining that the bitstream includes a virtual reference frame usage indicator; A device comprising: one or more processors configured to generate a decoded version of the image frame based on the virtual reference frame.
2. The device of claim 1 , wherein the synthesis support data includes facial landmark data, motion-based data, or a combination thereof.
3. The device of claim 1 , wherein the bitstream indicates a first set of reference candidates that includes the virtual reference frame.
4. The device of claim 3 , wherein the bitstream indicates one or more additional first sets of reference candidates including one or more additional virtual reference frames associated with one or more additional image frames of a sequence of image frames.
5. The device of claim 1 , wherein the bitstream further indicates a second set of reference candidates comprising one or more previously decoded image frames.
6. The device of claim 1 , wherein the bitstream includes a supplemental enhancement information (SEI) message indicating the synthesis support data.
7. 2. The device of claim 1, wherein the synthesis support data includes facial landmark data indicating locations of facial features, and the one or more processors are configured to generate the virtual reference frame based at least in part on previously decoded image frames and the locations of the facial features.
8. 2. The device of claim 1, wherein the synthesis support data includes motion-based data indicative of global motion, and the one or more processors are configured to generate the virtual reference frame based at least in part on a previously decoded image frame and the global motion.
9. 2. The device of claim 1, wherein the one or more processors are configured to warp previously decoded image frames using motion-based data to generate the virtual reference frame, and the synthesis support data includes the motion-based data.
10. 2. The device of claim 1 , wherein the one or more processors are configured to use a trained model to generate the virtual reference frame, inputs to the trained model including the synthetic support data and at least one previously decoded image frame.
11. The device of claim 1 , further comprising a modem configured to receive the bitstream from a second device.
12. The device of claim 1 , further comprising a display device configured to display the decoded version of the image frame.
13. obtaining, at the device, a bitstream corresponding to an encoded version of the image frame; generating a virtual reference frame based on synthesis support data included in the bitstream based on determining that the bitstream includes a virtual reference frame usage indicator; generating, at the device, a decoded version of the image frame based on the virtual reference frame.
14. obtaining synthetic support data associated with an image frame of the sequence of image frames; selectively generating a virtual reference frame based on the synthetic support data; 11. A device comprising: one or more processors configured to generate a bitstream corresponding to an encoded version of the image frame that is based at least in part on the virtual reference frame.
15. The device of claim 14 , wherein the synthesis support data includes facial landmark data, motion-based data, or a combination thereof.
16. The device of claim 14 , wherein the bitstream includes the compositing support data.
17. The device of claim 14 , wherein the one or more processors are configured to generate a first set of reference candidates that include the virtual reference frame.
18. The device of claim 17 , wherein the bitstream indicates the first set of reference candidates.
19. 20. The device of claim 17, wherein the one or more processors are configured to generate one or more additional first sets of reference candidates including one or more additional virtual reference frames associated with one or more additional image frames of the sequence of image frames.
20. 15. The device of claim 14, wherein the bitstream further indicates a second set of reference candidates including one or more previously decoded image frames, and the one or more processors are configured to generate the virtual reference frame based at least in part on determining that a count of reference frames in the second set of reference candidates is less than a threshold reference count of a coding configuration.
21. The device of claim 14 , wherein the one or more processors are configured to generate the virtual reference frame based at least in part on detecting a face in the image frame.
22. the one or more processors obtaining motion-based data associated with the image frames; The device of claim 14 , configured to generate the virtual reference frame based at least in part on a determination that the motion-based data exhibits global motion greater than a global motion threshold.
23. The device of claim 14 , wherein the synthesis support data includes facial landmark data indicating locations of facial features within the image frames.
24. The device of claim 14 , wherein the compositing support data includes motion sensor data indicative of motion of an image capture device associated with the image frame.
25. 25. The device of claim 24, wherein the image capture device comprises at least one of an augmented reality (XR) device, a vehicle, or a camera.
26. 15. The device of claim 14, wherein the one or more processors are configured to warp previously decoded image frames using motion-based data to generate the virtual reference frame, and the synthesis support data includes the motion-based data.
27. The device of claim 14 , wherein the bitstream includes a supplemental enhancement information (SEI) message indicating virtual reference frame usage to generate a decoded version of the image frame.
28. 15. The device of claim 14, wherein the one or more processors are configured to use a trained model to generate the virtual reference frame, inputs to the trained model including the synthetic support data and at least one previously decoded image frame.
29. 15. The device of claim 14, further comprising a modem configured to transmit the bitstream to a second device.
30. obtaining, at the device, synthesis support data associated with an image frame of the sequence of image frames; selectively generating a virtual reference frame based on the synthetic support data; generating, at the device, a bitstream corresponding to an encoded version of the image frame that is based at least in part on the virtual reference frame.