Method and apparatus for time resampling
The method and apparatus address the lack of temporal upsampling in SEI messages by employing neural networks for video coding, enhancing video coding efficiency and improving machine vision tasks.
Patent Information
- Application Number
- JP2025538745
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2023-12-26
- Publication Date
- 2026-01-16
AI Technical Summary
Existing video coding standards, such as VVC/H.266, do not consider temporal upsampling for machine vision tasks when generating Supplemental Enhancement Information (SEI) messages, necessitating a solution for realizing temporal upsampling based on SEI messages.
A method and apparatus for performing temporal resampling using neural networks based on SEI messages, involving decoding and encoding processes with a generator and encoder configured to utilize neural networks for temporal upsampling.
Enhances video coding efficiency by enabling effective temporal upsampling for machine vision applications, improving the quality and utility of decoded video data for tasks like object recognition and augmented reality.
Smart Images

Figure 2026501640000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to video processing, and more particularly to a method and apparatus for temporal resampling based on Supplemental Enhancement Information (SEI) messages.
[0002] (CROSS-REFERENCE TO RELATED APPLICATIONS) This disclosure claims the benefit of priority to U.S. Provisional Application No. 63 / 436,625, filed January 1, 2023, and claims the benefit of U.S. Patent Application No. 18 / 392,715, entitled "METHOD AND APPARATUS FOR TEMPORAL RESAMPLING," filed December 21, 2023. Both of the above applications are incorporated herein by reference in their entireties. [Background technology]
[0003] A video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is typically called encoding, and the decompression process is typically called decoding. A variety of video coding formats exist that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Standardization organizations have developed video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, that specify particular video coding formats. As more advanced video coding techniques are adopted into video standards, the coding efficiency of new video coding standards becomes higher. Summary of the Invention [Means for solving the problem]
[0004] Embodiments of the present disclosure provide a method and apparatus for performing temporal resampling for machine vision tasks based on SEI messages.
[0005] In a first aspect, one embodiment of the present disclosure provides a method for decoding video data, the method including: generating a reconstructed frame sequence based on compressed video; decoding an SEI message for the reconstructed frame sequence according to the compressed video; and performing temporal upsampling on the reconstructed frame sequence based on the SEI message using a neural network.
[0006] In a second aspect, one embodiment of the present disclosure provides a method for encoding video data, the method including compressing a frame sequence and encoding an SEI message for the frame sequence, the SEI message indicating whether to perform temporal upsampling on the frame sequence using a neural network.
[0007] In a third aspect, a decoding device is provided, including a generator configured to generate a reconstructed frame sequence based on compressed video, a decoder configured to decode an SEI message related to the reconstructed frame sequence according to the compressed video, and a processor configured to perform temporal upsampling on the reconstructed frame sequence based on the SEI message using a neural network.
[0008] In a fourth aspect, an embodiment of the present disclosure provides an encoding device including: a compression unit configured to compress a frame sequence; and an encoding unit configured to encode an SEI message for the frame sequence, wherein the SEI message indicates whether to perform temporal upsampling on the frame sequence using a neural network.
[0009] In a fifth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing a video bitstream, the non-transitory computer-readable storage medium causing the processor to perform a video data decoding method according to the first aspect when the video bitstream is decoded by the processor.
[0010] In a sixth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing a video bitstream, the non-transitory computer-readable storage medium causing the processor to perform the video data encoding method according to the second aspect when the video bitstream is encoded by the processor.
[0011] In a seventh aspect, an embodiment of the present disclosure provides an electronic device including: a memory having stored thereon a set of instructions; and one or more processors configured to execute the set of instructions to cause the one or more processors to perform a method for decoding video data according to the first aspect.
[0012] In an eighth aspect, an embodiment of the present disclosure provides an electronic device including: a memory having stored thereon a set of instructions; and one or more processors configured to execute the set of instructions to cause the one or more processors to perform a method for encoding video data according to the second aspect.
[0013] In a ninth aspect, embodiments of the present disclosure provide a computer program product comprising computer program instructions that enable a computer to perform the method for decoding video data according to the first aspect.
[0014] In a tenth aspect, embodiments of the present disclosure provide a computer program product comprising computer program instructions that enable a computer to perform the method for encoding video data according to the second aspect.
[0015] In an eleventh aspect, an embodiment of the present disclosure provides a computer program enabling a computer to perform the video data decoding method according to the first aspect.
[0016] In a twelfth aspect, an embodiment of the present disclosure provides a computer program enabling a computer to perform the video data encoding method according to the second aspect.
[0017] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and accompanying drawings, in which the various features shown are not drawn to scale. [Brief explanation of the drawings]
[0018] [Figure 1] FIG. 1 is a schematic diagram illustrating an exemplary system for pre-processing and encoding image data, according to some embodiments of the present disclosure. [Figure 2A] FIG. 2 is a schematic diagram illustrating an example encoding process for a hybrid video encoding system, consistent with embodiments of the present disclosure. [Figure 2B] FIG. 2 is a schematic diagram illustrating another exemplary encoding process for a hybrid video encoding system, consistent with embodiments of the present disclosure. [Figure 3A] FIG. 2 is a schematic diagram illustrating an example decoding process for a hybrid video coding system, consistent with embodiments of the present disclosure. [Figure 3B] FIG. 10 is a schematic diagram illustrating another example decoding process for a hybrid video coding system, consistent with embodiments of the present disclosure. [Figure 4] 1 is a block diagram of an exemplary apparatus for pre-processing or encoding image data, according to some embodiments of the present disclosure. [Figure 5] FIG. 10 is a schematic diagram illustrating a luma data channel with nnpfc_inp_order_idc equal to 3, according to some embodiments of the present disclosure. [Figure 6] 1 shows a flowchart of an exemplary method for encoding video data, according to some embodiments of the present disclosure. [Figure 7]FIG. 10 is a schematic diagram illustrating an example segment of an SEI message, in accordance with some embodiments of the present disclosure. [Figure 8] FIG. 10 is a schematic diagram illustrating another example segment of an SEI message, in accordance with some embodiments of the present disclosure. [Figure 9] FIG. 10 is a schematic diagram illustrating another example segment of an SEI message, in accordance with some embodiments of the present disclosure. [Figure 10] 1 shows a flowchart of an exemplary video data decoding method according to some embodiments of the present disclosure. [Figure 11] 11 is a flowchart illustrating substeps of the exemplary video data decoding method shown in FIG. 10, according to some embodiments of the present disclosure. [Figure 12] FIG. 1 is a schematic diagram illustrating an exemplary process for interpolating frames, consistent with embodiments of the present disclosure. [Figure 13] FIG. 10 is a schematic diagram illustrating an example segment of a decoded SEI message in accordance with some embodiments of the present disclosure. [Figure 14] FIG. 10 is a schematic diagram illustrating another example segment of a decoded SEI message in accordance with some embodiments of the present disclosure. [Figure 15] FIG. 10 is a schematic diagram illustrating another example segment of a decoded SEI message in accordance with some embodiments of the present disclosure. [Figure 16] 11 shows a flowchart of sub-steps of the exemplary video data decoding method shown in FIG. 10, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0019] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings in which like numbers in different figures, unless otherwise specified, represent the same or similar elements. The implementations set forth in the following description of the exemplary embodiments do not represent all implementations consistent with the present disclosure. Instead, the implementations are merely examples of apparatus and methods consistent with aspects related to the present disclosure as recited in the appended claims. Certain aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions set forth herein shall control.
[0020] The ITU-T Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) Joint Video Experts Team (JVET) have been developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 while using half the bandwidth.
[0021] The VVC standard has been progressing well since April 2018, and continues to incorporate more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.
[0022] Neural network post-processing filters are employed in image processing. Some existing techniques use SEI messages to specify the post-processing filter characteristics of neural network post-processing filters. However, temporal upsampling for machine vision is not considered when generating the SEI messages. Therefore, it is necessary to realize temporal upsampling for machine vision based on the SEI messages.
[0023] 1 is a block diagram illustrating a system 100 for preprocessing and encoding image data, according to some disclosed embodiments. The image data may include an image (also called a "picture" or "frame"), multiple images, or a video. An image is a still image. Multiple images may or may not be related spatially or temporally. A video is a set of images arranged in chronological order.
[0024] 1 , system 100 includes a source device 120 that provides encoded video data to be later decoded by a destination device 140. Consistent with embodiments disclosed herein, each of source device 120 and destination device 140 may include any of a wide range of devices, including a desktop computer, a notebook (e.g., laptop) computer, a server, a tablet computer, a set-top box, a mobile phone, a vehicle, a camera, an image sensor, a robot, a television, a camera, a wearable device (e.g., a smart watch or wearable camera), a display device, a digital media player, a video game console, a video streaming device, etc. Source device 120 and destination device 140 may be equipped for wireless or wired communication.
[0025] Referring to Figure 1, source device 120 may include an image / video preprocessor 122, an image / video encoder 124, and an output interface 126. Destination device 140 may include an input interface 142, an image / video decoder 144, and one or more machine vision applications 146. Image / video preprocessor 122 preprocesses image data, i.e., image(s) or video(s), to generate an input bitstream for image / video encoder 124. Image / video encoder 124 encodes the input bitstream and outputs an encoded bitstream 162 via output interface 126. Encoded bitstream 162 is transmitted over communication medium 160 and received by input interface 142. Image / video decoder 144 then decodes encoded bitstream 162 to generate decoded data usable by machine vision application(s) 146.
[0026] More specifically, source device 120 may further include various devices (not shown) for providing source image data to be preprocessed by image / video preprocessor 122. The devices for providing source image data may include an image / video capture device such as a camera, an image / video archive or storage device containing previously captured images / video, or an image / video distribution interface for receiving images / video from an image / video content provider.
[0027] Image / video encoder 124 and image / video decoder 144 may each be implemented as any of a variety of suitable encoder or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If encoding or decoding is implemented partially in software, image / video encoder 124 or image / video decoder 144 may store instructions for the software on a suitable non-transitory computer-readable medium and execute these instructions in hardware using one or more processors to perform techniques consistent with this disclosure. Each of image / video encoder 124 or image / video decoder 144 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) within the respective device.
[0028] Image / video encoder 124 and image / video decoder 144 may operate according to any video coding standard, such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), AOMedia Video 1 (AV1), Joint Photographic Experts Group (JPEG), Moving Picture Experts Group (MPEG), etc. Alternatively, image / video encoder 124 and image / video decoder 144 may be customized devices that do not conform to existing standards. Although not shown in FIG. 1 , in some embodiments, image / video encoder 124 and image / video decoder 144 may be integrated with an audio encoder and decoder, respectively, and may include appropriate MUX-DEMUX units or other hardware and software to handle the encoding of both audio and video in a common data stream or separate data streams.
[0029] Output interface 126 may include any type of medium or device capable of transmitting encoded bitstream 162 from source device 120 to destination device 140. For example, output interface 126 may include a transmitter or transceiver configured to transmit encoded bitstream 162 directly from source device 120 to destination device 140 in real time. Encoded bitstream 162 may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 140.
[0030] The communication medium 160 may include a transitory medium, such as a wireless broadcast or a wired network transmission. For example, the communication medium 160 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). The communication medium 160 may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. In some embodiments, the communication medium 160 may include routers, switches, base stations, or any other equipment that may be useful in facilitating communication from the source device 120 to the destination device 140. For example, a network server (not shown) may receive the encoded bitstream 162 from the source device 120 and provide such encoded bitstream 162 to the destination device 140, for example, via a network transmission.
[0031] Communication medium 160 may also be in the form of a storage medium (e.g., a non-transitory storage medium), such as a hard disk, a flash drive, a compact disc, a digital video disc, a Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded image data. In some embodiments, a computing device at a media production facility, such as a disc stamping facility, may receive the encoded image data from source device 120 and produce a disc containing such encoded video data.
[0032] Input interface 142 may include any type of medium or device capable of receiving information from communication medium 160. The received information includes encoded bitstream 162. For example, input interface 142 may include a receiver or transceiver configured to receive encoded bitstream 162 in real time.
[0033] The machine vision application 146 includes various hardware and / or software for utilizing the decoded image data generated by the image / video decoder 144. For example, the machine vision application 146 may include a display device that displays the decoded image data to a user, and may include any of various display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device. In another example, the machine vision application 146 may include one or more processors configured to use the decoded image data to perform various machine vision applications, such as object recognition and tracking, facial recognition, image matching, image / video retrieval, augmented reality, robotic vision and navigation, autonomous driving, 3D structure construction, stereo correspondence, motion tracking, etc.
[0034] Exemplary image data encoding and decoding techniques will now be described with reference to Figures 2A-2B and 3A-3B.
[0035] FIG. 2A shows a schematic diagram of an exemplary encoding process 200A consistent with embodiments of the present disclosure. For example, encoding process 200A may be performed by an encoder, such as image / video encoder 124 of FIG. 1. As shown in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to process 200A. Video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in chronological order. Each original picture of video sequence 202 may be divided into multiple basic processing units, basic processing subunits, or regions for processing by the encoder. In some embodiments, the encoder may perform process 200A at the level of the basic processing units for each original picture of video sequence 202. For example, the encoder may perform process 200A in an iterative manner, in which case the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for regions of each original picture of video sequence 202.
[0036] 2A , an encoder may provide a fundamental processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to a binary encoding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a “forward path.” During process 200A, after quantization stage 214, the encoder may provide quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to a prediction BPU 208 to generate a prediction reference 224, which is used in prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.
[0037] The encoder may iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and generate prediction reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.
[0038] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or any act by any method for inputting data.
[0039] In the prediction stage 204, in the current iteration, the encoder receives the original BPU and a prediction reference 224 and can perform a prediction operation to generate predicted data 206 and a predicted BPU 208. The prediction reference 224 can be generated from a reconstruction path in a previous iteration of the process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting predicted data 206, which can be used to reconstruct the original BPU from the predicted data 206 and the prediction reference 224 as a predicted BPU 208.
[0040] Ideally, predicted BPU 208 may be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 generally differs slightly from the original BPU. To record such differences, after generating predicted BPU 208, the encoder may subtract the predicted BPU 208 from the original BPU to generate residual BPU 210. For example, the encoder may subtract pixel values (e.g., grayscale or RGB values) of predicted BPU 208 from corresponding pixel values of the original BPU. As a result of such subtraction between corresponding pixels of the original BPU and pixels of predicted BPU 208, each pixel of residual BPU 210 may have a residual value. Compared to the original BPU, predicted data 206 and residual BPU 210 may require fewer bits, which can be used to reconstruct the original BPU without significant loss of quality. Thus, the original BPU is compressed.
[0041] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns," where each basis pattern may represent a variation frequency (e.g., luma variation frequency) component of the residual BPU 210. None of the basis patterns can be reconstructed from any combination (e.g., a linear combination) of any other basis patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is analogous to the discrete Fourier transform of a function, where the basis patterns are analogous to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are analogous to the coefficients associated with the basis functions.
[0042] Various transform algorithms can use various basis patterns. For example, various transform algorithms can be used in transform stage 212, such as a discrete cosine transform, a discrete sine transform, etc. The transform in transform stage 212 is invertible. That is, the encoder can recover residual BPU 210 by inversely operating the transform (called an "inverse transform"). For example, to recover pixels of residual BPU 210, the inverse transform can multiply the values of corresponding pixels in the basis pattern by their associated coefficients and add the products to generate a weighted sum. In video coding standards, both the encoder and the decoder can use the same transform algorithm (and therefore the same basis pattern). Therefore, the encoder can record only the transform coefficients, and the decoder can reconstruct residual BPU 210 from the transform coefficients without receiving the basis pattern from the encoder. Although the transform coefficients may have fewer bits than residual BPU 210, they can still be used to reconstruct residual BPU 210 without significant loss of quality. Therefore, the residual BPU 210 is further compressed.
[0043] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different fluctuation frequencies (e.g., luma fluctuation frequencies). Because the human eye is generally better at perceiving low-frequency fluctuations, the encoder can ignore high-frequency fluctuation information without significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization parameter") and rounding the quotient to the nearest integer. After such an operation, some transform coefficients of the high-frequency basis patterns can be converted to zero, and those of the low-frequency basis patterns can be converted to smaller integers. The encoder can ignore the zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process is also invertible, in which case the quantized transform coefficients 216 can be reconstructed into transform coefficients in an inverse operation of quantization (called "dequantization").
[0044] Because the encoder ignores the remainder of such a division in a rounding operation, the quantization stage 214 can be lossy. Typically, the quantization stage 214 can result in the greatest information loss in the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To achieve different levels of information loss, the encoder can use different values of the quantization parameter or any other parameter of the quantization process.
[0045] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as the prediction mode used in the prediction stage 204, parameters of the prediction operation, the type of transform in the transform stage 212, parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). The encoder may use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0046] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may apply the reconstructed residual BPU 222 to a prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.
[0047] It should be noted that other variations of process 200A can be used to encode video sequence 202. In some embodiments, the stages of process 200A can be performed by an encoder in a different order. In some embodiments, one or more stages of process 200A can be combined into a single stage. In some embodiments, a single stage of process 200A can be separated into multiple stages. For example, transform stage 212 and quantization stage 214 can be combined into a single stage. In some embodiments, process 200A can include additional stages. In some embodiments, process 200A can omit one or more stages in FIG. 2A.
[0048] 2B shows a schematic diagram of another exemplary encoding process 200B consistent with embodiments of the present disclosure. The encoding process 200B may be performed by an encoder such as the image / video encoder 124 of FIG. 1. The process 200B may be a modification of the process 200A. For example, the process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to the process 200A, the forward path of the process 200B further includes a mode decision stage 230, which divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of the process 200B further includes a loop filter stage 232 and a buffer 234.
[0049] In general, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can predict a current BPU using pixels from one or more neighboring BPUs already coded in the same picture. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce picture-specific spatial redundancy. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can predict a current BPU using regions from one or more already coded pictures. That is, the prediction reference 224 in temporal prediction can include coded pictures. Temporal prediction can reduce picture-specific temporal redundancy.
[0050] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. With respect to an original BPU of a picture being encoded, the prediction reference 224 may include one or more neighboring BPUs coded (in the forward path) and reconstructed (in the reconstruction path) within the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform extrapolation at the pixel level for each pixel of the predicted BPU 208, for example, by extrapolating the value of the corresponding pixel. The neighboring BPUs used for extrapolation can be positioned from various directions relative to the original BPU, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction defined in the video coding standard used. For intra prediction, the prediction data 206 may include, for example, the positions (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, etc.
[0051] In another example, in the temporal prediction stage 2044, the encoder may perform inter-prediction. For an original BPU of the current picture, the prediction reference 224 may include one or more pictures (called "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. Once all the reconstructed BPUs of the same picture have been generated, the encoder may generate the reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a certain range (called a "search window") of the reference picture. The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered in the reference picture at a location having the same coordinates as the coordinates of the original BPU in the current picture and extend over a predetermined distance. When the encoder identifies a region within the search window that is similar to the original BPU (e.g., by using a pixel recursion algorithm, a block matching algorithm, etc.), the encoder can determine such a region as a matching region. The matching region may have different dimensions than the original BPU (e.g., smaller than, equal to, larger than, or of a different shape). Because the reference picture and the current picture are temporally separated in the timeline, the matching region can be considered to "move" toward the position of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector." If multiple reference pictures are used, the encoder can search for the matching region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to pixel values of the matching region in each matching reference picture.
[0052] Motion estimation can be used to identify various types of motion, such as, for example, translation, rotation, zooming, etc. In inter prediction, the prediction data 206 may include, for example, the coincidence (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference pictures, the weights associated with the reference pictures, etc.
[0053] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., a motion vector) and the prediction reference 224. For example, the encoder may shift the matching region of the reference picture according to the motion vector, thereby allowing the encoder to predict the original BPU of the current picture. If multiple reference pictures are used, the encoder may shift the matching region of the reference picture according to each motion vector and average the pixel values of the matching region. In some embodiments, if the encoder assigns weights to the pixel values of the matching region of each matching reference picture, the encoder may add a weighted sum of the pixel values of the shifted matching region.
[0054] In some embodiments, inter prediction can be unidirectional or bidirectional. Unidirectional inter prediction can use one or more reference pictures in the same temporal direction relative to the current picture. Unidirectional inter prediction uses a reference picture preceding the current picture. Bidirectional inter prediction can use one or more reference pictures in both temporal directions relative to the current picture.
[0055] Continuing with reference to the forward path of process 200B, after spatial prediction 2042 and temporal prediction step 2044, in mode decision step 230, the encoder can select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder can perform a rate-distortion optimization technique, in which the encoder can select a prediction mode to minimize the value of a cost function depending on the bitrate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the encoder can generate a corresponding predicted BPU 208 and predicted data 206.
[0056] In the reconstruction path of process 200B, if an intra prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU coded and reconstructed in the current picture), the encoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If an inter prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current picture coded and reconstructed in all BPUs), the encoder can provide the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced by the inter prediction. The encoder can apply various loop filter techniques in the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The loop-filtered reference pictures may be stored in a buffer 234 (or a "decoded picture buffer") for later use (e.g., for use as inter-prediction reference pictures for future pictures in the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information in the binary encoding stage 226.
[0057] In some embodiments, the input video sequence 202 is processed block-by-block by the encoding process 200B. In VVC, a coding tree unit (CTU) is the largest block unit and may be approximately as large as 128x128 luma samples (plus corresponding chroma samples depending on the chroma format). A CTU may be further divided into multiple coding units (CUs) using a quadtree, binary tree, or ternary tree. The leaf nodes of the division structure transmit coding information such as the coding mode (intra mode or inter mode), motion information (reference index, motion vector difference, etc.) if inter-coded, and quantized transform coefficients 216. When intra prediction (also called spatial prediction) is used, spatially neighboring samples are used to predict the current block. When inter prediction (also called temporal prediction or motion-compensated prediction) is used, samples from previously coded pictures, called reference pictures, are used to predict the current block. In inter prediction, uni- or bi-prediction may be used. In uni-prediction, only one motion vector pointing to one reference picture is used to generate a prediction signal for the current block. In bi-prediction, two motion vectors, each pointing to a respective reference picture, are used to generate a prediction signal for the current block. The motion vectors and reference indexes are sent to the decoder to identify where the prediction signal(s) for the current block come from. After intra- or inter-prediction, a mode decision stage 230 selects an optimal prediction mode for the current block, for example, based on a rate-distortion optimization method. Based on the optimal prediction mode, a prediction BPU 208 is generated and subtracted from the input video block.
[0058] 2B , the prediction residual BPU 210 is sent to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The quantized transform coefficients 216 are then inverse quantized in an inverse quantization stage 218 and inverse transformed in an inverse transform stage 220 to obtain a reconstructed residual BPU 222. The prediction BPU 208 and the reconstructed residual BPU 222 are summed to form a prediction reference 224 before loop filtering, which is used to provide a reference sample for intra prediction. Loop filtering, such as deblocking, sample adaptive offset (SAO), or adaptive loop filter (ALF), may be applied to the prediction reference 224 in a loop filter stage 232 to form a reconstructed block. The reconstructed block is stored in a buffer 234 and used to provide a reference sample for inter prediction. The coding information generated in the mode decision stage 230, such as the coding mode (intra- or inter-prediction), intra-prediction mode, motion information, quantized residual coefficients, etc., is sent to the binary coding stage 226 to further reduce the bitrate before being packed into the output video bitstream 228.
[0059] FIG. 3A shows a schematic diagram of an exemplary decoding process 300A consistent with embodiments of the present disclosure. For example, decoding process 300A may be performed by a decoder such as image / video decoder 144 of FIG. 1. Process 300A may be a decompression process corresponding to compression process 200A of FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder (image / video decoder 144 of FIG. 1) can decode video bitstream 228 into video stream 304 according to process 300A. Video stream 304 may be very similar to video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 of FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. 2A-2B, the decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder decodes a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for regions (e.g., regions 114-118) of each picture encoded in video bitstream 228.
[0060] In FIG. 3A , a decoder may provide a portion of a video bitstream 228 associated with a basic processing unit (referred to as a “coded BPU”) of a coded picture to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may provide the prediction data 206 to a prediction stage 204 to generate a prediction BPU 208. The decoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may provide the prediction reference 224 to the prediction stage 204 for performing a prediction operation in a next iteration of the process 300A.
[0061] The decoder may iteratively perform process 300A to decode each coded BPU of the coded picture and generate a prediction reference 224 for coding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder may output the picture to the video stream 304 for display and proceed to decode the next coded picture in the video bitstream 228.
[0062] In binary decoding stage 302, the decoder may perform the inverse operation of the binary encoding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to prediction data 206 and quantized transform coefficients 216, the decoder may decode other information in binary decoding stage 302, such as, for example, a prediction mode, parameters of the prediction operation, a type of transform, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. In some embodiments, if video bitstream 228 is transmitted in the form of packets over a network, the decoder may depacketize video bitstream 228 before providing it to binary decoding stage 302.
[0063] 3B shows a schematic diagram of another exemplary decoding process 300B consistent with embodiments of the present disclosure. For example, the decoding process 300B may be performed by a decoder such as the image / video decoder 144 of FIG. 1. The process 300B may be a modification of the process 300A. For example, the process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to the process 300A, the process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and further includes a loop filter stage 232 and a buffer 234.
[0064] In process 300B, prediction data 206 decoded by the decoder from binary decoding stage 302 for a coded elementary processing unit (referred to as a "current BPU") of a coded picture being decoded (referred to as a "current picture") may include various types of data, depending on which prediction mode was used by the encoder to code the current BPU. For example, if intra prediction was used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation may include, for example, the positions (e.g., coordinates) of one or more neighboring BPUs used as references, sizes of the neighboring BPUs, parameters of extrapolation, directions of the neighboring BPUs relative to the original BPU, etc. In another example, if inter prediction was used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. Parameters for inter-prediction operations may include, for example, the number of reference pictures associated with the current BPU, weights associated with each of the reference pictures, the locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors associated with each of the matching regions, etc.
[0065] Based on the prediction mode indicator, the decoder may determine whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details regarding the performance of such spatial or temporal prediction are shown in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate the predicted BPU 208. As described in FIG. 3A, the decoder may add the predicted BPU 208 and the reconstructed residual BPU 222 to generate the prediction reference 224.
[0066] In process 300B, the decoder may provide the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 to perform a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may provide the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture from which all BPUs are decoded), the decoder may provide the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B . The loop-filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., for use as an inter-prediction reference picture for a future encoded picture of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, if the prediction mode indicator in the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data includes parameters of the loop filter.
[0067] Returning to FIG. 1 , each of the image / video preprocessor 122, image / video encoder 124, and image / video decoder 144 may be implemented as any suitable hardware, software, or combination thereof. FIG. 4 is a block diagram of an example apparatus 400 for processing image data consistent with embodiments of the present disclosure. For example, the apparatus 400 may be a preprocessor, an encoder, or a decoder. As shown in FIG. 4 , the apparatus 400 may include a processor 402. When the processor 402 executes the instructions described herein, the apparatus 400 may become a dedicated machine for preprocessing, encoding, and / or decoding image data. The processor 402 may be any type of circuitry capable of manipulating or processing information. For example, processor 402 may include any combination of any number of central processing units (i.e., "CPUs"), graphics processing units (i.e., "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems-on-chips (SoCs), application-specific integrated circuits (ASICs), etc. In some embodiments, processor 402 may also be a set of processors grouped as a single logical component. For example, as shown in FIG. 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0068] The device 400 may also include a memory 404 configured to store data (e.g., sets of instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for performing steps in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. The memory 404 may include high-speed random access storage or non-volatile storage. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAMs), read-only memories (ROMs), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, etc. Memory 404 may also be a collection of memories (not shown in FIG. 4) grouped together as a single logical component.
[0069] Bus 410 may be a communication device that transfers data between components within apparatus 400, such as an internal bus (e.g., a CPU memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like.
[0070] For ease of explanation and to avoid ambiguity, this disclosure will collectively refer to the processor 402 and other data processing circuitry as "data processing circuitry." The data processing circuitry may be implemented entirely in hardware or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry may be a single, independent module, or may be fully or partially combined within any other component of the device 400.
[0071] The device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communications network, etc.) In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.
[0072] In some embodiments, the device 400 may further include a peripheral interface 408 for providing connection to one or more peripheral devices, including, but not limited to, a cursor control device (e.g., a mouse, touchpad, or touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera, or an input interface coupled to a video archive), etc., as shown in FIG.
[0073] It should be noted that a video codec (e.g., a codec that performs process 200A, 200B, 300A, or 300B) may be implemented as any combination of software or hardware modules within device 400. For example, some or all of the stages of process 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instructions loadable into memory 404. In another example, some or all of the stages of process 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).
[0074] SEI messages are intended to be conveyed within coded video bitstreams in a manner specified in a video coding specification, or by other means determined by the specification of a system that uses such coded video bitstreams. Hereinafter, the video coding specification(s) will be referred to simply as the specification(s). SEI messages can contain various types of data that indicate the timing of video pictures, describe various characteristics of the coded video, or describe how the coded video can be used or enhanced. SEI messages are also defined as containing arbitrary user-defined data. SEI messages do not affect the core decoding process, but can indicate recommended ways in which the video should be post-processed or displayed.
[0075] In some embodiments consistent with this disclosure, a neural network post-filter (NNPF) is used to improve the quality of the decoded video. Therefore, an NNPF SEI message is used. The NNPF SEI message includes a neural network post-filter characteristic (NNPFC) and a neural network post-filter activation (NNPFA).
[0076] The semantic information of the NNPF SEI message and NNPFC is shown in Table 1 [NNPF SEI message syntax] and Table 2 [NNPF SEI message syntax].
[0077] [Table 1] [Table 2]
[0078] To use NNPF, the SEI message must define the following variables: The width and height of the cropped decoded output picture in units of luma samples, denoted herein by CroppedWidth and CroppedHeight respectively. - The luma sample array CroppedYPic and the chroma sample arrays CroppedCbPic and CroppedCrPic (if present) of the cropped decoded output picture for vertical coordinate y and horizontal coordinate x, where the top left corner of each sample array has coordinate y equal to 0 and coordinate x equal to 0. -BitDepthY, the bit depth of the luma sample array of the cropped decoded output picture. - The bit depth of the chroma sample array, if any, of the cropped decoded output picture, BitDepthC. A chroma format indicator, denoted herein by ChromaFormatIdc. - Quantized strength value StrengthControlVal when nnpfc_auxiliary_inp_idc is equal to 1.
[0079] When this NNPF SEI message specifies a neural network that can be used as a post-processing filter, the semantics specify the derivation of the luma sample array FilteredYPic[x][y] and chroma sample arrays FilteredCbPic[x][y] and FilteredCrPic[x][y], as indicated by the value of nnpfc_out_order_idc, that contain the output of the post-processing filter.
[0080] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc as specified by Table 3 below [SubWidthC and SubHeightC Values Derived from ChromaFormatIdc].
[0081] [Table 3]
[0082] nnpfc_id contains an identification number that can be used to identify a post-processing filter. The value of nnpfc_id can be in the range of 0 to 232-2 (inclusive). The nnpfc_id values of 256 to 511 (inclusive) and 231 to 232-2 (inclusive) are reserved for future use by ITU-T / ISO / IEC. Decoders with nnpfc_id values in the range of 256 to 511 (inclusive) or 231 to 232-2 (inclusive) can ignore it.
[0083] nnpfc_mode_idc equal to 0 specifies that the post-processing filter associated with the nnpfc_id value is determined by external means not specified in the current specification. nnpfc_mode_idc equal to 1 specifies that the post-processing filter associated with the nnpfc_id value is a neural network represented by the ISO / IEC 15938-17 bitstream included in this SEI message.
[0084] An nnpfc_mode_idc equal to 2 specifies that the post-processing filter associated with the nnpfc_id value is the neural network identified by the specified tag uniform resource identifier (URI) (nnpfc_uri_tag[i]) and neural network information URI (nnpfc_uri[i]).
[0085] The value of nnpfc_mode_idc can be in the range 0 to 255, inclusive. Values of nnpfc_mode_idc greater than 2 are reserved for future specifications by ITU-T|ISO / IEC and cannot be present in bitstreams conforming to the current specification. Decoders conforming to the current specification can ignore SEI messages containing reserved values of nnpfc_mode_idc.
[0086] nnpfc_purpose_and_formatting_flag equal to 0 specifies that syntax elements related to filter purpose, input formatting, output formatting, and complexity are not present. nnpfc_purpse_and_formatting_flag equal to 1 specifies that syntax elements related to filter purpose, input formatting, output formatting, and complexity are present.
[0087] nnpfc_purpose_and_formatting_flag may be equal to 1 when nnpfc_mode_idc is equal to 1 and the current CLVS (Coding Layer Video Sequence) does not contain a preceding neural network post filter characteristics SEI message in decoding order with a value of nnpfc_id equal to the value of nnpfc_id in this SEI message.
[0088] When the current CLVS contains a preceding neural network post-filter characteristics SEI message in decoding order with the same value of nnpfc_id equal to the value of nnpfc_id in this SEI message, at least one of the following conditions may apply: This SEI message has nnpfc_mode_idc equal to 1 and nnpfc_purpose_and_formatting_flag equal to 0 to provide neural network updates. This SEI message has the same content as the preceding Neural Network Post-Filter Characteristics SEI message.
[0089] When this SEI message is the first neural network post-filter characteristics SEI message in decoding order with a particular nnpfc_id value within the current CLVS, it specifies a base post-processing filter that pertains to the current decoded picture and all subsequent decoded pictures of the current layer in output order until the end of the current CLVS. When this SEI message is not the first neural network post-filter characteristics SEI message in decoding order with a particular nnpfc_id value within the current CLVS, this SEI message pertains to the current decoded picture and all subsequent decoded pictures of the current layer in output order until the end of the current CLVS, or the next neural network post-filter characteristics SEI message with that particular nnpfc_id value within the current CLVS in output order.
[0090] nnpfc_purpose indicates the purpose of the post-processing filter as specified in Table 4 [nnpfc_purpose definition]. Values of nnpfc_purpose can be in the range of 0 to 2-2 (inclusive). Values of nnpfc_purpose not listed in Table 4 are reserved for future specifications by ITU-T|ISO / IEC and cannot be present in bitstreams conforming to the current specification. Decoders conforming to the current specification can ignore SEI messages containing reserved values of nnpfc_purpose.
[0091] [Table 4]
[0092] Note that if reserved values of nnpfc_purpose are used in the future by ITU-T|ISO / IEC, the syntax of this SEI message may be extended with syntax elements whose presence is conditioned by nnpfc_purpose being equal to that value.
[0093] If SubWidthC is equal to 1 and SubHeightC is equal to 1, nnpfc_purpose cannot be equal to 2 or 4.
[0094] nnpfc_out_sub_c_flag equal to 1 specifies that outSubWidthC is equal to 1 and outSubHeightC is equal to 1. nnpfc_out_sub_c_flag equal to 0 specifies that outSubWidthC is equal to 2 and outSubHeightC is equal to 1. If nnpfc_out_sub_c_flag is not present, outSubWidthC is inferred to be equal to SubWidthC and outSubHeightC is inferred to be equal to SubHeightC. If SubWidthC is equal to 2 and SubHeightC is equal to 1, nnpfc_out_sub_c_flag cannot be equal to 0.
[0095] nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples specify the width and height, respectively, of the picture's luma sample array obtained by applying the post-processing filter identified by nnpfc_id to the cropped decoded output picture. If nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they are inferred to be equal to CroppedWidth and CroppedHeight, respectively.
[0096] nnpfc_component_last_flag equal to 0 specifies that the second dimension in the input tensor to the post-processing filter, inputTensor, and the resulting output tensor, outputTensor, is used for the channel. nnpfc_component_last_flag equal to 1 specifies that the last dimension in the input tensor to the post-processing filter, inputTensor, and the resulting output tensor, outputTensor, is used for the channel.
[0097] Note that the first dimension in the input and output tensors is used for the batch index, which is a practice in some neural network frameworks. Although the semantics of this SEI message uses a batch size equal to 1, it is up to the post-processing implementation to determine the batch size used as input to the neural network inference.
[0098] A color component is an example of a channel.
[0099] nnpfc_inp_format_flag indicates how to convert the sample values of the cropped decoded output picture into input values to the post-processing filter. If nnpfc_inp_format_flag is equal to zero, the input values to the post-processing filter are real numbers, and the functions InpY() and InpC() are specified as follows: InpY( x ) = x ÷ ( ( 1 << BitDepthY ) - 1 ) (Definition 1) InpC( x )= x ÷ ( ( 1 << BitDepthC ) - 1 ) (Definition 2)
[0100] If nnpfc_inp_format_flag is equal to 1, the input values to the post-processing filter are unsigned integers, and the functions InpY() and InpC() are specified as follows: shiftY = BitDepthY - inpTensorBitDepth if( inpTensorBitDepth >= BitDepthY) InpY( x ) = x << ( inpTensorBitDepth - BitDepthY ) (Definition 3) else InpY( x ) = Clip3(0, ( 1 << inpTensorBitDepth ) - 1, ( x + ( 1 << ( shiftY - 1 ) ) ) >> shiftY ) shiftC = BitDepthC - inpTensorBitDepth if( inpTensorBitDepth >= BitDepthC ) InpC( x ) = x << ( inpTensorBitDepth - BitDepthC ) (Definition 4) else InpC( x ) = Clip3(0, ( 1 << inpTensorBitDepth ) - 1, ( x + ( 1 << ( shiftC - 1 ) ) ) >> shiftC )
[0101] The variable inpTensorBitDepth is derived from the syntax element nnpfc_inp_tensor_bitdepth_minus8 as specified below.
[0102] nnpfc_inp_tensor_bitdepth_minus8 plus 8 specifies the bit depth of the luma sample values in the input integer tensor. The value of inpTensorBitDepth is derived as follows: inpTensorBitDepth = nnpfc_inp_tensor_bitdepth_minus8 + 8 (definition 5)
[0103] It is a bitstream conformance requirement that the value of nnpfc_inp_tensor_bitdepth_minus8 can be in the range of 0 to 24 (inclusive).
[0104] nnpfc_auxiliary_inp_idc not equal to 0 specifies that auxiliary input data is present in the neural network postfilter's input tensor. nnpfc_auxiliary_inp_idc equal to 0 indicates that auxiliary input data is not present in the input tensor. nnpfc_auxiliary_inp_idc equal to 1 specifies that auxiliary input data is derived as specified in Table 7 below. The value of nnpfc_auxiliary_inp_idc can be in the range of 0 to 255 (inclusive). Values of nnpfc_auxiliary_inp_idc greater than 1 are reserved for future specifications by ITU-T|ISO / IEC and cannot be present in bitstreams conforming to the current specification. Decoders conforming to the current specification can ignore SEI messages containing reserved values of nnpfc_auxiliary_inp_idc.
[0105] nnfpc_separate_colour_description_present_flag equal to 1 indicates that the separate combinations of primaries, transfer characteristics, and matrix coefficients for the picture resulting from the post-processing filter are specified in the SEI message syntax structure. nnpfc_separate_colour_description_present_flag equal to 0 indicates that the combinations of primaries, transfer characteristics, and matrix coefficients for the picture resulting from the post-processing filter are the same as those indicated in the VUI parameters for CLVS.
[0106] nnpfc_colour_primaries has the same semantics as specified for the vui_colour_primaries syntax element, except as follows: -nnpfc_colour_primaries specifies the picture primaries that result from applying the neural network postfilter specified in the SEI message, not the primaries used for CLVS. If -nnpfc_colour_primaries is not present in the Neural Network Postfilter Characteristics SEI message, the value of nnpfc_colour_primaries is inferred to be equal to vui_colour_primaries.
[0107] nnpfc_transfer_characteristics has the same semantics as specified for the vui_transfer_characteristics syntax element, except as follows: -nnpfc_transfer_characteristics specifies the transfer characteristics of the picture resulting from applying the neural network postfilter specified in the SEI message, rather than the transfer characteristics used for CLVS. If -nnpfc_transfer_characteristics is not present in the Neural Network Postfilter Characteristics SEI message, the value of nnpfc_transfer_characteristics is inferred to be equal to vui_transfer_characteristics.
[0108] nnpfc_matrix_coeffs has the same semantics as specified for the vui_matrix_coeffs syntax element, except as follows: -nnpfc_matrix_coeffs specifies the matrix coefficients of the picture resulting from applying the neural network postfilter specified in the SEI message, rather than the matrix coefficients used for CLVS. If -nnpfc_matrix_coeffs is not present in the Neural Network Postfilter Characteristics SEI message, the value of nnpfc_matrix_coeffs is inferred to be equal to vui_matrix_coeffs. The allowed values for -nnpfc_matrix_coeffs are not constrained by the chroma format of the decoded video picture indicated by the value of ChromaFormatIdc for the semantics of the VUI parameter. - If nnpfc_matrix_coeffs is equal to zero, nnpfc_out_order_idc cannot be equal to 1 or 3.
[0109] nnpfc_inp_order_idc indicates how the sample array of the cropped decoded output picture is ordered as input to the post-processing filter. Table 5 below [Informative Descriptions of nnpfc_inp_order_idc Values] contains informative descriptions of nnpfc_inp_order_idc values. The semantics of nnpfc_inp_order_idc, which range from 0 to 3 (inclusive), are specified in Table 7 below, which specifies the process for deriving the input tensor inputTensor for different values of nnpfc_inp_order_idc and for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft, which specify the top-left sample location of the patch of samples contained in the input tensor. If the chroma format of the cropped decoded output picture is not 4:2:0, nnpfc_inp_order_idc cannot be equal to 3. The value of nnpfc_inp_order_idc can be in the range from 0 to 255 (inclusive). Values of nnpfc_inp_order_idc greater than 3 are reserved for future specifications by ITU-T|ISO / IEC and MUST NOT be present in bitstreams conforming to the current specification. Decoders conforming to the current specification MAY ignore SEI messages containing reserved values of nnpfc_inp_order_idc.
[0110] [Table 5]
[0111] A patch is a rectangular array of samples from a component of a picture (eg, the luma or chroma component).
[0112] nnpfc_constant_patch_size_flag equal to 0 specifies that the post-processing filter accepts as input any patch size that is a positive integer multiple of the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. If nnpfc_constant_patch_size_flag equals zero, the patch size width can be less than or equal to CroppedWidth. If nnpfc_constant_patch_size_flag equals zero, the patch size height can be less than or equal to CroppedHeight. nnpfc_constant_patch_size_flag equal to 1 specifies that the post-processing filter accepts as input exactly the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1.
[0113] nnpfc_patch_width_minus1+1 specifies the horizontal sample count of the patch size required for input to the post-processing filter if nnpfc_constant_patch_size_flag is equal to 1. If nnpfc_constant_patch_size_flag is equal to zero, any positive integer multiple of (nnpfc_patch_width_minus1+1) can be used as the horizontal sample count of the patch size used for input to the post-processing filter. The value of nnpfc_patch_width_minus1 can be in the range of 0 to Min(32766,CroppedWidth-1), inclusive.
[0114] nnpfc_patch_height_minus1+1 specifies the vertical sample count of the patch size required for input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. When nnpfc_constant_patch_size_flag is equal to zero, any positive integer multiple of (nnpfc_patch_height_minus1+1) can be used as the vertical sample count of the patch size used for input to the post-processing filter. The value of nnpfc_patch_height_minus1 can be in the range of 0 to Min(32766,CroppedHeight-1), inclusive.
[0115] nnpfc_overlap specifies the overlapping horizontal and vertical sample count of adjacent input tensors of the post-processing filter. The value of nnpfc_overlap can be in the range of 0 to 16383 (inclusive).
[0116] The variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows: inpPatchWidth = nnpfc_patch_width_minus1 + 1 inpPatchHeight = nnpfc_patch_height_minus1 + 1 outPatchWidth = (nnpfc_pic_width_in_luma_samples * inpPatchWidth) / CroppedWidth outPatchHeight = ( nnpfc_pic_height_in_luma_samples * inpPatchHeight ) / CroppedHeight horCScaling = SubWidthC / outSubWidthC verCScaling = SubHeightC / outSubHeightC outPatchCWidth = outPatchWidth * horCScaling outPatchCHeight = outPatchHeight * verCScaling overlapSize = nnpfc_overlap (definition 6)
[0117] It is a bitstream conformance requirement that outPatchWidth*CroppedWidth can be equal to nnpfc_pic_width_in_luma_samples*inpPatchWidth and outPatchHeight*CroppedHeight can be equal to nnpfc_pic_height_in_luma_samples*inpPatchHeight.
[0118] nnpfc_padding_type specifies the padding process when referring to sample positions outside the boundaries of the cropped decoded output picture, as described in Table 6 [Informative Descriptions of nnpfc_padding_type Values]. The value of nnpfc_padding_type can be in the range of 0 to 15 (inclusive).
[0119] [Table 6]
[0120] nnpfc_luma_padding_val specifies the luma value used for padding when nnpfc_padding_type is equal to 4.
[0121] nnpfc_cb_padding_val specifies the Cb value used for padding when nnpfc_padding_type is equal to 4.
[0122] nnpfc_cr_padding_val specifies the Cr value used for padding when nnpfc_padding_type is equal to 4.
[0123] The function InpSampleVal(y, x, picHeight, picWidth, croppedPic) whose inputs are vertical sample position y, horizontal sample position x, picture height picHeight, picture width picWidth, and sample array croppedPic returns the value of sampleVal derived as follows: if( nnpfc_padding_type = = 0 ) if( y < 0 | | x < 0 | | y >= picHeight | | x >= picWidth ) sampleVal = 0 else sampleVal = croppedPic[ x ][ y ] else if( nnpfc_padding_type = = 1 ) sampleVal=croppedPic[Clip3(0, picWidth-1, x)][Clip3(0,picHeight-1, y)] else if( nnpfc_padding_type= =2) sampleVal=croppedPic[Reflect(picWidth-1, x)][Reflect(picHeight-1, y)] else if( nnpfc_padding_type = = 3 ) if( y >= 0 && y < picHeight ) sampleVal = croppedPic[ Wrap( picWidth − 1, x ) ][ y ] else if( nnpfc_padding_type = = 4 ) if( y < 0 | | x < 0 | | y >= picHeight | | x >= picWidth ) sampleVal[0]=npfc_luma_padding_val sampleVal
[0001] = nnpfc_cb_padding_val sampleVal
[0002] = nnpfc_cr_padding_val else sampleVal = croppedPic[ x ][ y ] (Definition 7)
[0124] The semantics of nnpfc_inp_order_idc, in the range 0 to 3 (inclusive), are specified in Table 7 below: [The process for deriving the input tensor inputTensor for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft that specify the top-left sample position of the patch of samples contained in the input tensor].
[0125] [Table 7]
[0126] An nnpfc_complexity_idc greater than 0 specifies that one or more syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id may be present. An nnpfc_complexity_idc equal to 0 specifies that there are no syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id. The value nnpfc_complexity_idc may be in the range of 0 to 255, inclusive. Values of nnpfc_complexity_idc greater than 1 are reserved for future specifications by ITU-T|ISO / IEC and MUST NOT be present in bitstreams conforming to the current specification. Decoders conforming to the current specification MAY ignore SEI messages containing reserved values of nnpfc_complexity_idc.
[0127] nnpfc_out_format_flag equal to 0 indicates that the sample values output by the post-processing filter are real numbers, and the functions OutY() and OutC() for converting the luma and chroma sample values output by the post-processing filter to integer values at bit depths BitDepthY and BitDepthC, respectively, are specified as follows: OutY(x) = Clip3(0, (1 << BitDepthY ) - 1, Round(x * ( ( 1 << BitDepthY ) - 1 ) ) ) (Definition 8) OutC(x)= Clip3(0, (1 << BitDepthC ) - 1, Round(x * ( ( 1 << BitDepthC ) - 1 ) ) ) (Definition 9)
[0128] nnpfc_out_format_flag equal to 1 indicates that the sample values output by the post-processing filter are unsigned integers and the functions OutY() and OutC() are specified as follows: shiftY = outTensorBitDepth - BitDepthY if( shiftY > 0 ) OutY( x ) = Clip3( 0, ( 1 << BitDepthY ) - 1, ( x + ( 1 << ( shiftY - 1 ) ) ) >> shiftY ) ) else OutY( x ) = x << ( BitDepthY − outTensorBitDepth ) (Definition 10) shiftC = outTensorBitDepth - BitDepthC if( shiftC > 0 ) OutC( x )= Clip3( 0, ( 1 << BitDepthC ) - 1, ( x + ( 1 << ( shiftC - 1 ) ) ) >> shiftC ) else OutC( x ) = x << ( BitDepthC − outTensorBitDepth ) (Definition 11)
[0129] The variable outTensorBitDepth is derived from the syntax element nnpfc_out_tensor_bitdepth_minus8, as described below.
[0130] nnpfc_out_tensor_bitdepth_minus8 plus 8 specifies the bit depth of the sample values in the output integer tensor. The value of outTensorBitDepth is derived as follows: outTensorBitDepth=nnpfc_out_tensor_bitdepth_minus8+8 (definition 12)
[0131] It is a bitstream conformance requirement that the value of nnpfc_out_tensor_bitdepth_minus8 can be in the range of 0 to 24 (inclusive).
[0132] nnpfc_out_order_idc indicates the output order of the samples resulting from the post-processing filter. Table 8 [Informative Descriptions of nnpfc_out_order_idc Values] contains informative descriptions of nnpfc_out_order_idc values. The semantics of nnpfc_out_order_idc, in the range 0 to 3 (inclusive), are specified in Table 9, which specifies different values of nnpfc_out_order_idc and the process for deriving sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft that specify the top-left sample location of a patch of samples contained in the input tensor. nnpfc_out_order_idc cannot be equal to 3 if nnpfc_purpose is equal to 2 or 4. The value of nnpfc_out_order_idc can be in the range 0 to 255 (inclusive). Values of nnpfc_out_order_idc greater than 3 are reserved for future specifications by ITU-T|ISO / IEC and cannot be present in bitstreams conforming to the current specification. Decoders conforming to the current specification can ignore SEI messages containing reserved values of nnpfc_out_order_idc.
[0133] [Table 8] [Table 9]
[0134] The base post-processing filter for the cropped decoded output picture picA is the filter identified by the first Neural Network Post-Filter Characteristics SEI message in decoding order that has a particular nnpfc_id value in the CLVS.
[0135] If there is another neural network post-filter characteristics SEI message related to picture picA that has the same nnpfc_id value, has nnpfc_mode_idc equal to 1, and has different content than the neural network post-processing filter SEI message that defines the base post-processing filter, then the base post-processing filter is updated by decoding the ISO / IEC 15938-17 bitstream in that neural network post-filter characteristics SEI message to obtain the post-processing filter PostProcessingFilter(). Otherwise, the post-processing filter PostProcessingFilter() is assigned to be the same as the base post-processing filter.
[0136] The following process is used to filter the cropped decoded output picture using the post-processing filter PostProcessingFilter() to generate a filtered picture containing Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc_out_order_idc. if( nnpfc_inp_order_idc = = 0 ) for( cTop = 0; cTop < CroppedHeight; cTop += inpPatchHeight ) for( cLeft = 0; cLeft < CroppedWidth; cLeft += inpPatchWidth ) { DeriveInputTensors( ) outputTensor = PostProcessingFilter( inputTensor ) StoreOutputTensors( ) } else if( nnpfc_inp_order_idc = = 1 ) for( cTop = 0; cTop < CroppedHeight / SubHeightC; cTop += inpPatchHeight ) for( cLeft = 0; cLeft < CroppedWidth / SubWidthC; cLeft += inpPatchWidth ) { DeriveInputTensors( ) outputTensor = PostProcessingFilter( inputTensor ) StoreOutputTensors( ) } else if( nnpfc_inp_order_idc = = 2 ) for( cTop = 0; cTop < CroppedHeight; cTop += inpPatchHeight) for( cLeft = 0; cLeft < CroppedWidth; cLeft += inpPatchWidth) { DeriveInputTensors( ) outputTensor = PostProcessingFilter( inputTensor ) StoreOutputTensors( ) } else if( nnpfc_inp_order_idc = = 3 ) for( cTop = 0; cTop < CroppedHeight; cTop += inpPatchHeight * 2 ) for( cLeft = 0; cLeft < CroppedWidth; cLeft += inpPatchWidth * 2 ) { DeriveInputTensors( ) outputTensor = PostProcessingFilter( inputTensor ) StoreOutputTensors( ) } (definition 13)
[0137] nnpfc_reserved_zero_bit can be equal to 0.
[0138] nnpfc_uri_tag[i] contains a NULL-terminated UTF-8 string that specifies a tag URI. The UTF-8 string may contain a URI with syntax and semantics as specified in IETF RFC 4151 that uniquely identifies the format and associated information about the neural network to be used as the post-processing filter specified by the nnrpf_uri[i] value.
[0139] Note that the nnrpf_uri_tag[i] element represents a "tag" URI that allows the format of the neural network data specified by the nnrpf_uri[i] value to be uniquely identified without the need for a central registration authority.
[0140] nnpfc_uri[i] may contain a null-terminated UTF-8 string as specified in ISO / IEC 10646. The UTF-8 string may contain a URI with syntax and semantics as specified in IETF Internet Standard 66 that identifies neural network information (e.g., data representation) to be used as a post-processing filter.
[0141] nnpfc_payload_byte[i] contains the i-th byte of a bitstream conforming to ISO / IEC 15938-17. The byte sequence nnpfc_payload_byte[i] for all possible values of i may be a complete bitstream conforming to ISO / IEC 15938-17.
[0142] nnpfc_parameter_type_idc equal to 0 indicates that the neural network uses integer parameters only. nnpfc_parameter_type_flag equal to 1 indicates that the neural network can use floating-point or integer parameters. nnpfc_parameter_type_idc equal to 2 indicates that the neural network uses binary parameters only. nnpfc_parameter_type_idc equal to 3 is reserved for future specifications by ITU-T|ISO / IEC and cannot be present in bitstreams conforming to the current specification. Decoders conforming to the current specification can ignore SEI messages containing reserved values of nnpfc_parameter_type_idc.
[0143] nnpfc_log2_parameter_bit_length_minus3 equal to 0, 1, 2, and 3 indicates that the neural network shall not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. If nnpfc_parameter_type_idc is present and nnpfc_log2_parameter_bit_length_minus3 is not present, the neural network shall not use parameters with bit lengths greater than 1.
[0144] nnpfc_num_parameters_idc indicates the maximum number of neural network parameters for the post - processing filter in units of powers of 2048. A nnpfc_num_parameters_idc equal to 0 indicates that the maximum number of neural network parameters is not specified. The value of nnpfc_num_parameters_idc can be within the range of 0 to 52 (including both end values). Values of nnpfc_num_parameters_idc greater than 52 are reserved for future specifications by ITU - T|ISO / IEC and cannot be present in a bit - stream conforming to the current specification. A decoder conforming to the current specification can ignore SEI messages containing reserved values of nnpfc_num_parameters_idc.
[0145] When the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters is derived as follows. maxNumParameters=(2048<<nnpfc_num_parameters_idc)-1 (Definition 14)
[0146] It is a bit - stream conformance requirement that the number of neural network parameters of the post - processing filter can be less than or equal to maxNumParameters.
[0147] nnpfc_num_kmac_operations_idc greater than 0 specifies that the maximum number of multiply - accumulate operations per sample of the post - processing filter is less than or equal to nnpfc_num_kmac_operations_idc * 1000. A nnpfc_num_kmac_operations_idc equal to 0 specifies that the maximum number of multiply - accumulate operations of the network is not specified. The value of nnpfc_num_kmac_operations_idc can be within the range of 0 to 232 - 1 (including both end values).
[0148] The SEI message and the semantic information of NNPFA are shown in Table 10 [NNPFA Syntax].
[0149] [Table 10]
[0150] This SEI message specifies the neural network post-processing filters that can be used for post-processing filtering of the current picture.
[0151] The neural network post-processing filter activation SEI message only lasts for the current picture.
[0152] Note that there may be several neural network post-processing filter activation SEI messages for the same picture, for example, when the post-processing filters are for different purposes or filter different color components.
[0153] nnpfa_id relates to the current picture and specifies that the neural network post-processing filters specified by one or more Neural Network Post-Processing Filter Characteristics SEI messages with nnpfc_id equal to nnfpa_id can be used for post-processing filtering of the current picture.
[0154] Despite the aforementioned features of the SEI design for NNPFC, existing SEIs do not consider temporal upsampling for machine vision as part of the neural network post-filter method. Therefore, according to some embodiments consistent with the present disclosure, a novel feature is added to the NNPFC message for temporal upsampling for machine vision.
[0155] FIG. 6 shows a flowchart of an exemplary video data encoding method 600 according to some embodiments of the present disclosure. For example, method 600 may be performed by one or more processors associated with an encoder, such as image / video encoder 124 (FIG. 1). In some embodiments, image / video encoder 124 may be integrated into apparatus 400 shown in FIG. 4, such that method 600 may be performed by apparatus 400. In other examples, method 600 may be performed by an encoder simulated by a general-purpose processing unit, optionally with auxiliary components. In this case, the encoder is implemented as a user-facing application or program. As shown in FIG. 6, method 600 includes the following steps 610 and 620:
[0156] In step 610, the encoder compresses the frame sequence. Specifically, the encoder may compress frames in the frame sequence according to the Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or other video standards. The compression scheme for the frame sequence may conform to specifications according to these standards. In some embodiments, if the encoder determines that the motion between adjacent frames is low, the frame sequence may be compressed at a lower frame rate.
[0157] In step 620, the encoder encodes a supplemental enhancement information (SEI) message for the frame sequence, where the SEI message indicates whether to perform temporal upsampling on the frame sequence using a neural network. Temporal upsampling, also known as frame interpolation, can be used to improve smoothness between adjacent frames. Upon receiving the SEI message indicating to perform temporal upsampling, the decoder can perform temporal upsampling on the reconstructed frames. In some embodiments, the decoder can determine how to perform temporal upsampling based on, for example, differences between adjacent frames. In some embodiments, the decoder can perform temporal upsampling based on a suggestion included in the SEI message from the encoder.
[0158] In some embodiments, Table 11 [NNPFC SEI Message Syntax for Machine Vision] and Table 12 [NNPFC SEI Message Syntax for Machine Vision] provide NNPFC SEI messages (or simply referred to as SEI messages) for temporal upsampling in machine vision that may be employed in an encoder.
[0159] [Table 11] [Table 12]
[0160] The semantics of the NNPFC SEI message for machine vision are described below. The SEI message specifies a neural network that can be used as a temporal resampling filter. The use of a temporal resampling filter in a particular video sequence is indicated with a neural network postfilter activation SEI message.
[0161] nnpfc_purpose indicates the purpose of the post-processing filter, as specified in Table 13 [nnpfc_purpose definition]. Values of nnpfc_purpose may range from 0 to 2-2 (inclusive). Values of nnpfc_purpose not shown in Table 13 may be reserved for future specifications and may not be present in bitstreams conforming to the current specification. Decoders conforming to the current specification may ignore SEI messages containing reserved values of nnpfc_purpose.
[0162] [Table 13]
[0163] 7 is a schematic diagram illustrating an example segment of an SEI message 700, according to some embodiments of the present disclosure. The SEI message 700 generated at the encoder includes three indicators: (1) nnpfc_purpose 701 for indicating whether to perform temporal upsampling on a frame sequence using a neural network; (2) nnpfc_temp_factor 702 for indicating the number of upsamplings; and (3) nnpfc_temp_strength 703 for indicating an upsampling strength that describes the relationship between the number of upsamplings and a reference range. In some embodiments, the reference range can be conveyed by nnpfc_temp_factor 702 and nnpfc_temp_strength 703.
[0164] With further reference to Tables 11 and 13, nnpfc_purpose in the SEI message (e.g., nnpfc_purpose 701 shown in FIG. 7) can be used to indicate whether to perform temporal upsampling on a frame sequence using a neural network. Specifically, an encoder may instruct a decoder to perform temporal upsampling by setting the value of the indicator nnpfc_purpose to 5. If the value of the indicator nnpfc_purpose is set to any other value (e.g., 1, 2, 3, or 4), the decoder does not perform temporal upsampling on the reconstructed frame sequence. In some embodiments, the value of the indicator nnpfc_purpose can be greater than 5, in which case nnpfc_purpose is used to indicate other purposes, not described here. In some embodiments, as shown in FIG. 7, nnpfc_purpose 701 can be encoded with 32 bits.
[0165] As also shown in Table 11, the SEI message indicates an upsampling number and a reference range that indicate how to perform temporal upsampling on a frame sequence using a neural network. At the decoder side, the neural network can interpolate the upsampling number of frames between two adjacent frames in the frame sequence by referencing frames in the reference range. nnpfc_temp_factor (e.g., nnpfc_temp_factor 702 shown in FIG. 7) indicates a temporal upsampling factor in the range of 1 to 2-1 (inclusive), which means the number of upsampled frames between the previous frame and the current decoded (reconstructed) frame at the decoder side. In some embodiments, as shown in FIG. 7, nnpfc_temp_factor 702 may be coded with 32 bits.
[0166] For example, if nnpfc_temp_factor is 2 and the current decoded (reconstructed) frame index is i, then upon receiving SEI message 700, the decoder can use a temporal upsampling model to interpolate two frames between decoded frame i-1 and decoded frame ii.
[0167] As also shown in Table 11, nnpfc_temp_strength (e.g., nnpfc_temp_strength 703 shown in FIG. 7) indicates the temporal resampling strength in the range of 0 to 15. The temporal resampling strength specifies the relationship between the temporal upsampling coefficient (e.g., nnpfc_temp_factor 702 shown in FIG. 7) and the input frame range of the temporal upsampling network, denoted as NumInputTem. Therefore, on the decoder side, the reference range can be determined based on nnpfc_temp_factor 702 and nnpfc_temp_strength 703. Specifically, if nnpfc_temp_strength is not equal to zero, NumInputTem = nnpfc_temp_strength * nnpfc_temp_factor. Alternatively, NumInputTem = (nnpfc_temp_strength + 1) * nnpfc_temp_factor, whereby nnpfc_temp_strength can be equal to 0. In some embodiments, nnpfc_temp_strength 703 may be encoded with 4 bits, as shown in FIG.
[0168] For example, if NumInputTem is 2 and the current frame index is i, in a bidirectional referencing scheme, the input frame indices for reference to the temporal upsampling network can be i-1, i-2, i, and i+1, where NumInputTem can be determined based on nnpfc_temp_factor 702 and nnpfc_temp_strength 703, and therefore it can be concluded that the three indicators in the SEI message 700 are sufficient to perform temporal upsampling on the reconstructed frame sequence.
[0169] 8 is a schematic diagram illustrating an example segment of an SEI message 800 according to some embodiments of the present disclosure. The SEI message 800 generated by the encoder includes four indicators: (1) nnpfc_purpose 801 for indicating whether to perform temporal upsampling on a frame sequence using a neural network; (2) nnpfc_temp_factor 802 for indicating the number of upsamplings; (3) nnpfc_temp_strength 803 for indicating an upsampling strength that describes the relationship between the number of upsamplings and the reference range; and (4) nnpfc_inp_range 804 for directly indicating the reference range. Thus, the reference range can be conveyed directly by (A) nnpfc_temp_factor 802 and nnpfc_temp_strength 803, or (B) nnpfc_inp_range 804.
[0170] The indicators nnpfc_purpose 801, nnpfc_temp_factor 802, and nnpfc_temp_strength 803 may be defined similarly to nnpfc_purpose 701, nnpfc_temp_factor 702, and nnpfc_temp_strength 703, respectively, shown in FIG.
[0171] In some embodiments, as also shown in Table 11, nnpfc_inp_range (e.g., nnpfc_inp_range 804 shown in FIG. 8) indicates the bidirectional input range of the reconstructed frame for the temporal upsampling network, in the range of 1 to 2-1 (inclusive). In some embodiments, the reference range can be determined at the decoder side based on nnpfc_temp_factor 802 and nnpfc_temp_strength 803. Specifically, if nnpfc_temp_strength is not equal to zero, NumInputTem = nnpfc_temp_strength * nnpfc_temp_factor. In some embodiments, nnpfc_inp_range can be used as NumInputTem only if nnpfc_temp_strength is equal to zero. In some embodiments, as shown in FIG. 8, nnpfc_inp_range 804 can be coded in 32 bits.
[0172] 9 is a schematic diagram illustrating an example segment of an SEI message 900 according to some embodiments of the present disclosure. The SEI message 900 generated by the encoder includes three indicators: (1) nnpfc_purpose 901 for indicating whether to perform temporal upsampling on a frame sequence using a neural network, (2) nnpfc_temp_factor 902 for indicating the number of upsamplings, and (3) nnpfc_inp_range 903 for indicating a reference range. Thus, the reference range can be directly conveyed by nnpfc_inp_range 903.
[0173] The indicators nnpfc_purpose 901, nnpfc_temp_factor 902, and nnpfc_inp_range 903 may be defined similarly to nnpfc_purpose 801, nnpfc_temp_factor 802, and nnpfc_inp_range 804, respectively, shown in FIG.
[0174] In some embodiments, nnpfc_inp_range (e.g., nnpfc_inp_range 903 shown in FIG. 9) indicates the bidirectional input range of the reconstructed frame for the temporal upsampling network, in the range of 1 to 2-1 (inclusive). In some embodiments, nnpfc_inp_range and nnpfc_temp_factor have the following constraint: 16≧nnpfc_inp_range / nnpfc_temp_factor≧1 / 16. Here, NumInputTem can be determined based on nnpfc_inp_range 903, and thus it can be concluded that the three indicators in the SEI message 900 are sufficient to perform temporal upsampling on the reconstructed frame sequence.
[0175] The encoded frame sequence according to one or more of the above embodiments can be stored or transmitted to a decoder side for processing. In some embodiments, the encoder may generate the frame sequence without considering how the decoder will be implemented. That is, the encoding side and the decoding side may be designed separately according to a common specification and may not impose restrictions on each other. Note that although the indicators in the SEI message are shown adjacent to each other, they may be included separately in the SEI message.
[0176] FIG. 10 shows a flowchart of an exemplary video data decoding method 1000 according to some embodiments of the present disclosure. For example, method 1000 may be performed by one or more processors, such as image / video decoder 144 (FIG. 1). In some embodiments, image / video decoder 144 may be integrated into device 400 shown in FIG. 4, such that method 1000 may be performed by device 400. In some other examples, method 1000 may be performed by a decoder simulated by a general-purpose processing unit, optionally with auxiliary components. In this case, the decoder is implemented as a user application or program. As shown in FIG. 10, method 1000 includes the following steps 1010-1030:
[0177] In step 1010, the decoder generates a reconstructed frame sequence based on the compressed video. The decoder may decode a compressed frame sequence transmitted from an encoder that conforms to a common specification used to encode the compressed frame sequence. For example, the decoder may obtain a reconstructed frame sequence conforming to a specification according to the Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or other video standard.
[0178] In step 1020, the decoder decodes a supplemental enhancement information (SEI) message for the reconstructed frame sequence according to the compressed video. As mentioned above, the SEI message can be included in the compressed video and can be used to indicate whether and how to perform temporal upsampling on the received compressed video.
[0179] In step 1030, the decoder performs temporal upsampling on the frame sequence reconstructed based on the SEI message using a neural network. Thus, the decoder can obtain a video segment with a higher frame rate. Also, the communication cost between the encoder and the decoder can be reduced to some extent. Figure 11 shows a flowchart of sub-steps of the exemplary video data decoding method 1000 shown in Figure 10 according to some embodiments of the present disclosure. As shown in Figure 11, step 1030 includes the following sub-steps 1110 and 1120.
[0180] In sub-step 1110, the decoder determines, based on the SEI message, the upsampling number and reference range for the first and second frames in the reconstructed frame sequence, where the second frame is adjacent to the first frame. Figure 12 is a schematic diagram illustrating an exemplary process for interpolating frames consistent with an embodiment of the present disclosure. As shown in Figure 12, four reconstructed frames 1230, 1210, 1220, and 1240 are sorted in order. If the number of frames determined based on the SEI message is two, two interpolated frames 1250 and 1260 may be interpolated between frames 1210 and 1220. Furthermore, if the determined reference ranges are two reconstructed frames located before and after the interpolation position, respectively, frames 1250 and 1260 are interpolated by referring to frames 1210, 1220, 1230, and 1240. Similarly, frames may be interpolated between frames 1230 and 1210 and between frames 1220 and 1240.
[0181] In sub-step 1120, the decoder interpolates an upsampled number of frames between the first and second frames by referencing the reference range through a neural network. The interpolation process can be realized by a conventional neural network. The decoder can input requirements (e.g., the upsampling number and the reference range) into the neural network to obtain an interpolated frame between the first and second frames.
[0182] 13 is a schematic diagram illustrating an example segment of a decoded SEI message 1300 according to some embodiments of the present disclosure. The decoded SEI message 1300, which corresponds to the SEI message 700 shown in FIG. 7, includes three indicators: (1) nnpfc_purpose 1301 for indicating whether to perform temporal upsampling on a frame sequence using a neural network; (2) nnpfc_temp_factor 1302 for indicating the number of upsamplings; and (3) nnpfc_temp_strength 1303 for indicating an upsampling strength that describes the relationship between the number of upsamplings and a reference range. The decoder can determine the reference range based on the nnpfc_temp_factor 1302 and the nnpfc_temp_strength 1303.
[0183] With further reference to Table 11, the nnpfc_purpose in the decoded SEI message (e.g., nnpfc_purpose 1301 shown in FIG. 13) can be used to indicate whether to perform temporal upsampling on the frame sequence using a neural network. Specifically, if temporal upsampling is expected at the decoder, the value of the indicator nnpfc_purpose can be set to 5. If the value of the indicator nnpfc_purpose is set to any other value (e.g., 1, 2, 3, or 4), the decoder does not perform temporal upsampling on the reconstructed frame sequence. In some embodiments, the value of the indicator nnpfc_purpose can be greater than 5, in which case nnpfc_purpose is used to indicate other purposes, not described here. In some embodiments, as shown in FIG. 13, nnpfc_purpose 1301 can be decoded as 32 bits.
[0184] As also shown in Table 11, the decoded SEI message may also indicate an upsampling number and a reference range that indicate how to perform temporal upsampling on a frame sequence using a neural network. The decoder's neural network can interpolate the upsampling number of frames between two adjacent frames in a frame sequence by referencing frames within the reference range. nnpfc_temp_factor (e.g., nnpfc_temp_factor 1302 shown in FIG. 13) indicates a temporal upsampling factor in the range of 1 to 2-1 (inclusive) and indicates the number of upsampled frames between the previous frame and the currently decoded (reconstructed) frame. In some embodiments, as shown in FIG. 13, nnpfc_temp_factor 1302 may be decoded as 32 bits.
[0185] For example, with further reference to FIG. 12, if nnpfc_temp_factor 1302 is decoded as 2 and the current decoded (reconstructed) frame index is 1220, upon receiving SEI message 1300, the decoder can interpolate two frames between decoded frames 1210 and 1220 using a temporal upsampling model.
[0186] As also shown in Table 11, nnpfc_temp_strength (e.g., nnpfc_temp_strength 1303 shown in FIG. 13) indicates the temporal resampling strength in the range of 0 to 15. The temporal resampling strength specifies the relationship between the temporal upsampling coefficient (e.g., nnpfc_temp_factor 1302 shown in FIG. 13) and the input frame range of the temporal upsampling network, denoted as NumInputTem. Thus, the decoder can determine the reference range based on nnpfc_temp_factor 1302 and nnpfc_temp_strength 1303. Specifically, if nnpfc_temp_strength is not equal to zero, NumInputTem = nnpfc_temp_strength * nnpfc_temp_factor. Alternatively, NumInputTem = (nnpfc_temp_strength + 1) * nnpfc_temp_factor, whereby nnpfc_temp_strength can be equal to 0. In some embodiments, as shown in FIG. 13, nnpfc_temp_strength 1303 may be decoded as 4 bits.
[0187] For example, with further reference to FIG. 12 , if nnpfc_temp_factor 1302 is decoded as 2 and nnpfc_temp_strength 1303 is decoded as 1, then NumInputTem can be derived as 2. With respect to the current frame index 1220, the input frame indices indicating the reference range for the temporal upsampling network can be 1230, 1210, 1220, and 1240.
[0188] 14 is a schematic diagram illustrating another example segment of a decoded SEI message 1400 according to some embodiments of the present disclosure. The decoded SEI message 1400, which corresponds to the SEI message 800 shown in FIG. 8, includes four indicators: (1) nnpfc_purpose 1401 for indicating whether to perform temporal upsampling on a frame sequence using a neural network; (2) nnpfc_temp_factor 1402 for indicating the number of upsamplings; (3) nnpfc_temp_strength 1403 for indicating an upsampling strength that describes the relationship between the number of upsamplings and the reference range; and (4) nnpfc_inp_range 1404 for indicating the reference range. The decoder side can directly determine the reference range based on (A) nnpfc_temp_factor 1302 and nnpfc_temp_strength 1303, or (B) nnpfc_inp_range 1304.
[0189] As also shown in Table 11, nnpfc_inp_range (e.g., nnpfc_inp_range 1404 shown in FIG. 14) indicates the bidirectional input range of the reconstructed frame for the temporal upsampling network, in the range of 1 to 2-1 (inclusive). In some embodiments, the reference range may be determined by the decoder based on nnpfc_temp_factor 1402 and nnpfc_temp_strength 1403. Specifically, if nnpfc_temp_strength is not equal to zero, then NumInputTem = nnpfc_temp_strength * nnpfc_temp_factor. In some embodiments, nnpfc_inp_range can be used as NumInputTem only if nnpfc_temp_strength is equal to zero. In some embodiments, as shown in FIG. 14, nnpfc_temp_strength 1404 can be coded in 32 bits.
[0190] Compared with the decoded SEI message 1300 of Figure 13, the decoded SEI message 1400 may include more range information, while communication overhead may increase. Figure 15 is a schematic diagram illustrating another example segment of a decoded SEI message 1500 according to some embodiments of the present disclosure. The decoded SEI message 1500, which corresponds to the SEI message 900 shown in Figure 9, includes three indicators: (1) nnpfc_purpose 1501 for indicating whether to perform temporal upsampling on the frame sequence using a neural network, (2) nnpfc_temp_factor 1502 for indicating the number of upsamplings, and (3) nnpfc_inp_range 1503 for indicating the reference range. A decoder can directly determine the reference range based on the nnpfc_inp_range 1503.
[0191] Figure 16 shows a flowchart of sub-steps of the exemplary video data decoding method shown in Figure 10, according to some embodiments of the present disclosure. As shown in Figure 16, step 1120 includes the following sub-steps S1610 and S1620.
[0192] In sub-step 1610, the decoder generates input tensors for the neural network based on the upsampling number and the reference range. The input tensors can be generated based on the following process:
[0193] The function InpSampleValTem(y, x, picHeight, picWidth, croppedSeq, picIdx) whose inputs are the vertical sample position y, the horizontal sample position x, the picture height picHeight, the picture width picWidth, and the sample array croppedSeq returns the value of sampleVal derived as follows (where croppedSeq is the sample array of a decoded video file that has been only spatially cropped before temporal upsampling, the decoded frame index denoted as picIdx is located in the highest dimension, and the shape of croppedSeq is [picWidth, picHeight, frame_num], meaning that frame_num is the number of frames in the decoded sequence): croppedPic= croppedSeq[:, :, picIdx] if( nnpfc_padding_type = = 0 ) if( y < 0 | | x < 0 | | y >= picHeight | | x >= picWidth ) sampleVal = 0 else sampleVal = croppedPic[ x ][ y ] else if( nnpfc_padding_type = = 1 ) sampleVal=croppedPic[Clip3(0, picWidth-1, x)][Clip3(0, picHeight-1, y)] else if( nnpfc_padding_type = = 2 ) sampleVal=croppedPic[Reflect(picWidth-1, x)][Reflect(picHeight-1, y)] else if( nnpfc_padding_type = = 3 ) if( y >= 0 && y < picHeight ) sampleVal = croppedPic[ Wrap( picWidth − 1, x ) ][ y ] else if( nnpfc_padding_type = = 4 ) if( y < 0 | | x < 0 | | y >= picHeight | | x >= picWidth ) sampleVal
[0000] = nnpfc_luma_padding_val sampleVal
[0001] = nnpfc_cb_padding_val sampleVal
[0002] = nnpfc_cr_padding_val else sampleVal = croppedPic[ x ][ y ] (Definition 15)
[0194] The input of the temporal upsampling network also needs to be specified by a process called DeriveInputTensorsTem() shown in Table 14 [DeriveInputTensorsTem() Syntax] as follows: CurPicIdx indicates the frame index of the current decoded (reconstructed) frame in the decoded video. Specifically, if nnpfc_purpose is equal to 5, the DeriveInputTensors() process can be replaced by DeriveInputTensorsTem().
[0195] [Table 14]
[0196] The process for deriving the sample values of the filtered output sample array is also specified in Table 15 [StoreOutputTensorsTem() Syntax]. Returning to Figure 16, in sub-step 1620, the decoder receives output tensors for the upsampled number of frames from the neural network. The output tensors may be generated according to the syntax shown in Table 15. Specifically, StoreOutputTensorsTem() defines the storage format of the output tensors for pictures that have undergone temporal upsampling between the current decoded picture and the previous picture. Similarly, if nnpfc_purpose is equal to 5, the StoreOutputTensors() process may be replaced with StoreOutputTensorsTem().
[0197] [Table 15]
[0198] In some embodiments, the solutions shown in Tables 11 and 12 can alternatively be implemented as Table 16 below [NNPFC SEI Message Syntax for Machine Vision]. Table 16 shows part of the syntax, but other parts of the syntax can be used as in Tables 11 and 12.
[0199] [Table 16]
[0200] 13 includes three indicators: nnpfc_purpose 1301, nnpfc_temp_factor 1302, and nnpfc_temp_strength 1303. The decoded SEI message 1300 may also conform to the syntax shown in Table 16.
[0201] As shown in Table 16, nnpfc_temp_factor (e.g., nnpfc_temp_factor 1302 shown in FIG. 13) indicates a temporal upsampling factor in the range of 1 to 2-1 (inclusive), which means the number of upsampled frames between the previous frame and the current decoded (reconstructed) frame.
[0202] For example, if nnpfc_temp_factor 1302 is decoded as 2 and the current decoded (reconstructed) frame index is i, then the temporal upsampling model should generate two frames between decoded frame i-1 and decoded frame i.
[0203] As also shown in Table 16, nnpfc_temp_strength (e.g., nnpfc_temp_strength 1303 shown in FIG. 13) indicates the temporal resampling strength in the range of 0 to 15. The temporal resampling strength defines the relationship between the temporal upsampling factor (e.g., nnpfc_temp_factor 1302 shown in FIG. 13) and the input frame range of the temporal upsampling network, denoted as NumInputTem. Specifically, NumInputTem = (nnpfc_temp_strength + 1) * nnpfc_temp_factor, whereby nnpfc_temp_strength can be equal to 0.
[0204] In some embodiments, the solutions shown in Tables 11 and 12 can be alternatively implemented as Table 17 below. Table 17 shows part of the syntax, but other parts of the syntax can be used as in Tables 11 and 12.
[0205] As described above, the decoded SEI message 1500 of Figure 15 includes three indicators: nnpfc_purpose 1501, nnpfc_temp_factor 1502, and nnpfc_inp_range 1503. The decoded SEI message 1500 may also conform to the syntax shown in Table 17 [NNPFC SEI Message Syntax for Machine Vision].
[0206] [Table 17]
[0207] As shown in Table 17, nnpfc_inp_range (e.g., nnpfc_inp_range 1503 shown in FIG. 15) indicates the bidirectional input range of the reconstructed frame for the temporal upsampling network, in the range of 1 to 2-1 (inclusive). In some embodiments, as shown in FIG. 15, nnpfc_inp_range 1503 may be decoded as 32 bits.
[0208] As shown in Table 17, nnpfc_temp_factor (e.g., nnpfc_temp_factor 1502 shown in FIG. 15) indicates a temporal upsampling factor in the range of 1 to 2-1 (inclusive), which means the number of upsampled frames between the previous frame and the current decoded (reconstructed) frame.
[0209] For example, if nnpfc_temp_factor 1502 is decoded as 2 and the current decoded (reconstructed) frame index is i, then the temporal upsampling model should produce two frames between decoded frame i-1 and decoded frame i.
[0210] In some embodiments, the solutions shown in Tables 11 and 12 can be alternatively implemented as Table 18 below. Table 18 shows part of the syntax, but other parts of the syntax can be used as in Tables 11 and 12.
[0211] As described above, the decoded SEI message 1500 includes three indicators. In some embodiments, the nnpfc_inp_range (e.g., nnpfc_inp_range 1503 shown in FIG. 15) and the nnpfc_temp_factor (e.g., nnpfc_temp_factor 1502 shown in FIG. 15) are restricted such that 16≧nnpfc_inp_range / nnpfc_temp_factor≧1 / 16. This restriction is a supplemental syntax in Table 18 [NNPFC SEI Message Syntax for Machine Vision].
[0212] [Table 18]
[0213] nnpfc_temp_factor (e.g., nnpfc_temp_factor 1502 shown in FIG. 15) indicates a temporal upsampling factor in the range of 1 to 2-1 (inclusive), which means the number of upsampled frames between the previous frame and the current decoded (reconstructed) frame.
[0214] For example, if nnpfc_temp_factor 1502 is decoded as 2 and the current decoded (reconstructed) frame index is i, then the temporal upsampling model should generate two frames between decoded frame i-1 and decoded frame i. As also shown in Table 18, nnpfc_inp_range (e.g., nnpfc_inp_range 1503 shown in FIG. 15) indicates the bidirectional input range of the reconstructed frame for the temporal upsampling network, ranging from 1 to 2-1 (inclusive). Specifically, the ratio of nnpfc_inp_range to nnpfc_temp_factor is constrained as 16≧nnpfc_inp_range / nnpfc_temp_factor≧1 / 16.
[0215] In some embodiments, a generator configured to generate a reconstructed frame sequence based on the compressed video; a decoder configured to decode an SEI message for the reconstructed frame sequence according to the compressed video; and a processing unit configured to perform temporal upsampling on the reconstructed frame sequence based on the SEI message using a neural network.
[0216] In one implementation, the processing unit: Determine an upsampling number and a reference range for a first frame and a second frame adjacent to the first frame in the reconstructed frame sequence based on the SEI message; The neural network is configured to interpolate an upsampled number of frames between the first frame and the second frame by referencing within the reference range.
[0217] In one implementation, the SEI message includes: a first indicator for indicating an upsampling number; and a second indicator for indicating an upsampling strength that describes a relationship between the upsampling number and a reference range; A reference range is determined based on the first indicator and the second indicator.
[0218] In one implementation, if the second indicator is not equal to zero, the reference range is the product of the upsampling number and the upsampling strength.
[0219] In one implementation, the SEI message further includes a third indicator for indicating a reference range, and when the second indicator is equal to zero, the reference range is determined based on the third indicator.
[0220] In one implementation, the reference range is the product of the upsampling strength plus one and the number of upsamplings.
[0221] In one implementation, the SEI message includes a first indicator for indicating an upsampling number and a third indicator for indicating a reference range; A reference range is determined based on the third indicator.
[0222] In one implementation, the ratio of the reference range to the upsampling number is within a predetermined range.
[0223] In one implementation, the predetermined range is [1 / 16,16].
[0224] In one implementation, the SEI message includes a fourth indicator to indicate whether to perform temporal resampling on the reconstructed frame sequence.
[0225] In one implementation, the processing unit: configured to generate an input tensor for the neural network based on the upsampling number and the reference range; The decoding device A receiver configured to receive an output tensor for the upsampled number of frames from the neural network is included.
[0226] In some embodiments, a compressor configured to compress the sequence of frames; an encoding unit configured to encode an SEI message for the frame sequence; An encoding device is provided in which the SEI message indicates whether to perform temporal upsampling on a frame sequence using a neural network.
[0227] In one implementation, the SEI message further indicates how to perform temporal upsampling on the frame sequence using a neural network.
[0228] In one implementation, the SEI message indicates the upsampling number and the reference range, so that the neural network interpolates the upsampling number of frames between two adjacent frames in the frame sequence by referencing frames within the reference range.
[0229] In one implementation, the SEI message includes: a first indicator for indicating an upsampling number; and a second indicator for indicating an upsampling strength that describes a relationship between the upsampling number and a reference range; A reference range is determined based on the first indicator and the second indicator.
[0230] In one implementation, the SEI message further includes a third indicator for indicating a reference range, and when the second indicator is equal to zero, the reference range is determined based on the third indicator.
[0231] In one implementation, the SEI message includes a first indicator for indicating an upsampling number and a third indicator for indicating a reference range; A reference range is determined based on the third indicator.
[0232] In some embodiments, an electronic device is provided that includes a memory having a set of instructions stored thereon, and one or more processors configured to execute the set of instructions to cause the one or more processors to perform a method for decoding video data in accordance with an embodiment of the method for decoding video data described above.
[0233] In some embodiments, an electronic device is provided that includes a memory having stored thereon a set of instructions and one or more processors configured to execute the set of instructions to cause the one or more processors to perform a method for encoding video data in accordance with an embodiment of the method for encoding video data described above.
[0234] Some embodiments also provide a non-transitory computer-readable storage medium having stored thereon a bitstream that can be encoded and decoded according to the above-mentioned NNPFC SEI messages for machine vision.
[0235] Some embodiments also provide a non-transitory computer-readable storage medium containing instructions that can be executed by a device (such as the encoders and decoders disclosed herein) to perform the methods described above. Common forms of non-transitory media include, for example, a floppy disk, a flexible disk, a hard disk, a solid-state drive, magnetic tape or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROM, EPROM, FLASH-EPROM or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge, and networked versions thereof. A device may include one or more processors (CPUs), input / output interfaces, a network interface, and / or memory.
[0236] In some embodiments, a computer program product is provided that includes computer program instructions that enable a computer to perform a method for decoding video data according to an embodiment of the above method for decoding video data.
[0237] In some embodiments, a computer program product is provided comprising computer program instructions that enable a computer to perform a method for encoding video data according to an embodiment of the above method for encoding video data.
[0238] In some embodiments, a computer program is provided that enables a computer to perform a method for decoding video data according to an embodiment of the above method for decoding video data.
[0239] In some embodiments, a computer program is provided that enables a computer to perform a method for encoding video data according to an embodiment of the above method for encoding video data.
[0240] The embodiments can be further described using the following clauses. 1. A method for decoding video data, comprising: generating a reconstructed frame sequence based on the compressed video; decoding an SEI message for the reconstructed frame sequence according to the compressed video; and performing temporal upsampling on the reconstructed frame sequence based on the SEI message using a neural network. 2. Performing temporal upsampling on the reconstructed frame sequence determining an upsampling number and a reference range for a first frame and a second frame adjacent to the first frame in the reconstructed frame sequence based on the SEI message; 2. The method of claim 1, comprising interpolating an upsampled number of frames between the first frame and the second frame by referencing within the reference range via a neural network. 3. The SEI message includes a first indicator for indicating an upsampling number and a second indicator for indicating an upsampling strength that describes a relationship between the upsampling number and a reference range; 3. The method of clause 2, wherein the reference range is determined based on the first indicator and the second indicator. 4. The method according to clause 3, wherein if the second indicator is not equal to zero, the reference range is the product of the upsampling number and the upsampling strength. 5. The method of clause 4, wherein the SEI message further includes a third indicator for indicating a reference range, and when the second indicator is equal to zero, the reference range is determined based on the third indicator. 6. The method according to clause 3, wherein the reference range is the product of the upsampling intensity plus 1 and the number of upsamplings. 7. The SEI message includes a first indicator for indicating an upsampling number and a third indicator for indicating a reference range; 3. The method of clause 2, wherein the reference range is determined based on a third indicator. 8. The method of clause 7, wherein the ratio of the reference range to the upsampling number is within a predetermined range. 9. The method of clause 8, wherein the predetermined range is [1 / 16,16]. 10. The method of any one of clauses 2 to 9, wherein the SEI message includes a fourth indicator for indicating whether or not to perform temporal resampling on the reconstructed frame sequence. 11. Interpolating an upsampled number of frames between the first frame and the second frame by referencing within the reference range; generating an input tensor for the neural network based on the upsampling number and the reference range; and receiving output tensors for the upsampled number of frames from the neural network. 12. A method for encoding video data, comprising: compressing the frame sequence; encoding an SEI message for the frame sequence; The method, wherein the SEI message indicates whether to perform temporal upsampling on the frame sequence using a neural network. 13. The method of clause 12, wherein the SEI message further indicates how to perform temporal upsampling on the frame sequence using a neural network. 14. The method of clause 13, wherein the SEI message indicates an upsampling number and a reference range, whereby the neural network interpolates the upsampling number of frames between two adjacent frames of the frame sequence by referring to frames within the reference range. 15. The SEI message includes a first indicator for indicating an upsampling number and a second indicator for indicating an upsampling strength that describes a relationship between the upsampling number and a reference range; 15. The method of clause 14, wherein the reference range is determined based on the first indicator and the second indicator. 16. The method according to clause 15, wherein if the second indicator is not equal to zero, the reference range is the product of the upsampling number and the upsampling strength. 17. The method of clause 16, wherein the SEI message further includes a third indicator for indicating the reference range, and when the second indicator is equal to zero, the reference range is determined based on the third indicator. 18. The method according to clause 15, wherein the reference range is the product of the upsampling intensity plus 1 and the number of upsamplings. 19. The SEI message includes a first indicator for indicating an upsampling number and a third indicator for indicating a reference range; 15. The method of clause 14, wherein the reference range is determined based on a third indicator. 20. The method of clause 19, wherein the ratio of the reference range to the upsampling number is within a predetermined range. 21. The method of clause 20, wherein the predetermined range is [1 / 16,16]. 22. The method of any one of clauses 14 to 21, wherein the SEI message includes a fourth indicator for indicating whether or not to perform temporal resampling on the reconstructed frame sequence. 23. Upsampling number of frames generating an input tensor for the neural network based on the upsampling number and the reference range; receiving an output tensor for an upsampled number of frames from the neural network; and 24. A non-transitory computer-readable storage medium storing a video bitstream having a frame sequence, the bitstream comprising: Contains an SEI message for a frame sequence, A non-transitory computer-readable storage medium comprising information describing an SEI message performing temporal upsampling on a sequence of frames using a neural network. 25. Performing temporal upsampling on the reconstructed frame sequence determining an upsampling number and a reference range for a first frame and a second frame adjacent to the first frame in the reconstructed frame sequence based on the SEI message; and interpolating, via a neural network, an upsampled number of frames between the first frame and the second frame by referencing within the reference range. 26. The SEI message includes a first indicator for indicating an upsampling number and a second indicator for indicating an upsampling strength that describes the relationship between the upsampling number and a reference range; 26. The non-transitory computer-readable storage medium of clause 25, wherein the reference range is determined based on the first indicator and the second indicator. 27. The non-transitory computer-readable storage medium of clause 26, wherein if the second indicator is not equal to zero, the reference range is the product of the upsampling number and the upsampling strength. 28. A non-transitory computer-readable storage medium as described in clause 27, wherein the SEI message further includes a third indicator for indicating a reference range, and when the second indicator is equal to zero, the reference range is determined based on the third indicator. 29. The non-transitory computer-readable storage medium of clause 26, wherein the reference range is the product of the upsampling intensity plus one and the number of upsamplings. 30. The SEI message includes a first indicator for indicating an upsampling number and a third indicator for indicating a reference range; 26. The non-transitory computer-readable storage medium of clause 25, wherein the reference range is determined based on a third indicator. 31. The non-transitory computer-readable storage medium of clause 30, wherein the ratio of the reference range to the upsampling number is within a predetermined range. 32. The non-transitory computer-readable storage medium of clause 31, wherein the predetermined range is [1 / 16, 16]. 33. A non-transitory computer-readable storage medium according to any one of clauses 25 to 32, wherein the SEI message includes a fourth indicator for indicating whether to perform temporal resampling on the reconstructed frame sequence. 34. Interpolating an upsampled number of frames between a first frame and a second frame by referencing within a reference range; generating an input tensor for the neural network based on the upsampling number and the reference range; and receiving output tensors for the upsampled number of frames from the neural network. 35. A non-transitory computer-readable storage medium storing a bitstream, the bitstream comprising: compressing the frame sequence; encoding an SEI message for the frame sequence; A non-transitory computer-readable storage medium generated by a method, wherein the SEI message indicates whether or not to perform temporal upsampling on a sequence of frames using a neural network. 36. The non-transitory computer-readable storage medium of clause 35, wherein the SEI message further indicates how to perform temporal upsampling on the frame sequence using a neural network. 37. A non-transitory computer-readable storage medium as described in clause 36, wherein the SEI message indicates an upsampling number and a reference range, whereby the neural network interpolates the upsampling number of frames between two adjacent frames in the frame sequence by referencing frames within the reference range. 38. The SEI message includes a first indicator for indicating an upsampling number and a second indicator for indicating an upsampling strength that describes the relationship between the upsampling number and a reference range; 38. The non-transitory computer-readable storage medium of clause 37, wherein the reference range is determined based on the first indicator and the second indicator. 39. The non-transitory computer-readable storage medium of clause 38, wherein if the second indicator is not equal to zero, the reference range is the product of the upsampling number and the upsampling strength. 40. A non-transitory computer-readable storage medium as described in clause 39, wherein the SEI message further includes a third indicator for indicating the reference range, and when the second indicator is equal to zero, the reference range is determined based on the third indicator. 41. The non-transitory computer-readable storage medium of clause 38, wherein the reference range is the product of the upsampling intensity plus one and the number of upsamplings. 42. The SEI message includes a first indicator for indicating an upsampling number and a third indicator for indicating a reference range; 38. The non-transitory computer-readable storage medium of clause 37, wherein the reference range is determined based on a third indicator. 43. The non-transitory computer-readable storage medium of clause 42, wherein the ratio of the reference range to the upsampling number is within a predetermined range. 44. The non-transitory computer-readable storage medium of clause 43, wherein the predetermined range is [1 / 16, 16]. 45. A non-transitory computer-readable storage medium according to any one of clauses 37 to 44, wherein the SEI message includes a fourth indicator for indicating whether to perform temporal resampling on the reconstructed frame sequence. 46. The number of frames upsampled is generating an input tensor for the neural network based on the upsampling number and the reference range; receiving an output tensor for an upsampled number of frames from the neural network; and interpolating between two adjacent frames by performing operations including: receiving an output tensor for an upsampled number of frames from the neural network; 47. A memory storing a set of instructions; and one or more processors configured to execute a set of instructions to cause the apparatus to perform the following operations, said operations comprising: generating a reconstructed frame sequence based on the compressed video; decoding an SEI message for the reconstructed frame sequence according to the compressed video; and performing temporal upsampling on the reconstructed frame sequence based on the SEI messages using a neural network. 48. Performing temporal upsampling on the reconstructed frame sequence determining an upsampling number and a reference range for a first frame and a second frame adjacent to the first frame in the reconstructed frame sequence based on the SEI message; and interpolating an upsampled number of frames between the first frame and the second frame by referencing within the reference range via a neural network. 49. The SEI message includes a first indicator for indicating an upsampling number and a second indicator for indicating an upsampling strength that describes the relationship between the upsampling number and a reference range; 49. The apparatus of clause 48, wherein the reference range is determined based on the first indicator and the second indicator. 50. The apparatus of clause 49, wherein if the second indicator is not equal to zero, the reference range is the product of the upsampling number and the upsampling strength. 51. The apparatus described in clause 50, wherein the SEI message further includes a third indicator for indicating a reference range, and when the second indicator is equal to zero, the reference range is determined based on the third indicator. 52. The apparatus of clause 49, wherein the reference range is the product of the upsampling intensity plus one and the number of upsamplings. 53. The SEI message includes a first indicator for indicating an upsampling number and a third indicator for indicating a reference range; 49. The apparatus of clause 48, wherein the reference range is determined based on a third indicator. 54. The apparatus of clause 53, wherein the ratio of the reference range to the upsampling number is within a predetermined range. 55. The apparatus of clause 54, wherein the predetermined range is [1 / 16, 16]. 56. The apparatus of any one of clauses 48 to 55, wherein the SEI message includes a fourth indicator for indicating whether to perform temporal resampling on the reconstructed frame sequence. 57. Interpolating an upsampled number of frames between a first frame and a second frame by referencing within a reference range; generating an input tensor for the neural network based on the upsampling number and the reference range; and receiving output tensors for the upsampled number of frames from the neural network. 58. A memory storing a set of instructions; and one or more processors configured to execute a set of instructions to cause the device to perform the following operations, wherein the operations include: compressing the frame sequence; encoding an SEI message for the frame sequence; 10. An image data processing apparatus, wherein the SEI message indicates whether to perform temporal upsampling on the frame sequence using a neural network. 59. The apparatus of clause 58, wherein the SEI message further indicates how to perform temporal upsampling on the frame sequence using a neural network. 60. The apparatus described in clause 59, wherein the SEI message indicates an upsampling number and a reference range, whereby the neural network interpolates the upsampling number of frames between two adjacent frames in the frame sequence by referencing frames within the reference range. 61. The SEI message includes a first indicator for indicating an upsampling number and a second indicator for indicating an upsampling strength that describes the relationship between the upsampling number and a reference range; 61. The apparatus of clause 60, wherein the reference range is determined based on the first indicator and the second indicator. 62. The apparatus of clause 61, wherein if the second indicator is not equal to zero, the reference range is the product of the upsampling number and the upsampling strength. 63. The apparatus described in clause 62, wherein the SEI message further includes a third indicator for indicating a reference range, and when the second indicator is equal to zero, the reference range is determined based on the third indicator. 64. The apparatus of clause 61, wherein the reference range is the product of the upsampling intensity plus one and the number of upsamplings. 65. The SEI message includes a first indicator for indicating an upsampling number and a third indicator for indicating a reference range; 61. The apparatus of clause 60, wherein the reference range is determined based on a third indicator. 66. The apparatus of clause 65, wherein the ratio of the reference range to the upsampling number is within a predetermined range. 67. The device according to clause 66, wherein the predetermined range is [1 / 16, 16]. 68. The apparatus of any one of clauses 60 to 67, wherein the SEI message includes a fourth indicator for indicating whether to perform temporal resampling on the reconstructed frame sequence. 69. Upsampling number of frames generating an input tensor for the neural network based on the upsampling number and the reference range; receiving an output tensor for an upsampled number of frames from the neural network; and
[0241] It should be noted that, as used herein, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another and do not require or imply any actual relationship or order among those entities or operations. Furthermore, the words "comprising," "having," "containing," and "including," and other similar forms, are intended to be equivalent and open-ended, in that the item or items following any of these words are not intended to be an exhaustive list of such item or items, nor are they intended to be limited to only the listed item or items.
[0242] As used herein, unless specifically stated otherwise, the term "or" includes all possible combinations unless impracticable. For example, if it is stated that a database may include A or B, then the database may include A, or B, or A and B, unless specifically stated otherwise or impracticable. As a second example, if it is stated that a database may include A, B, or C, then the database may include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C, unless specifically stated otherwise or impracticable.
[0243] It will be understood that the above-described embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the computer-readable medium described above. The software, when executed by a processor, can perform the methods disclosed herein. The arithmetic units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above-described modules / units can be combined into one module / unit, and that each of the above-described modules / units can be further divided into multiple sub-modules / sub-units.
[0244] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications of the above-described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the invention being indicated by the appended claims. It is also intended that the order of steps depicted in the figures is for illustrative purposes only and is not intended to be limited to any particular order of steps. Thus, one skilled in the art will recognize that steps can be performed in different orders while performing the same method.
[0245] In the drawings and specification, illustrative embodiments are disclosed. However, many variations and modifications to these embodiments are possible. Accordingly, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. A method for decoding video data, comprising: generating a reconstructed frame sequence based on the compressed video; decoding a supplemental enhancement information (SEI) message for the reconstructed frame sequence according to the compressed video; and performing temporal upsampling on the reconstructed frame sequence based on the SEI message using a neural network.
2. performing temporal upsampling on the reconstructed frame sequence; determining an upsampling number and a reference range for a first frame and a second frame adjacent to the first frame in the reconstructed frame sequence based on the SEI message; and interpolating the upsampled number of frames between the first frame and the second frame by referencing within the reference range via the neural network.
3. the SEI message includes a first indicator for indicating the upsampling number and a second indicator for indicating an upsampling strength that describes a relationship between the upsampling number and the reference range; The method of claim 2 , wherein the reference range is determined based on the first indicator and the second indicator.
4. The method of claim 3 , wherein if the second indicator is not equal to zero, the reference range is the product of the upsampling number and the upsampling strength.
5. The method of claim 4 , wherein the SEI message further includes a third indicator for indicating the reference range, and when the second indicator is equal to zero, the reference range is determined based on the third indicator.
6. The method of claim 3 , wherein the reference range is the product of the upsampling strength plus one and the upsampling number.
7. the SEI message includes a first indicator for indicating the upsampling number and a third indicator for indicating the reference range; The method of claim 2 , wherein the reference range is determined based on the third indicator.
8. The method of claim 7 , wherein the ratio of the reference range to the upsampling number is within a predetermined range.
9. The method of claim 8 , wherein the predetermined range is [1 / 16, 16].
10. The method of claim 2 , wherein the SEI message includes a fourth indicator for indicating whether to perform temporal resampling on the reconstructed frame sequence.
11. interpolating the upsampled number of frames between the first frame and the second frame by referencing within the reference range; generating an input tensor for the neural network based on the upsampling number and the reference range; and receiving an output tensor for the upsampled number of frames from the neural network.
12. A video data encoding method, comprising: compressing the frame sequence; encoding a supplemental enhancement information (SEI) message for the sequence of frames; The method, wherein the SEI message indicates whether to perform temporal upsampling on the sequence of frames using a neural network.
13. The method of claim 12 , wherein the SEI message further indicates how to perform temporal upsampling on the sequence of frames using the neural network.
14. 14. The method of claim 13, wherein the SEI message indicates an upsampling number and a reference range, whereby the neural network interpolates the upsampling number of frames between two adjacent frames of the frame sequence by referring to frames within the reference range.
15. the SEI message includes a first indicator for indicating the upsampling number and a second indicator for indicating an upsampling strength that describes a relationship between the upsampling number and the reference range; The method of claim 14 , wherein the reference range is determined based on the first indicator and the second indicator.
16. 16. The method of claim 15, wherein the SEI message further includes a third indicator for indicating the reference range, and when the second indicator is equal to zero, the reference range is determined based on the third indicator.
17. the SEI message includes a first indicator for indicating the upsampling number and a third indicator for indicating the reference range; The method of claim 14 , wherein the reference range is determined based on the third indicator.
18. a generator configured to generate a reconstructed frame sequence based on the compressed video; a decoder configured to decode an SEI message for the reconstructed frame sequence according to the compressed video; a processing unit configured to perform temporal upsampling on the reconstructed frame sequence based on the SEI message using a neural network.
19. The processing unit determining an upsampling number and a reference range for a first frame and a second frame adjacent to the first frame in the reconstructed frame sequence based on the SEI message; 20. The decoding device of claim 18, configured to interpolate the upsampled number of frames between the first frame and the second frame by referencing within the reference range via the neural network.
20. the SEI message includes a first indicator for indicating the upsampling number and a second indicator for indicating an upsampling strength that describes a relationship between the upsampling number and the reference range; The decoding device according to claim 19 , wherein the reference range is determined based on the first indicator and the second indicator.
21. 21. The decoding device of claim 20, wherein if the second indicator is not equal to zero, the reference range is the product of the upsampling number and the upsampling strength.
22. 22. The decoding device of claim 21, wherein the SEI message further includes a third indicator for indicating the reference range, and when the second indicator is equal to zero, the reference range is determined based on the third indicator.
23. The decoding device according to claim 20 , wherein the reference range is a product of the upsampling strength plus 1 and the number of upsamplings.
24. the SEI message includes a first indicator for indicating the upsampling number and a third indicator for indicating the reference range; 20. The decoding device of claim 19, wherein the reference range is determined based on the third indicator.
25. The decoding device according to claim 24 , wherein a ratio of the reference range to the upsampling number is within a predetermined range.
26. 26. The decoding device according to claim 25, wherein the predetermined range is [1 / 16, 16].
27. 27. The decoding device according to claim 19, wherein the SEI message includes a fourth indicator for indicating whether to perform temporal resampling on the reconstructed frame sequence.
28. The processing unit configured to generate an input tensor for the neural network based on the upsampling number and the reference range; The decoding device 28. A decoding device according to any one of claims 19 to 27, comprising a receiver configured to receive output tensors for the upsampled number of frames from the neural network.
29. a compressor configured to compress the sequence of frames; an encoder configured to encode an SEI message for the sequence of frames; An encoding apparatus, wherein the SEI message indicates whether to perform temporal upsampling on the frame sequence using a neural network.
30. 30. The encoding device of claim 29, wherein the SEI message further indicates how to perform temporal upsampling on the frame sequence using the neural network.
31. The encoding device of claim 30, wherein the SEI message indicates an upsampling number and a reference range, whereby the neural network interpolates a frame of the upsampling number between two adjacent frames of the frame sequence by referring to frames within the reference range.
32. the SEI message includes a first indicator for indicating the upsampling number and a second indicator for indicating an upsampling strength that describes a relationship between the upsampling number and the reference range; 32. The encoding device of claim 31, wherein the reference range is determined based on the first indicator and the second indicator.
33. The encoding device of claim 32, wherein the SEI message further includes a third indicator for indicating the reference range, and when the second indicator is equal to zero, the reference range is determined based on the third indicator.
34. the SEI message includes a first indicator for indicating the upsampling number and a third indicator for indicating the reference range; 32. The encoding device of claim 31, wherein the reference range is determined based on the third indicator.
35. 12. A video data decoding device comprising: a memory storing a set of instructions; and one or more processors configured to execute the set of instructions to cause the one or more processors to perform the video data decoding method of any one of claims 1 to 11.
36. 18. An apparatus for encoding video data, comprising: a memory having stored thereon a set of instructions; and one or more processors configured to execute the set of instructions to cause the one or more processors to perform the method for encoding video data according to any one of claims 12 to 17.
37. 12. A non-transitory computer-readable storage medium storing a video bitstream, the non-transitory computer-readable storage medium causing a processor to perform the video data decoding method of any one of claims 1 to 11 when the video bitstream is decoded by the processor.
38. 18. A non-transitory computer-readable storage medium storing a video bitstream, the non-transitory computer-readable storage medium causing a processor to perform the video data encoding method of any one of claims 12 to 17 when the video bitstream is encoded by the processor.
39. A computer program product comprising computer program instructions that enable a computer to carry out the method for decoding video data according to any one of claims 1 to 11.
40. A computer program product comprising computer program instructions that enable a computer to carry out the method for encoding video data according to any one of claims 12 to 17.
41. A computer program product that enables a computer to carry out the method for decoding video data according to any one of claims 1 to 11.
42. A computer program product enabling a computer to carry out the method for encoding video data according to any one of claims 12 to 17.