System and method for adaptive decoder-side padding in video region packing

Through the adaptive video encoding and decoding system, the area extraction, transformation and packaging technology is used to solve the problem of inefficient video encoding consumed by machines, and more efficient video compression and decoding are achieved, which enhances the context provision of machine tasks.

CN120266475APending Publication Date: 2025-07-04OP SOLUTIONS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380081603.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-27
Filing Date
2023-09-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently process video data consumed by machines, and fails to fully utilize the characteristics of machine processing images and videos, resulting in inefficient encoding.

Method used

Adaptive video encoding and decoding systems are adopted to generate compressed bit streams through area extraction, transformation and packaging, and adaptively fill on the decoder side, providing selectively filled areas to enhance machine task evaluation.

Benefits of technology

Improves video encoding efficiency, reduces the number of transmitted bits, and provides more context information to support the effective execution of machine tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120266475A_ABST
    Figure CN120266475A_ABST
Patent Text Reader

Abstract

Systems and methods for machine consumed video encoding and decoding are disclosed. A decoder is provided for decoding a bitstream encoded with a packed frame having at least one region of interest defined therein and encoding region parameters associated with the region of interest. The decoder includes a video decoder that receives the bitstream and extracts packed frames and region parameters. The region unpacking module receives the packed frame and the region parameters, and reconstructs the unpacked frame with the at least one region of interest. A region filling module is provided in the decoder that applies at least one filling parameter to at least one dimension of the region of interest in the unpacked frame. The region fill may be fixed or dynamic.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority of U.S. Provisional Patent Application No. 63 / 410,285, entitled "Systems and Methods for Adaptive Decoder - side Padding in Video Region Packing", filed on September 27, 2023, the entire content of which is incorporated herein by reference. Technical Field

[0003] The present invention generally relates to the field of video encoding and decoding. In particular, the present invention relates to Video Coding for Machines (VCM) encoders, VCM decoders, and VCM bitstreams. Background Art

[0004] Latest trends in robotics, surveillance, monitoring, Internet of Things, etc. have introduced usage scenarios where a large portion of all images and videos recorded on - site are only consumed by machines and do not reach human eyes. These machines process images and videos to complete tasks such as object detection, object tracking, segmentation, event detection, etc. Recognizing that this trend is widespread and will only accelerate in the future, international standard - setting organizations are working to standardize image and video coding that is optimized mainly for machine consumption. For example, in addition to already established standards such as "Compact Descriptors for Visual Search" and "Compact Descriptors for Video Analytics", standards like "JPEG AI" and "Video Coding for Machines" have been enabled. Compared with traditional image and video coding technologies, solutions that can improve efficiency are needed. A such solution is presented herein. Summary of the Invention

[0005] Various objects, features, aspects, and advantages of the subject matter of the present invention will become more apparent from the following detailed description of the embodiments and the accompanying drawings, in which like numerals represent like components.

[0006] Throughout this specification, the word "comprising" or variations thereof, such as "comprises" or "comprising", shall be understood to mean including the stated element, integer, or step, or group of elements, integers, or steps, but not excluding any other element, integer, or step, or group of elements, integers, or steps.

[0007] The present invention relates to a video encoding and decoding system for machine consumption and its video encoding and decoding method.

[0008] According to one aspect of the present invention, a video encoding system for machine consumption includes: an encoder having an inference of a region extractor, adapted to receive a video input and generate a first encoded output; a region transformer and a packer, adapted to receive the first encoded output and generate a second encoded output with region packing; an adaptive video encoder, adapted to receive the second encoded output and generate a compressed bitstream, wherein the compressed bitstream includes the encoded packed regions and the parameters and padding information required to reconstruct and reposition each region in the decoded frame.

[0009] According to another aspect of the present invention, a video decoding system for machine consumption includes: a decoder having an adaptive video decoder, adapted to receive a compressed bitstream including the encoded packed regions and the parameters required to reconstruct and reposition each region in the decoded frame and provide a first decoded output; an unpacker and a de-transformer, adapted to receive the first decoded output and provide an output video with adaptive padding transformation in region packing, wherein the unpacker and the de-transformer receive the unpacked reconstructed video and provide selectively padded regions for machine task evaluation.

[0010] In one embodiment, a decoder for decoding a bitstream encoded with packed frames is provided, the packed frames having at least one region of interest defined therein and encoded region parameters associated with the region of interest. The decoder includes a video decoder that receives the bitstream and extracts the packed frames and region parameters from the bitstream. A region unpacking module is provided that receives the packed frames and region parameters and reconstructs an unpacked frame with the at least one region of interest. The decoder further includes a region padding module that receives the unpacked frame and region parameters and applies at least one padding parameter to at least one dimension of the region of interest in the unpacked frame.

[0011] In certain embodiments, the region padding module may receive adaptive padding parameters, and the applied padding parameters are determined at least in part based on the adaptive padding parameters.

[0012] In certain embodiments, the region of interest is bounded by a rectangular bounding box, and the applied padding parameters are pixels of a predetermined color added to at least one boundary of the region bounding box. Alternatively, the applied padding parameters may be pixels of an average color determined by the pixels within the region, which are added to at least one boundary of the bounding box.

[0013] In some embodiments, the applied padding parameters are a fixed number of pixels. Alternatively, the region padding module also receives adaptive padding parameters, and the applied padding parameters may be a variable number of pixels determined at least in part by the adaptive padding parameters.

[0014] In some cases, the padding parameter can take the form of repeating pixels at the edge of the region of interest. Additionally or alternatively, the padding value can be signaled in the bitstream, and the padding parameter is determined at least in part based on the padding value.

[0015] In one embodiment, a packed object frame is sent to an adaptive video encoder that is adapted to receive and process the packed object frame to produce a compressed bitstream, where the adaptive video encoder is adapted to signal additional parameters for use in the reconstruction process on the decoder side, including signaling which pixels should be used for decoder-side padding.

[0016] In one embodiment, an adaptive video decoder receives the compressed bitstream and decodes it to produce a packed region frame and its signaled region information.

[0017] In one embodiment, the signaled region information includes the parameters needed for the reconstructed frame and any additional parameters to be incorporated for padding transformation of the unpacked frame.

[0018] In one embodiment, the unpacker and de-transformer are adapted to process each unpacked frame based on the region parameters and one or more adaptive parameters to expand the region pixels, thereby providing more context for the endpoint machine task evaluation of the machine task system.

[0019] According to another aspect of the present invention, a video encoding method for machine consumption includes: receiving a video input via inference with a region extractor and generating a first encoded output; receiving the first encoded output via a region transformer and a packer and generating a second encoded output with region packing; receiving the second encoded output via an adaptive video encoder and generating a compressed bitstream, where the compressed bitstream contains the encoded packed regions and the parameters and padding information needed to reconstruct and reposition each region in the decoded frame.

[0020] According to another aspect of the present invention, a video decoding method for machine consumption includes: receiving, via an adaptive video decoder, a compressed bitstream containing the encoded packed regions and the parameters needed to reconstruct and reposition each region in the decoded frame and providing a first decoded output; receiving the first decoded output via an unpacker and a de-transformer and providing an output video with adaptive padding transformation in region packing, where the unpacker and de-transformer receive the unpacked reconstructed video as part of the first decoded output and provide selectively padded regions in the output video for machine task evaluation.

[0021] According to one aspect of the present invention, a method for video compression in a machine-based video processing system includes the steps of: using an encoder-side region detector module to identify meaningful regions within a frame or image; using an extraction module to extract the identified meaningful image regions; using a region packing module to tightly pack the meaningful image regions into a single frame; processing the packed regions through a video encoder to generate a compressed bitstream, wherein the compressed bitstream includes the encoded packed regions and the parameters and padding information required for reconstructing and repositioning each region in the decoded frame.

[0022] According to one aspect of the present invention, a method for video decoding in a machine-based video processing system includes the steps of: using a video decoder to decode the compressed bitstream to generate a frame of packed regions; using a region unpacking module to unpack the region frames and restore them to their original positions in the context of the original video frames; applying an adaptive padding transform to the unpacked frames; and using a region padding module to provide additional context for machine task evaluation.

[0023] According to one aspect of the present invention, a method for padding in video compression includes the steps of: including padding information in the compressed bitstream to be added to the decoder; determining the padding type and padding size, wherein the padding size can vary on each side of a rectangular region; using different padding types, including repeating edge pixels, using the average color of pixels in the region, or including padding pixel values in the bitstream; and adaptively determining the padding based on the characteristics of the objects in the region.

[0024] According to one aspect of the present invention, a method for video encoding and decoding based on machine consumption includes: using an inference with a region extractor to encode a video input to generate a first encoded output; further encoding the first encoded output using a region transformer and a packer to generate a second encoded output with region packing; encoding the second encoded output using an adaptive video encoder to generate a compressed bitstream with padding information; using an adaptive video decoder to decode the compressed bitstream to provide a first decoded output; using an unpacker and a de-transformer to unpack and de-transform the first decoded output to provide an output video with adaptive padding transform in region packing, wherein the unpacker and the de-transformer receive the unpacked reconstructed video and provide selectively padded regions for machine task evaluation.

[0025] A computer-readable medium is also disclosed, having instructions stored thereon that, when executed by a processor, cause the processor to perform the following steps including: encoding a video input using an inference with a region extractor to generate a first encoded output; further encoding the first encoded output using a region transformer and a packer to generate a second encoded output with region packing; encoding the second encoded output using an adaptive video encoder to generate a compressed bitstream, wherein the compressed bitstream includes encoded packed regions and parameters and padding information required to reconstruct and reposition each region in a decoded frame.

[0026] On the other hand, a computer-readable medium having instructions stored thereon that, when executed by a processor, cause the processor to perform the steps including: decoding the compressed bitstream using an adaptive video decoder to provide a first decoded output; unpacking and inverse-transforming the first decoded output using an unpacker and an inverse-transformer to provide an output video with adaptive padding transformation in region packing, wherein the unpacker and the inverse-transformer receive the unpacked reconstructed video and provide selectively padded regions for a machine evaluation task. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a block diagram of a system for encoding and decoding video for a machine with region packing (e.g., in a machine video coding (VCM) system) according to an embodiment of the present disclosure.

[0028] Figure 2 is for encoding and decoding video for a machine according to an embodiment of the present disclosure Figure 1 detailed block diagram of the system.

[0029] Figure 3 is a graphical representation illustrating adaptive region filling according to the present disclosure. DETAILED DESCRIPTION

[0030] Some embodiments of the present invention that illustrate all of its features will now be discussed in detail. The words "comprising", "having", "including", and their other forms are intended to be equivalent in meaning and are open-ended, because one or more items following any of these words do not mean an exhaustive list of such one or more items, nor are they limited to the listed one or more items.

[0031] Figure 1 is a block diagram of an exemplary embodiment of a VCM coding system that includes an encoder, a decoder, and a bitstream well-suited for machine-based video consumption (e.g., as envisioned in applications for machine video coding (VCM)). As used herein, the term VCM is not limited to a specific proposed protocol, but more generally includes all systems for encoding and decoding video for machine consumption. Although Figure 1It has been simplified to depict the components used for encoding for machine consumption, but it can be understood that the present system and method can be applied to hybrid systems that also encode, transmit, and decode video for human consumption. Systems for encoding / decoding video using various protocols (such as HEVC, VVC, AV1, etc.) are well known in the art.

[0032] Reference Figure 1 , which illustrates an exemplary embodiment of a VCM encoding system 100 including a VCM encoder 105, a VCM bitstream 155, and a VCM decoder 130.

[0033] Further reference Figure 1 , which illustrates an exemplary embodiment of a machine video coding (VCM) encoder. The VCM encoder 105 can be implemented using any circuit including but not limited to digital circuits and / or analog circuits; the VCM encoder 105 can be configured using a hardware configuration, a software configuration, a firmware configuration, and / or any combination thereof. The VCM encoder 105 can be implemented as a computing device and / or a component of a computing device, which can include but not limited to any of the computing devices described below. In one embodiment, the VCM encoder 105 can be configured to receive an input video 102 and generate an output bitstream 155. The reception of the input video 102 can be accomplished in any of the ways described below. The bitstream can include but not limited to any bitstream known in the art for advanced codecs (CODECS) (such as HEVC, AV1, or VVC or as described below).

[0034] The VCM encoder 105 can include but not limited to an inference 110 with a region extractor, a region transformer and packer 115, a packed picture converter and shifter 120, and / or an adaptive video encoder 125.

[0035] The packed picture converter and shifter 120 processes the packed images to remove more redundant information before encoding. Examples of the conversion are conversions of color spaces (such as from RGB to grayscale), quantization of pixel values (such as reducing the range of represented pixel values and thus reducing the contrast), and other conversions to remove redundancy in the sense of the machine model. Shifting requires reducing the range of represented pixel values by a direct right shift operation (for example, shifting the pixel value right by 1 is equivalent to dividing all values by 2). The conversions and shifts on the decoder side are reversed by block 140 using mathematical operations opposite to those used in 120.

[0036] Further reference Figure 1 , the adaptive video encoder 125 can include but not limited to any video encoder known in the art for advanced CODEC standards, such as HEVC, AV1, VVC, etc., or as described in further detail below.

[0037] Still referring to Figure 1 , which illustrates an exemplary embodiment of a VCM decoder. The VCM decoder 130 can be implemented using any circuitry including but not limited to digital circuitry and / or analog circuitry; the VCM decoder 130 can be configured using a hardware configuration, a software configuration, a firmware configuration, and / or any combination thereof. The VCM decoder 130 can be implemented as a computing device and / or a component of a computing device, which can include but not be limited to any of the computing devices described below. In one embodiment, the VCM decoder 130 can be configured to receive an input bitstream 155 and generate an output video 147. The reception of the bitstream 155 can be accomplished in any of the ways described below. The bitstream can include but not be limited to any of the bitstreams described below.

[0038] Continuing to refer to Figure 1 , the machine model 160 can exist in the VCM encoder 105 or be sent to the VCM encoder 105 in an online or offline mode using an available communication channel. The machine model 160 contains information that fully describes the requirements of the machine 150 for completing the task. This information can be used by the VCM encoder 105 and, in some embodiments, specifically by the region transformer and packer 115.

[0039] Given a frame of a video or image, efficient compression of such media can be achieved by detecting and extracting its significant regions and packing them into a single frame. At the same time, the system discards any detected regions that are not of interest. These packed frames serve as the input to the encoder to produce a compressed bitstream. The generated bitstream contains the encoded packed regions as well as the parameters required to reconstruct and reposition each region in the decoded frame. The machine task system 150 can perform specified machine tasks on the reconstructed video frames, such as computer vision-related functions.

[0040] Such a video compression system can be improved by performing additional processing on the unpacked frames. The improved system provides a decoder-side region filling method to provide better context for end-point machine task evaluation. Figure 2 Illustrated is the proposed video compression system 100, which includes an encoder-side module and a decoder-side module, including decoder-side region filling.

[0041] Figure 2 is a detailed block diagram showing the sub-components of a Figure 1 VCM encoding system according to an embodiment of the present disclosure. The system (previously referred to herein by reference numeral 200) shows an encoder 208 and a decoder 236. The encoder 208 includes a region detection block 212, a region extractor block 216, a region packing block 220, a region parameter block 224, and a video encoder block 228, which cooperate to generate a compressed bitstream 232.

[0042] The decoder 236 includes a video decoder 240 that receives the compressed bitstream 232, and a region unpacking module 244 coupled to a region parameter module 248, which generates an unpacked reconstructed video frame 252. A region filling module 256 coupled to the adaptive filling parameter 264 receives the unpacked reconstructed video from the region unpacking module 244 and provides a selectively filled region for the machine task system 260.

[0043] Region Detection and Extraction

[0044] The encoder-side region detector module 212 is used to identify meaningful frames or image regions, which generates the coordinates of the detected objects. A saliency-based detection method using video motion can also be employed to identify important regions. The obtained coordinates are used to determine the regions to be packed and can identify the pixels that are considered unimportant to the detection module. Such unimportant regions can be discarded and not used in the packing. The region extraction module 216 extracts the pixels of the image regions identified by 212 and prepares the coordinates for the remainder of the pipeline. The extraction module can output additional parameters to be encoded in 228. This can include information about which packed regions should receive decoder-side filling and the amount of filling.

[0045] Region Packing

[0046] The extracted region box coordinates returned from the extraction module 216 are used as the input to the region packing system 220. The region packing module extracts the meaningful image regions and tightly packs them into a single frame. The region packing module 220 generates the packing parameters to be signaled in the bitstream 232, such as in the header information or supplementary information signaled in the bitstream.

[0047] Video Encoding

[0048] The packed object frame is processed by the video encoder 228 to generate the compressed bitstream 232. The compressed bitstream includes the encoded packed regions and the parameters 224 required to reconstruct and relocate each region in the decoded frame. Additional parameters can be signaled for use in the decoder-side 236 reconstruction process, such as signaling which pixels should be used for decoder-side filling. The encoding is not limited to any specific standard and can substantially conform to known CODEC protocols, such as HEVC, AV1, VVC, etc. or their variants.

[0049] Video Decoding

[0050] The compressed bitstream 232 is decoded by the video decoder 240 to produce the packed region frames and the region parameter information signaled thereby by the region parameter module 248. Such region information generally includes the parameters required to reconstruct the frame (e.g., region coordinates, object type, etc.), as well as any additional parameters to be incorporated for padding transformation of the unpacked frame.

[0051] Region unpacking

[0052] The decoded parameters are used to unpack the region frames via the region unpacking module 244. Each region of interest identified by the encoder 208 is restored to its position within the context of the original video frame. The resulting unpacked frames include only the pixels of the regions of interest determined by the region detection system 212 and not the discarded pixels.

[0053] The proposed padding module 256 considers any adaptive parameters in the region parameters 248 and 264 to apply further processing to the unpacked frames. Such further processing focuses on expanding the region pixels in order to provide more context for the endpoint machine task evaluation by the module 260.

[0054] Padding methods can use various techniques to expand the region boundaries. This includes directly expanding the edge pixels or filling the edge pixels with a specified color. Such a specified color can be signaled by the decoded parameters 248, or such a specified color can simply be applied using the average of the pixel colors found in the region box. Additionally, prediction techniques (e.g., inpainting) can be applied to reconstruct the edge pixels around the region. Similarly, a specified pixel block can be signaled in the bitstream and tiled across the region edge to create a new texture around the region. Such transformations applied to the unpacked frames are used to provide additional context for the machine-related tasks by the machine task system 260. Applying padding to the reconstructed frames can obviate any need for region padding on the encoder side, ultimately reducing the number of bits transmitted.

[0055] The compressed bitstream can include padding information to be added to the decoder. Such padding descriptions can be included in the region description information. The padding information preferably contains information including the padding type and the padding size. The padding size can be different for each side of the rectangular region. The padding type signals to the decoder how to obtain the pixels for padding. In one signaled padding type, the edge pixels of the region are repeated to the padding size. In another signaled padding type, the average color of the pixels within the region is used for padding. In yet another signaled padding type, the pixel values for padding are included in the bitstream. Padding can also be determined adaptively based on the characteristics of the objects in the region.

[0056] Figure 3An example 300 of the extended region boundary performed by the padding module 256 is shown. In this diagram, the image 308 is the decoder-side padded version of the image in 304. Small patches of extended region pixels are represented by the numeral 316. The numeral 312 shows the same patch of pixels without padding taken from the corresponding region in the image. The extension of the edge pixels in the previous case can well reconstruct the region boundary without introducing any new artifacts or textures that may have a negative impact on machine analysis.

[0057] Machine tasks

[0058] The reconstructed, unpacked, and padded video frame 252 is used as an input to the machine task system 260, which can perform functions related to computer vision. The performance of machine tasks on the unpacked and padded frames can be analyzed and used to determine techniques for applying padding techniques on a per-frame basis. The optimization parameters 264 can be updated and signaled to the decoder-side pipeline to effectively pad regions in the unpacked frames.

[0059] Some embodiments may include a non-transitory computer program product (i.e., a physically embodied computer program product) storing instructions that, when executed by one or more data processors of one or more computing systems, cause at least one data processor to perform the operations herein.

[0060] Embodiments may include circuitry configured to implement any of the operations described above in any embodiment in any order and any degree of repetition. For example, a module (e.g., an encoder or decoder) may be configured to repeat a single step or sequence until a desired or commanded result is achieved; the repetition of steps or sequences of steps may be performed iteratively and / or recursively using the output of a previous repetition as the input to a subsequent repetition, aggregating the repeated inputs and / or outputs to produce an aggregated result, decrementing or decaying one or more variables (e.g., global variables), and / or dividing a larger processing task into a set of smaller processing tasks to be solved iteratively. An encoder or decoder may perform any step or sequence of steps described in this disclosure in parallel, e.g., performing the step two or more times simultaneously and / or substantially simultaneously using two or more parallel threads, processor cores, etc.; the division of tasks between parallel threads and / or processes may be performed according to any protocol suitable for dividing tasks between iterations. Those skilled in the art, after reviewing the entire disclosure, will recognize the various ways in which iterative, recursive, and / or parallel processing can be used to subdivide, share, or otherwise process steps, sequences of steps, processing tasks, and / or data.

[0061] A non-transitory computer program product (i.e., a physically embodied computer program product) can store instructions that, when executed by one or more data processors of one or more computing systems, cause at least one data processor to perform the operations and / or steps described in this disclosure, including but not limited to any of the operations described above and / or any operations that a decoder and / or encoder can be configured to perform. Similarly, a computer system is also described that can include one or more data processors and a memory coupled to the one or more data processors. The memory can temporarily or permanently store instructions that cause at least one processor to perform one or more of the operations described herein. Additionally, the methods can be implemented by one or more data processors within a single computing system or distributed between two or more computing systems. Such computing systems can be connected and can exchange data and / or commands or other instructions, etc., via one or more connections, including connections via a network (e.g., the Internet, wireless wide area network, local area network, wide area network, wired network, etc.), direct connections between one or more of the multiple computing systems, etc.

[0062] It should be noted that any one or more aspects and embodiments described herein can be conveniently implemented using one or more machines (e.g., one or more computing devices used as user computing devices for electronic documents, one or more server devices (such as document servers), etc.) programmed according to the teachings of this specification, as this will be obvious to those of ordinary skill in the computer art. Based on the teachings of this disclosure, a skilled programmer can easily prepare the appropriate software code, as this will be obvious to those of ordinary skill in the software art. The aspects and implementations using software and / or software modules discussed above can also include appropriate hardware for assisting in implementing the software and / or software modules.

[0063] Such software can be a computer program product employing a machine-readable storage medium. A machine-readable storage medium can be any medium that is capable of storing and / or encoding a sequence of instructions executable by a machine (e.g., a computing device) and causing the machine to perform any of the methods and / or embodiments described herein. Examples of machine-readable storage media include but are not limited to magnetic disks, optical disks (e.g., CD, CD-R, DVD, DVD-R, etc.), magneto-optical disks, read-only memory (ROM) devices, random access memory (RAM) devices, magnetic cards, optical cards, solid-state memory devices, EPROM, EEPROM, and any combination thereof. As used herein, a machine-readable medium is intended to include a single medium as well as a collection of physically separate media, such as a collection of optical disks or one or more hard disk drives combined with computer memory. As used herein, a machine-readable storage medium does not include transitory forms of signal transmission.

[0064] Such software may also include information (e.g., data) carried on a data carrier (e.g., a carrier wave) in the form of a data signal. For example, machine-executable information may be included as a data-carrying signal embodied in a data carrier, in which the signal encodes a sequence of instructions or portions thereof to be executed by a machine (e.g., a computing device), and any associated information (e.g., data structures and data) that causes the machine to perform any one of the methods and / or embodiments described herein.

[0065] Examples of computing devices include, but are not limited to, e-reader devices, computer workstations, terminal computers, server computers, handheld devices (e.g., tablets, smart phones, etc.), network appliances, network routers, network switches, network bridges, any machine capable of executing a sequence of instructions specifying actions to be taken by that machine, and any combination thereof. In one example, a computing device may include and / or be included in a kiosk.

[0066] It must also be noted that, as used herein and in the appended claims, the singular forms "a", "an", and "the" include plural references unless the context clearly dictates otherwise. Although any systems and methods similar or equivalent to those described herein may be used in practicing or testing embodiments of the present invention, preferably, the systems and methods are now described.

[0067] In the foregoing description, some terms have been used for the sake of brevity, clarity, and understanding. Except as required by the prior art, no unnecessary limitations should be implied therefrom, as such terms are used for descriptive purposes and are intended to be broadly construed. Accordingly, the present invention is not limited to the specific details, representative embodiments, and illustrative examples shown and described. Therefore, this application is intended to cover alterations, modifications, and variations that fall within the scope of the appended claims.

[0068] Furthermore, although the present invention has been described in detail with its advantages, it should be understood that various changes, substitutions, and alterations can be made herein without departing from the invention as defined by the appended claims. Moreover, the scope of this application is not intended to be limited to the specific embodiments of the processes, machines, manufactures, compositions of matter, devices, methods, and steps described in the specification. From this disclosure, one will readily recognize that processes, machines, manufactures, compositions of matter, devices, methods, or steps, existing or later developed, that perform substantially the same function or achieve substantially the same result as the corresponding embodiments described herein can be used. Accordingly, the appended claims are intended to embrace such processes, machines, manufactures, compositions of matter, devices, methods, or steps within their scope.

[0069] The foregoing description has been presented with reference to various embodiments. Those skilled in the art relevant to this application will recognize that modifications and changes can be made to the described structures and methods of operation without meaningfully departing from the principles and scope.

Claims

1. A decoder for decoding a bitstream encoded with packed frames, wherein, The packed frame has at least one region of interest defined therein and coded region parameters associated with the region of interest, and the decoder comprises: a video decoder that receives the bitstream and extracts the packed frame and the region parameters from the bitstream; a region unpacking module that receives the packed frame and the region parameters and reconstructs an unpacked frame with the at least one region of interest; and a region filling module that receives the unpacked frame and the region parameters and applies at least one filling parameter to at least one dimension of the region of interest in the unpacked frame.

2. The decoder according to claim 1, wherein, The region filling module further receives adaptive filling parameters, and the applied filling parameter is determined at least in part based on the adaptive filling parameters.

3. The decoder according to claim 1, wherein, The region of interest is delimited by a rectangular bounding box, and the applied filling parameter is pixels of a predetermined color added to at least one boundary of the region bounding box.

4. The decoder according to claim 1, wherein, The region of interest is delimited by a rectangular bounding box, and the applied filling parameter is pixels of an average color determined by pixels within the region, the pixels being added to at least one boundary of the bounding box.

5. The decoder according to claim 3 or 4, wherein, The applied filling parameter is a fixed number of pixels.

6. The decoder according to claim 3 or 4, wherein, The region filling module further receives adaptive filling parameters, and the applied filling parameter includes a variable number of pixels determined at least in part by the adaptive filling parameters.

7. The decoder according to claim 3 or 4, wherein, The filling parameter includes repeating pixels at the edge of the region of interest.

8. The decoder according to claim 3 or 4, wherein, A filling value is signaled in the bitstream, and the filling parameter is determined at least in part based on the filling value.

9. A method for decoding a bitstream having a packed frame, the packed frame having at least one region of interest defined therein and coded region parameters associated with the region of interest, comprising: receiving the bitstream; extracting the packed frame and the region parameters from the bitstream; reconstructing an unpacked frame with the at least one region of interest based on the packed frame and the region parameters; and applying at least one filling parameter to at least one dimension of the region of interest in the unpacked frame.

10. The method according to claim 9, further comprising receiving adaptive filling parameters, and the applied filling parameter is determined at least in part based on the adaptive filling parameters.

11. The method according to claim 9, wherein, The region of interest is delimited by a rectangular bounding box, and the applied filling parameter is pixels of a predetermined color added to at least one boundary of the region bounding box.

12. The method according to claim 9, wherein, The region of interest is delimited by a rectangular bounding box, and the applied filling parameter is pixels of an average color determined by pixels within the region, the pixels being added to at least one boundary of the bounding box.

13. The method according to claim 9, wherein The applied filling parameter is a fixed number of pixels.

14. The method according to claim 9, further comprising receiving adaptive filling parameters, and the applied filling parameter includes a variable number of pixels determined at least in part by the adaptive filling parameters.

15. The method according to claim 9, wherein, The filling parameter includes repeating pixels at the edge of the region of interest.

16. The method according to claim 9, wherein, A filling value is signaled in the bitstream, and the filling parameter is determined at least in part based on the filling value.

17. A computer-readable medium having instructions stored thereon that, when executed by a processor, cause the processor to perform steps including the following: Decode a compressed bitstream using an adaptive video decoder to provide a first decoded output; Unpack and inverse-transform the first decoded output using an unpacker and an inverse-transformer to provide an output video with adaptive padding transformation in region packing, wherein, The unpacker and the de-transformer receive the unpacked reconstructed video and provide a selectively padded region for machine task evaluation.