System and method for object boundary merging, segmentation, transformation and background processing in video encapsulation

Through top-down area extraction and packaging technology, the regions of interest in video are identified and processed, and the problems of low video encoding efficiency and insufficient machine task requirements in the prior art are solved, and efficient video data compression and encoding are achieved.

CN119948870APending Publication Date: 2025-05-06OP SOLUTIONS
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380068539.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-12
Filing Date
2023-09-27
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively optimize video data for machine consumption in video encoding and decoding, especially in improving efficiency and adapting to machine task requirements.

Method used

Using the top-down area extraction and packaging method, the region of interest is identified through the area detector module, and the top-down area extractor module merges and transforms to form a more efficient encapsulated frame and encodes it into an encoded bitstream.

Benefits of technology

It realizes efficient compression and encoding of video data, reduces unnecessary pixels, improves encoding efficiency, and optimizes the performance of machine task system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948870A_ABST
    Figure CN119948870A_ABST
Patent Text Reader

Abstract

Systems and methods for encoding and decoding video content for machine consumption with an enhanced region encapsulation policy. The encoder includes a region detector module that receives a source video and identifies a region of interest therein. A top-down region extractor module receives the identified regions of interest and generates a modified set of regions of interest that can be more efficiently encapsulated in the frame. A region encapsulation module receives the modified set of regions of interest and arranges the modified set of regions of interest into an encapsulation frame in which pixels outside the modified regions of interest are substantially excluded. The video encoder encodes the encapsulated frame and the region parameter into an encoded bitstream. A compatible decoder provides complementary processing to reconstruct a frame having a region of interest as arranged in a source frame.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of priority to U.S. Provisional Application Serial No. 63 / 410,251, filed on September 27, 2023, entitled “System and Method for Top-Down Object Boundary Merging and Splitting in Video Packing,” and also claims the benefit of priority to U.S. Provisional Application Serial No. 63 / 410,266, filed on September 27, 2023, entitled “System and Methods for Region Transformations in Video Box Packing,” and also claims the benefit of priority to U.S. Provisional Application Serial No. 63 / 410,266, filed on September 27, 2023, entitled “System and Methods for Adaptive Frame Reconstruction for Video No. 63 / 410,272, filed on October 12, 2022, entitled “Systems and Methods for Video Packing, Encoding and Decoding for Machine-based Applications,” and also claims the benefit of priority to U.S. Provisional Application Serial No. 63 / 415,376, filed on October 12, 2022, entitled “Systems and Methods for Video Packing, Encoding and Decoding for Machine-based Applications,” the disclosures of each of the above applications being incorporated herein by reference in their entirety. Technical Field

[0003] The present disclosure generally relates to the field of video encoding and decoding. In particular, the present disclosure relates to encoding and decoding of video for machines. Background Art

[0004] Recent trends in robotics, surveillance, monitoring, IoT, etc., have introduced use cases where a large portion of all images and videos recorded in the field are consumed only by machines and never reach human eyes. These machines process images and videos with the goal of completing tasks such as object detection, object tracking, segmentation, event detection, etc. Recognizing that this trend is pervasive and will only accelerate in the future, international standardization bodies have made efforts to standardize image and video coding that are primarily optimized for machine consumption. For example, standards such as JPEG AI and video coding for machines have been initiated in addition to already established standards such as Compact Descriptors for Visual Search and Compact Descriptors for Video Analytics. Solutions that improve efficiency compared to classical image and video coding techniques are needed. One such solution is proposed here. Summary of the invention

[0005] In one embodiment, an encoder for video for machine consumption is provided, comprising a region detector module that receives a source video and identifies regions of interest therein. A top-down region extractor module receives the identified regions of interest and generates a modified set of regions of interest that can be more efficiently packed into a frame, the modified regions of interest being defined at least in part by region parameters. A region packing module receives the modified set of regions of interest and arranges the modified set of regions of interest into a packed frame in which pixels outside the modified regions of interest are substantially excluded. A video encoder receives the packed frame and the region parameters and encodes the packed frame and the region parameters into an encoded bitstream.

[0006] The region of interest may be defined by a bounding box, such as a rectangular bounding box, and the region parameters may include coordinates of the region bounding box within the source video frame.

[0007] The top-down region extractor module may also provide processing of the detected regions of interest to form a merge of regions of interest, wherein at least one of adjacent and overlapping regions of interest are combined. In some embodiments, the processing further includes aligning the regions of interest from the merge process with a predetermined grid, slicing the aligned regions of interest along the grid partitions, and reattaching the slices to form modified regions of interest. The top-down region extractor module may then provide coordinates of the modified regions of interest.

[0008] In some embodiments, the predetermined grid is selected to align the region of interest with a boundary of a coding tree unit in a packed frame. In some embodiments, the predetermined grid is a 16x16 pixel grid.

[0009] In addition, the encoder may include a region transform module inserted between the top-down region extractor module and the region encapsulation module, the region transform module receiving the coordinates of the modified regions of interest and applying at least one transformation from the group including scaling, rotation, and translation to at least one modified region of interest and providing the transformed coordinates of the modified region of interest. In some embodiments, the region transform module receives adaptive transformation parameters related to the machine process and applies the selected transformation based in part on these parameters. In another embodiment, the region transform module applies the transformation based on at least one characteristic of the region of interest including object confidence, object category, region area, or coding unit parameters.

[0010] In another embodiment, an encoder for encoding video for machine consumption includes a region detector module that receives a source video and identifies a region of interest therein, the region of interest being defined in part by coordinates of a bounding box in a frame of the source video. A region transform module receives the coordinates of the region of interest and applies at least one transform from the group consisting of scaling, rotation, and translation to at least one modified region of interest to improve region packing efficiency and provide the transformed coordinates of the modified region of interest. The region packing module receives a set of regions of interest and the transformed regions of interest and arranges the regions of interest into a packed frame in which pixels outside the regions of interest are substantially excluded. A video encoder receives the packed frame and the coordinates of the regions of interest and encodes the packed frame and region parameters into an encoded bitstream.

[0011] A method for encoding a source video for machine consumption is also provided. The method preferably includes receiving a source video and identifying a region of interest therein. The detected regions of interest may be processed to form a merge of regions of interest, wherein at least one of adjacent and overlapping regions of interest are combined. The method also includes aligning the regions of interest from the merge process with a predetermined grid, slicing the aligned regions of interest along the grid partitions, and reattaching the slices to form a modified region of interest. Coordinates of the modified region of interest are provided, and the modified region of interest is arranged into a packed frame in which pixels outside the modified region of interest are substantially excluded. Region parameters and the packed frame containing the coordinates of the modified region are encoded in the encoded bitstream.

[0012] Decoders and decoding methods are provided. Preferably, the decoder is configured to receive an encoded bitstream encoded by any encoder and encoding method described herein, and comprises circuitry configured to decode the bitstream and reconstruct a frame having a region of interest from a source video while excluding pixels outside the region of interest.

[0013] In one embodiment, a decoder for decoding a coded bitstream of a packed frame having a region of interest is characterized by enhanced processing of background pixels. The decoder includes a video decoder that receives a coded bitstream and decompresses the bitstream to identify a region of interest and region parameters therefrom. A region decapsulation module receives the decoded packed frame and region parameters and arranges the decoded region of interest in a reconstructed frame having a size, position and orientation corresponding to the original frame. Pixels in the reconstructed frame outside the arranged region of interest are considered to be background pixels. A background processing module is provided and sets parameters of background pixels to optimize the performance of a machine task system that receives the reconstructed frame.

[0014] In some embodiments, the background processing module receives at least one adaptive fill parameter indicative of a performance metric of the machine task system based on at least one parameter of the background pixels. The parameter of the background pixels may be a fixed color or an average color of the pixels in the region of interest. In some cases, the adaptive fill parameter indicates a parameter of the background pixels in which object detection by the machine task system is optimized.

[0015] These and other aspects and features of non-limiting embodiments of the present invention will become apparent to those skilled in the art from the following description of specific non-limiting embodiments of the invention read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a simplified block diagram of a system for encoding and decoding video for a machine, such as in a system for video coding for machine (VCM).

[0017] Figure 2 is a simplified block diagram of a system for encoding and decoding video for a machine, such as a video coding for a machine ("VCM"), with region packing with top-down region extraction in accordance with the present disclosure.

[0018] Figure 3 is a block diagram illustrating the structure and operation of a top-down region extractor according to the present disclosure.

[0019] Figure 4A is a graphical representation of an exemplary image, where the inference prediction is given as Figure 3 The input of the top-down region extractor.

[0020] Figure 4B Shown in Figure 3 The merging of the inference predictions generated in process 308 in .

[0021] Figure 4C Shown in Figure 3 The polygon transformation process 312 is transformed into a 16x16 grid alignment Figure 4B Merging of inference predictions.

[0022] Figure 4D Description from Figure 3 The bounding boxes of the aligned polygons are further processed in the polygon transformation process.

[0023] Figure 4E Shown in Figure 3 After the region segmentation process 316 Figure 4D The bounding box of

[0024] Figure 4F The original image with processed inferences after the reattachment slice process 320 is shown.

[0025] Figure 5 is a simplified block diagram of a system for encoding and decoding video for a machine, such as a video coding machine ("VCM"), utilizing region packing according to the present method.

[0026] Fig. 6A and Figure 6B is a diagram illustrating examples of region packed frames with and without transforms applied, respectively, according to the present systems and methods.

[0027] Figure 7 is a simplified block diagram of an alternative embodiment of a decoder video for machine, such as a video coding for machine ("VCM"), for decoding a bitstream with region packing according to the present method.

[0028] Fig. 8A is a schematic diagram showing an example of a decapsulated frame with a black background, and Figure 8B is the same schematic diagram with an "average" background in accordance with the present systems and methods.

[0029] The drawings are not necessarily drawn to scale and may be illustrated by phantom lines, diagrammatic representations, and partial views. In certain instances, details that are not necessary for understanding the embodiments or that render other details difficult to perceive may have been omitted. DETAILED DESCRIPTION

[0030] Figure 1 is a block diagram illustrating an exemplary embodiment of a system well suited for machine-based video consumption, including an encoder, a decoder, and a bitstream, such as contemplated in applications for video coding for machines ("VCM"). Figure 1It has been simplified to depict the components used in encoding for machine consumption, but it should be understood that the present systems and methods are also applicable to hybrid systems that also encode, transmit and decode video for human consumption. Such systems for encoding / decoding video of various protocols such as HEVC, VVC, AV1, etc. are well known in the art.

[0031] Reference now Figure 1 , shows an exemplary embodiment of an encoding system 100 including an encoder 105 that generates an encoded bitstream 155 that is sent to a decoder 130 via a communication channel.

[0032] Further references Figure 1 , the encoder 105 may be implemented using any circuit, including but not limited to digital and / or analog circuits; the encoder 105 may be configured using hardware configuration, software configuration, firmware configuration, and / or any combination thereof. The encoder 105 may be implemented as a computing device and / or a component of a computing device, which may include but not limited to any computing device as described below.

[0033] The encoder 105 may include, but is not limited to, an inference module 110 having a region extractor, a region transformer and packer module 115 , a packed picture converter and shifter module 120 , and / or an adaptive video encoder module 125 .

[0034] Further references Figure 1 The adaptive video encoder 125 may be a standard encoder for generating a bitstream conforming to a known CODEC standard such as HEVC, AV1, VVC, etc., and may include, but is not limited to, any video encoder as described in further detail below.

[0035] Still reference Figure 1, an exemplary embodiment of a decoder 130 is shown. The decoder 130 can be implemented using any circuit, including but not limited to digital and / or analog circuits; the decoder 130 can be configured using hardware configuration, software configuration, firmware configuration, and / or any combination thereof. The decoder 130 can be implemented as a computing device and / or a component of a computing device, which can include but is not limited to any computing device as described below. In an embodiment, the decoder 130 can be configured to receive an encoded bitstream 155 and generate an output video 147 suitable for machine consumption and / or a video for human consumption in a hybrid system. The reception of the bitstream 155 can be implemented in any manner described below. The bitstream can include but is not limited to any bitstream as described below. The present system and method are not limited to a specific CODEC standard and are applicable to current standards such as AV1, HEVC, VVC, etc. and their variants and improvements. It should be understood that the specific structure and operation of the encoder 105 and the decoder 130 will depend in part on the specific coding standard used in the deployed system, and such encoders and decoders are known in the art.

[0036] Continue to refer Figure 1 , the machine model 160 may reside in the encoder 105 or otherwise be provided to the encoder 105 in an online or offline mode using an available communication channel. The machine model 160 is application / task specific and generally contains information sufficient to describe the requirements of the machine 150 to complete the task. The machine 150 may provide periodic updates to the machine model based on system updates or data related to processing performance. This information may be used by the encoder 105, and in some embodiments specifically by the region transformer and encapsulator 115.

[0037] Given a frame of video or image, efficient compression of such media for machine consumption can be achieved by detecting and extracting its important regions and packing them into a single frame. At the same time, the system discards any detected regions of no interest. These packed frames are then used as input to an encoder to produce a compressed bitstream. The resulting bitstream 155 contains the encoded packed regions and the parameters required to reconstruct and reposition each region in its original position within the decoded frame. The machine task system can perform tasks such as specified computer vision related functions on the reconstructed video frames.

[0038] The present method for video encoding for machine consumption is a region-based system that divides the input image into regions of interest that are retained in the encoded bitstream and regions that are not of interest and are discarded at the encoder. This video compression system can be improved by applying the post-inference region extraction techniques described herein to create new region boxes. These new regions can be derived from the boxes output by the region detection module. Some extraction modules may perform this type of post-processing on a per-box basis; however, the improved system and method described herein processes predictions in a top-down approach. This processing step can eliminate the occurrence of repeated pixels within the region image and can also give more object context, which is generally beneficial for endpoint machine task execution.

[0039] Figure 2 2 is a simplified block diagram of a system for region encapsulation according to the present disclosure, which shows a proposed CODEC system 200, which includes an encoder 208 and a decoder 232. The encoder includes a region detector module 212, which receives a source video 204 and identifies regions of interest and / or objects therein. The encoder also includes a top-down region extractor module 256, a region encapsulation module 216, a region parameter module 220, and a video encoder 224 that cooperate to generate a compressed bitstream 228. Preferably, adaptive extraction parameters 260 are provided to the top-down region extractor 256.

[0040] Decoder 232 includes a video decoder 236 that receives compressed bitstream 228 , a region decapsulation module 240 , and a region parameter module 244 that cooperate to decode the bitstream and generate decapsulated reconstructed video frames 248 for use by machine task system 252 .

[0041] Area Detection

[0042] Significant frame or image regions are identified using an encoder-side region detector module 212, which generates coordinates of the discovered objects. Typically, a region is defined at least in part by a bounding box, and the coordinates are the vertices of the bounding box. The bounding box can be rectangular and can be defined in any way sufficient to recreate the region of interest in the original frame, for example, by the coordinates of one corner and the width and height of the bounding box, the coordinates of the diagonally opposite corner, etc. The region detector module can also output category information for each identified object and a confidence score. In one embodiment, the Yolov7 object detection neural network can be used with a 0.0 confidence threshold. The identified object coordinates can be extended with region padding to give additional context to each detection. Filling inference objects can improve endpoint machine task performance. In one example, the inference boundary can be extended by 15 pixels in two dimensions. However, it should be understood that other region padding strategies can be adopted, including dynamic region padding, where the amount of padding in each dimension of the bounding box can depend on criteria such as object type or region size.

[0043] A saliency-based detection method using video motion can also be used to identify regions of interest. For example, uniform motion detected across consecutive frames can be designated as salient. In another example, any motion that lasts for a long period of time (e.g., 100 frames) in a continuous trajectory can be designated as salient. In another example, motion detected at the same coordinates where an object is detected can be designated as salient. The spatial coordinates of the salient region are used to determine the region for encapsulation and enable identification of pixels that are considered unimportant to the detection module. Such unimportant regions can be discarded and not used for encapsulation.

[0044] Top-down region extraction

[0045] The structure and operation of the top-down region extractor module 256 are described in detail in Figure 3 Further shown in the reference Figure 3 , the top-down region extractor module operates to provide Figure 4B The merge of detected regions 308 is shown, as Figure 4C and Figure 4D As shown, polygon transformation is performed on region 312, region segmentation is performed 316, and Figure 4E and Figure 4F Reattachment of slice 320 is shown being performed.

[0046] The top-down region extractor module 256 receives the object coordinates from the region detector module 212. The top-down region extractor module 256 examines the predictions in a top-down approach by merging all region detections into one or more polygons. Depending on the overlap that occurs, a merge of the inferred predictions 308 is taken to create new region polygons. These new region polygons may have irregular shapes and therefore require further processing in order to be used as Figure 2 The input of the encapsulation module 216 in.

[0047] The merged extrapolated prediction region identified in 308 may be expanded by the polygon transformation module 312 to reduce the number of irregular edges (ie, vertices outside the overall rectangular bounding coordinates), such as Figure 4C As shown. For example, polygonal regions can use a 16x16 grid to fill and expand the merged inferred region. Alignment with such a 16x16 grid can also result in better alignment with coding units of a video encoder, such as coding tree units (CTUs) in a later encoding stage, i.e., for video encoder module 224. This alignment can improve compression efficiency because prediction and residual coding are done within the coding block boundaries of each block. The alignment is further disclosed in PCT application PCT / US22 / 47829, entitled “Systems and Methods for Object and Event Detection and Feature Based Rate Distortionfor Video Coding,” filed on October 26, 2022, the disclosure of which is incorporated herein by reference in its entirety.

[0048] The resulting transformed polygon from the polygon transformation module 312 can be segmented into rectangular sub-boxes, such as Figure 4E . The partitioning performed by the region segmentation module 316 is accomplished by slicing the input polygon based on non-rectangular vertices. That is, vertices outside the minimum and maximum x, y coordinates are used as reference points for slicing. Vertical and horizontal lines are drawn outward from the reference points to intersect the edges of the polygon. These lines form the new boundaries of the sliced ​​regions. The resulting fragments can be processed directly by the region packing system or can undergo additional processing.

[0049] Slice regions may be greedily and recursively reattached by the reattachment module 320 based on shared edges to create better boxes for encapsulation. Reattachment may be performed by merging two boxes with shared edges to form a new region. This reattachment is considered "greedy" because it may take into account the characteristics of the box regions to determine the order in which the regions are processed. For example, slice regions may be sorted by area, category, and confidence so that reattachment of certain boxes takes precedence over other boxes. In some cases, reattaching boxes with larger areas first may help preserve regional characteristics. Similarly, confidence-based reattachment may help better preserve objects with high inference scores. Reattaching boxes may help reduce the number of rectangles sent to the encapsulation module 216 and may generally help provide more contextual space reservation for the endpoint machine task system 252. Figure 4F Shows Figure 4A An example of a sample image where the reattached slice forms a new region of interest.

[0050] 4A to 4F A pictorial example of a top-down region extractor process is shown. Figure 4A An image with inference predictions is shown. Figure 4B The merging of the detected extrapolated coordinates from block 308 is shown in FIG. Figure 4C As further shown in , padding and alignment are applied to polygonal regions using a 16x16 grid at 312 . Figure 4D The shape obtained in can be obtained by module 316 ( Figure 4E ) segmentation, wherein reattachment is performed by reattachment module 320, such as Figure 4F At 324, the final coordinates are overlaid on the source image.

[0051] Table 1 compares the top-down region extraction approach with the per-object merge-split region extraction approach. The merge-split approach considers the occurrence of overlapping predictions to be merged together or split into separate smaller regions on a per-prediction basis. According to the results, the top-down approach produces more efficient bit reduction and additionally maintains higher machine task performance. This approach provides better area for packing, thereby improving overall pipeline efficiency.

[0052]

[0053] Table 1. Top-down region extraction

[0054] Regional Encapsulation

[0055] The newly identified region box coordinates 324 returned from the extraction module 356 are used as input to the region packing module 216. The region packing module 216 extracts the regions of interest and packs them tightly into a single frame. The region packing module 216 generates packing parameters that are preferably signaled later in the bitstream 228.

[0056] Video Encoding

[0057] refer to Figure 2 , the encapsulated object frame containing the processed regions from the top-down region extractor 256 is input to the video encoder 224, which produces a compressed bitstream 228. It should be understood that the video encoder 224 can take the form of any suitable encoder known in the art for advanced compression standards (such as AV1, HEVC, VVC, etc. and / or their variants). The compressed bitstream 228 typically contains the encoded encapsulated regions and the parameters 220 required to reconstruct and reposition each region in the decoded frame. Optionally, the original region coordinates, that is, those derived from the region detector module 212 before the top-down region extraction 256, can be signaled in the bitstream for use by the decoder side. Signaling can include providing data in the bitstream header, SPS, PPS, or supplementary information (e.g., SEI), and can vary depending on the CODEC standard in which the present system and method are deployed.

[0058] Video Decoding

[0059] The compressed bitstream 228 is decoded using a video decoder 236 to produce a packed region frame and its signaled region information 244. Such region information preferably includes parameters sufficient to reconstruct the frame and may be combined with the original object coordinate information from the region detector 212 and the region information identified in the top-down region extractor 256.

[0060] Regional decapsulation

[0061] The decoded parameters 244 are used to depacketize the packed region frame via the region depacketization module 240. Each frame is preferably returned to its position within the context of the original video frame. The resulting depacketized frame includes the detected regions of interest placed at their original positions in the frame prior to packing, and includes only the salient regions determined by the region detection system 212 after the extractor module 256, and preferably does not include pixels outside the regions of interest.

[0062] Machine Tasks

[0063] The depacketized and reconstructed video frames 248 are used as input to a machine task system 252, which can perform machine tasks such as computer vision related functions. The machine task performance on the regions determined by the merge-segment extraction module of the top-down region extractor 156 can be analyzed to determine the best extraction actions. The optimized region extraction parameters 260 can be updated and signaled to the encoder-side pipeline to effectively unify and segment the prediction regions.

[0064] Figure 5 5 is a simplified block diagram of an alternative embodiment of an encoder with region encapsulation according to the present disclosure. The encoder 508 includes a region detection block 512 that receives a source video 504 and identifies regions of interest therein. The encoder also includes a region extractor module 516, a region transform module 560, a region encapsulation module 520, region parameters 524, and a video encoder 528 to generate a compressed bitstream 532. The adaptive transform parameters 564 are preferably provided to the region transform module 560.

[0065] Region Detection and Extraction

[0066] The encoder side region detector module 512 is used to identify significant frames or image regions. The region of interest is usually defined by a bounding box (such as a rectangular bounding box), and the region detector module 512 generates the coordinates of the bounding box. A detection method based on saliency using video motion can also be used to identify important regions. For example, uniform motion detected across consecutive frames can be designated as significant. In another example, any motion that lasts for a long period of time (e.g., 100 frames) in a continuous trajectory can be designated as significant. In another example, motion detected at the same coordinates where an object is detected can be designated as significant. The spatial coordinates of the significant region are used to determine the region for encapsulation and enable identification of pixels that are considered unimportant to the detection module. Such unimportant regions can be discarded and not used for encapsulation. The extraction module 516 extracts the image region identified by the region detector module 512 and prepares the coordinates to be used by the rest of the pipeline.

[0067] Region Transformation

[0068] The region transform module 560 receives the object coordinates from the region extractor module 516. The region transform module 560 can adaptively apply transformations such as scaling, rotation, and / or translation on a per-region basis. Internal decisions made within the module can apply these actions using confidence-based, category-based, region-based, or coding unit-based approaches.

[0069] The purpose of applying the transform is to reduce the bitrate budget required to encode frames containing the transformed regions. In some cases, the transform can be applied to improve the detection accuracy on a machine.

[0070] Fig. 6A shows a packed region frame of samples to which no transform is applied, and Figure 6B The illustration shows the same frame with the transformation applied. Figure 6B It is illustrated that the applied transformation not only reduces the overall frame size, in this example from 338x224 pixels to 160x142 pixels, but may also additionally affect and improve the packing arrangement.

[0071] The class-based transform may consider classes of objects that quickly degrade the machine task performance metric when scaling is applied. That is, of the classes of objects present in the video frame, the module should consider which objects can be scaled and by how much. Adaptive transform parameters 564 may be used to indicate classes of objects that may be scaled and classes of objects that should be preserved in size. This may be based in whole or in part on performance metrics from the machine task system 252. The detected objects from the region detector 512 within the extracted regions may be used to identify which classes are present in each region.

[0072] The input transformation parameters 564 can be determined by examining the machine task performance across different scaling factors on a per-class basis. This can be done by taking video or image samples of a type similar to (or from) the type of data seen at the source video 504 and evaluating its behavior at different scales. Additionally, the scaling parameters 564 for each class can be determined by analyzing the data on which the endpoint machine task system 556 was trained. That is, it may be beneficial to identify the characteristics of the objects in the training dataset. For example, a machine task system trained on a dataset with small objects can implement more aggressive scaling from the region transformation module.

[0073] Area-based scaling provides an alternative solution to determine the scaling factor of a region box. In class-unknown scenarios, the relative size of the extracted regions (and / or the size of the objects present in the regions) can be used to perform scaling. For example, a box containing relatively large objects can be scaled more than a region with smaller objects.

[0074] Scaling can be applied using a coding tree unit (CTU) aware approach so that each frame can be encoded more efficiently. Scaling each region box to better align with the coding tree unit reduces the frame size while also optimizing the packed frame for encoding.

[0075] A rotational transformation may be applied to create an optimal box for packing. This may be performed on a per-region basis based on the characteristics of the frame and the objects present. That is, the box may be rotated to create a better packed frame that may be encoded more efficiently.

[0076] Such a region transformation is intended to reduce the number of bits in the compressed bitstream 532 while ensuring that the machine task 252 ( Figure 2) system performance is improved or maintained. Table 1 compares a region packing system without any region transform applied to a system using a class-based transform approach. The results show that this transform approach improves overall pipeline performance. In this case, the class-based technique significantly reduces the number of bits per pixel while maintaining performance for the machine task.

[0077]

[0078] Table 1. Results of the region encapsulation system for class-based transformations

[0079] Regional Encapsulation

[0080] The transformed region box coordinates returned from the transform module 560 are used as input to the region packing system 520. The region packing module extracts the salient image regions and packs them tightly into a single frame. The region packing module 520 generates packing parameters which are preferably signaled in the bitstream 532.

[0081] Video Encoding

[0082] The encapsulated object frame containing the transformed regions from the region transform module 560 is processed by a video encoder 528 to produce a compressed bitstream 532. It should be understood that the video encoder 528 can take the form of any suitable encoder known in the art for advanced compression standards (such as AV1, HEVC, VVC, etc. and their variants). The compressed bitstream contains the encoded encapsulated regions and the parameters 524 required to reconstruct and reposition each region in the decoded and reconstructed frames. The original region sizes (i.e., those derived from the inference 512 before the region transform 560) are signaled in the bitstream 132 for use by the decoder side. Signaling can include providing data in the bitstream header, SPS, PPS, or supplementary information (e.g., SEI), and can vary depending on the CODEC standard in which the present systems and methods are deployed.

[0083] For decoding Figure 5 The encoder encodes the bit stream and the decoder Figure 2 The decoder described in is substantially the same. In 240, the compressed bitstream 532 is decoded using the video decoder 236 to produce a packed region frame and the region information it signals. Such region information 248 includes the parameters required to reconstruct the frame. This includes any transform information signaled from the transform module 560.

[0084] The decoded parameters 248 are used to depacketize the region frame via the region depacketization module 244. Each box is returned to its position within the context of the original video frame. The resulting depacketized frame includes only the salient regions determined by the region detection system 512 and does not include discarded pixels. The transformation performed by the transformation module 560 is undone by the complementary process in the depacketization stage. That is, each box is returned to its original size and orientation before being placed in the depacketized frame.

[0085] The depacketized and reconstructed video frames 248 are used as input to the machine task system 252, which can perform tasks such as computer vision related functions. The machine task performance on the region determined by the detection module 512 can be analyzed to determine the optimal transformation parameters. The optimized region transformation parameters 564 can be updated and signaled to the pipeline to effectively apply the transformation to the encoder side region box.

[0086] Figure 7 7 is a block diagram of an alternative embodiment of a decoder according to the present disclosure. The decoder 732 includes a video decoder 736 that receives a compressed bitstream 728, a region decapsulation module 740, and region parameters 744 that cooperate to generate a decapsulated reconstructed video frame 748. The decoder 732 also includes a background processing module 750 that receives the decapsulated reconstructed video frame and the adaptive padding parameters 764. The background processing module 750 is coupled to a machine task system 752.

[0087] Region Detection and Extraction

[0088] Significant frames or image regions are identified using an encoder-side region detector module 512, which produces the coordinates of the objects found. A saliency-based detection method using video motion can also be used to identify important regions. The resulting coordinates are used to determine regions for encapsulation and enable identification of pixels that are considered unimportant to the detection module. Such unimportant regions can be discarded and not used for encapsulation. An extraction module 516 extracts the image regions identified by 512 and prepares the coordinates to be used by the rest of the pipeline. The extraction module can output additional parameters and regions to be encoded in 528. This can include one or more small patches of background pixels used later in the decapsulation module 740.

[0089] Video Encoding

[0090] The packed object frames are processed by a video encoder 528 to produce a compressed bitstream 532. The compressed bitstream contains the encoded packed regions and the parameters needed to reconstruct and reposition each region in the decoded frame 524. Additional parameters may be signaled for use in the decoder side 536 reconstruction process, such as signaling which pixels belong to the background and which pixels contain objects.

[0091] Video Decoding

[0092] The compressed bitstream 532 is received by decoder 732 and decoded by video decoder 736 to produce a packed region frame and its region information signaled in 748. The region information includes the parameters needed to reconstruct the frame and any additional parameters incorporated to perform background transformation on the depacketized frame.

[0093] Regional decapsulation

[0094] The decoded parameters are used to depacketize the packed region frame via the region depacketization module 740. Each box is returned to its position within the context of the original video frame. The resulting depacketized frame 748 includes only the salient regions determined by the region detection system 512 and does not include discarded pixels.

[0095] The proposed background processing module 750 takes into account the region parameters 744 and any adaptive parameters 764 to apply further processing to the decapsulated frame. Such further processing focuses on reconstructing the discarded pixels from the region detector 512. This includes applying different background filling techniques to provide more context for the machine task system 752.

[0096] refer to Fig. 8A , the default black background color in the decapsulated frame may degrade machine task system performance. Figure 8B As shown, alternative background filling methods can use the average color of some (or all) pixels within the area box to create a new background in the decapsulated frame. Background prediction techniques, such as repair, can also be applied to reconstruct background pixels. Optionally, bitstream parameters 744 can be used to signal a specific fill color to be applied to the decapsulated frame. Similarly, a specified tile of background pixels can be signaled in the bitstream and tiled across black areas to create a new background texture. This transformation applied to the decapsulated frame is used to provide additional context for machine-related tasks performed by 752.

[0097] Background color may have an impact on machine task prediction and performance. Fig. 8A , using a black background in 804 causes some false positive inference predictions to occur. The false predictions can be minimized by replacing background pixels with the average background color, such as Figure 8BSuch false positive predictions may affect performance; therefore, it is important to consider the sensitivity of such systems to background pixels and colors.

[0098] Machine Tasks

[0099] The depacketized and reconstructed video frames 748 are used as input to a machine task system 752, which can perform machine tasks, such as computer vision related functions. The machine task performance on the depacketized frames containing the filled background can be analyzed and used to determine the background filling technique on a per-frame basis. The optimized parameters 764 can be updated and signaled to the decoder side pipeline to effectively fill the background pixels in the depacketized frames.

[0100] In general, the processed frames of the method disclosed herein are encoded using a standard CODEC process (e.g., VVC compliant with the VTM12 encoding process). The CODEC can be modified to accept input region parameters as disclosed herein. Such region parameters are preferably included in the compressed bitstream and decoded by the corresponding video decoder. The original bounding box size and the encapsulated box size are preferably recorded to track any applied transformations. The box parameters are defined as the original box coordinates generated by the top-down extractor and the corresponding encapsulation position. The box parameter encoding can be improved by utilizing a block alignment process (such as 16×16 block alignment) and CTU scaling so that the encoded box size and position are in units of 16. The present implementation encodes the parameters in a simplified form. Given a box parameter p, the simplified form p' is defined as follows:

[0101]

[0102] In some embodiments, the frame parameters (position and size of the original frame and the packed frame) can be encoded in the slice header of the frame. The bitstream decoder obtains the frame parameters and decodes the packed frame. The frame parameters are used to decapsulate and reconstruct the frame for machine processing.

[0103] Some embodiments may include a non-transitory computer program product (ie, a physically embodied computer program product) storing instructions that, when executed by one or more data processors of one or more computing systems, cause at least one data processor to perform the operations herein.

[0104] Embodiments may include circuits configured to implement any operation as described in any embodiment above in any order and with any degree of repetition. For example, a module such as an encoder or decoder may be configured to repeatedly perform a single step or sequence until a desired or commanded result is achieved; the output of a previous repetition may be used as an input to a subsequent repetition to iteratively and / or recursively perform repetitions of a step or sequence of steps, aggregate the input and / or output of the repetition to produce an aggregated result, reduce or decrement one or more variables such as global variables, and / or divide a larger processing task into a set of iteratively addressed smaller processing tasks. The encoder 500 may perform any step or sequence of steps as described in the present disclosure in parallel, such as using two or more parallel threads, processor cores, etc. to perform steps twice or more simultaneously and / or substantially simultaneously; the task division between parallel threads and / or processes may be performed according to any protocol suitable for dividing tasks between iterations. Those skilled in the art will know various ways in which steps, sequence of steps, processing tasks, and / or data may be subdivided, shared, or otherwise processed using iteration, recursion, and / or parallel processing when reading the entire contents of the present disclosure.

[0105] A non-transient computer program product (i.e., a physically embodied computer program product) may store instructions that, when executed by one or more data processors of one or more computing systems, cause at least one data processor to perform the operations and / or steps thereof described in the present disclosure, including but not limited to any of the operations described above and / or any operations that a decoder and / or encoder may be configured to perform. Similarly, a computer system that may include one or more data processors and a memory coupled to the one or more data processors is also described. The memory may temporarily or permanently store instructions that cause at least one processor to perform one or more operations described herein. In addition, the method may be implemented by one or more data processors within a single computing system or distributed between two or more computing systems. Such computing systems may be connected via one or more connections, including connections through a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), via direct connections between one or more of a plurality of computing systems, etc., and may exchange data and / or commands or other instructions, etc.

[0106] It should be noted that any one or more aspects and embodiments described herein can be conveniently implemented using one or more machines (e.g., one or more computing devices used as user computing devices for electronic documents, one or more server devices, such as document servers, etc.) programmed according to the teachings of this specification, which is obvious to those of ordinary skill in the computer arts. Based on the teachings of this disclosure, a skilled programmer can easily prepare appropriate software coding, which is obvious to those of ordinary skill in the software arts. The aspects and implementations using software and / or software modules discussed above may also include appropriate hardware for assisting in implementing the machine executable instructions of the software and / or software modules.

[0107] Such software can be a computer program product using a machine-readable storage medium. A machine-readable storage medium can be any medium capable of storing and / or encoding a sequence of instructions executed by a machine (e.g., a computing device) and causing the machine to perform any of the methods and / or embodiments described herein. Examples of machine-readable storage media include, but are not limited to, disks, optical disks (e.g., CDs, CD-Rs, DVDs, DVD-Rs, etc.), magneto-optical disks, read-only memory "ROM" devices, random access memory "RAM" devices, magnetic cards, optical cards, solid-state memory devices, EPROMs, EEPROMs, and any combination thereof. As used herein, machine-readable media is intended to include a single medium and a collection of physically separated media, such as, for example, a collection of optical disks or one or more hard disk drives combined with a computer memory. As used herein, a machine-readable storage medium does not include signal transmission in transient form.

[0108] Such software may also include information (e.g., data) carried as a data signal on a data carrier (such as a carrier wave). For example, machine executable information may be included as a data-bearing signal embodied in a data carrier, wherein the signal encodes a sequence of instructions or a portion thereof executed by a machine (e.g., a computing device) and any related information (e.g., data structures and data) that causes the machine to perform any of the methods and / or embodiments described herein.

[0109] Examples of computing devices include, but are not limited to, electronic book reading devices, computer workstations, terminal computers, server computers, handheld devices (e.g., tablet computers, smart phones, etc.), network devices, network routers, network switches, bridges, any machine capable of executing a sequence of instructions specifying actions to be taken by the machine, and any combination thereof. In one example, the computing device may include and / or be included in an all-in-one machine.

Claims

1. An encoder for video for machine consumption, comprising: a region detector module that receives a source video and identifies a region of interest therein; a top-down region extractor module that receives the identified regions of interest and generates a modified set of regions of interest that can be more efficiently packed into a frame, the modified regions of interest being defined at least in part by region parameters; a region packing module, the region packing module receiving the modified set of regions of interest and packing the modified set of regions of interest into a packing frame, in which pixels outside the modified regions of interest are substantially excluded; as well as A video encoder receives the packed frame and region parameters and encodes the packed frame and region parameters into a coded bit stream.

2. The encoder according to claim 1, wherein The top-down region extractor module further comprises: processing the detected regions of interest to form a merge of regions of interest, wherein at least one of adjacent and overlapping regions of interest are combined; aligning the region of interest from the merging process to a predetermined grid; Slicing the aligned region of interest along the grid partitions; reattaching the slices to form a modified region of interest; and The coordinates of the modified region of interest are provided.

3. The encoder according to claim 2, wherein: The region of interest is defined by a rectangular bounding box, and wherein the region parameters include coordinates of the region bounding box within the source video frame.

4. The encoder according to claim 2, wherein: The predetermined grid is selected to align the region of interest with boundaries of coding tree units in the encapsulated frame.

5. The encoder according to claim 2, wherein: The predetermined grid is a 16×16 pixel grid.

6. The encoder according to claim 3 further includes a region transformation module inserted between the top-down region extractor module and the region encapsulation module, the region transformation module receiving the coordinates of the modified region of interest, and applying at least one transformation from the group including scaling, rotation and translation to at least one modified region of interest, and providing the transformed coordinates of the modified region of interest.

7. The encoder according to claim 6, wherein The region transformation module receives adaptive transformation parameters associated with a machine process and applies a selected transformation based in part on the parameters.

8. The encoder according to claim 6, wherein: The region transform module applies a transform based on at least one characteristic of the region of interest, the at least one characteristic of the region of interest comprising an object confidence, an object category, a region area, or a coding unit parameter.

9. An encoder for video for machine consumption, comprising: a region detector module that receives a source video and identifies a region of interest therein, the region of interest being defined in part by coordinates of a bounding box in a frame of the source video; a region transformation module receiving the coordinates of the region of interest and applying at least one transformation from the group consisting of scaling, rotation and translation to at least one modified region of interest and providing the transformed coordinates of the modified region of interest; a region packing module, the region packing module receiving the set of regions of interest and the transformed regions of interest, and arranging the regions of interest into a packing frame in which pixels outside the regions of interest are substantially excluded; as well as A video encoder receives the packed frame and the coordinates of the region of interest and encodes the packed frame and region parameters into a coded bit stream.

10. The encoder according to claim 9, wherein: The region transformation module receives adaptive transformation parameters associated with a machine process and applies a selected transformation based in part on the parameters.

11. The encoder according to claim 9, wherein: The region transformation module selectively applies a transformation to the region of interest based on at least one characteristic of the region of interest, the at least one characteristic of the region of interest comprising at least one of an object confidence, an object category, a region area, and a coding unit parameter.

12. A method for encoding source video for machine consumption, comprising: Receiving the source video and identifying a region of interest therein; processing the detected regions of interest to form a merge of regions of interest, wherein at least one of adjacent and overlapping regions of interest are combined; aligning the region of interest from the merging process to a predetermined grid; Slicing the aligned region of interest along the grid partitions; reattaching the slices to form a modified region of interest; providing coordinates of the modified region of interest; receiving the modified regions of interest and arranging the set of modified regions of interest into a packed frame in which pixels outside the modified regions of interest are substantially excluded; as well as The encapsulated frame and region parameters are encoded into a coded bitstream.

13. The method of claim 12, further comprising receiving coordinates of the modified regions of interest, and applying at least one transformation from the group consisting of scaling, rotation, and translation to at least one modified region of interest, and providing the transformed coordinates of the modified region of interest.

14. A decoder configured to receive an encoded bitstream encoded by an encoder according to any one of claims 1-13, and the decoder comprises a circuit configured to decode the bitstream and reconstruct a frame using a region of interest from a source video while excluding pixels outside the region of interest.

15. A decoder for decoding a coded bit stream of encapsulated frames having a region of interest, the decoder comprising: a video decoder receiving the encoded bitstream and decompressing the bitstream to identify a region of interest and region parameters therefrom; a region depackaging module, the region depackaging module receiving the decoded packed frame and region parameters, and arranging the decoded region of interest in a reconstructed frame having a size, position and orientation corresponding to the original frame, the pixels in the reconstructed frame outside the arranged region of interest being background pixels; A background processing module is provided to set parameters of the background pixels to optimize the performance of a machine task system receiving the reconstructed frame.

16. The decoder according to claim 15, wherein: The background processing module receives at least one adaptive fill parameter indicating a performance metric of the machine task system based on at least one parameter in the background pixels.

17. The decoder according to claim 16, wherein: The parameter of the background pixel is a fixed color.

18. The decoder according to claim 16, wherein: The parameter of the background pixel is the average color of the pixel in the region of interest.

19. The decoder according to claim 16, wherein: The adaptive fill parameters indicate parameters of the background pixels in which object detection by the robotic task system is optimized.

Citation Information

Cited By

  • Video frame streaming transmission methods in image system operation and maintenance

    CN122578855A