System and method for region detection and region encapsulation in video encoding and decoding of machine
By introducing area detection and packaging modules into the video encoder, identifying and extracting the regions of interest in the video for packaging, the problem of inefficient video encoding in the prior art is solved, and efficient video data encoding and machine task performance optimization is achieved.
Patent Information
- Application Number
- CN202380068866.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-26
- Filing Date
- 2023-09-26
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is difficult to efficiently process machine-specific video data in video encoding, resulting in inefficiency and waste of resources.
By introducing the area detector selection module, the area extractor module, the area packaging module and the area parameter module in the video encoder, the area of interest in the video is identified and extracted, and the area packaging and parameter provision are performed to generate the encoded bitstream.
It realizes efficient encoding of video data, reduces unnecessary pixel processing, improves encoding efficiency, and optimizes the performance of machine tasks.
Smart Images

Figure CN119968658A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Application Serial No. 63 / 409,843, filed on September 26, 2023, entitled “System and Method for Adaptive Region Detection and Region Packing,” and also claims priority to U.S. Provisional Application Serial No. 63 / 409,847, filed on September 26, 2023, entitled “System and Method for Extending Predicted Object Boundaries in a Video Packing System,” and also claims priority to U.S. Provisional Application Serial No. 63 / 409,851, filed on September 26, 2023, entitled “Systems and Methods for Merge and Split Region Extraction in Video Region Packing,” and the disclosures of each of the above applications are incorporated herein by reference in their entirety. Technical Field
[0003] The present disclosure relates generally to the field of video encoding and decoding, and in particular to encoding and decoding video and other data for machines. Background Art
[0004] Recent trends in robotics, surveillance, monitoring, IoT, etc. have introduced use cases where a large portion of all images and videos recorded in the field are consumed only by machines, without ever reaching the human eye. These machines process images and videos with the goal of completing specific tasks such as object detection, object tracking, segmentation, event detection, etc. Recognizing that this trend is pervasive and will only accelerate in the future, international standardization bodies have begun efforts to standardize image and video coding that is optimized primarily for machine consumption. For example, in addition to already established standards such as Compact Descriptors for Visual Search and Compact Descriptors for Video Analytics, standards like JPEG AI and video coding for machines are also ongoing efforts. Solutions that improve efficiency compared to classical image and video coding techniques are needed and are presented in this article. Summary of the invention
[0005] In one embodiment, a video encoder for encoding data for machine consumption is provided. The video encoder includes a region detector selection module that receives a source video and detector selection parameters and selects an object detector model. The region detection module applies the selected model to the source video to identify a region of interest in the source video. A region extractor module extracts pixels of the identified region from the source video. A region encapsulation module receives regions extracted from the source video and encapsulates the regions into encapsulated frames in which pixels outside the region of interest are omitted. A region parameter module receives the identified region from the region extractor and provides parameters for placing the region of interest in a reconstructed video frame. The video encoder receives the encapsulated frames from the region encapsulation module and receives the region parameters from the region parameter module and generates an encoded bitstream.
[0006] In some embodiments, the region detector selection module selects one of the plurality of models based on detector selection parameters from the machine task system.The detection selection parameters from the machine task system may be updated based on the performance of the machine task system for the encoded bitstream.
[0007] In some embodiments, the detector module may include at least one of a RetinaNet model and a Yolov7 model.
[0008] The region detection module may define each detected region at least in part by a rectangular bounding box. In some embodiments, the encoder may include a region filling module that adds filling parameters to one or more dimensions of the bounding box of the detected region. Each detected region may have an associated region type, and the filling parameters may be determined at least in part based on the object type. Alternatively or additionally, the filling parameters may be determined at least in part based on the region size and / or the bounding box size.
[0009] In another embodiment, the encoder may include a merged segmented region extractor module that further processes the detected regions and performs at least one of selectively merging regions having substantial overlap and selectively segmenting the regions to optimize packaging performance. The merged segmented region extractor module may receive adaptive extraction parameters from the machine task system and dynamically adjust the merging and segmenting parameters based on the parameters.
[0010] In some embodiments, an encoder may include both a region filling module and a merged segmented region extractor module.
[0011] A method of encoding video data for machine processing consumption is provided, the method comprising the steps of: receiving a source video; identifying at least one region of interest in the source video, each region of interest being defined by an associated bounding box; extracting the identified content of the region of interest within the associated bounding box from the source video; packing the extracted region into a packed video frame in which pixels outside the region of interest are omitted; providing region parameters for the bounding box sufficient to reconstruct the region of interest in a reconstructed video frame; and generating an encoded bitstream including the packed frames and the associated region parameters.
[0012] In some cases, the method may further include, for at least one region of interest, applying region padding to at least one dimension of an associated bounding box. The method may also include a merge segmentation process including at least one of selectively merging regions of interest having substantial overlap and selectively segmenting regions to optimize packing performance. The region of interest may have an associated object type, and the region padding may be determined based at least in part on the object type. In some embodiments, the region of interest has an associated bounding box size, and the region padding is determined based at least on the bounding box size.
[0013] In some embodiments, the method may include receiving performance data from a machine system located at a decoder site that receives the encoded bitstream, and the region filling is determined based at least in part on the received performance data.
[0014] The present disclosure also includes a video decoder including a circuit configured to receive and decode a coded bitstream generated by the above-mentioned encoder and coding method. The present disclosure also discloses an embodiment of a computer-readable medium storing a coded bitstream thereon, the coded bitstream being generated by any encoder and coding method described herein.
[0015] These and other aspects and features of non-limiting embodiments of the present invention will become apparent to those skilled in the art upon reading the following description of specific non-limiting embodiments in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] For the purpose of illustrating the present invention, the accompanying drawings show aspects of one or more embodiments of the present invention. However, it should be understood that the present invention is not limited to the precise arrangements and instrumentalities shown in the accompanying drawings, in which:
[0017] Figure 1 is a simplified block diagram of a system for encoding and decoding video for a machine, such as in a system for video coding for a machine (VCM).
[0018] Figure 2is a simplified block diagram of a system for encoding and decoding video for a machine, such as a video coding for a machine ("VCM"), using region packing in accordance with the present method.
[0019] Figure 3 is a schematic diagram comparing inference predictions using two different object detection networks;
[0020] Figure 4 is a simplified block diagram of an alternative embodiment of an encoder for video of a machine utilizing both region packing and region padding according to the present method;
[0021] Figure 5 is a schematic diagram showing an example of area filling according to the present disclosure;
[0022] Figure 6 is a simplified block diagram of an alternative embodiment of an encoder for encoding video content for a machine, such as a video coding for a machine ("VCM"), utilizing region packing and merging and segmentation extraction processing in accordance with the present method;
[0023] Figure 7 It further shows Figure 6 A simplified block diagram of a merge / split region extractor module of an encoder of ; and
[0024] Figure 8 is a schematic diagram illustrating extraction of identified objects using merged / segmented regions according to the disclosed systems and methods.
[0025] The drawings are not necessarily drawn to scale and may be illustrated by phantom lines, diagrammatic representations, and partial views. In certain instances, details that are not necessary for understanding the embodiments or that render other details difficult to perceive may have been omitted. DETAILED DESCRIPTION
[0026] Figure 1 is a block diagram illustrating an exemplary embodiment of a system well suited for machine-based consumption of video and related data, including encoders, decoders, and bitstreams, such as contemplated in the application of video coding for machines ("VCM"). Figure 1 It has been simplified to depict the components used in encoding for machine consumption, but it should be understood that the present systems and methods are also applicable to hybrid systems that also encode, transmit and decode video for human consumption. Such systems for encoding / decoding video of various protocols such as HEVC, VVC, AV1, etc. are well known in the art.
[0027] Reference now Figure 1 , an exemplary embodiment of an encoding system 100 including an encoder 105, a bitstream 155, and a decoder 130 is shown.
[0028] Further references Figure 1 , an exemplary embodiment of an encoder for encoding video for a machine is shown. The encoder 105 can be implemented using any circuit, including but not limited to digital and / or analog circuits; the encoder 105 can be configured using hardware configuration, software configuration, firmware configuration and / or any combination thereof. The encoder 105 can be implemented as a computing device and / or a component of a computing device, which can include but is not limited to any computing device described below. In an embodiment, the encoder 105 can be configured to receive input image or video data 102 and generate an output bitstream 155. The reception of the input video 102 can be accomplished in any manner described below or known in the art. The bitstream can include but is not limited to header content, payload content, supplemental signaling information, etc., and can include any bitstream described below.
[0029] The encoder 105 may include, but is not limited to, an inference module having a region extractor 110 , a region converter and packer 115 , a packing picture converter and shifter 120 , and / or an adaptive video encoder 125 .
[0030] The packed picture converter and shifter 120 processes the packed image so that additional redundant information can be removed before encoding. Examples of conversions are conversion of color space (e.g., conversion from RGB to grayscale), quantization of pixel values (e.g., reducing the range of pixel values represented and thus reducing contrast), and other conversions that remove redundancy in the sense of the machine model. Shifting requires reducing the range of pixel values represented by a direct right shift operation (e.g., shifting the pixel values right by 1 is equivalent to dividing all values by 2). Using the inverse mathematical operation used in 120, the conversion and shifting process is reversed on the decoder side by box 140.
[0031] Further references Figure 1 , the adaptive video encoder 125 may include, but is not limited to, any video encoder for encoding video to an advanced CODEC standard (such as HEVC, AV1, VVC, etc.) as described in further detail below or known in the art.
[0032] Still reference Figure 1, an exemplary embodiment of a decoder 130 is shown. The decoder 130 can be implemented using any circuit, including but not limited to digital and / or analog circuits; the decoder 130 can be configured using hardware configuration, software configuration, firmware configuration and / or any combination thereof. The decoder 130 can be implemented as a computing device and / or a component of a computing device, which can include but is not limited to any computing device described below. In an embodiment, the decoder 130 can be configured to receive an input bitstream 155 and generate an output video 147 suitable for machine consumption. The reception of the bitstream 155 can be accomplished in any manner described below or known in the art. The bitstream can include but is not limited to a compatible bitstream provided by the encoder 105 with, for example, header information, payload data, enhanced or supplemental signaling, etc., and / or can include any bitstream described below.
[0033] Continue to refer Figure 1 , the machine model 160 may be present in the encoder 105 or otherwise provided to the encoder 105 in an online or offline mode using an available communication channel. The machine model 160 is application / task specific and generally contains information sufficient to describe the requirements of the machine 150 to complete the task at the decoder site. This information may be used by the encoder 105, and in some embodiments specifically by the region converter and encapsulator 115.
[0034] Given a frame of video or image, efficient compression of such media can be achieved by detecting and extracting its important regions and packing them into a single frame. At the same time, the system discards any detected regions of no interest. These packed frames are used as input to an encoder to generate a compressed bitstream. The generated bitstream 155 contains the encoded packed regions and the parameters required to reconstruct and reposition each region in the decoded frame. The machine task system 150 can perform tasks such as specified computer vision related functions on the reconstructed video frames.
[0035] Such video compression systems can be improved by implementing adaptive selection of region detection methods. Adaptively selecting which encoder-side region detection system to use is advantageous in supporting endpoint target machine tasks.
[0036] Figure 2 is a simplified block diagram of a system for machine video encoding and decoding using region packing according to the present disclosure. Figure 2, the system 200 includes an encoder 208 and a decoder 236. The encoder includes a detector selection block 260 that receives a source video 204 and an adaptive selection parameter 264 for a machine-task system. The encoder also includes a region detection module 212, a region extractor module 216, a region encapsulation module 220, region parameters 224, and a video encoder 228, which cooperate to generate a compressed bitstream 232, which is delivered to a decoder site via a communication channel, and the decoder site decodes the bitstream for processing by the machine task system 256.
[0037] The decoder 236 includes a video decoder 240 that receives the compressed bitstream 232 , a region decapsulation module 244 , and region parameters 248 that generate a decapsulated reconstructed video frame 252 for a machine task system 256 .
[0038] Encoder site area detection
[0039] Significant frame or image regions are identified using the detection system 212 that generates the coordinates of the discovered objects. The resulting coordinates are used to determine regions for encapsulation and enable identification of pixels that are considered unimportant to the detection module. Such unimportant regions can be discarded and are not required for encapsulation.
[0040] Improvements to the region detection module 212 within the compression pipeline are intended to better support target machine task performance. The adaptive selection method described herein provides that an encoder-side detection algorithm can be selected based on specified characteristics of the endpoint evaluation network. For example, a neural network can be selected that matches a neural network with similar characteristics used by the machine, such as a convolutional neural network with a similar number of layers and input and output dimensions. In some examples, the same algorithm can be used. If information about the detection algorithm used by the machine 150 is not available, or does not include a detailed description, or the algorithm itself is not available for implementation on the encoder side, a similar algorithm can be used. In some cases, a similar but more recent algorithm can be used as an encoder-side alternative to allow faster operation.
[0041] Figure 3 A and Figure 3 B shows the same example image with inferred predictions created by two different object detection networks. Figure 3 In A, the inferred prediction 304 is derived from the RetinaNet model of Detectron2, while Figure 3B shows the 308 inferred predicted coordinates output from the Yolov7 network. The RetinaNet model is described in Lin, TY, Goyal, P., Girshick, R., He, K., & Dollár, P. (2017), Focal Loss for Dense Object Detection, published in Proceedings of the IEEE International Conference on Computer Vision (pp. 2980-2988). The Yolov7 network is described in Wang, CY, Bochkovskiy, A., & Liao, HYM (2022), YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-art for Real-time Object Detectors, arXiv preprint arXiv:2207.02696. It should be understood that these are merely examples of suitable detection networks, and the present systems and methods are not limited to these particular detection networks.
[0042] Those skilled in the art will appreciate that, when applying the proposed method, one particular detection network may outperform other detection networks based on the specific target machine task. Figure 2 , the currently disclosed CODEC pipeline preferably sends the adaptive selection parameters 264 of the target machine task to the detector selection module 260. The detector selection module 260 interprets the received information and uses it to select the best available detection network. The preferred detection network is the detection network that provides the best region detection to support the endpoint computer vision related tasks in 256. For example, for an endpoint neural network task that is highly sensitive to object context, the system may generate a selection of the RetinaNet detection method because more background pixels are included in the image. As shown in FIG. Figure 3 A and Figure 3 As shown in B, in contrast, a system aiming to reduce the size of the transmitted bitstream 232 may select the Yolov7 model because it contains fewer predictions.
[0043] The selected detection method can be included in the output bitstream 232 along with information characterizing the detection method. Such information can be signaled in the bitstream header or provided in the bitstream as supplemental information. In some embodiments, this can be included in sequence parameter set ("SPS") data that typically remains unchanged for a sequence of frames, or in picture parameter set ("PPS") data that can change from frame to frame. The detection method information includes the purpose of the detection method, the version number of the detection method, the training data used in the detection method, performance parameters such as the minimum and maximum detection confidence of the detection in the frame, and the object category detected in the frame. The detection confidence for each object category can also be included in the bitstream. Other parameters characterizing the detection performance can be determined and included. At the decoder, the model parameters extracted from the bitstream can be used to select or adapt the machine / algorithm for the machine task.
[0044] Object detector information
[0045] The following is a description of exemplary object detection semantics that may be encoded in the bitstream:
[0046]
[0047]
[0048] Object detector information semantics
[0049] object_detector_ID - The ID of the object detector. This ID can come from a known detector registry, or can be configured and agreed upon between the encoding and decoding systems.
[0050] object_detector_version - The version of the object detector. The object detector version can be used to identify how a particular detector was trained. Additional information, such as the number of classes the detector can handle, can be obtained based on the version number. This bitstream field can be extended to include a list of classes that the detector can detect.
[0051] object_detector_name_length - Number of bytes to use for the object detector name
[0052] object_detector_name[object_detector_name_length] - The name of the object detector. This is typically a displayable string.
[0053] object_classes_detected – The number of object classes detected in this frame.
[0054] min_detection_condfience - Confidence is described as a number between 0 and 100. 100 is 100% confidence and 0 is 0% confidence.
[0055] max_detection_condfience - Confidence is described as a number between 0 and 100. 100 is 100% confidence and 0 is 0% confidence.
[0056] object_class_name_length - the number of bytes used for the object class name
[0057] object_class_name [object_class_name_length] – name of the object class
[0058]
[0059] object_detector_information_present - A one-bit field, when set to 1, signals the presence of object_detector_information in a sequence parameter set ("SPS").
[0060]
[0061] object_detector_information_present - A one-bit field, when set to 1, signals the presence of object_detector_information in the picture parameter set ("PPS").
[0062] Similarly, this object detector information may be expanded or reduced and signaled elsewhere in the video bitstream, such as in a slice header. In some cases, such object detection information can be signaled as supplemental enhancement information (SEI) data that may be associated with a frame and signaled in a separate information packet that is not included in the video bitstream.
[0063] The selected method for object detection is used as input to the region detector module 212 along with any specification information. The signaled method from the machine system 256 and any selection parameters 264 are used to perform inference and identify regions in 212. Additional region extraction methods can be paired with the proposed adaptive system. These steps can be used to isolate the best possible combination of region detection predictions derived from specified characteristics. That is, based on the selected detection network, different thresholds for box selection can be applied based on the characteristics of the network. This includes thresholding the box confidences to keep high / low / all confidence predictions, or otherwise using a class-based approach to select regions based on their class inferred from the network.
[0064] Table 1 shows the performance difference between the aforementioned RetinaNet detection network and the Yolov7 detection network. From the table, it can be concluded that such a region encapsulation system may be affected by the inferred predictions derived from 212. Therefore, including the selection module 260 is beneficial to the system because it allows for improved performance through adaptive selection of detection methods.
[0065]
[0066] Table 1 Comparison of detection methods in regional packaging systems
[0067] Region extraction and encapsulation
[0068] The resulting coordinates from the region detection system 212 are used as input to the region extraction system 216. The extraction module 216 extracts the important image regions and prepares the coordinates for the packing module 220. The packing module 220 receives the extracted regions and packs them tightly into a single frame. This module additionally outputs packing parameters that will be signaled later in the bitstream 232.
[0069] Video Encoding
[0070] The encapsulated object frames are processed by a video encoder 228 to generate a compressed bitstream 232. The compressed bitstream contains the encoded encapsulated regions and the parameters 224 required to reconstruct and reposition each region in the decoded frame. The video encoder 228 can take the form of any known advanced video encoder known for coding standards such as HEVC, AV1, and VVC or variants thereof on such known encoders. Optionally, any detection thresholds applied with the network selected by the detection module can be signaled in the bitstream 232 for reconstruction on the decoder 236 side.
[0071] Video Decoding
[0072] The compressed bitstream 232 is decoded using a video decoder 240 to generate a packed region frame and its signaled region information. The video decoder 240 will typically take a form complementary to the selected video encoder 228 and may take the form of any known advanced video decoder known for conventional codec standards such as HEVC, AV1, and VVC or variants of such known standards. The signaled region information includes the parameters required for reconstruction of the frame and may be combined with the signaled detection threshold and method for each region applied in the encoder 208.
[0073] Regional decapsulation
[0074] The region parameter module 248 provides the decoded region parameters via the region decapsulation module 244 to decapsulate the region frames. During region decapsulation, each box is returned to its position within the context of the original video frame. The resulting decapsulated frames include only the salient regions determined by the region detection system 212 and do not include discarded pixels. These decapsulated regions include predictions made by the adaptively selected detection network in the encoder 208.
[0075] Machine Tasks
[0076] The depacketized and reconstructed video frames 252 are used as input to a machine task system 256, which performs a specified machine task, such as a computer vision related function. The machine task performance on the regions selected by the selection module 260 can be analyzed to determine the best box selection method and inference threshold on a scene-by-scene basis. The optimized region detection parameters 264 can be changed and signaled to the encoder-side pipeline to more effectively select the region detection method.
[0077] Figure 4 is a simplified block diagram of an encoder for region packing with region filling in a system for video encoding of a machine according to the present disclosure. Region filling is the process of extending the boundaries of detected regions to improve system performance. Encoder 408 includes region detector 412 and region filling module 460, which receives source video 404 and adaptive filling parameters 464 for the machine task system from region detector 412. Encoder 408 also includes region extractor module 416, region packing module 420, region parameters 424 and video encoder 428, which cooperate to generate a compressed bitstream 432, each of which is substantially similar in structure and operation to the above-mentioned combined system except as noted below. Figure 2 Those corresponding components described.
[0078] Area Detection
[0079] Significant frames or image regions are identified using an encoder side region detector module 412, which typically generates the coordinates of the found objects in the form of a rectangular bounding box around the detected object. A saliency-based detection method using video motion can also be used to identify significant regions. For example, uniform motion detected across consecutive frames can be designated as significant. In another example, any motion that lasts for a long period of time (e.g., 100 frames) in a continuous trajectory can be designated as significant. In another example, motion detected at the same coordinates where an object is detected can be designated as significant. The spatial coordinates of the significant region can be used to determine the region for encapsulation, and enable identification of pixels that are considered unimportant to the detection module. Such unimportant regions can be discarded and not used for encapsulation.
[0080] Area Fill
[0081] The discovered objects and regions from the region detector 412 may be additionally processed prior to the extraction and packing stages to achieve more efficient compression and / or endpoint machine task performance. A region padding module 460 may be used to extend the detected object boundaries. Padding is the extension of a bounding box beyond the minimum area detected by one or more pixels in one or more directions or dimensions. Padding may be applied uniformly around the bounding box, e.g., the same number of pixels in each dimension, or dynamically where padding varies in different dimensions.
[0082] The padding size can be determined based on an internal decision made by the module with or without receiving adaptive padding parameters 464. This can include applying padding based on object category, object size, and / or object confidence. For example, the decision to enable padding and the amount of padding can be calculated using an optimization search in the inference space that can compare the detection accuracy of boxes with or without padding, and various amounts of padding are applied. It should be understood that not all object classes and instances need to be evaluated. A representative sample of categories with similar characteristics (such as size, orientation, and color) can be used to assign padding decisions to all objects represented by the exemplary object.
[0083] Application of the filling module 460 can provide better context for the post-compression machine task evaluation of the machine system at the decoder site. Each endpoint machine system may have a different sensitivity to background pixel information; therefore, expansion of the initial prediction region can help improve evaluation accuracy. The object filling described herein is preferably performed by expanding each dimension relative to the overall image boundaries to include additional pixels for context. Such additional pixels are pixels that reside outside of the original coordinates output by the region detection module.
[0084] Prediction box expansion can be performed using a fixed amount of padding, or can be adaptively determined on a box-by-box basis. Adaptive padding can be performed using characteristics of the detected object including object category, inference confidence / score, and / or object size. Additionally, padding can be skipped based on the type of region detection method selected or using a similar box-by-box. The padding size can be determined using supplementary input adaptive padding parameters 464 based on machine task feedback.
[0085] Region extraction and encapsulation
[0086] The extended / filled coordinates obtained from the padding module 460 are used as input to the region extraction module 416. The region extraction module 416 extracts the image regions and prepares the coordinates for the packing module 420. The region packing module 420 takes the extracted regions and packs them tightly into a single frame. The packing module 420 additionally outputs packing parameters that will be signaled later in the bitstream 432.
[0087] Video Encoding
[0088] The encapsulated object frames are processed by a video encoder 428 to generate a compressed bitstream. The video encoder 428 can take the form of any known advanced video encoder, such as for encoding standards such as HEVC, AV1, and VVC, or a variant of such a known standard for machine use. The compressed bitstream contains the encoded encapsulated regions and the parameters 424 required to reconstruct and reposition each region in the decoded frame. Optionally, the padding size for each in the frame can be signaled in the bitstream for decoder-side reconstruction and for data collection. Such signaling may include signaling within the header, SPS, PPS, or auxiliary signaling such as supplemental enhancement information (SEI).
[0089] Although various functional modules in the encoder 408 have been described as different functional modules, it should be understood that these functional modules can be further divided into sub-modules or functional combinations without departing from the intent of the embodiments described herein.
[0090] Figure 5 is a graphical representation of the current region fill, where a source frame 504 is processed to provide four inferred predictions or objects in 506. A 15 pixel padding parameter is applied 508 to the four objects to provide the filled inferred predictions shown in 510. While in this example, a fixed padding is applied to each detected object, it should be appreciated that in different examples, adaptive padding may be applied, where different padding parameters are applied to different detected objects.
[0091] Video Decoding
[0092] Except as described below, the structure and operation of the decoder for a bitstream with region filling is similar to that of Figure 2 When region padding is used during encoding, the compressed bitstream 432 is decoded by the video decoder 240 to generate a packed region frame and the region information it signals. Such region information preferably includes region parameters 248 useful for reconstruction of the frame, and may be combined with the padding size applied to each region determined by the encoder-side process.
[0093] Regional decapsulation
[0094] refer to Figure 2 , the decoded parameters are used to decapsulate the encapsulated region frame via the region decapsulation module 244. Each frame is returned to its position within the context of the original video frame. The resulting decapsulated frame 252 includes only the significant areas determined by the region detection system 412 - post-filling - and preferably does not include discarded pixels outside the padding object boundaries. The padding size sent with the signal from the region parameters 248 can be used in the decapsulation stage to determine background pixels. These background pixels can be used to improve the object context and can help fill in the empty spaces within the decapsulated frame 252.
[0095] Machine Tasks
[0096] The depacketized and reconstructed video frames 252 are used as input to a machine task system 256, which can perform machine tasks, such as computer vision related functions. The machine task performance on the padding area can be analyzed and used to determine the optimal amount of padding on a box-by-box or object-by-object basis. The optimized padding parameters 464 can be updated and signaled to the encoder-side pipeline to effectively extend the object boundaries.
[0097] Figure 6 6 is a simplified block diagram of an alternative embodiment of an encoder with region packing with merge / split region extraction for video encoding of a machine according to the present disclosure. The encoder 600 processes a received source video 604 and generates a compressed bitstream 628. The encoder 600 includes a region detection block 612, a merge / split region extractor module 656, a region packing module 616, a region parameter module 620, and a video encoder 624, which cooperate to generate the compressed bitstream 628. The components of the encoder 600 are structurally and operationally combined with the above Figure 2 The description is similar except as follows.
[0098] Merge segmentation region extraction
[0099] The merged segmented region extraction module 656 receives object coordinates from the region detection system 612. These coordinates may contain multiple predictions within the same region and / or may consist of overlapping regions with redundant pixels. Figure 7 Further shown in FIG. 6 is a merge / segment extraction module 656 that creates new region boxes based on given predictions.
[0100] refer to Figure 7 , a merged segmented region extractor 756 receives region detection coordinates 704 from the region detection system 612. The merged segmented region extractor 756 identifies overlapping regions 708 in the detected regions and may then apply merge actions 712 and split actions 716 to merge multiple region predictions into a single box and / or split the regions into new smaller boxes defined by new region coordinates 720. These actions are performed on a per-object (or per-prediction) basis.
[0101] The merged segmented region extractor module 756 checks which regions are close and identifies them as candidates for further processing. The decision to merge and the decision to split are primarily based on rate savings considerations, and secondarily on considerations of expected detection accuracy on the machine. By merging region predictions, a more compact and continuous spatial structure can be obtained, which may be more suitable for predicting hybrid video and image coding. By splitting region predictions, smaller geometric structures are obtained, such as smaller rectangles, which can potentially be packed in a more spatially optimal manner. For example, in some cases, machine detection performance can be improved if the bits saved by improved object packing are spent on more accurate texture representation (e.g., retaining more high-frequency components).
[0102] Different criteria can be used to determine appropriate scenarios for merging and splitting actions. For example, inferred prediction boxes that overlap more than a determined threshold can be merged to form a single new region box. Here, the previous inferred box can be discarded and replaced with a unified new box. Segmentation can be performed on boxes that overlap with a smaller threshold. In this case, the segmented box is retained, while the original inferred box can be discarded.
[0103] One or more region detections may be grouped into local clusters. Machine learning methods may also be used to determine the optimal number of clusters based on given inference parameters and image characteristics. For example, each region detection may be assigned as a single instance in a k-means clustering algorithm - a cost function that minimizes the bit budget for encoding the frame and / or a cost function that maximizes the detection accuracy in the inference model may be used as the objective function for the k-means algorithm.
[0104] Clusters of boxes can be merged to form regions containing multiple inferred objects. In addition, various segmentation methods can be used on a case-by-case basis. This includes horizontal segmentation, vertical segmentation, and / or some combination of the two. For example, overlapping inferred boxes containing vertically oriented objects can be segmented vertically, while horizontal objects can be segmented horizontally.
[0105] Figure 8 An example of the merging and splitting actions applied is shown. The predicted objects in the original frame 804 are used as the basis for determining the new region box. The predicted regions 808, 812, 816, and 820 overlap each other by a significant amount and share many of the same image pixels. These objects are merged to form a new region box. The resulting merged coordinates still overlap with the existing inferred prediction box 824. Since the overlapping area is not as large as the previously described overlap, the two regions are vertically split to create a new smaller region box. Prediction 828 remains intact because it does not overlap with any other region. The final box can be tightly packed as shown in 804.
[0106] Merging inferred boxes can benefit the encoding and machine task process. The merged boxes can help reduce the number of individual boxes sent to the packaging system and can therefore improve the spatial protection of image objects. Alternatively, segmenting overlapping boxes helps reduce the appearance of repeated pixels propagated throughout the pipeline. The region extraction step performed by the merged segmented region extractor 656 can include additional pixels beyond those determined to be important by the region detection module 612. It can also change which pixels are to be discarded. The system outputs new region coordinates 720.
[0107] Regional Encapsulation
[0108] The newly identified region box coordinates 720 returned from the merged segmentation extraction module 656 are used as input to the region packing system 616. The region packing module 616 extracts the salient image regions and packs them tightly into a single frame. The region packing module generates packing parameters that will be signaled later in the encoded bitstream 628.
[0109] Video Encoding
[0110] The packed object frame containing the processed regions from 656 is input to a video encoder 624 that generates a compressed bitstream 628. The compressed bitstream contains the encoded packed regions and the parameters needed to reconstruct and reposition each region in the decoded frame 620. Optionally, the original region coordinates (i.e., those derived from the extrapolation from the region detector 612 before merging the segmented region extraction 656) can be signaled in the bitstream for use at the decoder side.
[0111] Video Decoding
[0112] Use basically Figure 2 The video decoder 236 described in the above decodes the compressed bitstream 628 to generate a packed region frame and the region information it signals. This region information includes the parameters required to reconstruct the frame and can be combined with the original object coordinate information from 612 and the new coordinates and region information 720 provided by the merge and segment region extractor 656.
[0113] Regional decapsulation
[0114] The decoded region parameters 248 are used by the region depacketization module 244 to depacketize the packed region frame. Each box is returned to its position within the context of the original video frame. The resulting depacketized frame includes only the salient regions determined by the region detection system 112 after the merge segmentation extractor module 656, and preferably does not include discarded pixels.
[0115] Machine Tasks
[0116] The depacketized and reconstructed video frames 252 are used as input to a machine task system 256, which can perform machine tasks, such as computer vision related functions. The machine task performance on the regions determined by the merge-segment extraction module 656 can be analyzed to determine the best extraction actions. The optimized region extraction parameters 660 can be updated and signaled to the encoder-side pipeline to effectively merge and segment the inferred boxes to determine the regions.
[0117] Some embodiments may include a non-transitory computer program product (ie, a physically embodied computer program product) storing instructions that, when executed by one or more data processors of one or more computing systems, cause at least one data processor to perform the operations herein.
[0118] Embodiments may include circuits configured to implement any operation as described in any embodiment above in any order and with any degree of repetition. For example, a module such as an encoder or decoder may be configured to repeatedly perform a single step or sequence until the desired or commanded result is achieved; the output of a previous repetition may be used as the input of a subsequent repetition to iteratively and / or recursively perform the repetition of a step or sequence of steps, aggregate the input and / or output of the repetition to generate an aggregated result, reduce or decrement one or more variables such as global variables, and / or divide a larger processing task into a set of iteratively addressed smaller processing tasks. The encoders and decoders described herein may perform any step or sequence of steps as described in the present disclosure in parallel, such as using two or more parallel threads, processor cores, etc. to perform steps twice or more simultaneously and / or substantially simultaneously; the task division between parallel threads and / or processes may be performed according to any protocol suitable for dividing tasks between iterations. Those skilled in the art will know various ways in which steps, sequence of steps, processing tasks and / or data may be subdivided, shared or otherwise processed using iteration, recursion and / or parallel processing when reading the entire contents of the present disclosure.
[0119] A non-transient computer program product (i.e., a physically embodied computer program product) may store instructions that, when executed by one or more data processors of one or more computing systems, cause at least one data processor to perform the operations and / or steps thereof described in the present disclosure, including but not limited to any of the operations described above and / or any operations that a decoder and / or encoder may be configured to perform. Similarly, a computer system that may include one or more data processors and a memory coupled to the one or more data processors is also described. The memory may temporarily or permanently store instructions that cause at least one processor to perform one or more operations described herein. In addition, the method may be implemented by one or more data processors within a single computing system or distributed between two or more computing systems. Such computing systems may be connected via one or more connections, via direct connections between one or more of a plurality of computing systems, etc., and may exchange data and / or commands or other instructions, etc., wherein the connection includes a connection through a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.).
[0120] It should be noted that any one or more aspects and embodiments described herein can be conveniently implemented using one or more machines (e.g., one or more computing devices used as user computing devices for electronic documents, one or more server devices, such as document servers, etc.) programmed according to the teachings of this specification, which is obvious to those of ordinary skill in the computer arts. Based on the teachings of this disclosure, a skilled programmer can easily prepare appropriate software coding, which is obvious to those of ordinary skill in the software arts. The aspects and implementations using software and / or software modules discussed above may also include appropriate hardware for assisting in implementing the machine executable instructions of the software and / or software modules.
[0121] Such software can be a computer program product using a machine-readable storage medium. A machine-readable storage medium can be any medium capable of storing and / or encoding a sequence of instructions executed by a machine (e.g., a computing device) and causing the machine to perform any of the methods and / or embodiments described herein. Examples of machine-readable storage media include, but are not limited to, disks, optical disks (e.g., CDs, CD-Rs, DVDs, DVD-Rs, etc.), magneto-optical disks, read-only memory "ROM" devices, random access memory "RAM" devices, magnetic cards, optical cards, solid-state memory devices, EPROMs, EEPROMs, and any combination thereof. As used herein, machine-readable media is intended to include a single medium and a collection of physically separated media, such as, for example, a collection of optical disks or one or more hard disk drives combined with a computer memory. As used herein, a machine-readable storage medium does not include signal transmission in transient form.
[0122] Such software may also include information (e.g., data) carried as a data signal on a data carrier (such as a carrier wave). For example, machine executable information may be included as a data-bearing signal embodied in a data carrier, wherein the signal encodes a sequence of instructions or a portion thereof executed by a machine (e.g., a computing device) and any related information (e.g., data structures and data) that causes the machine to perform any of the methods and / or embodiments described herein.
[0123] Examples of computing devices include, but are not limited to, electronic book reading devices, computer workstations, terminal computers, server computers, handheld devices (e.g., tablet computers, smart phones, etc.), network devices, network routers, network switches, bridges, any machine capable of executing a sequence of instructions specifying actions to be taken by the machine, and any combination thereof. In one example, the computing device may include and / or be included in an all-in-one machine.
[0124] The foregoing is a detailed description of an illustrative embodiment of the present invention. Various modifications and additions may be made without departing from the spirit and scope of the present invention. The features of each of the various embodiments described above may be appropriately combined with the features of the other described embodiments so as to provide a variety of feature combinations in associated new embodiments. In addition, although a plurality of separate embodiments have been described above, what has been described herein is merely an illustration of the application of the principles of the present invention. In addition, although the specific methods herein may be shown and / or described as being performed in a particular order, the order is highly variable within the ordinary technician to implement the methods, systems and software according to the present disclosure. Therefore, this description is intended to be merely an example, and not to otherwise limit the scope of the present invention.
[0125] Exemplary embodiments have been disclosed above and shown in the accompanying drawings. It will be appreciated by those skilled in the art that various changes, omissions and additions may be made to the specific disclosure herein without departing from the spirit and scope of the present invention.
Claims
1. A video encoder for encoding data for consumption by a machine, comprising: a region detector selection module that receives a source video and a detector selection parameter and selects an object detector model; a region detection module that receives the selected object detector model and applies the model to the source video to identify regions of interest in the source video; a region extractor module that extracts the identified region from the source video; A region encapsulation module, the region encapsulation module receiving the region extracted from the source video and encapsulating the region into an encapsulation frame; a region parameter module that receives the identified region from the region extractor and provides parameters for placing the region in a reconstructed video frame; A video encoder receives the encapsulated frames from the region encapsulation module and the region parameters from the region parameter module and generates an encoded bitstream.
2. The encoder according to claim 1, wherein The regional detector selection module selects one of a plurality of models based on a detector selection parameter from a machine task system.
3. The encoder according to claim 2, wherein: Detection selection parameters from the machine task system are updated based on the performance of the machine task system on the encoded bitstream.
4. The encoder according to claim 2, wherein: The multiple models include at least one of a RetinaNet model and a Yolov7 model.
5. The encoder according to claim 1, wherein The region detection module defines each detected region at least in part by a rectangular bounding box, and the encoder further includes a region padding module that adds padding parameters to one or more dimensions of the bounding box of the detected region.
6. The encoder according to claim 5, wherein: Each detected region has an associated region type, and the fill parameter is determined based at least in part on the object type.
7. The encoder according to claim 5, wherein: The fill parameter is determined based at least in part on the region size.
8. The encoder of claim 1, further comprising a merged segmented region extractor module that processes the detected regions for further processing and performs at least one of selectively merging regions having substantial overlap and selectively segmenting the regions to optimize packing performance.
9. The encoder according to claim 8, wherein: The merge-segment region extractor module receives adaptive extraction parameters from a machine task system and dynamically adjusts merge and segment parameters based on the parameters.
10. The encoder according to claim 1, wherein: Each detection region is defined by a rectangular bounding box, and the encoder also includes: a region filling module that adds padding parameters to one or more dimensions of a bounding box of a detected region; and A merged segmented region extractor module processes the detected regions for further processing and performs at least one of selectively merging regions having substantial overlap and selectively segmenting the regions to optimize packaging performance.
11. A method of encoding video data for consumption by machine processing, the method comprising: Receive source video; identifying at least one region of interest in the source video, each region of interest being defined by an associated bounding box; extracting, from the source video, identification content of the region of interest within the associated bounding box; Packing the extracted region into a packed video frame, wherein pixels outside the region of interest are omitted in the packed video frame; providing region parameters for a bounding box sufficient to reconstruct the region of interest in a reconstructed video frame; and An encoded bitstream is generated including the encapsulated frames and associated region parameters.
12. The encoding method according to claim 11, further comprising: For at least one region of interest, applying region padding to at least one dimension of an associated bounding box; as well as A merge-segment process is applied to optimize packaging performance, the merge-segment process including at least one of selectively merging regions of interest having substantial overlap and selectively segmenting regions.
13. The encoding method according to claim 12, wherein: The region of interest has an associated object type, and the region fill is determined based at least in part on the object type.
14. The encoding method according to claim 12, wherein: The region of interest has an associated bounding box size, and the region fill is determined based at least on the bounding box size.
15. The encoding method according to claim 12, further comprising: Performance data is received from a machine system located at a decoder site that receives the encoded bitstream, and wherein the region filling is determined based at least in part on the received performance data.
16. A video decoder comprising circuitry configured to receive and decode an encoded bitstream generated by any one of claims 1-15.
17. A machine-readable medium having stored thereon a coded bit stream, the coded bit stream being generated by any one of claims 1-15.