Region-based picture enhancement method and device

By dividing the video picture into macroblocks and predicting and splicing importance, the problems of high computing costs and reduced accuracy in video analysis are solved, efficient regional enhancement is achieved, and the accuracy and throughput of video analysis are improved.

CN120263999APending Publication Date: 2025-07-04TSINGHUA UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510404555.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the calculation cost of full-frame enhancement methods in video analysis is high, and the selective enhancement methods lead to a decrease in accuracy, and the area of ​​beneficial analysis tasks cannot be accurately identified, resulting in wasted computing resources.

Method used

The video picture is divided into macroblocks, and importance prediction is made through the pre-trained segmentation model, the areas to be enhanced are identified, and spliced ​​and enhanced processing is performed to reduce the amount of calculation and improve resource utilization.

Benefits of technology

Fast and accurate area enhancement is achieved, improving the accuracy and throughput of video analysis and reducing computational costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263999A_ABST
    Figure CN120263999A_ABST
Patent Text Reader

Abstract

The invention provides a region-based picture enhancement method and device, and the method comprises the steps: dividing a to-be-enhanced video picture by taking macro blocks as basic units, and carrying out the importance prediction of each divided macro block; based on the importance prediction result of each macro block, obtaining a plurality of to-be-enhanced areas of the to-be-enhanced video picture; the plurality of areas to be enhanced are spliced; and inputting the splicing result into the enhancement model for enhancement processing. According to the method, important regions can be quickly and accurately identified, and sparsely distributed regions are spliced into dense tensors and are efficiently enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and edge computing, and in particular, to a region-based video enhancement method and apparatus. Background Art

[0002] Video analysis is widely used in society, such as traffic control, school security, criminal investigation, etc. With the development of deep neural networks (DNNs), AI-driven automatic video analysis has become possible. However, due to outdated camera hardware and bandwidth limitations, the video quality is restricted, resulting in a reduction in analysis accuracy. Content enhancement technologies enhance the information details of video frames before inputting them into the final analysis model through neural enhancement models (such as super-resolution, generative adversarial networks, etc.), thereby improving accuracy and saving bandwidth.

[0003] Traditional frame-level content enhancement methods enhance every frame in the video, regardless of whether the frame contains regions beneficial to the analysis task. Since enhancement models (such as super-resolution, generative adversarial networks, etc.) need to process a large amount of pixel data. For example, for a 1080p video frame, the enhancement model needs to process approximately 2 million pixels. Therefore, this full-frame enhancement method has extremely high computational costs, which will cause significant delays in real-time video analysis. Moreover, both the enhancement model and the final analysis model (such as object detection, semantic segmentation, etc.) require a large amount of computing resources. The full-frame enhancement method will occupy too much computing resources, leading to resource competition with the analysis model, thereby reducing the overall system throughput.

[0004] Existing selective enhancement methods (such as NeuroScaler and Nemo) improve throughput by only enhancing some frames (called anchor frames), but this method will lead to a decrease in accuracy. Because selective enhancement methods need to make a trade-off between enhanced frames and unenhanced frames, the cumulative error of unenhanced frames will result in inaccurate analysis results. Traditional frame-level enhancement methods cannot accurately identify which regions are truly beneficial to the analysis task, resulting in a waste of a large amount of computing resources on regions that are not beneficial to the analysis. For example, in a video frame containing multiple objects, only some regions (such as the bounding boxes of target objects) may be truly important for the analysis task, while the full-frame enhancement method will enhance the entire frame, including irrelevant regions such as the background. Summary of the Invention

[0005] The present invention provides a region-based video enhancement method and apparatus to solve the defects of extremely high computational costs caused by enhancing each frame of video in the prior art and significant accuracy degradation caused by selective enhancement methods, and to achieve high accuracy and high throughput in content-enhanced video analysis.

[0006] The present invention provides a region-based video enhancement method, including the following steps.

[0007] Divide the video frame to be enhanced into macroblocks as the basic unit, and predict the importance of each divided macroblock; Based on the importance prediction results of each of the macroblocks, obtain multiple regions to be enhanced of the video frame to be enhanced; Stitch the multiple regions to be enhanced; Input the stitching result into an enhancement model for enhancement processing.

[0008] According to the region-based video frame enhancement method provided by the present invention, the step of dividing the video frame to be enhanced into macroblocks as the basic unit and predicting the importance of each divided macroblock specifically includes: Divide the video frame to be enhanced into macroblocks of a predetermined size; Input each of the divided macroblocks into a pre-trained segmentation model to obtain the importance prediction result of each of the macroblocks output by the segmentation model, where the segmentation model is trained based on a sample video frame, the enhanced result of the sample video frame, and the calculation result of the importance index of each sample macroblock in the sample video frame.

[0009] According to the region-based video frame enhancement method provided by the present invention, the training process of the segmentation model specifically includes: Perform enhancement processing on each frame in the sample video frame using a pre-trained super-resolution model; Input the enhanced frame and the original frame of the sample video frame into a preset analysis model respectively, and perform forward propagation to obtain an analysis result; Input the enhanced frame and the original frame of the sample video frame into the analysis model respectively, and perform backward propagation to calculate gradients; Calculate the contribution of any macroblock MB after enhancement in the divided sample video frame to the accuracy of the analysis task, and obtain the importance index of each macroblock MB. The importance index calculation formula is: , where, i are the pixels within the macroblock, is the L1 norm; represents the analysis task I (⋅) accuracy, SR(⋅) represents super-resolution, IN(⋅) represents bilinear interpolation with the same magnification factor, represents the set of all macroblocks in the video frame to be enhanced; Store the calculated importance index in the set where the set represents a matrix or tensor corresponding to the macroblock division of the video frame, and the set Each element in corresponds to a macroblock MB, and the value of each element corresponds to the contribution degree of the macroblock MB to the accuracy of the analysis task after enhancement; Taking the value of each element in the set as the label of the importance level of the corresponding macroblock, and using the enhanced video frame and the corresponding as training data to train the segmentation model using the cross-entropy loss function.

[0010] According to the region-based video enhancement method provided by the present invention, taking the value of each element in the set as the label of the importance level of the corresponding macroblock, and using the enhanced video frame and the corresponding as training data to train the segmentation model using the cross-entropy loss function specifically includes: S1. Initialize a preset segmentation model; S2. Input the original frame of the sample video into the segmentation model to obtain the importance prediction value of each macroblock; S3. Use the cross-entropy loss function to calculate the loss between the importance prediction value of each macroblock and the true label corresponding to each macroblock; S4. Update the parameters of the segmentation model through backpropagation to minimize the preset loss function; S5. Use a preset optimizer to update the parameters of the segmentation model; S6. Repeat steps S2 - S5 until the segmentation model converges or reaches a predetermined number of training epochs.

[0011] According to the region-based video enhancement method provided by the present invention, initialize the preset segmentation model using a pre-trained MobileNetV2 backbone network; The preset optimizer includes an Adam optimizer or an SGD optimizer.

[0012] According to the region-based video enhancement method provided by the present invention, obtaining multiple regions to be enhanced of the video to be enhanced based on the importance prediction results of each macroblock specifically includes: Construct a global queue, and sort all macroblocks in the video to be enhanced in descending order of importance; According to the estimated resource and performance goals, select the regions corresponding to a predetermined number of macroblocks with higher importance rankings as the regions to be enhanced.

[0013] According to the region-based video enhancement method provided by the present invention, the method further includes: In the subsequent frames of any continuous frame, if the change in the importance of each macroblock compared to the importance of each macroblock in the previous frame of the continuous frame is less than a preset threshold, then reuse the importance prediction results of each macroblock in the previous frame of the continuous frame; Among them, the change in the importance of each macroblock is obtained by accumulating the feature changes of each macroblock in each frame. The specific formula is: , Among them, S is the cumulative result of the feature changes of each macroblock in each frame, Norm(·) is L1 normalization, ∆Φ(ResY i ) = Φ(ResY i+1 )−Φ(ResY i ), ResY i , ResY i+1 respectively represent the features of the macroblocks in the previous frame and the subsequent frame of the continuous frame.

[0014] According to the region-based video enhancement method provided by the present invention, splicing the multiple regions to be enhanced specifically includes: Calculate the connected components of the macroblocks in each region to be enhanced, and obtain multiple irregular regions based on the connected components; Use the minimum rectangle to enclose each irregular region in a corresponding rectangular box respectively; Segment the rectangular box with a size greater than the preset threshold, and replace the corresponding rectangular box with the segmented sub-rectangular boxes; Calculate the average importance of each macroblock in each rectangular box; Based on the calculation results of the average importance of each macroblock, sort the rectangular boxes in descending order; Use the greedy algorithm to pack the rectangular boxes with higher rankings into boxes of a preset size, and maximize the utilization rate of the boxes to generate a dense tensor after splicing the regions to be enhanced.

[0015] According to the region-based video enhancement method provided by the present invention, inputting the splicing result into the enhancement model for enhancement processing specifically includes: Use a pre-trained super-resolution model to enhance the dense tensor to obtain the enhanced region; Splice the enhanced region back to the corresponding position of the original frame to form the enhanced frame of the video picture.

[0016] The present invention also provides a region-based video enhancement device, including: An importance prediction module, configured to divide the video picture to be enhanced into basic units of macroblocks, and perform importance prediction on each divided macroblock; An enhanced region acquisition module, configured to obtain multiple regions to be enhanced of the video frame to be enhanced based on the importance prediction results of the macroblocks; A splicing module, configured to splice the multiple regions to be enhanced; An enhancement module, configured to input the splicing result into an enhancement model for enhancement processing.

[0017] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for region-based video frame enhancement as described in any one of the above is implemented.

[0018] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for region-based video frame enhancement as described in any one of the above is implemented.

[0019] The present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for region-based video frame enhancement as described in any one of the above is implemented.

[0020] The method and device for region-based video frame enhancement provided by the present invention divide the video frame to be enhanced into macroblocks as basic units, predict the importance of each divided macroblock; obtain multiple regions to be enhanced of the video frame to be enhanced based on the importance prediction results of the macroblocks; splice the multiple regions to be enhanced; input the splicing result into an enhancement model for enhancement processing, which can quickly and accurately identify important regions, splice sparsely distributed regions into a dense tensor and enhance it efficiently. Description of the Drawings

[0021] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 is a flowchart of the method for region-based video frame enhancement provided by the present invention.

[0023] Figure 2 is a schematic overview diagram of the regions to be enhanced provided by the present invention.

[0024] Figure 3 is a heat map of the macroblock importance provided by the present invention.

[0025] Figure 4 is a classification of the importance levels of the macroblocks provided by the present invention.

[0026] Figure 5 It is a schematic diagram of model selection for macroblock-based region importance prediction provided by the present invention.

[0027] Figure 6 It is a schematic diagram of the correlation coefficient in macroblock importance reuse across frames provided by the present invention.

[0028] Figure 7 It is a schematic diagram of frame selection based on the cumulative distribution function (CDF) in macroblock importance reuse across frames provided by the present invention.

[0029] Figure 8 It is a schematic diagram of the operation of each component in the region-based video enhancement method provided by the present invention.

[0030] Figure 9 It is a schematic diagram of the code of the region-aware bin-packing algorithm provided by the present invention.

[0031] Figure 10 It is a specific embodiment of the region-aware bin-packing algorithm provided by the present invention.

[0032] Figure 11 It is a schematic diagram of the process of the execution method based on a configuration file provided by the present invention.

[0033] Figure 12 It is a specific embodiment of the execution method based on a configuration file provided by the present invention.

[0034] Figure 13 It is a comparison of the accuracy and throughput when the region-based video enhancement method provided by the present invention performs object detection tasks on different devices.

[0035] Figure 14 It is a comparison of the accuracy and throughput when the region-based video enhancement method provided by the present invention performs semantic segmentation tasks on different devices.

[0036] Figure 15 It is the trade-off relationship between the throughput (TPT) and accuracy (ACC) of the region-based video enhancement method provided by the present invention on different devices.

[0037] Figure 16 It is the accuracy performance when the region-based video enhancement method provided by the present invention performs object detection tasks under different numbers of video streams.

[0038] Figure 17 It is the precision performance of the region-based video enhancement method provided by the present invention under different resources.

[0039] Figure 18 It is the throughput of the region prediction provided by the present invention.

[0040] Figure 19 This is the GPU resource usage provided by the present invention.

[0041] Figure 20 This is the performance of different packing strategies provided by the present invention.

[0042] Figure 21 This is the accuracy gain of cross-stream macroblock selection provided by the present invention.

[0043] Figure 22 This is the accuracy gain of macroblock selection provided by the present invention.

[0044] Figure 23 This is the execution plan of different workloads provided by the present invention.

[0045] Figure 24 This is the usage of GPU and CPU provided by the present invention.

[0046] Figure 25 This is the schematic structural diagram of the region-based picture enhancement device provided by the present invention.

[0047] Figure 26 This is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0048] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention fall within the protection scope of the present invention.

[0049] The present invention will be specifically described below with reference to the accompanying drawings of the specification. The specific operation methods in the method embodiments can also be applied to the device embodiments or system embodiments. In the description of the present invention, unless otherwise specified, "at least one" includes one or more. "Multiple" means two or more. For example, at least one of A, B, and C includes: A alone, B alone, A and B existing simultaneously, A and C existing simultaneously, B and C existing simultaneously, and A, B, and C existing simultaneously. In the present invention, " / " means "or". For example, A / B may represent A or B; "and / or" herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A existing alone, A and B existing simultaneously, and B existing alone.

[0050] The present invention will be specifically described below in combination with the specific implementation manners.

[0051] In some specific embodiments of the present invention, as Figure 1 shown, this solution provides a region-based video frame enhancement method, including: Step 100: Divide the video frame to be enhanced into macroblocks as basic units, and perform importance prediction on each divided macroblock; Step 200: Based on the importance prediction results of each of the macroblocks, obtain multiple regions to be enhanced of the video frame to be enhanced; Step 300: Stitch the multiple regions to be enhanced; Step 400: Input the stitching result into an enhancement model for enhancement processing.

[0052] It should be noted that existing video frame enhancement solutions cannot perform region-by-region execution, and enhancing each frame of the video, although it can significantly improve the accuracy of the analysis task, has extremely high computational costs. Selective enhancement methods (such as NeuroScaler and Nemo) improve throughput by only enhancing some frames, but this method will lead to a significant decrease in accuracy. Because selective enhancement methods need to make a trade-off between enhanced frames and non-enhanced frames, resulting in insufficient quality of some frames, affecting the final analysis result. And enhancing each frame takes a long time to process, resulting in low throughput. For example, when performing frame-by-frame enhancement on high-resolution videos (such as 1080p), the processing speed may not meet the real-time requirements. Although selective enhancement methods improve throughput, they still cannot reach the throughput level of only performing the analysis task (without enhancement). For example, when selective enhancement methods process multiple video streams, due to the need to allocate resources between different frames, the overall throughput is still low.

[0053] Therefore, by dividing the video frame to be enhanced into macroblocks and predicting the importance of the macroblocks, the present invention can quickly and accurately identify the regions in the video that are most helpful for the analysis task. At the same time, it can stitch the regions to be enhanced composed of scattered macroblocks into a dense tensor, and then perform efficient enhancement, reducing the computational amount of the enhancement process and also improving resource utilization.

[0054] In some possible implementation manners of the present invention, the dividing the video frame to be enhanced into macroblocks as basic units and performing importance prediction on each divided macroblock specifically includes: Dividing the video frame to be enhanced into macroblocks of a predetermined size; Inputting each of the divided macroblocks into a pre-trained segmentation model to obtain the importance prediction results of each of the macroblocks output by the segmentation model, where the segmentation model is trained based on a sample video frame, the enhanced result of the sample video frame, and the calculation results of the importance indicators of each sample macroblock in the sample video frame.

[0055] Specifically, this embodiment provides an implementation manner for predicting the importance of each divided macroblock. By dividing the video picture to be enhanced into macroblocks of a predetermined size, the importance of each divided macroblock is predicted by a pre-trained segmentation model.

[0056] To quickly and accurately identify the enhancement regions (Eregions), we propose a spatial importance prediction method based on macroblocks (Macroblock, MB). This method includes two key parts: estimating the importance of macroblocks in each frame (spatial dimension), and reusing the importance of macroblocks between consecutive frames (temporal dimension).

[0057] It can be understood that the regions to be enhanced can be of any shape. Therefore, we need to determine the granularity for constructing these regions. A natural method is to use pixel granularity. Although the pixel granularity method can provide the most accurate importance prediction, the computational cost is too high, as Figure 2 shown. Figure 2 It shows that by selectively enhancing the important regions in the video frame (instead of the whole frame), the processing efficiency can be significantly improved and the waste of computing resources can be reduced. Using the per-frame enhancement method, the enhanced area is the largest and the latency is the longest. An ideal region selection method (Oracle) can perfectly identify the regions in the video frame that need to be enhanced, thus achieving optimal performance. The region selection obtained by the region detection (Rol Detc.) method is greatly reduced, and the latency can also be greatly shortened. The impact of the region selection method on the efficiency and performance of the enhancement process in video analysis. By selectively enhancing the important regions (Eregions) in the video frame, such as Figure 2 the red region (Region Slct) in, rather than enhancing the whole frame, the processing time and the consumption of computing resources can be significantly reduced. This method improves the overall throughput of the system by reducing unnecessary enhancement operations. The region selection method can save significant time costs (e.g., a 2.4-fold improvement), significantly improve the processing speed and system throughput while maintaining a high analysis accuracy. Through selective enhancement, the system can reduce the processing of unimportant regions, thereby reducing the computational cost. Although only part of the regions are enhanced, through precise region selection, the system can still maintain a high analysis accuracy.

[0058] In view of this, inspired by video coding knowledge, we believe that using macroblocks (MBs) as the basic units of the regions to be enhanced is both efficient and accurate. A macroblock is the basic unit in video coding to which a quantization parameter (QP) is applied to control the video quality compression level. For example, in H.264 coding, a frame is divided into an array of 16×16 pixel macroblocks, and each macroblock is assigned a different QP value so that more bits are allocated to regions that require higher visual quality, while fewer bits are allocated to less important regions. After using macroblocks as the basic units, our problem is transformed into: Given a video frame containing a set of macroblocks f , select some macroblocks for enhancement to maximize the accuracy of downstream analysis tasks: ACC (I( SR( )) (1), where ACC ( is the accuracy of the analysis task I( such as, the F1 score of object detection), and are super-resolution and bilinear interpolation with the same magnification factor respectively, represents the unselected macroblock MB.

[0059] In some possible embodiments of the present invention, the training process of the segmentation model specifically includes: Performing enhancement processing on each frame in the sample video picture using a pre-trained super-resolution model; Inputting the enhanced frame and the original frame of the sample video picture into a preset analysis model respectively for forward propagation to obtain analysis results; Inputting the enhanced frame and the original frame of the sample video picture into the analysis model respectively for backward propagation to calculate gradients; Calculating the contribution of any macroblock MB after enhancement in the divided sample video picture to the accuracy of the analysis task to obtain the importance index of each macroblock MB, and the importance index calculation formula is: (2), where i is the pixel within the macroblock, is the L1 norm; represents the accuracy of the analysis task I (⋅), SR(⋅) represents super-resolution, IN(⋅) represents bilinear interpolation with the same magnification factor, represents the set of all macroblocks in the picture to be enhanced; Storing the calculated importance index in the set Among them, the set represents a matrix or tensor corresponding to the macroblock partitioning of the video frame, and each element in the set corresponds to a macroblock MB, and the value of each element corresponds to the contribution degree of the macroblock MB to the accuracy of the analysis task after enhancement; Taking the value of each element in the set as the label of the importance level of the corresponding macroblock, and using the enhanced video frame and the corresponding as training data, the segmentation model is trained using the cross-entropy loss function.

[0060] Specifically, this embodiment provides an implementation manner for training the segmentation model. In order to select beneficial macroblocks MB, by setting importance metrics, an importance score is assigned to each macroblock MB, Figure 3 which shows the MB importance heatmap. Their similarity indicates that MB importance is a good representation of the area to be enhanced.

[0061] Furthermore, in order to select beneficial macroblocks, we need to quantify their importance. Ideally, the MBs that cause greater changes in the inference results and greater changes in pixel values after enhancement are more important. Based on this, we use the above "importance" metrics such as formula (2) to observe the impact of pixel changes of each macroblock MB on the accuracy gradient, and the magnitude of pixel value changes due to enhancement.

[0062] However, the above calculation process needs to be based on the already enhanced frames, that is to say, in fact, the importance of MB cannot be directly calculated according to the above importance metrics. It is necessary to predict the importance of macroblocks on the original frames. Therefore, a learning-based method can be used to predict the MB importance in the original frames. We construct a training set by enhancing all video frames and calculating the importance metrics of each macroblock, through one forward and backward propagation of the final analysis model. The importance value of each macroblock is an entry in the mask set for each frame.

[0063] On this basis, the problem can be regarded as a segmentation task and inspired by their model design. MB importance prediction is similar to the image segmentation problem. Image segmentation aims to semantically segment an image by assigning a predefined label to each pixel, while MB importance prediction assigns an importance score to each MB. In view of this, MB importance prediction can be approximated as an image segmentation problem by simplifying the importance values into multiple importance levels. Assuming that ten levels are set in the embodiment, the goal of MB importance prediction is to assign an importance level to each MB, similar to classification in image segmentation. Figure 4Shows that this approximation can yield good performance. This observation enables us to utilize various techniques tailored for MB importance predictors, including learning-based semantic segmentation. Figure 4 In Figure 4 , Acc-Model is a model for predicting macroblock importance that can provide more accurate importance values but has a high computational complexity. In contrast, embodiments of the present invention reduce the computational complexity by classifying importance values into discrete levels while achieving a comparable accuracy in practical applications.

[0064] Different from saliency maps in computer vision that capture which pixel values have a greater impact on the DNN output, our loss function captures how specifically enhancing or not enhancing an MB (with its content quality changing) alters the DNN inference accuracy.

[0065] To accurately predict the importance of MBs and achieve high throughput, we trained a super-lightweight segmentation model using the above importance metrics. We retrained six models using cross-entropy loss with segmented (as importance levels) to support MB importance prediction. The six models are a super-lightweight model MobileSeg (with two backbone networks), two lightweight models AccModel and HarDNet, and two heavyweight models FCN and DeepLabV3. As Figure 5 shown, MobileSeg is a lightweight segmentation model using MobileNetV2 as the backbone network, AccModel is a model for predicting the specific importance value of macroblocks, HarDNet is a lightweight model, FCN (Fully Convolutional Network) is a heavier model commonly used for image segmentation tasks, and DeepLabV3 is another heavier model also commonly used for image segmentation tasks. Figure 5 The performance of these models was compared, mainly focusing on two aspects: accuracy: the accuracy of the model in predicting the importance of macroblocks; throughput: the processing speed of the model in actual operation, usually measured by the number of frames processed per second (fps). The super-lightweight model provides almost the same accuracy as the heavyweight model while providing 4 - 18 times the throughput. This can be attributed to the significantly reduced complexity of MB-level segmentation compared to traditional image segmentation. The predefined MB size in video coding supports this. The 16×16-MB coding in H.264 reduces the labels output by traditional image segmentation models from 1920×1080 to 120×68. Given the best performance of MobileSeg, we selected it as our MB importance predictor. In the offline stage, RegenHance uses the Fine-tune the predictor. RegenHance is an efficient video analysis system that focuses on achieving high-accuracy and high-throughput video content enhancement on edge devices. It significantly improves the efficiency and performance of video analysis by enhancing only the regions (Eregions) in the video that are most helpful for the analysis task, rather than the entire frame.

[0066] In some possible embodiments of the present invention, taking the value of each element in the set as the label of the importance level of the corresponding macroblock, using the enhanced video frame and the corresponding as training data, and training the segmentation model using the cross-entropy loss function, specifically including: S1. Initialize a preset segmentation model; S2. Input the original frame of the sample video into the segmentation model to obtain the importance prediction value of each macroblock; S3. Use the cross-entropy loss function to calculate the loss between the importance prediction value of each macroblock and the true label corresponding to each macroblock; S4. Update the parameters of the segmentation model through backpropagation to minimize the preset loss function; S5. Use a preset optimizer to update the parameters of the segmentation model; S6. Repeat steps S2 - S5 until the segmentation model converges or reaches a predetermined number of training epochs.

[0067] Specifically, this embodiment provides an implementation of training a segmentation model using the cross-entropy loss function. Utilizing the core role of the cross-entropy loss function in training the segmentation model, by measuring the difference between the model output and the true label, guiding the model to learn the optimal parameters. Through the above steps, a segmentation model with good performance can be effectively trained.

[0068] In some possible embodiments of the present invention, the method further includes: In the subsequent frame of any continuous frame, if the change in the importance of each macroblock is less than a preset threshold compared to the importance of each macroblock in the previous frame of the continuous frame, then reuse the importance prediction results of each macroblock in the previous frame of the continuous frame; Among them, the change in the importance of each macroblock is obtained by accumulating the feature changes of each macroblock in each frame. The specific formula is: (3), where S is the cumulative result of the feature changes of each macroblock in each frame, is the Y channel of the residual of each frame, is L1 normalization, .

[0069] Specifically, this embodiment provides an implementation method for predicting the importance index of macroblocks in consecutive frames, and only predicts the importance of macroblocks in frames with significant changes.

[0070] It can be understood that it is a common practice to reuse the DNN output between consecutive frames. Although reusing the content of enhanced frames in region enhancement will result in a significant loss of accuracy, the macroblock importance values are reusable. Therefore, predict the importance of macroblocks for a group of frames and reuse their outputs on other frames to provide the best approximation of the macroblock importance prediction for each frame.

[0071] In a possible embodiment, our key choice is to use an ultra-lightweight operator to represent the change in macroblock importance, and then only predict the importance of macroblocks in frames with significant changes. We compared many lightweight features and proposed a 1 / Area operator. The Area operator captures large blocks in the image, while 1 / Area captures the changes of small objects, as Figure 3 required for the importance of macroblocks. Statistical analysis such as Figure 6 shows that Figure 6 illustrates how to reuse the importance prediction results of macroblocks (Macroblock, MB) in the time dimension in the RegenHance system to further improve the efficiency and throughput of the system. This reuse mechanism utilizes the temporal correlation between video frames and reduces the computational cost of independent prediction for each frame. The RegenHance system uses a lightweight feature operator, the 1 / Area Operator, to capture the change in macroblock importance. This operator can effectively detect the changes of small objects and is very suitable for estimating the importance of macroblocks. Statistical analysis shows that the correlation between the 1 / Area Operator and the change in macroblock importance is 0.91, which is a good estimation metric and can estimate the changes.

[0072] Denote the 1 / Area operator as Φ, and select the frames to be predicted according to the cumulative distribution function (CDF) of ΔΦ between consecutive frames in each block. It first accumulates the feature changes using the following function when decoding an n-frame block, as shown in Equation (3).

[0073] Then, select N frames according to the CDF M calculated from S, where = 1. As Figure 7 shown, the y-axis is divided into N equally spaced intervals; in each interval, it selects a value, such as , and then its corresponding frame index, such as , is the selected frame. The prediction results of this frame are reused for other frames in each interval. In multiple streams, for a given stream jThe selected number of frames is allocated by a ratio, and its total number is determined by a profile-based execution plan. The system selects frames for importance prediction through the cumulative distribution function (CDF). Specifically, when decoding a video stream, the system calculates the feature changes of each frame and selects some frames for importance prediction according to the CDF. By this method, the system can significantly reduce the number of frames to be processed while ensuring prediction accuracy, thus improving the overall efficiency. In a multi-stream processing scenario, the system allocates prediction tasks according to the importance changes of macroblocks in each stream. By reasonably allocating prediction tasks, the system can further optimize resource utilization and improve the overall throughput.

[0074] From Figure 6 , 7 it can be seen that by reusing the importance prediction results of macroblocks in the time dimension, the system reduces the computational cost of independent prediction for each frame. By selectively performing importance prediction on some frames and applying these prediction results to other frames, the system significantly improves the throughput. Although the number of prediction tasks is reduced, by reasonably selecting the frames to be predicted, the system can still maintain high accuracy. This method is applicable to multi-stream processing scenarios and can significantly improve the overall performance of the system.

[0075] Figure 8Shows the main components of the RegenHance system during operation and their interactions, demonstrating how the system works together during actual operation to achieve efficient region selection, enhancement processing, and analysis tasks. Based on the importance prediction of macroblocks, frames are first selected, and then the importance of their macroblocks is predicted. Decoder: Responsible for decoding the compressed video stream into RGB frames for subsequent processing. MB-based Region Importance Prediction is used to perform region importance prediction at the macroblock level on the decoded frames to determine which macroblocks are most important for the analysis task. Region Selection: Selects the regions to be enhanced based on the results of the macroblock importance prediction. Region-aware Enhancement is used to enhance the selected regions to improve the image quality of these regions. Analytical Model: Performs analysis tasks on the enhanced frames, such as object detection, semantic segmentation, etc. The decoder passes the decoded frames to the macroblock importance prediction component. The macroblock importance prediction component analyzes the macroblocks in the frames, generates the importance scores for each macroblock, and passes the results to the region selection component. The region selection component selects the regions to be enhanced based on the importance scores and passes the information of these regions to the region enhancement component. The region enhancement component enhances the selected regions and passes the enhanced frames to the analytical model. The analytical model performs analysis tasks on the enhanced frames and outputs the final analysis results. By only enhancing the important regions, the system can efficiently utilize limited computing resources and improve processing efficiency; by optimizing the execution order and resource allocation of each component, the system can quickly complete analysis tasks in real-time video streams and is suitable for application scenarios such as real-time monitoring; although only some regions are enhanced, through precise region selection and enhancement processing, the system can still maintain a high analysis accuracy.

[0076] In some possible embodiments of the present invention, a pre-trained MobileNetV2 backbone network is used to initialize a preset segmentation model; The preset optimizer includes an Adam optimizer or an SGD optimizer.

[0077] Specifically, this embodiment provides an implementation manner for initializing a preset segmentation model and a preset optimizer.

[0078] In some possible embodiments of the present invention, obtaining multiple regions to be enhanced of the video picture to be enhanced based on the importance prediction results of each of the macroblocks specifically includes: Construct a global queue and sort all macroblocks in the to-be-enhanced video frame in descending order of importance; Select the regions corresponding to the top predetermined number of macroblocks in the importance sorting as the to-be-enhanced regions according to the estimated resources and performance goals.

[0079] Specifically, this embodiment provides an implementation manner for obtaining the to-be-enhanced regions. To maximize the overall accuracy improvement, select the macroblocks (MBs) that can provide the highest total accuracy from all video streams. As Figure 7 shown, the system constructs a global queue for aggregating and sorting the macroblocks in all streams according to the importance (level) in the macroblock index, where the macroblock index may include {streamid, frameid, locx, locy, importance}, and locx and locy represent the coordinate positions of the macroblock in the frame.

[0080] On this basis, select the top N macroblocks in importance and pass their indices to the region-aware bin-packing algorithm for further processing. The number of selected macroblocks is estimated according to the following formula: (4), where is the size of the macroblock (for example, 16×16 in H.264), and H, W, and B are the preset best height, width, and batch size of the enhancement model, respectively, and these parameters are determined in the subsequent execution plan.

[0081] In some possible implementation manners of the present invention, splicing the multiple to-be-enhanced regions specifically includes: Calculate the connected components of the macroblocks in each to-be-enhanced region, and obtain multiple irregular regions based on the connected components; Use the minimum rectangle to limit each irregular region within the corresponding rectangular frame; Split the rectangular frames with sizes greater than the preset threshold, and replace the corresponding rectangular frames with the split sub-rectangular frames; Calculate the average importance of each macroblock within each rectangular frame; Sort the rectangular frames in descending order according to the calculated results of the average importance of each macroblock; Use the greedy algorithm to pack the top-ranked rectangular frames into a preset-size box and maximize the utilization rate of the box to generate a dense tensor after splicing the to-be-enhanced regions.

[0082] Specifically for region awareness, this embodiment provides an implementation manner for splicing the multiple to-be-enhanced regions. This method is called the region-aware bin-packing algorithm, and the region-aware bin-packing algorithm will be described in detail through specific embodiments below.

[0083] Considering the sparse distribution characteristics of the selected macroblocks, the unique requirements of the enhancement model (requiring rectangular inputs, with the latency being proportional to the input size and independent of pixel values), and the complexity of batch execution, we propose a region-aware binning algorithm to construct the selected macroblocks into irregular regions and stitch them into dense tensors.

[0084] This problem is modeled as a two-dimensional binning problem, with the goal of packing as many selected macroblocks as possible into the given bins. The input includes macroblock indices, the number of bins B, and the size of the bins H×W, and the output is the binning plan for the macroblocks. We deal with macroblock indices instead of the actual image data to avoid frequent memory I / O operations. This problem is known to be an NP-hard problem. Previous methods cannot handle irregular regions because they mainly perform batch processing for standard rectangular DNN inputs.

[0085] To achieve a better balance between bin utilization and algorithm efficiency, we propose a region-aware binning algorithm (see Algorithm 1 in Figure 9 ), and its key design choices include: 1) bounding the irregular regions with rectangular boxes for efficient search of the binning plan (lines 3 - 5); 2) cutting large bins into small bins and sorting the bins according to the average importance (i.e., importance density) of all macroblocks in the bins to achieve high bin utilization (line 6). Figure 10 shows how this prioritization leads to a higher (13%) accuracy improvement, while the traditional "large item first" (largest area first) strategy only brings a 6% improvement. The region-aware binning algorithm first constructs regions by computing the connected components of the selected macroblocks (line 3), and then bounds each region with an extended rectangular box (line 4), such as ①②③ in Figure 10 . Next, if the size of a bin exceeds a preset value, it is split into smaller bins to avoid introducing too many unimportant macroblocks (line 5), such as splitting bin ① in Figure 10 into ① and ① . Then, the bins are sorted according to the priority order we proposed, that is, not in the order of the largest area first strategy (①, ②, ③), but in the order of (②, ③, ① ) (line 6). Finally, the algorithm iteratively packs the bins into the bins. In each iteration (lines 7 - 11), it rotates and packs a bin into the free area in the bin, updates the placement position of the bin and the list of free areas, and then deletes the bin from the bin list. For example, the "L" - shaped region in Figure 8 and ① in Figure 10 in The box will be rotated and loaded into the box.

[0086] In some possible embodiments of the present invention, inputting the splicing result into the enhancement model for enhancement processing specifically includes: Using a pre-trained super-resolution model to enhance the dense tensor to obtain an enhanced region; Splicing the enhanced region back to the corresponding position of the original frame to form an enhanced frame of the video picture.

[0087] Specifically, this embodiment provides an implementation manner of inputting the splicing result into the enhancement model for enhancement processing, and uses a super-resolution model for enhancement.

[0088] In possible embodiments, according to the execution plans of different devices, except for Jetson AGX Orin equipped with unified memory, other devices need to transfer the actual frame from the main memory to the GPU memory. To save time, we hide this transfer process while performing macroblock selection and binning. This design is feasible because after the region importance prediction processing based on macroblocks, as Figure 2 shown, all modules only process macroblock indices before super-resolution processing. Therefore, we splice the actual content region into a tensor (box) on the GPU according to the binning plan, then perform enhancement processing on it, and splice the enhanced region back into the non-enhanced region of bilinear interpolation for final analysis.

[0089] When splicing the enhanced content back into the bilinear interpolation frame, jagged edges and blocky artifacts may occur.

[0090] In some specific implementation manners of the present invention, as Figure 11 shown, this solution provides an execution method based on a configuration file for the region-based picture enhancement method described in any one of the above, including: Step 1110, obtaining an edge server equipped with R computing resources and a user's video analysis job; Step 1120, parsing the analysis task to obtain a data flow graph, where the data flow graph is used to describe the data flow and dependency relationships between components in the video analysis task; each node in the data flow graph represents a component, and the edge represents the data flow between components; Step 1130, running a workload on the edge server to analyze the performance of each component on different hardware resources; Step 1140, collecting the performance information of each component on different hardware, and generating a performance configuration file for each component based on the performance information; Step 1150: Use a dynamic programming algorithm to generate an execution plan by recursively calculating the optimal resource allocation for each node. Step 1160: According to the generated execution plan, load each component onto the specified hardware and configure the corresponding resources. Step 1170: During runtime, schedule the execution of each component according to the execution plan to ensure the efficient operation of the system. Step 1180: During operation, monitor the performance of the system in real time and dynamically adjust the execution plan according to the actual load and resource usage. Step 1190: If there is a deviation between the performance of the system and the expected goal, automatically adjust the resource allocation and execution plan to ensure the optimal overall performance of the system.

[0091] Specifically, during the actual execution of any of the above embodiments, given an edge server equipped with computing resources R (i.e., the processing time of the processor at 100% utilization) and the user's analysis task, the execution plan based on the configuration file aims to allocate the most suitable resources for each component to maximize the end-to-end throughput while meeting the performance goals. The problem can be formulated as follows: (5), where, G is the dataflow graph (DFG) of the component, u is the graph G in the node, represents the resources allocated to the node u , represents the end-to-end throughput.

[0092] Simplify the steps in Figure 11 as shown in Figure 12 , we propose an execution plan based on the configuration file, which includes the following steps: Step 1210: Parse the dataflow graph (DFG) of the analysis task specified by the user.

[0093] Step 1220: Run the workload (e.g., the video stream specified by the user) on all components (including the model uploaded by the user) to analyze their capacities on all available hardware.

[0094] Step 1230: Generate an execution plan that meets the performance goals specified by the user.

[0095] Step 1240: Load each component into the corresponding hardware.

[0096] Our key choice is to allocate the size of the input tensor (e.g., batch size) for each component for resource allocation. Batch execution (i.e., combining input matrices together) is a common method for DNNs to achieve high processor utilization, and it also allows the inference engine to achieve different throughputs by adjusting different batch sizes.

[0097] Since the data flow graph (DFG) of any job is naturally a directed acyclic graph (DAG), we can use dynamic programming to solve this optimization problem. Define as the maximum throughput of node r and its subtree within the resource budget u in the graph G , and its value is equal to the minimum node on each path. For non-leaf nodes u , the algorithm allocates a resource budget u ′ for node r , allocates at most r − r ′ to the subtree, and then enumerates all r ≤ R to find the optimal allocation and allocate the optimal batch size b for each node. Formally, for all ∀(u, v) ∈ E(G), we have: (6), where is the resource cost of node u at the b batch size. Allocate the least resources for the analysis model to meet the user's latency target, and then allocate the batch size for other components according to the above formula.

[0098] Generally speaking, the optimal solution always converges to an allocation that is not bottlenecked by any node; in other words, each node in the graph produces the same throughput. Therefore, when the user's registration changes frequently, the execution plan must be able to be generated quickly. For this purpose, we will explore relevant methods in future work, such as online deep reinforcement learning and combinatorial optimization-related methods, such as local search.

[0099] To further verify the effectiveness of the present invention, RegenHance is evaluated on five heterogeneous devices using two video analysis tasks. The evaluation results are as follows: The present invention improves the accuracy by 10 - 19%, and achieves a 2 - 3 times end-to-end throughput improvement compared with existing frame-based enhancement methods; The present invention shows strong effectiveness on various devices, which have different computing resources, analysis tasks and models, diverse workloads and performance goals, and different resolutions.

[0100] The present invention brings significant improvements in accuracy and throughput based on macroblock region importance prediction and region-aware enhancement, while the profile-based execution plan greatly improves resource utilization and throughput.

[0101] In a specific embodiment, we implemented the enhanced region on commercial frameworks, including FFmpeg (v4.4.2), Pytorch (v1.8.2), Paddleseg (v2.7.0), ONNX, OpenVINO (v2023.0.1), TensorRT (v8.4.2.4), and the code was written in Python (v3.8). The implementation of the macroblock-based region importance prediction and region-aware enhancer is as follows: Retrain the macroblock importance predictor using MobileSeg with a MobileNetV2 backbone network and prune 50% of its parameters using an L1 norm pruner. Then, further export the model to the ONNX version (using the paddle.onnx.export API) for efficient operation using OpenVINO runtime on Intel CPUs and export it to the TensorRT version for NVIDIA GPUs (using the trtexec library). Unless otherwise specified, all TensorRT models are set to the FP16 dynamic shape version, and the engine files are exported and inferred using PyCUDA (v2022.2).

[0102] Furthermore, modify the ff_h264_idct_add API in FFmpeg to extract the residuals for temporal macroblock importance reuse. Use a pre-trained super-resolution model as the enhancer and convert it to the ONNX version and TensorRT version to improve efficiency.

[0103] Loading the model into memory may take several hundred milliseconds to several seconds. Therefore, at runtime, pre-load the DNN into the specified processor and then call it. It contains approximately 5.3K lines of code.

[0104] For specific experimental setups, downstream tasks, and datasets. As summarized in Table 1, shown in Figure 1, we selected two downstream analysis tasks, object detection and semantic segmentation, to evaluate the performance of RegenHance because they play a central role in various high-level tasks. Object detection aims to identify objects of interest (i.e., location and category) in each video frame; its accuracy is measured by the average F1 score in each stream with an Intersection over Union (IoU) threshold of 0.5. Semantic segmentation assigns a category to each pixel, and we use the mean Intersection over Union (mIoU) to measure its accuracy. We also evaluated their throughput, i.e., the number of video streams that can be processed in real time.

[0105] Table 1

[0106] These two tasks were tested on two models (lightweight and heavyweight) and two video sets. For object detection, we used the Yoda dataset and collected 120 video clips from YouTube containing various scenarios with different characteristics in terms of time, lighting, object density and speed, and road type. The labels used to train the importance predictor were generated using Mask R-CNN (Swin backbone) on single-frame enhancement because it has State-of-the-Art (SOTA) performance; if users provide their own customized analysis models, we will use their outputs as labels. Here, we take YOLO as an example. If this paper is accepted, we will make this dataset publicly available. For semantic segmentation, we used the BDD100K and Cityscape public datasets. We re-encoded these datasets into H.264 videos with a resolution of 360P, 30fps, and a bitrate of 1024kbps to avoid the influence of different video encoders.

[0107] For the settings of the device, we conducted experiments on five heterogeneous devices, which were divided into four categories. For comparison, we deployed and tested RegenHance on a cloud server equipped with NVIDIA A100 GPUs and Intel(R) i9-12900K CPUs. As one of the most popular configurations for edge servers, we conducted experiments on an edge device equipped with NVIDIA Tesla T4 and Intel i7-8700 CPUs. To explore gaming graphics cards and their generational differences, we tested on NVIDIA RTX4090 and NVIDIA RTX3090Ti equipped with Intel i9-13900K respectively. For embedded edge devices, we used NVIDIA Jetson AGX Orin 64GB as the platform baseline. To demonstrate the advantages brought by the present invention, RegenHance, we compared it with the following four different baselines: Only infer: Directly apply the analytical DNN on each original frame without enhancement.

[0108] NeuroScaler: This is the state-of-the-art frame-based enhancement method. It first enhances the anchor frames and multiplexes their quality gains on non-anchor frames, and then infers on all frames. It quickly selects anchor frames in a heuristic way.

[0109] Nemo: This method also only enhances the anchor frames and multiplexes their quality gains on non-anchor frames, and then infers on all frames; but it iteratively selects the best anchor frames based on the enhancement results.

[0110] Regarding the end-to-end performance, the present invention presents the end-to-end performance of RegenHance on various devices. Unless otherwise specified, all results were tested on RTX4090 with a target latency of 1 second, an accuracy of 90% for object detection, and an accuracy of 88% for semantic segmentation (a 10% improvement compared to only infer). Performance tests were conducted on different devices, such as Figure 13 and Figure 14As shown, RegenHance achieves high accuracy and high throughput simultaneously on various devices. Of course, it cannot reach the throughput of inference-only because the additional enhancement processing increases the time cost. However, compared with the existing state-of-the-art methods (NEMO and NeuroScale), it provides a significant throughput improvement. In object detection, RegenHance is on average 12 times and 2.1 times higher in throughput than theirs; in semantic segmentation, it is 11 times and 1.9 times higher respectively. Semantic segmentation shows a more significant improvement in accuracy because it is more sensitive to visual details. RegenHance can maintain this advantage on all five heterogeneous devices because the profile-based execution plan can always generate an optimal throughput plan.

[0111] Regarding the trade-off between accuracy and throughput, RegenHance creates a trade-off space between accuracy and throughput, as Figure 15 shown; it can be served on edge servers with heterogeneous resources. Edge servers equipped with higher resources can generate a larger trade-off space. For example, RegenHance supports object detection on RTX4090 or A100 GPUs, processing ten video streams (300fps) with an accuracy of 91%; if the accuracy requirement is more stringent, RegenHance will make corresponding adjustments and serve up to six streams with an accuracy of 95%. On devices with fewer resources, such as NVIDIA T4 and Jetson AGX Orin, although the maximum frame rate decreases, RegenHance can still provide significant throughput under different accuracy targets.

[0112] Performance improvement in multi-video stream scenarios. As the number of competing video streams increases, RegenHance can always achieve higher accuracy than the other three frame-based enhancement methods. For example, in the Figure 16 tests on RTX4090 in [reference], compared with the selective enhancement method, RegenHance improves the accuracy by 8-14% in six video streams. This is because in the case of limited resources, each video stream is allocated limited resources, and our method can always enhance the most valuable regions; in contrast, the selective baseline and frame-by-frame baseline waste too many resources on unimportant content. RegenHance significantly improves the throughput of content-enhanced video analysis.

[0113] Frame latency at different batch sizes. The definition of latency is the same as in previous studies (e.g., DDS, AWStream, and Reducto), that is, from encoding a 1-second video block (30 frames) by the camera to the inference results of all 30 frames on the edge device Based on any of the above embodiments, during the operation of the region-based video enhancement method proposed by the present invention, in-depth performance analysis was conducted on each component of RegenHance to evaluate the contribution of each component to the overall system performance.

[0114] On the one hand, macroblock-based region importance prediction is one of the key steps to obtain the enhanced region. It can significantly improve the throughput and accuracy of the system by quickly and accurately identifying the regions (Eregions) in the video frame that are most valuable for improving the analysis accuracy.

[0115] Specifically, considering accuracy, Figure 17 shows the accuracy improvement obtained by different enhancement methods when performing object detection tasks on six video streams with the same computational resource allocation. Compared with frame-based enhancement methods (such as Nemo and NeuroScaler), the region enhancement method of RegenHance can achieve higher accuracy improvement, which is 3 - 4% and 4 - 8% higher respectively. This is because our predictor can more accurately identify the regions that are most valuable for improving the analysis accuracy.

[0116] Considering throughput, Figure 18 shows that our lightweight MB importance predictor can run at a speed of 30 frames per second on a single i7 - 8700 CPU core, which is more than 60 times faster than the RPN-based RoI selection method in DDS. On the GPU, it can reach 973 frames per second, which is more than 12 times faster than DDS. In addition, through the time multiplexing mechanism, the throughput of our method is further increased by 2 times.

[0117] Considering resource savings, Figure 19 shows that compared with the frame-by-frame enhancement method based on frames, Nemo, and NeuroScaler, the method provided by the present invention reduces the GPU usage by 77%, 28%, and 20% respectively when enhancing a single video stream (30 FPS) to achieve an accuracy of more than 90%. Compared with the RoI selection method of DDS, this method saves 37% of the GPU usage because it can more accurately identify the regions that are most valuable for improving the analysis accuracy.

[0118] On the other hand, region-aware enhancement is the second key step to obtain the enhanced region. It significantly improves the throughput and overall accuracy of the system through efficient region binning and enhancement strategies.

[0119] Specifically, considering throughput, the region-aware binning algorithm provided by the present invention maximizes the throughput of RegenHance while maintaining high stability and occupancy rate, thereby effectively reducing the overhead of super-resolution (SR). We conducted 1000 experiments by randomly shuffling the order of six video streams and compared with the classic Guillotine strategy and Block strategy (i.e., MB binning) to evaluate the occupancy rate (i.e., the proportion of selected MBs in all enhanced content). Figure 20 Shows the occupancy rate differences at the mean, 90th percentile, and 95th percentile; our binning strategy achieved the highest occupancy rate of 75%, which is 13%, 9%, and 9% higher than the comparison methods respectively.

[0120] Considering accuracy, cross-stream MB selection, especially our custom sorting order, achieved a significant accuracy improvement by considering the heterogeneity of region accuracy improvement between different streams. We compared our MB selection method with the uniform allocation method (Uniform), which allocates the same number of MBs to each stream, and the threshold method (Threshold), which sets a fixed threshold of 0.5 for all streams to select the importance of MBs. As Figure 21 shown, our method achieved an accuracy improvement of 8 - 12% and 2 - 3% higher than the uniform allocation method and the threshold method respectively. Figure 22 Further demonstrates the superiority of our custom sorting priority. Compared with the traditional maximum-item-first (maximum region first in our context) binning strategy, our method achieved a 50% accuracy improvement.

[0121] Thirdly, the performance-analysis-based execution plan is the third key step to obtain the enhanced region. It maximizes the system throughput by performing performance analysis on different devices and allocating devices and resources to different components and models according to the analysis results.

[0122] Specifically, considering resource allocation, the performance-analysis-based execution plan enables RegenHance to achieve the best performance on the given devices and workloads, avoiding any component becoming a bottleneck. Figure 23 Visualizes the allocation of computing resources for two different object detection models on an i9-13900K + RTX4090 to meet the Figure 13 (a) shown performance requirements. Given the different workloads of YOLOv5s (16.9 GFLOPs) and Mask R-CNN (Swin backbone, 267 GFLOPS), it allocates more resources to the analysis task.

[0123] Considering resource usage, Figure 24Shows the real-time utilization of the CPU and GPU when performing object detection on six video streams. We used Nvidia Nsight to monitor the GPU utilization and HTOP to monitor the CPU utilization. The GPU can reach full load about 95 - 99% of the time, while the CPU can reach high load about 81% of the time. This result indicates that the execution plan achieves efficient GPU-CPU cooperation. Compared with the region-independent sketch method, the method provided by the present invention achieves a throughput improvement of 2.3 times.

[0124] Considering scalability and initialization, there are two types of preparation times in our system. If only the input video stream changes, the initialization takes about 0.6 - 2 seconds, depending on the requirements of specific devices and models. When given a new device, it takes about 1 - 3 minutes to execute the performance analysis-based execution plan.

[0125] The present invention provides a content enhancement technology in video analysis applications, which solves the problem that frame-based content enhancement methods waste too much computing resources on irrelevant images. And a region-based content enhancement technology, a region-aware resource scheduler matching it, and a region enhancement system are proposed. In the evaluation using five heterogeneous devices, we demonstrated that region enhancement can achieve an order-of-magnitude improvement in performance compared to frame-based content enhancement video analysis methods.

[0126] In some specific embodiments of the present invention, as Figure 25 shown, the present solution provides a region-based picture enhancement device, including: An importance prediction module 2501, configured to divide the video picture to be enhanced into macroblocks as basic units, and perform importance prediction on each divided macroblock; A region to be enhanced acquisition module 2502, configured to obtain multiple regions to be enhanced of the video picture to be enhanced based on the importance prediction results of the respective macroblocks; A splicing module 2503, configured to splice the multiple regions to be enhanced; An enhancement module 2504, configured to input the splicing result into an enhancement model for enhancement processing.

[0127] The region-based picture enhancement device provided by the embodiments of the present invention has the same implementation principle and beneficial effects as those of the region-based picture enhancement method shown in the above embodiments. For the implementation principle and beneficial effects of the region-based picture enhancement method shown in the above embodiments, reference can be made, and details are not described herein again.

[0128] Figure 26 Illustrates a schematic diagram of the physical structure of an electronic device, as Figure 26As shown, the electronic device may include: a processor 2610, a communications interface 2620, a memory 2630, and a communication bus 2640. Among them, the processor 2610, the communications interface 2620, and the memory 2630 communicate with each other via the communication bus 2640. The processor 2610 may invoke the logical instructions in the memory 2630 to execute a region-based picture enhancement method, which includes: dividing the video picture to be enhanced into macroblocks as basic units, and performing importance prediction on each divided macroblock; based on the importance prediction results of each of the macroblocks, obtaining multiple regions to be enhanced in the video picture to be enhanced; splicing the multiple regions to be enhanced; and inputting the splicing result into an enhancement model for enhancement processing.

[0129] In addition, when the logical instructions in the above-mentioned memory 2630 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs, Read-Only Memories), random access memories (RAMs, Random Access Memories), magnetic disks, or optical discs that can store program codes.

[0130] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the region-based picture enhancement method provided by the above-mentioned methods. The method includes: dividing the video picture to be enhanced into macroblocks as basic units, and performing importance prediction on each divided macroblock; based on the importance prediction results of each of the macroblocks, obtaining multiple regions to be enhanced in the video picture to be enhanced; splicing the multiple regions to be enhanced; and inputting the splicing result into an enhancement model for enhancement processing.

[0131] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a region-based video enhancement method provided by the above-mentioned various methods. The method includes: dividing the video frame to be enhanced into macroblocks as basic units, and performing importance prediction on each divided macroblock; obtaining a plurality of regions to be enhanced of the video frame to be enhanced based on the importance prediction results of the macroblocks; splicing the plurality of regions to be enhanced; and inputting the splicing result into an enhancement model for enhancement processing.

[0132] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0133] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0134] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A region-based image enhancement method, characterized in that, Including: Dividing the video frame to be enhanced into macroblocks as basic units, and predicting the importance of each divided macroblock; Based on the importance prediction results of each of the macroblocks, obtaining multiple regions to be enhanced in the video frame to be enhanced; Stitching the multiple regions to be enhanced; Inputting the stitching result into an enhancement model for enhancement processing.

2. The region-based image enhancement method according to claim 1, wherein The step of dividing the video frame to be enhanced into macroblocks as basic units and predicting the importance of each divided macroblock specifically includes: Dividing the video frame to be enhanced into macroblocks of a predetermined size; Inputting each of the divided macroblocks into a pre-trained segmentation model to obtain the importance prediction results of each of the macroblocks output by the segmentation model, where the segmentation model is trained based on a sample video frame, the enhanced result of the sample video frame, and the calculation results of the importance indicators of each sample macroblock in the sample video frame.

3. The region-based picture enhancement method according to claim 2, wherein The training process of the segmentation model specifically includes: Performing enhancement processing on each frame in the sample video frame using a pre-trained super-resolution model; Inputting the enhanced frame and the original frame of the sample video frame into a preset analysis model respectively for forward propagation to obtain analysis results; Inputting the enhanced frame and the original frame of the sample video frame into the analysis model respectively for backward propagation to calculate gradients; Calculating the contribution of any macroblock MB after enhancement in the divided sample video frame to the accuracy of the analysis task to obtain the importance indicators of each of the macroblocks MB, and the importance indicator calculation formula is: , Among them, i is the pixel within the macroblock, is the L1 norm; represents the analysis task I (⋅) represents the accuracy, SR(⋅) represents super-resolution, and IN(⋅) represents bilinear interpolation with the same magnification factor, represents the set of all macroblocks in the picture to be enhanced; Store the calculated importance indicators in the set where the set represents a matrix or tensor corresponding to the macroblock partitioning of the video frame, and each element in the set corresponds to a macroblock MB, and the value of each element represents the contribution degree of the macroblock MB to the accuracy of the analysis task after enhancement; Take the value of each element in the set as the label of the importance level of the corresponding macroblock, and use the enhanced video frame and the corresponding as training data to train the segmentation model using the cross-entropy loss function.

4. The region-based image enhancement method according to claim 3, wherein The step of taking the value of each element in the set as the label of the importance level of the corresponding macroblock, and using the enhanced video picture frame and the corresponding as training data, and training the segmentation model by using the cross-entropy loss function specifically includes: S1. Initializing a preset segmentation model; S2. Inputting the original frame of the sample video frame into the segmentation model to obtain the importance prediction value of each macroblock; S3. Using a cross-entropy loss function to calculate the loss between the importance prediction value of each macroblock and the true label corresponding to each of the macroblocks; S4. Updating the parameters of the segmentation model through backward propagation to minimize a preset loss function; S5. Using a preset optimizer to update the parameters of the segmentation model; S6. Repeating steps S2 - S5 until the segmentation model converges or reaches a predetermined number of training epochs.

5. The region-based image enhancement method according to claim 4, wherein Initializing a preset segmentation model using a pre-trained MobileNetV2 backbone network; The preset optimizer includes an Adam optimizer or an SGD optimizer.

6. The method for enhancing a picture based on a region according to claim 4, wherein The step of obtaining multiple regions to be enhanced in the video frame to be enhanced based on the importance prediction results of each of the macroblocks specifically includes: Constructing a global queue, and sorting the importance of all macroblocks in the video frame to be enhanced in descending order; Selecting the regions corresponding to a predetermined number of macroblocks with higher importance ranking as the regions to be enhanced according to the estimated resource and performance targets.

7. The region-based picture enhancement method according to any one of claims 1, wherein The method further includes: In the subsequent frame of any continuous frame, if the change in the importance of each macroblock is less than a preset threshold compared to the importance of each macroblock in the previous frame of the continuous frame, then reusing the importance prediction results of each macroblock in the previous frame of the continuous frame; Wherein, the change in the importance of each macroblock is obtained by accumulating the feature changes of each macroblock in each frame, and the specific formula is: , Among them, S is the cumulative result of the feature changes of each macroblock in each frame, Norm(·) is L1 normalization, and ∆Φ(ResY i ) = Φ(ResY i+1 ) − Φ(ResY i ), where ResY i and ResY i+1 respectively represent the features of the macroblocks in the previous frame and the subsequent frame in consecutive frames.

8. The region-based picture enhancement method according to any one of claims 1-7, characterized in that The step of stitching the multiple regions to be enhanced specifically includes: Calculate the connected components of the macroblocks in each of the regions to be enhanced, and obtain multiple irregular regions based on the connected components; Use the minimum rectangle to enclose each irregular region within a corresponding rectangular box; Segment the rectangular boxes with dimensions greater than a preset threshold, and replace the corresponding rectangular boxes with the segmented sub-rectangular boxes; Calculate the average importance of each macroblock within each rectangular box; Based on the calculation results of the average importance of each macroblock, sort the rectangular boxes in descending order; Use the greedy algorithm to pack the rectangular boxes with higher rankings into boxes of a preset size, and maximize the utilization rate of the boxes to generate a dense tensor after splicing the regions to be enhanced; 9. The region-based picture enhancement method according to claim 8, wherein The step of inputting the splicing result into an enhancement model for enhancement processing specifically includes: Use a pre-trained super-resolution model to enhance the dense tensor to obtain an enhanced region; Splice the enhanced region back to the corresponding position of the original frame to form an enhanced frame of the video picture.

10. A region-based image enhancement device, characterized in that It includes: An importance prediction module for dividing the video picture to be enhanced into macroblocks as basic units and predicting the importance of each divided macroblock; An enhanced region acquisition module for obtaining multiple regions to be enhanced of the video picture to be enhanced based on the importance prediction results of each macroblock; A splicing module for splicing the multiple regions to be enhanced; An enhancement module for inputting the splicing result into an enhancement model for enhancement processing.

Citation Information

Cited By

  • Video picture optimization method and device, equipment and medium

    CN121644889A