Real-time target detection system for super-resolution video
By adopting optimized background modeling and distance threshold-based target segmentation technology in the super-resolution video object detection system, the problems of low detection accuracy and high resource consumption in the prior art are solved, and the target detection effect with higher accuracy and lower resource consumption is achieved.
Patent Information
- Application Number
- CN202510043584.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
AI Technical Summary
The existing video object detection method based on deep learning has low detection accuracy and high hardware resource consumption when processing super-resolution video.
A real-time object detection system for super-resolution video is proposed. Through optimized background modeling and target segmentation based on distance threshold, the impact of background information changes is reduced, and the multi-objective segmentation effect with lower resource consumption and higher algorithm timing is achieved.
Lower latency and higher precision object detection effects are achieved in super-resolution video, with average accuracy (mAP) improved by about 15% and computational complexity reduced by about 65%.
Smart Images

Figure CN119964056A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of image processing, in particular to a real-time target detection system for super-resolution videos. Background Art
[0002] Existing deep learning-based video object detection methods usually combine object detection networks (such as YOLO, FasterR-CNN) with object tracking algorithms (such as SORT, DeepSORT), and have been widely used in industries such as autonomous driving and video surveillance. However, these methods will cause a significant drop in detection accuracy when processing ultra-high definition video (UHD) such as 4K (3840*2160 pixels) and 8K (7680*4320 pixels) videos. At the same time, the large amount of resource consumption caused by the deployment of neural network models for object detection on hardware platforms has also become an important reason for limiting super-resolution video object detection. Summary of the invention
[0003] In view of the shortcomings of the prior art in super-resolution video detection, such as low detection accuracy and large consumption of hardware resources, the present invention proposes a real-time target detection system for super-resolution video. Through optimized background modeling and target segmentation based on distance threshold, the influence of background information changes caused by environmental factors is avoided while achieving a multi-target segmentation effect with lower resource consumption and higher algorithm timing, and a lower latency and higher accuracy target detection effect for super-resolution video.
[0004] The present invention is achieved through the following technical solutions:
[0005] The present invention relates to a real-time target detection system for super-resolution video, comprising: a background modeling module, an image morphology operation module, a target segmentation module and a neural network detection reasoning module, wherein: the background modeling module performs background modeling and updates on an input video frame, and performs difference between the input video frame and the background frame to obtain a foreground frame; the image morphology operation module performs corrosion and expansion operations on the foreground frame; the target segmentation module extracts a target area from the foreground frame after image morphology processing according to a target segmentation algorithm based on a distance threshold; and the neural network detection reasoning module detects an image of the extracted target area and restores the result to a super-resolution, and then outputs an image detection result of each frame of the super-resolution video.
[0006] The background modeling module comprises: a weighting unit, a storage unit and a difference unit, wherein: the weighting unit performs weighted processing according to the scaled image pixel data and the background pixel data stored in the memory through the multiplier and adder logic resources on the FPGA to obtain the updated background pixel data; the storage unit saves the updated background pixel data in the memory according to the weighted calculation, and reads the background pixel data of the previous frame from the memory according to the effective signal of the image pixel data transmitted by the previous module; the difference unit performs a difference operation through the adder according to the image pixel data signal transmitted by the previous module and the background pixel data of the previous frame read from the memory to obtain the differenced foreground pixel data of each pixel, and performs a binarization operation on the differenced result through the comparator according to the set binarization threshold to obtain the binarized pixel data.
[0007] The image morphological operation module includes: an erosion unit and a dilation unit, wherein: the erosion unit uses the binary pixel data calculated by the background modeling module to cache the pixel data rows of the image by the hardware logic resource FIFO on the FPGA and uses the register to cache the pixels, and finally obtains the erosion / dilation value of the pixel point through a logical AND / OR operation.
[0008] The target segmentation module includes: a pixel counting unit (Pixel Counter Unit) and a target frame merging unit (Merge Box Unit), wherein: the pixel counting unit calculates the row and column position information corresponding to each pixel on the image according to the binary image pixel signal calculated by the image morphological operation module and the width and height information of the image, and the column address and row address are stored in the register group, and the target frame information corresponding to each pixel is generated in real time according to the position information and stored in the register group; the target frame merging unit determines whether the target frame generated by each pixel needs to be merged with the existing target frame group and updates and stores it in the register group.
[0009] The judgment refers to whether the coordinate position of the pixel, ie, the column address and the row address, falls within any target frame stored in the target frame diagram and is merged with the target frame that meets the conditions. Technical Effects
[0010] The present invention proposes a two-stage real-time target detection system for super-resolution video, which realizes the extraction and segmentation of multiple targets in binary images and a configurable motion region fusion strategy based on the target segmentation technology of the distance threshold, and is used to process the extracted target area to reach the image size required by the neural network and use the target detection model for detection, so as to reduce the complexity of the neural network reasoning calculation part in the target detection model. Compared with the simple one-stage detection scheme of directly scaling the super-resolution image to the size required by the neural network model, the average precision (mAP) of the present invention on the data set is improved by about 15%, and compared with the one-stage target detection scheme of cutting the original super-resolution image or video into multiple small images and detecting them separately, and finally merging the detection results and restoring them to the original super-resolution video, the amount of calculation is reduced by 65%. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a flow chart of the present invention;
[0012] Figure 2 (a) is the background modeling process flow chart;
[0013] Figure 2 (b) is the image corrosion flow chart;
[0014] Figure 2 (c) is the target segmentation flow chart based on distance threshold;
[0015] Figure 3 (a)-(f) are process flow diagrams of the embodiment;
[0016] Figure 4 This is a physical diagram of the embodiment system. DETAILED DESCRIPTION
[0017] like Figure 1 As shown, this embodiment relates to a real-time target detection system for super-resolution video, including: a background modeling module, an image morphological operation module, a target segmentation module and a neural network detection and reasoning module, wherein: the background modeling module performs background modeling and updates the input video frame, and differentiates the input video frame from the background frame to obtain a foreground frame; the image morphological operation module performs corrosion and expansion operations on the foreground frame; the target segmentation module extracts the target area from the foreground frame after image morphological processing according to a target segmentation algorithm based on a distance threshold; the neural network detection and reasoning module detects the image of the extracted target area and restores the result to super-resolution, and then outputs the image detection result of each frame of the super-resolution video.
[0018] like Figure 2As shown in (a), the background modeling refers to: performing a weighted summation of the current input video frame and the calculated background frame to obtain the background frame updated at the current moment, specifically: B n (x,y)=BG-K·B n-1 (x,y)+IM-K·I n (x,y), where: B n-1 (x, y) is the background frame data of the previous moment, I n (x, y) is the input video frame data at the current moment, B n (x, y) is the calculated updated background frame data, n represents the current moment, n-1 represents the previous moment, (x, y) represents the position coordinates of a pixel point on the image; IM_K is the weight parameter of the input video frame, and BG_K is the weight parameter of the background frame. In principle, the two parameters need to satisfy the condition IM_K+BG_K=1 to prevent pixel overflow in the calculated background frame (grayscale value exceeds 255 or is less than 0).
[0019] This embodiment stores the calculated background frame in a memory (BRAM) to achieve fast data reading and writing.
[0020] The difference refers to: differentiating the input video frame stream from the background frame to obtain a foreground frame stream containing target information, specifically: F t (x,y)=|I t (x,t)-B(x,y)|, where: I t (x,y), B(x,y), and F t (x, y) respectively represent the current input video frame data, the current calculated background frame data and the foreground frame data after differential calculation. According to the set binarization threshold, the foreground frame stream is binarized to obtain a binarized target frame stream.
[0021] like Figure 2 As shown in (a), the erosion operation means that at the coordinate (x, y), the value of the erosion operation is to compare each point (s, t) of the structural element B with the pixel value of the corresponding position in the image I and take the minimum value, specifically: The expansion operation means that at the coordinate (x, y), the value of the expansion operation is to compare each point (s, t) of the structural element B with the pixel value of the corresponding position in the image I and take the maximum value, specifically:
[0022] like Figure 3As shown in (c), the target segmentation refers to: after setting the distance threshold according to the image size and the specific environment, the distance between each adjacent target pixel is calculated and filtered using the distance threshold, and the entire image is traversed, and finally multiple non-overlapping target areas can be obtained, that is, the function of multi-target segmentation of the binary image is realized, specifically: [(x0, y0), (x1, y1)] = [(min(x2, x4), min(y2, y4)), (max(x3, x5), max(y3, y5))], where: (x0, y0), (x1, y1) are the upper left corner and lower right corner coordinates of the merged target frame, (x2, y2), (x3, y3) and (x4, y4), (x5, y5) are the upper left corner and lower right corner coordinates of the two target frames to be merged, and min and max are the minimum and maximum calculation operations, respectively. After the processing of the above modules, the rectangular frame positions of all target areas of the frame image can be obtained in the target frame diagram.
[0023] like Figure 3 As shown in the figure, the process effect of the system processing super-resolution images, where (a) represents the original super-resolution image with a pixel size of 3840*2160, (b) represents the image after preliminary scaling with a pixel size of 2048*2048, (c) is the binary image after background modeling and difference, and the red color in the figure is the target box, (d) is the binary image after the corrosion and expansion modules, (e) is the image after the target segmentation module, and compared with (d), the position coordinates of the rectangular box of the target area are additionally obtained, and (f) is a schematic diagram of the configurable motion area fusion strategy. By fusing and cropping the above target area rectangular boxes, a series of fixed-width target areas are obtained to meet the image size requirements of the target detection module.
[0024] The neural network detection and reasoning module in this embodiment is implemented by deploying a neural network accelerator of the YOLOv3 model on an FPGA.
[0025] After specific practical experiments, a software system with the same algorithm flow was built in the Python environment and three network models, YOLOv3, YOLOv4 and YOLOv5, were deployed to calculate and compare the actual performance of the system. Deployed in a dedicated processor in the field of road traffic monitoring, the overall system module construction and functional verification have been completed in the FPGA environment, which can realize real-time super-resolution video target detection from the camera sensor to the FPGA end. The FPGA board chip model is XilinxXCZU15EG. The specific resource usage is shown in the following table. quantity Chip resources Occupancy (%) Look-up Tables (LUTs) 213187 341280 62.5 Flip Flops (FFs) 194456 682560 28.5 Block Random-Access Memory (BRAMs) 530 744 71.2 Digital Signal Processor (DSPs) 1442 3528 40.9
[0026] Compared with the scaling one-stage detection scheme, since this system does not directly scale the original super-resolution video and send it to the neural network for detection, but first undergoes target segmentation processing to obtain the target area, and then merges all the target areas and sends them to the neural network for detection, compared with this technology, the present invention improves the target average detection accuracy (mAP) by about 15%; at the same time, compared with the cropping one-stage detection scheme, that is, the original super-resolution picture or video is cropped into multiple small pictures and detected separately, and finally the detection results are merged and restored to the original super-resolution video, the present invention does not need a neural network to infer all the image data of the original super-resolution, but the motion area therein, so the computational complexity of the neural network can be reduced by about 65% compared with this type of existing technology. The specific data are shown in the following table.
[0027] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principle and purpose of the present invention. The protection scope of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. Each implementation scheme within its scope shall be subject to the constraints of the present invention.
Claims
1. A real-time object detection system for super-resolution video, characterized in that: include: Background modeling module, image morphological operation module, target segmentation module and neural network detection and reasoning module, wherein: the background modeling module performs background modeling and updates the input video frame, and obtains a foreground frame by performing a difference between the input video frame and the background frame; the image morphological operation module performs corrosion and expansion operations on the foreground frame; the target segmentation module extracts the target area from the foreground frame after image morphological processing according to a target segmentation algorithm based on a distance threshold; The neural network detection and reasoning module detects the extracted image of the target area and restores the result to super-resolution, then outputs the image detection result of each frame of the super-resolution video.
2. The real-time object detection system for super-resolution video according to claim 1, characterized in that: The background modeling module comprises: a weighting unit, a storage unit and a difference unit, wherein: the weighting unit performs weighted processing according to the scaled image pixel data and the background pixel data stored in the memory through the multiplier and adder logic resources on the FPGA to obtain the updated background pixel data; the storage unit saves the updated background pixel data in the memory according to the weighted calculation, and reads the background pixel data of the previous frame from the memory according to the effective signal of the image pixel data transmitted by the previous module; the difference unit performs a difference operation through the adder according to the image pixel data signal transmitted by the previous module and the background pixel data of the previous frame read from the memory to obtain the differenced foreground pixel data of each pixel, and performs a binarization operation on the differenced result through the comparator according to the set binarization threshold to obtain the binarized pixel data.
3. The real-time object detection system for super-resolution video according to claim 1, characterized in that: The image morphological operation module includes: an erosion unit and a dilation unit, wherein: the erosion unit uses the binary pixel data calculated by the background modeling module to cache the pixel data rows of the image by the hardware logic resource FIFO on the FPGA and uses the register to cache the pixels, and finally obtains the erosion / dilation value of the pixel point through a logical AND / OR operation.
4. The real-time object detection system for super-resolution video according to claim 1, characterized in that: The target segmentation module includes: a pixel counting unit and a target frame merging unit, wherein: the pixel counting unit calculates the row and column position information of each pixel on the image according to the binary image pixel signal calculated by the image morphological operation module and the width and height information of the image, and stores the column address and row address in a register group, and generates a target frame corresponding to each pixel in real time according to the position information, and stores it in the register group; the target frame merging unit determines whether the target frame generated by each pixel needs to be merged with the existing target frame group and updates and stores it in the register group.
5. The real-time object detection system for super-resolution video according to claim 4 is characterized in that: The judgment refers to whether the coordinate position of the pixel, ie, the column address and the row address, falls within any target frame stored in the target frame diagram and is merged with the target frame that meets the conditions.
Citation Information
Cited By
Lightweight method for reducing calculated amount in background schlieren imaging of high-speed aircraft
CN120765679A
A lightweight method for reducing the amount of calculation in high-speed aircraft background schlieren imaging
CN120765679B