Object Detection System and Method Based on Super-Resolution Reconstruction
Through the end-to-end object detection method, combined with the multi-frame super-resolution module and the learnable data division module, the problems of high computational complexity and poor small object detection in image super-resolution reconstruction and object detection are solved, and the rapid and accurate detection of small targets in video is achieved, which is suitable for real-time detection in multiple scenarios.
Patent Information
- Application Number
- CN202010052220.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-01-17
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-01-17
AI Technical Summary
When processing multi-frame images, the existing image super-resolution reconstruction and object detection methods have problems such as high computational complexity, large model parameters, slow detection speed and poor detection effect on small objects, especially in videos, the detection accuracy of small objects is insufficient.
The end-to-end object detection method is adopted, and the scale invariance is introduced through the multi-frame super-resolution module, and a learningable data division module is designed to adaptively crop the image after super-resolution reconstruction, maintain the integrity of small objects in the sub-image, and combine the spatiotemporal sub-pixel convolution network and SSD algorithm for detection.
It realizes fast and accurate detection of small and medium-sized target detection in video, improves detection accuracy, and has a flexible model structure, which is suitable for real-time detection tasks in various scenarios.
Smart Images

Figure CN113139896B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing. Specifically, the present invention relates to an object detection system and method based on super-resolution reconstruction. Background Art
[0002] Image super-resolution technology is a signal processing technology that improves the spatial resolution of images or targets based on existing imaging devices. This technology solves the problem of too low imaging resolution of scenes or targets that may exist in some video and image-based applications. Image super-resolution technology includes: single-frame image super-resolution technology, which only uses a single image itself to improve its resolution, such as SRCNN and EDSR, etc.; and multi-frame image super-resolution technology, which uses adjacent multi-frame images to improve the image resolution of a specific frame, such as sub-pixel convolutional neural network and ESPCN, etc. In addition, image super-resolution technology also involves image quality assessment algorithms, which mainly include: image quality assessment algorithms using convolutional neural networks; and image quality assessment algorithms using image gradient features. The present invention mainly relates to image super-resolution technology for super-resolution reconstruction through multiple image frames.
[0003] In the process of video super-resolution processing, three main issues are concerned: 1) how to make full use of the correlation information between multiple frames; 2) how to effectively fuse image details into high-resolution images; and 3) how to improve the calculation speed. In the process of video super-resolution processing, sometimes it is necessary to first map the low-resolution image to a high-definition grid through upsampling, but this operation will increase the computational complexity. In response to the above problems, in the article "Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network." (2016): 1874-1883 published by Shi, Wenzhe et al. in 2016, a real-time image super-resolution algorithm ESPCN (Efficient Sub-Pixel Convolutional Neural Network) was proposed, and upsampling was performed before the network output layer. The previous layers are all ordinary (integer pixel) convolutional layers with activation function layers set, and a new sub-pixel convolutional layer is set in the last layer of the network to rearrange the pixels according to channels instead of convolution operations, that is, H×W×C×r 2The feature map is rearranged into (r×H)×(r×W)×C as the output. To further improve the speed, Caballero, Jose et al. published "Real-Time Video Super-Resolution with Spatio-Temporal Networks and Motion Compensation" in 2016, (2016):2848-2857. In this article, an end-to-end jointly trained motion compensation and video super-resolution algorithm was proposed, introducing a spatio-temporal sub-pixel convolutional network to achieve real-time video image super-resolution. This network mainly uses early fusion and slow fusion to process the time dimension, then establishes a motion compensation framework based on spatial transformation, and combines it with the ESPCN spatio-temporal network to achieve real-time calculation of video super-resolution reconstruction. In the existing technologies for realizing video super-resolution reconstruction, there are the following problems: (1) Using traditional super-resolution reconstruction (interpolation) and object detection methods, the performance has been surpassed by deep learning-based methods. The reconstruction quality of traditional super-resolution reconstruction (interpolation) and object detection methods is low, and the description ability of template-based detection methods is limited, with little semantic information that can be described; (2) Dependence on prior knowledge: The algorithm depends on the accuracy of prior knowledge (target image template). When the actual application scenario does not match the introduced prior knowledge, the accuracy of the algorithm will decrease. To solve the above problems of traditional super-resolution reconstruction (interpolation) and object detection methods, methods for improving the accuracy of object detection by increasing context information have been proposed. DSSD attempts to improve the performance of SSD by adding context. However, this method also has the following problems: (1) Large parameter (computation) amount and slow algorithm speed; (2) The large number of parameters makes the model occupy a large storage space.
[0004] As an important field of image processing using artificial intelligence, object detection technology is essentially the localization of multiple objects. That is, object detection is a combination of classification tasks and localization tasks. Its task is to accurately find the location (coordinates) of objects in a given picture and label the categories of the objects. The main performance indicators of object detection models are detection accuracy and speed. Currently, the mainstream object detection algorithms are mainly based on deep learning models, which can be divided into two categories: (1) Two-stage detection algorithms, which divide the detection problem into two stages. First, candidate regions are generated, and then the candidate regions are classified. The typical representatives of this type of algorithm are the R-CNN series algorithms based on region proposal; (2) One-stage detection algorithms, which do not require the region proposal stage and directly generate the class probability and position coordinate values of the object. Typical algorithms include YOLO and SSD.
[0005] Scale (scaling) invariance is crucial for object recognition and localization. Since the deeper layers of modern CNNs have a large stride (32 pixels), this results in a very rough representation of the input image, making small object detection very challenging. When faced with the small object problem (essentially a scale invariance problem), the detection results of the above-mentioned methods all have deficiencies. The reason for this problem is that there are the following contradictions in the convolutional network structure: the feature maps in the shallow layers of the network are large but the semantic information is insufficient, and the semantic information in the deep layers of the network is sufficient but the feature maps are too small. To detect objects of multiple scales, various solutions have been proposed, such as:
[0006] (1) Using dilated / atrous convolution to increase the resolution of the feature map, which retains the weights and receptive fields of the pre-trained network and is not affected by the performance degradation of large objects;
[0007] (2) Based on the fact that the shallow and deep layers contain complementary information, fusing the shallow features and deep features (context information) for prediction;
[0008] (3) Making predictions directly and independently on the feature maps of the shallow and deep layers of the network respectively;
[0009] (4) Upsampling the network input image during training.
[0010] Therefore, how to perform multi-frame image super-resolution reconstruction to obtain a higher-resolution image, while maintaining the integrity of small targets in the image and achieving fast and accurate target detection is a very worthy research issue. Summary of the Invention
[0011] An embodiment of the present invention provides an end-to-end target detection method. Scale invariance is introduced through a multi-frame super-resolution module, and a learnable data partitioning module is designed to adaptively crop the image after super-resolution reconstruction to maintain the integrity of small targets in the sub-images. Finally, the cropped images are input into the target detection module for detection, improving the detection effect of small-sized targets in the application scenario of detecting targets in a video (multi-frame).
[0012] According to one aspect of the embodiments of the present invention, there is provided a target detection system based on super-resolution reconstruction, characterized in that the system includes: a data acquisition module configured to acquire image data to be detected; a super-resolution reconstruction module configured to receive the image data acquired by the data acquisition module and perform super-resolution reconstruction processing on the image data; a target detection module configured to perform target detection on the image data that has undergone super-resolution reconstruction processing; and a partitioning and fusion module configured to crop the image data that has undergone target detection into multiple sub-image data and map the detection results of each sub-image data to the combined image data for coordinate fusion, thereby obtaining a target detection result.
[0013] In the target detection module, the single-shot multi-box detection (SSD) algorithm is used.
[0014] In the partitioning and fusion module, the step size for cropping the image data that has undergone target detection is a value obtained based on edge detection.
[0015] The super-resolution reconstruction module is a trained spatio-temporal sub-pixel convolutional network, and the spatio-temporal sub-pixel convolutional network includes a motion estimation part and a super-resolution part. Among them, the spatio-temporal sub-pixel convolutional network is trained through the following processing:
[0016] The Loss formula of the super-resolution part is as follows:
[0017]
[0018] The Loss formula of the motion estimation part is as follows:
[0019]
[0020] Where is approximately
[0021]
[0022] ε = 0.01
[0023] The total Loss formula of the spatio-temporal sub-pixel convolution network during end-to-end training is as follows:
[0024]
[0025] where θ Δ are the parameters of the motion estimation part, and θ are the parameters of the super-resolution part. represents an image frame, represents the image frame after warping processing.
[0026] In the target detection module, the target loss function L Det , the target loss function L Det is obtained through the following equation:
[0027]
[0028] where: N is the number of default boxes matching the ground truth boxes, L loc is the smooth L1-norm loss function in Fast R-CNN, L conf is the Softmax Loss, c is the confidence of each class, and α is the weight term and is set to 1.
[0029] According to another aspect of the embodiments of the present invention, there is also provided a target detection method based on super-resolution reconstruction, characterized in that the method includes the following steps: a data acquisition step of acquiring image data to be detected; a super-resolution reconstruction step of receiving the acquired image data and performing super-resolution reconstruction processing on the image data; a target detection step of performing target detection on the image data after super-resolution reconstruction processing; and a division and fusion step of cropping the image data after target detection into multiple sub-image data and mapping the detection results of each sub-image data to the combined image data for coordinate fusion, so as to obtain the target detection result.
[0030] In the target detection step, the single-shot multi-box detection (SSD) algorithm is used.
[0031] In the division and fusion step, the step size for cropping the image data after target detection is a value obtained based on edge detection.
[0032] In the super-resolution reconstruction step, a trained spatio-temporal sub-pixel convolution network is used. The spatio-temporal sub-pixel convolution network includes a motion estimation part and a super-resolution part. Among them, the spatio-temporal sub-pixel convolution network is trained through the following processing:
[0033] The Loss formula of the super-resolution part is as follows:
[0034]
[0035] The Loss formula of the motion estimation part is as follows:
[0036]
[0037] Among them, is approximated as
[0038]
[0039] ε = 0.01
[0040] The total Loss formula of the spatio-temporal sub-pixel convolution network during end-to-end training is as follows:
[0041]
[0042] Among them, θ Δ is the parameter of the motion estimation part, and θ is the parameter of the super-resolution part. represents the image frame, represents the image frame after warping processing.
[0043] In the target detection step, the target loss function L Det is adopted, and the target loss function is obtained through the following equation:
[0044]
[0045] Among them: N is the number of default boxes matching the ground truth box, L loc is the smooth L1-norm loss function in Fast R-CNN, L conf is the Softmax Loss, c is the confidence of each class, and α is the weight term and is set to 1. Description of the Drawings
[0046] The drawings described herein are used to provide a further understanding of the present invention, form a part of the present invention, and the schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention. In the drawings:
[0047] Figure 1 shows a schematic diagram of the target detection system based on super-resolution reconstruction according to an embodiment of the present invention.
[0048] Figure 2 shows a schematic diagram of the super-resolution reconstruction module in the target detection system based on super-resolution reconstruction according to an embodiment of the present invention.
[0049] Figure 3 Shows a schematic diagram of the target detection module in the target detection system based on super-resolution reconstruction according to an embodiment of the present invention.
[0050] Figure 4 Shows a schematic diagram of the division and fusion module in the target detection system based on super-resolution reconstruction according to an embodiment of the present invention.
[0051] Figure 5 Shows a flowchart of the target detection method based on super-resolution reconstruction according to an embodiment of the present invention. Detailed implementation manners
[0052] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0053] It should be noted that the terms "including" and "having" in the description and claims of the present invention and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or modules or units does not necessarily have to be limited to those steps or modules or units clearly listed, but may include other steps or modules or units not clearly listed or inherent to these processes, methods, products or devices.
[0054] For the convenience of describing the technical solution of the present invention below, several basic concepts are first described.
[0055] Deep neural network: A type of neural network, belonging to a branch of machine learning.
[0056] Feature: A representation method of an image. Traditional methods use pixels in the RGB three channels to represent an image. In order to better utilize a computer for recognition, it is necessary to filter out redundant information in the RGB and extract more semantic features. An image is often represented as a vector, and this vector becomes a feature. Image features contain some prominent information in the image, such as contour edges, colors, etc.
[0057] Scale: The size of an image.
[0058] Super-Resolution: Super-Resolution is a method to improve the resolution of the original image through hardware or software. The process of obtaining a high-resolution image from a series of low-resolution images is super-resolution reconstruction.
[0059] Object Detection: Object detection is the combination of classification task and localization task. Its task is to accurately find the location (coordinates) of the objects in the given picture and label the categories of the objects.
[0060] Scar Detection: It is a specific application scenario of the object detection problem. Given an image of a certain material surface, it identifies the damages and their categories (the categories include damaged rivets, scratches, cracks, paint peeling, etc.) in the picture.
[0061] Scale Invariance: It means that after the system undergoes a scale transformation, a certain property of it remains unchanged.
[0062] The object detection system and method based on super-resolution reconstruction proposed in the present invention can be used in the scenario of detecting (small) objects in a video in practical applications. Its core is to introduce scale invariance in the object detection scenario of the video to improve the detection effect of small objects. As long as there are sufficient training samples, the learned algorithm model will have excellent resolution ability and strong robustness. This algorithm can learn the features of various images and can be widely applied to scenarios such as the identification of fine scars on the surface of various materials. For example, the detection of rivet damages on the surface materials of airplanes.
[0063] Figure 1 FIG. shows a schematic diagram of an object detection system based on super-resolution reconstruction according to an embodiment of the present invention. As Figure 1 shown, the object detection system 100 based on super-resolution reconstruction includes: a data acquisition module 102 configured to acquire image data to be detected; a super-resolution reconstruction module 104 configured to receive the image data acquired by the data acquisition module 102 and perform super-resolution reconstruction processing on the image data; an object detection module 106 configured to perform object detection on the image data that has undergone super-resolution reconstruction processing; and a division and fusion module 108 configured to crop the image data that has undergone object detection into multiple sub-image data and map the detection results of each sub-image data to the combined image data for coordinate fusion, so as to obtain an object detection result.
[0064] Figure 2 FIG. shows a schematic diagram of the super-resolution reconstruction module in the object detection system based on super-resolution reconstruction according to an embodiment of the present invention. Figure 2After being trained, the super-resolution reconstruction module 200 shown in Figure 1 can be used as the super-resolution reconstruction module 104 shown in Figure 2 . Through training, the super-resolution reconstruction module can be made into a trained spatio-temporal sub-pixel convolutional network. As shown in Figure 2 , the spatio-temporal sub-pixel convolutional network includes a motion estimation part 202 and a super-resolution part 204. The multi-frame super-resolution module of the network structure is based on real-time video super-resolution with a spatio-temporal network and motion compensation. This network can handle video image super-resolution and achieve real-time speed. An algorithm that combines motion compensation and video super-resolution is also proposed and can be trained end-to-end. Compared with the single-frame model, the spatio-temporal network can reduce calculations and maintain the output quality. As shown in
[0065] , the spatio-temporal sub-pixel convolutional network is trained through the following process: and The difference between and is that and are two different frames, and the position of the object in the image may have changed. By warping, the position of the object in
[0066] can be made almost the same (with slight differences), and then it is sent into the super-resolution part 204.
[0067]
[0068] During training, the loss of the motion estimation part 202 is the MSE loss plus the Huber loss. The Huber Loss is added to make the optical flow smooth in space. The loss formula is as follows:
[0069]
[0070] where ε = 0.01
[0071] The loss formula of the super-resolution part 204 is as follows:
[0072]
[0073] Finally, after passing through the motion estimation part 202 and the super-resolution part 204, when performing end-to-end training, the overall loss is
[0074]
[0075] where θ Δ is the parameter of the motion estimation part 202, and θ is the parameter of the super-resolution part 204.
[0076] Figure 3 It is a schematic diagram showing the target detection module in the target detection system based on super-resolution reconstruction according to an embodiment of the present invention. Figure 3 After being trained, the target detection module 300 shown in Figure 1 can be used as the target detection module 106 shown in
[0077] The target detection module 300 includes: an image input unit 302, a first set of convolutional layers 304, a second set of convolutional layers 306, and a detection output unit 308. Among them, in the second set of convolutional layers 306, a feature pyramid structure is adopted for detection, that is, when detecting, feature maps of different sizes such as conv4-3, conv-7 (FC7), conv6-2, conv7-2, conv8_2, conv9_2 are utilized, and object category classification and location regression are performed simultaneously on multiple feature maps. The initial part of the SSD model, which is called the base network in the text (VGG-16 is used in this article, and lightweight base networks such as MobileNet and ShuffleNet are used to improve the speed of the algorithm), is a commonly used network for image classification. After the base network, additional auxiliary network structures are added, and these auxiliary structures mainly include the following three parts: (1) Multi-scale feature maps for detection: After the base network structure, additional convolutional layers are added, and the sizes of these convolutional layers decrease layer by layer, enabling prediction at multiple scales. (2) Convolutional predictors for detection: Each newly added layer (or the feature layer in the base network structure) can use a series of convolutional kernels to generate a series of predictions of a fixed size. (3) Default boxes and aspect ratios: The position of each box relative to its corresponding feature map cell is fixed. In each feature map cell, it is necessary to predict the offset between the predicted box and the default box, as well as the score of the object contained in each box. The predicted box actually predicts the offset (offsets) relative to the default box.
[0078] The training objective of SSD can handle multiple target categories. Using Indicates that the \(i\)-th default box matches the \(j\)-th ground truth box of class \(p\). If not, then According to this matching strategy, there must be which means that for the \(j\)-th ground truth box, there may be multiple default boxes that match it. The overall objective loss function is obtained by the weighted sum of the localization loss (loc) and the confidence loss (conf):
[0079]
[0080] The meanings of the parameters are explained below, where: \(N\) is the number of default boxes that match the ground truth boxes. The localization loss (loc) is the Smooth L1 Loss in Fast R-CNN, which is used for the parameters of the predicted box (\(l\)) and the ground truth box (\(g\)) (i.e., the center coordinate position, width, and height). The confidence loss (conf) is the Softmax Loss, and the input is the confidence \(c\) of each class. The weight term \(\alpha\) is set to 1.
[0081] After generating a series of predictions, there will be many predicted boxes that match the ground truth boxes, but at the same time, there are also many predicted boxes that do not match the ground truth boxes, and the number of negative samples is much larger than that of positive samples. This makes it difficult to converge during training, and we need to perform Hard negative mining. Therefore, in this paper, first, the boxes corresponding to the predicted boxes (default boxes) at each object position that are negative are sorted according to the confidence of the default boxes. Select the top few to ensure that the ratio of negative samples to positive samples is finally 3:1. (This algorithm is in the function MineHardExamples in bbox_util.cpp) The author of this paper found through experiments that such a ratio can be optimized faster and the training is more stable. Data augmentation is performed on the training data during the training process. To make the model more robust to the scale and size of the target, data augmentation is performed on the training images. Each training image is randomly generated by the following methods: (1) Use the original image. (2) Sample a patch, and the minimum Jaccard overlap with the object is: 0.1, 0.3, 0.5, 0.7, 0.9. (3) Randomly sample a patch.
[0082] Figure 4Shows a schematic diagram of the partitioning and fusion module in the object detection system based on super-resolution reconstruction according to an embodiment of the present invention. Figure 4 The partitioning and fusion module 400 shown in Figure 1 can be used as the partitioning and fusion module 108 shown in Figure 4 The upper part of shows the data partitioning unit 402 in the partitioning and fusion module 400, which is a learnable module for dividing an image into n sub-images using a sliding window. The step size of each step is variable. The purpose of cropping is to retain the resolution after reconstruction and adapt to the fixed input of the object detection module. The purpose of the variable step size is to maintain the integrity of the object to be detected in each segmented sub-image. In the data partitioning unit 402, the step size for cropping the image data after object detection is based on the value obtained from edge detection. Therefore, the step size for cropping is learnable. Specifically, a value is predicted through network learning, and this value is used to specify the cropping size. The goal is to retain the complete detection result (sub-image) in the image, and its learning label is the edge coordinates of the object in each image (predicting this value can retain the relatively complete object), while also retaining the benefits brought by super-resolution (higher resolution).
[0083] Figure 4 The lower part of shows the data fusion unit 404 in the partitioning and fusion module 400. The data fusion unit 404 is responsible for recombining the segmented and detected images into one image, and mapping the detection results of each sub-image to the combined large image for coordinate fusion to obtain the final detection result. In the data fusion unit 404, the process of restoring the bounding box coordinates is shown taking the example of dividing into four sub-images with a fixed step size. Finally, after completing data fusion in the data fusion unit 404, the object detection result is output.
[0084] Figure 5 Shows a flowchart of the object detection method based on super-resolution reconstruction according to an embodiment of the present invention. The object detection method based on super-resolution reconstruction includes the following steps: a data acquisition step S502 of acquiring image data to be detected; a super-resolution reconstruction step S504 of receiving the acquired image data and performing super-resolution reconstruction processing on the image data; an object detection step S506 of performing object detection on the image data after super-resolution reconstruction processing; and a partitioning and fusion step S508 of cropping the image data after object detection into multiple sub-image data and mapping the detection results of each sub-image data to the combined image data for coordinate fusion to obtain the object detection result.
[0085] Through the object detection technology and method based on super-resolution reconstruction provided by the present invention, the following effects can be achieved:
[0086] 1) Accuracy: In the scenario of detecting targets in a video, the super-resolution module is used to improve the resolution of the input image, thereby introducing scale invariance, enhancing the detection results for small targets, and overall improving the accuracy.
[0087] 2) Flexibility: It is an end-to-end network as a whole and is easy to train. The lightweight basic network and hyperparameters in the model sub-module can be replaced according to the needs of users.
[0088] 3) Wide application range: It can be applied to real-time detection tasks with small-sized targets in various scenarios, and has a very wide application range.
[0089] 4) Strong generalization ability: As long as there are enough training samples, the learned algorithm model has excellent accuracy (generalization ability) in practical applications.
Claims
1. A target detection system based on super-resolution reconstruction, characterized in that, The system includes: a data acquisition module configured to acquire image data to be detected; a super-resolution reconstruction module configured to receive the image data acquired by the data acquisition module and perform super-resolution reconstruction processing on the image data; an object detection module configured to perform object detection on the image data that has undergone super-resolution reconstruction processing; and a division and fusion module configured to crop the image data that has undergone object detection into multiple sub-image data and map the detection results of each sub-image data to the combined image data for coordinate fusion, thereby obtaining an object detection result, wherein the super-resolution reconstruction module is a trained spatio-temporal sub-pixel convolutional network, and the spatio-temporal sub-pixel convolutional network includes a motion estimation part and a super-resolution part, and wherein the spatio-temporal sub-pixel convolutional network is trained through the following processing: The Loss formula of the super-resolution part is as follows: The Loss formula of the motion estimation part is as follows: Among them, is approximately ε = 0.01 The total Loss formula of the spatio-temporal sub-pixel convolutional network during end-to-end training is as follows: where θ Δ is a parameter of the motion estimation part, and θ is a parameter of the super-resolution part, represents an image frame, represents the image frame after warping processing.
2. The object detection system based on super-resolution reconstruction according to claim 1, characterized in that, In the object detection module, the single-shot multibox detector (SSD) algorithm is used.
3. The object detection system based on super-resolution reconstruction according to claim 1, wherein In the division and fusion module, the step size for cropping the image data that has undergone object detection is a value obtained based on edge detection.
4. The object detection system based on super-resolution reconstruction according to claim 1, characterized in that, Adopt the target loss function L in the target detection module Det , the target loss function L Det is obtained through the following equation: Where: N is the number of default boxes matching the ground truth box, L loc is the smooth L1-norm loss function in Fast R-CNN, L conf is the Softmax Loss, c is the confidence of each class, α is the weight term and is set to 1, x is the position of the object in the image, l is the predicted box, and g is the ground truth box.
5. A target detection method based on super-resolution reconstruction, characterized in that, The method includes the following steps: a data acquisition step of acquiring image data to be detected; a super-resolution reconstruction step of receiving the acquired image data and performing super-resolution reconstruction processing on the image data; an object detection step of performing object detection on the image data that has undergone super-resolution reconstruction processing; and a division and fusion step of cropping the image data that has undergone object detection into multiple sub-image data and mapping the detection results of each sub-image data to the combined image data for coordinate fusion, thereby obtaining an object detection result, wherein, in the super-resolution reconstruction step, a trained spatio-temporal sub-pixel convolutional network is used, and the spatio-temporal sub-pixel convolutional network includes a motion estimation part and a super-resolution part, and wherein the spatio-temporal sub-pixel convolutional network is trained through the following processing: The Loss formula of the super-resolution part is as follows: The Loss formula of the motion estimation part is as follows: Among them, is approximately ε = 0.01 The total Loss formula of the spatio-temporal sub-pixel convolutional network during end-to-end training is as follows: where θ Δ is a parameter of the motion estimation part, and θ is a parameter of the super-resolution part, represents an image frame, and represents the image frame after warping processing.
6. The object detection method based on super-resolution reconstruction according to claim 5, wherein, In the object detection step, the single-shot multibox detector (SSD) algorithm is used.
7. The object detection method based on super-resolution reconstruction according to claim 5, characterized in that, In the division and fusion step, the step size for cropping the image data that has undergone object detection is a value obtained based on edge detection.
8. The object detection method based on super-resolution reconstruction according to claim 5, characterized in that, In the target detection step, a target loss function L is adopted Det , and the target loss function is obtained by the following equation: Where: N is the number of default boxes matching the ground truth box, L loc is the smooth L1 norm loss function in Fast R-CNN, L conf is the Softmax Loss, c is the confidence of each class, α is the weight term and is set to 1, x is the location of the object in the image, l is the predicted box, and g is the ground truth box.
Citation Information
Patent Citations
Video super-resolution reconstruction method and device
CN106254722A
Multi-scale label-based sub-pixel convolution image super-resolution reconstruction method
CN108734659A