Target detection precision improvement method and system based on joint training
Through the joint training of super-resolution and object detection model super-resolution training at the image level, the problems of poor super-segment images and strong model coupling in the existing technology are solved, and the target detection accuracy is significantly improved and the flexible application of the model is achieved.
Patent Information
- Application Number
- CN202311785262.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, the joint training method of image super-resolution and object detection model has the problem that unsupervised training leads to poor consistency between super-score images and high-score images, and the super-resolution model and object detection model are strongly coupled, and flexibility is limited.
Supersolution is carried out at the image level by supersolution, and through joint training of supersolution model and object detection model, the object detection model supersolution optimization is allowed to supersolution model, so that the image generated by supersolution model is consistent with the characteristics of real high-score images.
Without retraining the object detection model, the super-segment images generated by the super-resolution model significantly improve the object detection accuracy, and the super-resolution model is relatively independent of the object detection model, suitable for a variety of different model structures and has strong flexibility.
Smart Images

Figure CN120198635A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection in image processing, and particularly relates to a joint training method and system for super-resolution and object detection models aiming at improving object detection accuracy. Background Art
[0002] Image super-resolution is currently an important research direction in the field of image processing. Compared with traditional methods such as bicubic interpolation, deep learning super-resolution models have more significant image detail restoration capabilities and can provide better visual perception capabilities. Therefore, they are also used in various downstream tasks, such as object detection and semantic segmentation, to improve object detection accuracy.
[0003] Generally, image super-resolution and object detection are two independent image processing tasks. The goal of image super-resolution is to reconstruct image details, and object detection realizes the recognition of objects based on object features. Although image super-resolution can provide more object detail information for object detection and is beneficial to improving object detection performance, there are few cases of jointly learning these two tasks to improve object detection accuracy.
[0004] Some traditional joint training methods for super-resolution and object detection focus on improving the reconstruction effect of super-resolution images. Most of these methods train super-resolution models in an unsupervised or self-supervised manner. To obtain sufficient additional supervision information and improve the super-resolution image effect, usually, an object detection or semantic segmentation model is used to identify the super-resolution image, and the loss is calculated based on the recognition result and annotation information. The loss function is used to guide the super-resolution model to optimize its generated result. Although this type of method can generate images with good quality without supervision, the generated super-resolution images often do not perform well in object detection because the unsupervised training method has deficiencies in maintaining consistency with high-resolution images.
[0005] Another part of the joint training methods for super-resolution and object detection focus on improving the object feature extraction ability of the object detection model. These methods use the super-resolution model in the feature processing part of the object detection model, perform super-resolution on object features using the super-resolution model, and enable the object detection model to extract object features at different scales to address the problem of insufficient object information in small object detection, thereby improving the small object detection accuracy. This type of method can improve the performance of object detection. However, since this method performs super-resolution at the object feature level, the super-resolution model and the object detection model are tightly coupled, and heterogeneous model structures (such as convolutional neural network models and Transformer models) need to significantly change the input and network structure to be used, so the flexibility is limited.
[0006] In summary, the existing technologies are divided into two categories, among which: the object detection model adopted by the unsupervised super-resolution training method does not need to be retrained. However, this type of method is aimed at the problem of lacking high-low resolution image pair data, while the present invention aims to solve the problem with image pair data. For the other type of method with high-low resolution image pairs, in the prior art, after super-resolution is achieved at the feature level, the features are used as the input of the object detection model (while the input of the object detection model should originally be a natural image). In order to adjust the parameters of the object detection model according to the input features so that the correct object detection results can be output finally, the object detection model needs to be retrained. Since the method of the present invention does not perform super-resolution at the feature level but only at the image level, it ensures that the input of the object detection model remains a natural image. Therefore, it does not need to be retrained, and the generation result of the super-resolution model can be optimized by means of the loss function provided by the object detection model, thereby improving the accuracy of object detection. Summary of the Invention
[0007] In view of the above problems, aiming at the requirement of improving the accuracy of object detection, the present invention adopts a supervised training method to perform super-resolution at the image level. Through the joint training of the super-resolution model and the object detection model, the object detection model is used to supervise and optimize the super-resolution model, so that the characteristics of the image generated by the super-resolution model are consistent with those of the real high-resolution image. Without the need to retrain the object detection model, the accuracy of the object detection model can be greatly improved only by using the super-resolved image generated by the super-resolution model. Moreover, the super-resolution model and the object detection model of this method are relatively independent and can be disassembled for use, which is applicable to a variety of different model structures and has strong flexibility.
[0008] Specifically, the present invention proposes a method for improving the accuracy of object detection based on joint training, which includes:
[0009] Step 1: Obtain the high-resolution image with object annotations, perform degradation processing on the high-resolution image to obtain a low-resolution image, and use it as a high-low resolution image pair;
[0010] Step 2: Use the low-resolution image in the high-low resolution image pair as training data and the high-resolution image as the training target to train the super-resolution model to obtain the super-resolution model parameters;
[0011] Step 3: Use the high-resolution image as training data and the object annotation of the high-resolution image as the training target to train the object detection model to obtain the object detection model parameters;
[0012] Step 4: By cascading the target detection model behind the super-resolution network model, a joint training model is constructed. Fix the parameters of the target detection model in the joint training model to the parameters of the target detection model, and initialize the parameters of the super-resolution model in the joint training model to the parameters of the super-resolution model.
[0013] Step 5: Input the low-resolution image into the super-resolution network model in the joint training model to obtain a super-resolved image. The target detection model in the joint training model obtains a target detection result based on the super-resolved image. Construct a super-resolution loss according to the super-resolved image and the high-resolution image corresponding to the low-resolution image, and construct a target detection loss according to the target detection result and the target annotation of the high-resolution image corresponding to the low-resolution image. After weighted addition of the super-resolution loss and the target detection loss, perform backpropagation to update and train the parameters of the super-resolution model in the joint training model.
[0014] Step 6: Input the image to be target-recognized into the trained joint training model to obtain a target recognition result including the target position and category.
[0015] The method for improving target detection accuracy based on joint training, wherein the degradation processing includes downsampling and Gaussian blur.
[0016] The method for improving target detection accuracy based on joint training, wherein the super-resolution model in the joint training model independently performs the super-resolution task of the image to obtain a super-resolution image.
[0017] The method for improving target detection accuracy based on joint training, wherein in step 1, after degradation, the scale of the low-resolution image is aligned with the high-resolution image by an interpolation method.
[0018] The present invention also proposes a system for improving target detection accuracy based on joint training, which includes:
[0019] Module 1: Obtain a high-resolution image with target annotation, perform degradation processing on the high-resolution image to obtain a low-resolution image, and use it as a high-low resolution image pair.
[0020] Module 2: Use the low-resolution image in the high-low resolution image pair as training data and the high-resolution image as the training target to train the super-resolution model to obtain the parameters of the super-resolution model.
[0021] Module 3: Use the high-resolution image as training data and the target annotation of the high-resolution image as the training target to train the target detection model to obtain the parameters of the target detection model.
[0022] Module 4. After cascading the target detection model behind the super-resolution network model, a joint training model is constructed. The parameters of the target detection model in the joint training model are fixed to the parameters of the target detection model, and the parameters of the super-resolution model in the joint training model are initialized to the parameters of the super-resolution model;
[0023] Module 5. Input the low-resolution image into the super-resolution network model in the joint training model to obtain a super-resolved image. The target detection model in the joint training model obtains a target detection result based on the super-resolved image; a super-resolution loss is constructed according to the super-resolved image and the high-resolution image corresponding to the low-resolution image, and a target detection loss is constructed according to the target detection result and the target annotation of the high-resolution image corresponding to the low-resolution image; after weighted addition of the super-resolution loss and the target detection loss, backpropagation is performed to update and train the parameters of the super-resolution model in the joint training model;
[0024] Module 6. Input the image to be target-recognized into the trained joint training model to obtain a target recognition result including the target position and category.
[0025] In the target detection accuracy improvement system based on joint training, the degradation processing includes downsampling and Gaussian blur.
[0026] In the target detection accuracy improvement system based on joint training, the super-resolution model in the joint training model independently performs the super-resolution task of the image to obtain a super-resolution image.
[0027] In the target detection accuracy improvement system based on joint training, after degradation, module 1 aligns the scale of the low-resolution image with the high-resolution image by an interpolation method.
[0028] The present invention also proposes a storage medium for storing a program for executing any target detection accuracy improvement method based on joint training.
[0029] The present invention also proposes a client for any target detection accuracy improvement system based on joint training.
[0030] As can be seen from the above solutions, the advantages of the present invention are as follows:
[0031] 1. The present invention adopts a supervised training method, which can ensure the consistency between the super-resolved generated image and the real high-resolution image. At the same time, the generated super-resolution image can significantly improve the small target detection accuracy compared with the low-resolution image;
[0032] 2. The super-resolution model and the object detection model of the present invention are independent of each other, and do not require any modification to the super-resolution model and the object detection model. They can be directly cascaded for training and can be disassembled for use, and can be flexibly applied to various super-resolution models and object detection models. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flowchart of the present invention;
[0034] Figure 2 is a schematic diagram of the joint training of the models of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The object of the present invention is to overcome the problems of poor consistency between the super-resolution image and the high-resolution image caused by unsupervised learning in the above-mentioned prior art, and the limited flexibility of the algorithm caused by the strong coupling between the super-resolution model and the object detection model. A joint training method for super-resolution and object detection models for improving object detection accuracy is proposed. In order to achieve the above technical effects, the present invention proposes the following key technical points:
[0036] Key point 1: Collect a high-resolution object detection data set containing object annotations, and divide the data set into a training set and a test set; construct a training set and a test set for high-low resolution image pairs to support the pre-training of the object detection model.
[0037] Key point 2: Generate corresponding low-resolution images for the images in the high-resolution object detection data set through a degradation method, and construct a training set and a test set containing high-low resolution image pairs; support the pre-training of the super-resolution model, the joint training and test verification of the super-resolution model and the object detection model.
[0038] Key point 3: Train the super-resolution model on the high-low resolution image pairs in the training set to obtain a set of better super-resolution model parameters; provide the initialization parameters of the super-resolution model for the joint training of the super-resolution model and the object detection model.
[0039] Key point 4: Train the object detection model on the high-resolution images in the training set, and fix a set of better parameters for subsequent supervised training of the super-resolution model; provide the object detection model parameters for the joint training of the super-resolution model and the object detection model.
[0040] Key point 5: Cascade the object detection model behind the super-resolution network model, fix the parameters of the object detection model, load the super-resolution model parameters obtained in Key point 3, train the super-resolution model based on the high-low resolution image pairs of the training set, add the training loss of the super-resolution model and the loss of the object detection model after weighting, and perform backpropagation to optimize the parameters of the super-resolution model; Technical effect: Through the joint training of the super-resolution model and the object detection model, let the object detection model supervise and optimize the super-resolution model, so that the characteristics of the image generated by the super-resolution model are consistent with those of the real high-resolution image. Without the need to retrain the object detection model, the accuracy of the object detection model can be greatly improved only by using the super-resolution image generated by the super-resolution model;
[0041] Key point 6: Split the super-resolution model and the object detection model, load the optimized super-resolution model parameters, and then super-resolution images that can significantly improve the object detection accuracy can be generated. The super-resolution model and the object detection model are relatively independent and can be disassembled and used, which is applicable to a variety of different model structures and has strong flexibility.
[0042] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are hereinafter given and detailed descriptions are made in conjunction with the accompanying drawings of the specification as follows.
[0043] The present invention discloses a joint training method for super-resolution and object detection models for improving object detection accuracy. The training process is as shown in the appendix Figure 1 as follows.
[0044] Generally speaking, the present invention proposes a joint training method for super-resolution and object detection models for improving object detection accuracy. First, a high-resolution object detection training set and test set, as well as a training set and test set containing high-low resolution image pairs are constructed. Then, an image super-resolution model is pre-trained based on the high-resolution-low resolution image pair training set, and an object detection model is pre-trained based on the high-resolution object image training set. Then, the parameters of the object detection model are fixed, and the object detection model is cascaded behind the super-resolution model. The object detection model supervises the training of the super-resolution model. Finally, the super-resolution model generates super-resolution images that better conform to the high-resolution target features, so as to obtain higher object detection accuracy on the generated super-resolution images. The specific implementation steps and methods are as follows:
[0045] Step 1: Select a high-resolution image dataset for object detection (such as DOTA, LEVIR, etc.), and divide the dataset into a training set and a test set according to a certain ratio (preferably, 80% of the dataset is used as the training set, and the remaining 20% is used as the test set) for subsequent model training and testing. It should be noted that the dataset should contain the target category and location annotation information in each image;
[0046] Step 2: Construct a low-resolution image dataset corresponding to the high-resolution image dataset through degradation methods such as downsampling and Gaussian blur to form a training set and a test set of high-low resolution image pairs. To ensure the correct execution of object detection, after degradation, the scale of the low-resolution image should be aligned with that of the high-resolution image through interpolation methods such as bilinear and bicubic interpolation, so that the annotation information of the object category and location can also be applied to the low-resolution image;
[0047] Step 3: Based on the high-low resolution image pair training set constructed in Steps 1 and 2, first train the image super-resolution model. During the training process, use the pixel L1 Loss as the loss function to optimize the super-resolution model, and select a set of parameters with the smallest loss value within a certain number of rounds (200 rounds are used in this embodiment) as the super-resolution model parameters to be used subsequently. The calculation formula of the pixel L1 Loss is as follows:
[0048]
[0049] where, is the super-resolution image, I is the high-resolution image (reference image), h is the image height, w is the image width, c is the number of image channels, i represents the horizontal position of the current pixel being calculated, j represents the vertical position of the current pixel being calculated, and k represents the channel of the current pixel being calculated.
[0050] Optionally, the loss function here can be replaced with pixel L2 Loss, texture Loss, etc., and should be selected according to the actually used super-resolution model;
[0051] Step 4: Train the object detection model based on the high-resolution images and object annotation information in the high-low resolution image pair training set constructed in Steps 1 and 2. During the training process, calculate the loss according to the object category and location errors between the prediction results of the object detection model and the annotations, optimize the object detection model, and select a set of parameters with the smallest loss value within a certain number of rounds (50 rounds are used in this embodiment) as the object detection model parameters to be used subsequently. For the Deformable DETR object detection model used in this embodiment, its loss function is defined as follows:
[0052]
[0053] where, σ is the current matching strategy (or the initialized matching strategy), c i is the i-th category, b i is the i-th bounding box in the annotation box, is the probability of a certain category corresponding to the prediction box pointed to by σ(i), is the prediction box corresponding to the element pointed to by σ(i).
[0054]
[0055] Among them, is the GIoU loss, and λ iou , λ L1 are the weights corresponding to the IoU loss and the L1 loss respectively. The essence of DETR optimization is to find the optimal matching strategy between the predicted bounding boxes and the labeled bounding boxes. Therefore, the loss function of DETR is the matching loss
[0056] Step 5: As shown in the appendix Figure 2 , cascade the super-resolution model trained in Step 3 and the object detection model trained in Step 4, fix the parameters of the object detection model, and at the same time optimize the parameters of the super-resolution model using the super-resolution loss and the object detection loss. The specific implementation steps are as follows:
[0057] Step 51: Cascade the super-resolution model and the object detection model according to the structure of the appendix Figure 2 , load the model parameters trained in Steps 3 and 4, and fix the parameters of the object detection model (not participating in the backpropagation of the loss value);
[0058] Step 52: Use the high-low resolution image pair training set constructed in Steps 1 and 2 and the corresponding object category and location annotations to train the cascaded network model. Load the low-resolution image, obtain the super-resolution image through the forward propagation of the super-resolution model, calculate the super-resolution loss with the corresponding high-resolution image (in this embodiment, the pixel L1 loss is used), and input the super-resolution image into the object detection model to regress the object position and category, and calculate the object detection loss in combination with the annotation information (in this embodiment, the matching loss is used);
[0059] Step 53: Add the super-resolution loss and the object detection loss weighted to obtain the total loss value, and backpropagate this loss value. Since the parameters of the object detection model are fixed, the backpropagation will only optimize the parameters of the super-resolution model. Select a set of model parameters with the minimum object detection loss as the final saved super-resolution model parameters.
[0060] Step 6: After the training in Step 5, split the super-resolution model and the object detection model. At this time, load the super-resolution model parameters saved in Step 53 to generate a super-resolution image that can significantly improve the object detection accuracy.
[0061] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.
[0062] The present invention also provides a system for improving the accuracy of object detection based on joint training, which includes:
[0063] Module 1: Obtain a high-resolution image with object annotations, perform degradation processing on the high-resolution image to obtain a low-resolution image, and use it as a high-low resolution image pair;
[0064] Module 2: Use the low-resolution image in the high-low resolution image pair as training data and the high-resolution image as the training target to train a super-resolution model and obtain super-resolution model parameters;
[0065] Module 3: Use the high-resolution image as training data and the object annotation of the high-resolution image as the training target to train an object detection model and obtain object detection model parameters;
[0066] Module 4: After cascading the object detection model behind the super-resolution network model, construct a joint training model, fix the parameters of the object detection model in the joint training model to the object detection model parameters, and initialize the parameters of the super-resolution model in the joint training model to the super-resolution model parameters;
[0067] Module 5: Input the low-resolution image into the super-resolution network model in the joint training model to obtain a super-resolved image. The object detection model in the joint training model obtains an object detection result based on the super-resolved image; construct a super-resolution loss according to the super-resolved image and the high-resolution image corresponding to the low-resolution image, and construct an object detection loss according to the object detection result and the object annotation of the high-resolution image corresponding to the low-resolution image; after weighted addition of the super-resolution loss and the object detection loss, perform backpropagation to update and train the super-resolution model parameters in the joint training model;
[0068] Module 6: Input the image to be object-recognized into the trained joint training model to obtain an object recognition result including the object position and category.
[0069] In the system for improving the accuracy of object detection based on joint training, the degradation processing includes downsampling and Gaussian blur.
[0070] In the system for improving the accuracy of object detection based on joint training, the super-resolution model in the joint training model independently performs the super-resolution task of the image to obtain a super-resolution image.
[0071] In the system for improving the accuracy of object detection based on joint training, after degradation, Module 1 aligns the scale of the low-resolution image with the high-resolution image by an interpolation method.
[0072] The present invention also provides a storage medium for storing a program for executing any of the above-described methods for improving the accuracy of object detection based on joint training.
[0073] The present invention also provides a client for any of the systems for improving the accuracy of object detection based on joint training.
Claims
1. A method for improving the accuracy of object detection based on joint training, characterized in that Including: Step 1: Obtain a high-resolution image with target annotations. By performing degradation processing on this high-resolution image, a low-resolution image is obtained and used as a high-low resolution image pair. Step 2: Use the low-resolution image in the high-low resolution image pair as training data and the high-resolution image as the training target to train a super-resolution model and obtain super-resolution model parameters. Step 3: Use this high-resolution image as training data and the target annotation of this high-resolution image as the training target to train a target detection model and obtain target detection model parameters. Step 4: By cascading the target detection model behind the super-resolution network model, a joint training model is constructed. Fix the parameters of the target detection model in the joint training model to the target detection model parameters, and initialize the parameters of the super-resolution model in the joint training model to the super-resolution model parameters. Step 5: Input the low-resolution image into the super-resolution network model in the joint training model to obtain a super-resolved image. The target detection model in the joint training model obtains a target detection result based on this super-resolved image. Construct a super-resolution loss according to the super-resolved image and the high-resolution image corresponding to the low-resolution image, and construct a target detection loss according to the target detection result and the target annotation of the high-resolution image corresponding to the low-resolution image. After weighted addition of the super-resolution loss and the target detection loss, perform backpropagation to update and train the super-resolution model parameters in the joint training model. Step 6: Input the image to be target-recognized into the trained joint training model to obtain a target recognition result including the target position and category.
2. The method for improving the target detection accuracy based on joint training according to claim 1, wherein The degradation processing includes downsampling and Gaussian blur.
3. The method for improving the object detection accuracy based on joint training according to claim 1, wherein, Use the super-resolution model in the joint training model alone to perform the super-resolution task of the image and obtain a super-resolution image.
4. The method for improving the target detection accuracy based on joint training according to claim 1, wherein, In Step 1, after degradation, the scale of the low-resolution image is aligned with the high-resolution image through an interpolation method.
5. A target detection accuracy improvement system based on joint training, characterized in that, Including: Module 1: Obtain a high-resolution image with target annotations. By performing degradation processing on this high-resolution image, a low-resolution image is obtained and used as a high-low resolution image pair. Module 2: Use the low-resolution image in the high-low resolution image pair as training data and the high-resolution image as the training target to train a super-resolution model and obtain super-resolution model parameters. Module 3: Use this high-resolution image as training data and the target annotation of this high-resolution image as the training target to train a target detection model and obtain target detection model parameters. Module 4: By cascading the target detection model behind the super-resolution network model, a joint training model is constructed. Fix the parameters of the target detection model in the joint training model to the target detection model parameters, and initialize the parameters of the super-resolution model in the joint training model to the super-resolution model parameters. Module 5: Input the low-resolution image into the super-resolution network model in the joint training model to obtain a super-resolved image. The object detection model in the joint training model is based on the super-resolved image to obtain an object detection result; construct a super-resolution loss according to the super-resolved image and the high-resolution image corresponding to the low-resolution image, and construct an object detection loss according to the object detection result and the object annotation of the high-resolution image corresponding to the low-resolution image; perform weighted addition on the super-resolution loss and the object detection loss and then perform backpropagation to update and train the super-resolution model parameters in the joint training model. Module 6: Input the image to be object-recognized into the trained joint training model to obtain an object recognition result including the object position and category.
6. The system for improving the target detection accuracy based on joint training according to claim 5, characterized in that The degradation process includes downsampling and Gaussian blur.
7. The target detection accuracy improvement system based on joint training according to claim 5, characterized in that, Use the super-resolution model in the joint training model to independently perform the super-resolution task of the image to obtain a super-resolved image.
8. The target detection accuracy improvement system based on joint training according to claim 5, wherein After degradation, Module 1 aligns the scale of the low-resolution image with that of the high-resolution image through an interpolation method.
9. A storage medium for storing a program for executing any one of the methods for improving object detection accuracy based on joint training as described in claims 1 to 4.
10. A client for any one of the systems for improving object detection accuracy based on joint training as described in claims 5 to 8.