A target detection method based on deep learning

By adopting deep learning methods in the object detection algorithm, optimizing the training set and anchor frame, enhancing image processing, and using combined object sets for comparison learning, the difficulty of taking into account the speed and accuracy of the object detection algorithm in the prior art is solved, and a more efficient and accurate object detection effect is achieved.

CN114926704BActive Publication Date: 2025-05-23NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210449818.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-26
Publication Date
2025-05-23
Estimated Expiration
2042-04-26

AI Technical Summary

Technical Problem

Existing object detection algorithms are difficult to balance between speed and accuracy. Single-stage algorithms are fast but have low accuracy, while two-stage algorithms are time-consuming and prone to overfitting problems.

Method used

Using deep learning-based object detection method, through steps such as creating training sets, target embedding and reconstructing training sets, model training and loss function calculation, the anchor box and enhance image are optimized, and the combined target set is used for comparison learning to improve the model's positioning ability and boundary clarity.

Benefits of technology

It improves the accuracy and speed of target detection, enhances the accuracy of positioning and obvious boundary distinction of the model for target candidate box areas, and reduces the risk of overfitting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114926704B_ABST
    Figure CN114926704B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer technology, and in particular to a target detection method based on deep learning. A target embedding method is used to refer to a target of a detected original image candidate frame as an original image, and the target is combined with a reconstructed image to form a combined target set. A failed image with an iou value lower than 0.2 in a training set is used as an extended image, and a part of the extended image in the system is replaced with an image of the combined target set to form a new image, thereby obtaining a larger data set, which becomes very effective when the original data set is small. Since a neural network is more sensitive to these successfully detected images, areas outside the target are replaced multiple times, so that the model can more accurately locate the area of ​​the target candidate frame when performing target detection, and the boundary distinction of the candidate frame is clearer, thereby enhancing the positioning capability. The present invention only uses anchor frames with an iou value greater than 0.5, and performs non-maximum suppression, so that the spatial positioning capability is stronger.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a target detection method based on deep learning. Background Art

[0002] With the rapid development of computer technology, target detection in computer vision is being used in more and more places. Target detection algorithm refers to an algorithm that can obtain a rectangular box to detect the target and predict the object in the image by inputting one or more images and performing operations such as convolutional layers and pooling layers. With the widespread application of deep learning, there are more and more target detection algorithms, which can be roughly divided into two types: single-stage target detection and two-stage target detection. Single-stage target detection algorithms have low accuracy but high speed, such as Yolo and SSD algorithms, which increase the running speed of the neural network algorithm by reducing the number of layers and candidate regions of the convolutional neural network. Two-stage target detection algorithms are mostly optimized based on R-CNN. First, the candidate box in the image is determined by a certain algorithm, and then the candidate box is classified and regressed through spatial pyramids, anchor boxes, support vector machines, etc. to make predictions. The accuracy of the algorithm prediction is increased by increasing the scale of the algorithm and building a deep neural network for deep learning.

[0003] Whether it is a single-stage or two-stage algorithm, there are always more or less defects. Although the single-stage algorithm is faster, its accuracy is lower, and the two-stage algorithm takes too long and loses the timeliness of the data. The accuracy of the target detection algorithm is difficult to improve, and when the training model is too long, it often overfits with the training set. Summary of the invention

[0004] The purpose of the present invention is to provide a target detection method based on deep learning to solve the problems raised in the above background technology.

[0005] The technical solution of the present invention is: a target detection method based on deep learning, comprising the following steps:

[0006] S1. Create a training set and initialize training: including model initialization, initial training and anchor box optimization;

[0007] S2, target embedding, reconstruction training set: including image enhancement and target embedding reconstruction;

[0008] S3, training the model, calculating the loss function, and updating the model parameters: including retraining the model, calculating the loss function, and performing deep learning;

[0009] S4. Repeat S3.

[0010] Preferably, in S1, creating a training set and initializing training includes the following steps:

[0011] S11. Model initialization:

[0012] Use the moco-v2 model to randomly initialize and input the initial image. The dataset can be PascalVOC, COCO, etc. The learning rate is set to 0.05, and the iteration is 10,000 times. The anchor box is initially set to a rectangular box with 25 positions, aspect ratios, and scales.

[0013] The loss function for the initial training is:

[0014]

[0015] Where q is a query representation, k+ is a positive sample of the key sample, τ is a temperature hyperparameter, and N is the number of samples;

[0016] S12. Optimize anchors and save successfully detected images:

[0017] According to the ground-truth value of the input image in the dataset, the anchor boxes with Iou values ​​greater than 0.5 are retained and the rest are discarded; Iou refers to the result of dividing the overlapping part of two regions by the combined part of the two regions.

[0018]

[0019] Overlap represents the overlapping area, and Union represents the union area of ​​two areas;

[0020] S13, extraction target:

[0021] All the targets of the candidate boxes in the ground-truth of these successfully detected images are cropped out.

[0022] Preferably, in S2, training the model, calculating the loss function, and updating the parameters of the model include the following steps:

[0023] S21. Image enhancement:

[0024] S211, flipping the cropped object;

[0025] S212, performing random color dithering to randomly increase a value in each color channel;

[0026] S213, divided into 4 or 9 equal parts, each part is called a part, each

[0027] The parts are randomly rotated 10° to 30° and the positions are disrupted;

[0028] S214, using an encoder to extract the features of each part, integrate them into one, and then synthesize them into a new image, which is called a reconstructed image;

[0029] S22, target embedding and reorganization: The target of the detected candidate box of the original image is called the original image, and is combined with the reconstructed image to form a combined target set. The original image is called l and the reconstructed image is called p.

[0030] Preferably, in S3, contrast learning is performed on the images of the combined target set, wherein the images include original images and original images, original images and reconstructed images, and reconstructed images and reconstructed images.

[0031] Preferably, the contrast loss function of the target embedding includes the following four:

[0032] Original image and original image:

[0033] Original image and reconstructed image:

[0034]

[0035] Reconstruct image and reconstruct image:

[0036] The present invention provides a target detection method based on deep learning through improvement, which has the following improvements and advantages compared with the prior art:

[0037] The present invention uses a target embedding method, and the target of the detected original image candidate box is called the original image, and is combined with the reconstructed image to form a combined target set; the failed image in the training set with an iou lower than 0.2 is used as an extended image, and the part of the extended image in the system is replaced with the image of the combined target set to be combined into a new image to obtain a larger data set, which becomes very effective when the original data set is small; because the neural network is more sensitive to these images that have been successfully detected, the area outside the target is replaced multiple times, so that the model can more accurately locate the area of ​​the target candidate box when performing target detection, and the boundary distinction of the candidate box is clearer, thereby enhancing the positioning ability; the present invention only uses anchor boxes with an iou value greater than 0.5, and performs non-maximum suppression, so that the spatial positioning ability is stronger. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The present invention will be further explained below in conjunction with the accompanying drawings and embodiments:

[0039] Figure 1 It is a flow chart for implementing the method of the present invention;

[0040] Figure 2 It is a model structure diagram of the present invention. DETAILED DESCRIPTION

[0041] The present invention is described in detail below, and the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0042] The present invention provides a target detection method based on deep learning through improvement. The technical solution of the present invention is:

[0043] like Figure 1 As shown, a target detection method based on deep learning includes the following steps:

[0044] S1. Create a training set and initialize training: including model initialization, initial training and anchor box optimization;

[0045] Specifically, the following steps are included:

[0046] S11. Model initialization:

[0047] Use as Figure 2 The moco-v2 model shown is first randomly initialized, and the initial image is input. The data set can be PascalVOC, COCO, etc. The learning rate is set to 0.05, and the iteration is 10,000 times. The anchor box is initially set to a rectangular box with 25 positions, aspect ratios, and scales; Figure 2 The moco-v2 model shown first generates two different sets of images from the original data in the dataset, then processes the images through the encoder and extracts the features in the images, and then merges the two sets of data to generate the loss function loss;

[0048] The loss function for the initial training is:

[0049]

[0050] Where q is a query representation, k+ is a positive sample of the key sample, τ is a temperature hyperparameter, and N is the number of samples;

[0051] S12. Optimize anchors and save successfully detected images:

[0052] According to the ground-truth value of the input image in the dataset, the anchor boxes with Iou values ​​greater than 0.5 are retained and the rest are discarded; Iou refers to the result of dividing the overlapping part of two regions by the combined part of the two regions.

[0053]

[0054] Overlap represents the overlapping area, and Union represents the union area of ​​two areas;

[0055] S13, extraction target:

[0056] All the objects in the candidate boxes in the ground-truth of these successfully detected images are cropped out and called indicators. These images are detected, indicating that the neural network is more sensitive to these images;

[0057] S2, target embedding, reconstruction training set: including image enhancement and target embedding reconstruction;

[0058] Specifically, the following steps are included:

[0059] S21, Image Enhancement: Since the model can only detect the target, the image that has been successfully detected by the model is enhanced to strengthen its detection ability. The specific steps are as follows:

[0060] S211, flipping the cropped object;

[0061] S212, performing random color dithering to randomly increase a value in each color channel;

[0062] S213, divided into 4 or 9 equal parts, each part is called a part, and each part is randomly rotated 10° to 30° to shuffle;

[0063] S214, using an encoder to extract the features of each part, integrate them into one, and then synthesize them into a new image, which is called a reconstructed image;

[0064] S22, target embedding and reorganization: The target detected in the original image candidate box is called the original image, and it is combined with the reconstructed image to form a combined target set. The failed images in the training set with iou lower than 0.2 are used as extended images. The image of the combined target set is used to replace the same position part in the extended image in the system to form a new image. The images of a combined target set will be combined with two or more different images to obtain a new training set that is more than four times the original one. The original image is called l and the reconstructed image is called p.

[0065] S3, training the model, calculating the loss function, and updating the model parameters: including retraining the model, calculating the loss function, and performing deep learning;

[0066] Specifically, contrast learning is performed on the images of the combined target set, where the images include original images and original images, original images and reconstructed images, and reconstructed images and reconstructed images.

[0067] Among them, the contrast loss functions of target embedding include the following four:

[0068] Original image and original image:

[0069] Original image and reconstructed image:

[0070]

[0071] Reconstruct image and reconstruct image:

[0072] Among them, the symbol It represents the extraction of features from an image, after operations such as convolution and pooling;

[0073] S4. Repeat S3.

[0074] The present invention uses a target embedding method, and the target of the detected original image candidate box is called the original image, and is combined with the reconstructed image to form a combined target set; the failed image in the training set with an iou lower than 0.2 is used as an extended image, and the part of the extended image in the system is replaced with the image of the combined target set to be combined into a new image to obtain a larger data set, which becomes very effective when the original data set is small; because the neural network is more sensitive to these images that have been successfully detected, the area outside the target is replaced multiple times, so that the model can more accurately locate the area of ​​the target candidate box when performing target detection, and the boundary distinction of the candidate box is clearer, thereby enhancing the positioning ability; the present invention only uses anchor boxes with an iou value greater than 0.5, and performs non-maximum suppression, so that the spatial positioning ability is stronger.

[0075] The above description enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A target detection method based on deep learning, Features: The following steps are involved: S1. Create a training set and initialize training: including model initialization, initial training and anchor box optimization, including: S11. Model initialization: Use the moco-v2 model to randomly initialize and input the initial image. The dataset can be Pascal VOC or COCO. The learning rate is set to 0.05, and the iteration is 10,000 times. The anchor box is initially set to a rectangular box with 25 positions, aspect ratios, and scales. The loss function for the initial training is: Where q is a query representation, k+ is a positive sample of the key sample, τ is a temperature hyperparameter, and N is the number of samples; S12. Optimize anchors and save successfully detected images: According to the ground-truth value of the input image in the dataset, the anchor boxes with Iou values ​​greater than 0.5 are retained and the rest are discarded; S13, extraction target: All the targets of the candidate boxes in the ground-truth of these successfully detected images are cropped out; S2, target embedding, reconstructing training set: including image enhancement and target embedding reconstructing, training model, calculating loss function, updating model parameters including the following steps: S21. Image enhancement: S211, flipping the cropped object; S212, performing random color dithering to randomly increase a value in each color channel; S213, divided into 4 or 9 equal parts, each part is called a part, and each part is randomly rotated 10° to 30° to shuffle; S214, using an encoder to extract the features of each part, integrate them into one, and then synthesize them into a new image, which is called a reconstructed image; S22, target embedding and reorganization: The target of the detected original image candidate box is called the original image, and is combined with the reconstructed image to form a combined target set. The original image is called l and the reconstructed image is called p; S3, training the model, calculating the loss function, and updating the model parameters: including retraining the model, calculating the loss function, and performing deep learning; S4. Repeat S3.

2. According to the target detection method based on deep learning in claim 1, Features: In S12, Iou refers to the result obtained by dividing the overlapping part of the two regions by the combined part of the two regions. Overlap represents the overlapping area, and Union represents the union area of ​​two areas.

3. According to the target detection method based on deep learning in claim 1, Features: In S3, contrast learning is performed on the images of the combined target set, where the images include original images and original images, original images and reconstructed images, and reconstructed images and reconstructed images.

4. According to claim 3, a target detection method based on deep learning, Features: The contrast loss functions of the target embedding include the following four: Original image and original image: Original image and reconstructed image: Reconstruct image and reconstruct image: ; Among them, the symbol represents taking features of the image.

Citation Information

Patent Citations

  • Semi-supervised target detection method based on consistency constraint

    CN112926673A

  • Target detection method and device based on multi-scale image

    CN113221925A