Small target detection method and system

By improving the threshold calculation method of sample screening and positive and negative sample feature learning, the feature extraction problem in small object detection is solved, the accuracy and accuracy of the detection are improved, and the robustness and adaptability of the model are enhanced.

CN119992129AActive Publication Date: 2025-05-13UNIV OF SCI & TECH BEIJING
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510205028.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-13
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

Existing small object detection methods are difficult to capture sufficient effective information during feature extraction, resulting in low detection accuracy and accuracy, especially in complex backgrounds and multi-objective scenarios.

Method used

By improving the threshold calculation method for dynamic sample screening, samples suitable for small object detection are selected, and the quality of small object feature representation is improved through positive and negative sample feature learning, and the accuracy and accuracy of detection are improved.

Benefits of technology

It effectively reduces the interference of background noise, improves the robustness and adaptability of the model in complex environments, and enhances the detection effect of small targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992129A_ABST
    Figure CN119992129A_ABST
Patent Text Reader

Abstract

The invention provides a small target detection method and system, and belongs to the field of computer vision. The method comprises the following steps: firstly, acquiring a historical image and a real labeling sample, wherein the sample category comprises an extremely small target, a relatively small target, a common small target and a normal size target; constructing a small target detection model; inputting the historical image into a convolutional neural network to extract a feature map; inputting the feature map into an improved CRPN, obtaining candidate samples based on different sample categories, mapping the candidate samples on the feature map to obtain a region of interest and features, and obtaining prediction categories and regression positions of the samples based on the features; meanwhile, positive samples and typical negative samples are screened out from the features of the region of interest, a small target detection model is trained based on a prediction category, a regression position, a positive sample teacher set and a typical negative sample teacher set, finally, a current to-be-detected image is input into the trained small target detection model, and a detection result is output. According to the invention, the precision and accuracy of small target detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical field:

[0001] The present invention belongs to the field of computer vision, and in particular relates to a small target detection method and system. Background technology:

[0002] With the rapid development of computer vision technology, object detection is increasingly used in daily life. Small object detection (SOD) is an important branch of object detection tasks, which aims to identify and locate objects that occupy a small pixel area from images or videos. Small object detection tasks are crucial in many practical applications. For example, in unmanned driving, the accurate identification of small objects such as pedestrians and traffic signs is critical to driving safety; in security monitoring, the detection of suspicious behaviors in a small range depends on the precise positioning of small objects in the distance; in remote sensing images, the detection of small buildings, ships and vehicles is also a typical application scenario of small object detection tasks. However, compared with conventional object detection, the size of small objects is usually smaller and the number of pixels they occupy in the image is limited. This characteristic causes the features of small objects to be easily lost during the convolution process, and the detection difficulty increases significantly. In addition, the appearance features of small objects are fuzzy, the information is scarce, and they are greatly interfered by noise. Especially in complex backgrounds and multi-object scenes, traditional object detection methods are difficult to obtain ideal detection results. How to detect small objects efficiently and accurately has become an important research direction in the field of object detection.

[0003] In the prior art, the mainstream methods of small target detection can be divided into one-stage detection methods and two-stage detection methods. These two methods have different focuses on the feature extraction, sample allocation and detection accuracy improvement of small targets. Among them, the single-stage detector in the one-stage target detection method usually adopts an end-to-end architecture, directly performs dense predictions on the input image, and does not need to generate candidate regions. The main advantage is that the detection speed is fast and it is suitable for application scenarios with high real-time requirements. At the same time, due to its simple structure, it is easy to implement and optimize; while the two-stage target detection method adopts a more detailed detection process, which is usually divided into two stages: candidate region generation and candidate region classification. The main advantage is high detection accuracy, especially when dealing with complex scenes and multi-scale targets. It performs well and can more effectively capture the detailed features of the target. However, due to the small size of small targets, there are problems such as insufficient training samples and the inability to unify the screening standards for targets of different sizes, which makes it difficult for the above methods to capture enough effective information in the feature extraction process; in addition, the amount of information carried by small targets is limited, the feature discrimination is low, and they are often easily confused with the background or other objects, which greatly increases the risk of model misjudgment, resulting in low precision and accuracy of small target detection. Summary of the invention:

[0004] In order to solve the above problems, the present invention provides a small target detection method and system, which improves the threshold calculation method of dynamic sample screening, which is more suitable for sample screening of small-sized targets. At the same time, high-quality positive samples and typical negative sample features in training are collected for positive and negative sample feature learning, which serve as teacher features for positive and negative sample feature learning, thereby improving the quality of small target feature representation, thereby improving the accuracy and precision of small target detection, and improving the detection effect of small target tasks.

[0005] In order to achieve the above object, the technical solution adopted by the embodiment of the present invention is as follows:

[0006] In a first aspect, an embodiment of the present invention provides a small target detection method, the method comprising the following steps:

[0007] Step S1, obtaining historical images and real annotated samples in the historical images as training data sets, and the training data sets include categories of real annotated samples divided by target size, including extremely small targets eS, relatively small targets rS, ordinary small targets gS, and normal-sized targets N;

[0008] Step S2, constructing a small target detection model, the model includes a convolutional neural network, an improved thick and thin pipeline network CRPN, a mapping module, a classification head, a detection head and a loss function module;

[0009] Step S3, inputting the historical image into a convolutional neural network to extract a feature map of the image, wherein the feature map contains real annotated samples and category information;

[0010] Step S4, inputting the feature map into the improved CRPN, generating candidate samples through a coarse pipeline, optimizing the candidate samples through a fine pipeline, and calculating the classification score of each candidate sample;

[0011] Step S5, mapping the optimized candidate samples on the feature map to obtain the region of interest, pooling and aligning the original features of the region of interest to obtain the features of the region of interest;

[0012] Step S6, input the features of the region of interest into the classification head and the detection head to obtain the predicted categories and regression positions of all real labeled samples; at the same time, screen out positive samples and typical negative samples from the features of the region of interest to establish a positive sample teacher set and a typical negative sample teacher set;

[0013] Step S7, based on the predicted category, regression position, positive sample teacher set and typical negative sample teacher set, calculate the classification loss, regression loss, positive sample contrast loss of feature imitation and negative sample contrast loss, calculate the comprehensive loss through the four losses, and complete the training of the small target detection model based on the comprehensive loss;

[0014] Step S8, inputting the current image to be detected into the trained small target detection model, and outputting all target detection results, wherein the target detection results include small target detection results.

[0015] As a preferred embodiment of the present invention, step S4 generates candidate samples through a thick pipeline, specifically including:

[0016] Step 41, determine the categories of all real labeled samples in the feature map; for extremely small targets eS, execute step S42; for relatively small targets rS and ordinary small targets gS, execute step S43; for normal size targets N, execute step S44;

[0017] Step S42, calculating the IoU values ​​of a predetermined number of M anchor boxes corresponding to each real labeled sample belonging to eS, sorting the anchor boxes in descending order of IoU values, and selecting the first Q anchor boxes as candidate samples, with Q≤M;

[0018] Step S43, calculating the IoU values ​​of a predetermined number of M anchor boxes corresponding to each real annotated sample belonging to rS and gS; constructing a nonlinear function of the area of ​​each real annotated sample to calculate the IoU threshold; taking the anchor boxes with IoU values ​​greater than the IoU threshold T as candidate samples to implement the screening of samples in the rS and gS categories;

[0019] Step S44, calculate the IoU values ​​of a predetermined number M of anchor boxes corresponding to each real annotated sample belonging to N; preset an IoU threshold, and take the anchor boxes with IoU values ​​greater than the IoU threshold as candidate samples to implement the screening of samples in N categories.

[0020] As a preferred embodiment of the present invention, the nonlinear function of the area of ​​each real labeled sample constructed is as follows:

[0021] T=min(T max , max(αC, βC+δγ×f(wh))) (1)

[0022] In formula (1), T max represents the upper limit of the IoU threshold, C represents the IoU standard value, α represents the lower limit coefficient of the IoU threshold, β represents the starting coefficient of the nonlinear function, w and h represent the width and height of the real labeled sample, f(wh) represents the nonlinear function of the product of the width and height of the real sample, γ represents the slope of the nonlinear function, and δ represents the slope coefficient.

[0023] As a preferred embodiment of the present invention, the convolutional neural network in step S2 adopts a backbone network and a fast regional convolutional neural network Faster RCNN model.

[0024] As a preferred embodiment of the present invention, the improved fine pipeline network described in step S2 is constructed based on the dynamic threshold method, including a coarse pipeline network and a fine pipeline network, which respectively generate and optimize candidate regions for the input image, and improve the CRPN by optimizing the threshold calculation process of the coarse pipeline and the fine pipeline.

[0025] As a preferred embodiment of the present invention, step S6 selects positive samples and typical negative samples from the features of the region of interest, and specifically includes the following steps:

[0026] Get the IoU value and classification score of the candidate samples corresponding to the features of the region of interest;

[0027] Define IQ_H = IoU × classification score, and preset the positive sample threshold T 正 , IQ_H>T 正 The features of are taken as positive samples;

[0028] Define IQ_L = (1-IoU) × classification score, and preset the negative sample threshold T 负 , IQ_L>T 负 The features of are taken as typical negative samples.

[0029] As a preferred embodiment of the present invention, the positive sample contrast loss in step S7 adopts the following loss function:

[0030]

[0031] In formula (2), ρ represents the candidate sample set, represents the current candidate sample, υ p represents the current positive sample, represents the samples in the positive sample set, and τ represents the temperature coefficient for controlling contrastive learning.

[0032] As a preferred embodiment of the present invention, the typical negative sample contrast loss in step S7 adopts the following loss function:

[0033]

[0034] In formula (3), ρ represents the candidate sample set, represents the current candidate sample, v n represents the current negative sample, represents the samples in the negative sample set, and τ represents the temperature coefficient for controlling contrastive learning.

[0035] In a second aspect, an embodiment of the present invention further provides a small target detection system, the system comprising: a data acquisition module, a model construction module, a feature map extraction module, a CRPN processing module, a mapping feature map module, a classification and detection module, a positive and negative sample set construction module, a model training module and a result output module; wherein,

[0036] The data acquisition module is used to acquire historical images and real annotated samples in the historical images as training data sets, and input them into the feature map extraction module; the training data sets include categories of real annotated samples divided by target size, including extremely small targets eS, relatively small targets rS, ordinary small targets gS and normal-sized targets N; and is also used to acquire the current image to be detected, and input it into the model construction module after the model training is completed;

[0037] The model building module is used to build a small target detection model, which includes a convolutional neural network, an improved thick and thin pipeline network CRPN, a mapping module, a classification head, a detection head and a loss function module; and is also used to input the current image to be detected into the trained small target detection model after the model training is completed, and send the detection result to the result output module;

[0038] The feature map extraction module is used to input the historical image into a convolutional neural network to extract a feature map of the image, wherein the feature map contains real labeled samples and category information;

[0039] The CRPN processing module is used to input the feature map into the improved CRPN, generate candidate samples through a thick pipeline, optimize the candidate samples through a thin pipeline, and calculate the classification score of each candidate sample;

[0040] The mapping feature map module is used to map the optimized candidate samples on the feature map to obtain the region of interest, and to pool and align the original features of the region of interest to obtain the features of the region of interest;

[0041] The classification and detection module is used to map the optimized candidate samples on the feature map to obtain the region of interest, and to pool and align the original features of the region of interest to obtain the features of the region of interest;

[0042] The positive and negative sample set construction module is used to screen out positive samples and typical negative samples from the features of the region of interest, and establish a positive sample teacher set and a typical negative sample teacher set;

[0043] The model training module is used to calculate the classification loss, regression loss, positive sample contrast loss of feature imitation and negative sample contrast loss based on the predicted category, regression position, positive sample teacher set and typical negative sample teacher set, and obtain the comprehensive loss through the four loss calculations, and complete the training of the small target detection model based on the comprehensive loss;

[0044] The result output module is used to output all target detection results, including small target detection results.

[0045] The solution of the embodiment of the present invention has the following beneficial effects:

[0046] The small target detection method and system provided by the embodiment of the present invention can more accurately screen out samples suitable for small target detection by improving the dynamic threshold calculation method of sample screening, especially for targets with smaller size and weaker features, and select a fixed number of qualified candidate samples for training, which can effectively reduce the interference of background noise and improve the robustness and adaptability of the model in complex environments; and by introducing high-quality positive samples and typical negative sample features for learning during the training process, it can further effectively reduce the interference of background noise, reduce the probability of false detection, and improve the reliability of the detection results. The present invention reduces the negative impact of redundant samples on training by optimizing sample selection and feature extraction, further improves the discrimination ability of the model in the small target detection process, enhances the sensitivity of the model to small target features, and improves the detection effect of small targets.

[0047] Of course, it is not necessary to achieve all of the advantages described above at the same time to implement any product or method of the present invention. Description of the drawings:

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0049] Figure 1 is a flow chart of a small target detection method according to an embodiment of the present invention;

[0050] Figure 2 This is a comparison diagram of the effects of the small target detection method of the present invention and the existing detection method. Specific implementation method:

[0051] After discovering the above problems, the inventors of this application conducted a detailed study on the existing small target detection methods. The study found that, in response to the problems existing in small target detection, some scholars have proposed corresponding improvement methods, but there are still certain deficiencies. For example, in response to the problem of insufficient training samples, the IoU-based training sample screening method can improve the sample screening effect; however, for small targets, IoU will change greatly due to the slight difference between the sample and the basic true value, making the calculated threshold far away from the common threshold range, resulting in the obtained training samples filtering out most of the positive samples; some scholars have also adopted the method of lowering the threshold to obtain more samples of the same size as small targets, but drastically lowering the threshold is not a good solution. Although the number of training samples for small targets will increase, the training sample set for normal-sized targets will introduce a large number of low-quality samples, and the overall sample quality will decrease. It is also impossible to take into account the detection effects of small targets and large-sized targets, and it is impossible to further improve the problems of the inability to unify the screening standards for targets of different sizes and the serious shortage of small target training samples. In addition,

[0052] Since small targets are small in size and carry limited feature information, the feature representations obtained during feature extraction are often rough and fuzzy. In this case, it is difficult for the model to effectively capture the unique features of small targets, reducing its ability to identify small targets. In addition, existing solutions do not improve the quality of small target feature representation. The features of small targets are often easily confused with the features of the background or other objects, further increasing the risk of misjudgment in classification and positioning. Since small targets have fewer pixels in the image, their edge information and detail features are not obvious, making it difficult for the model to extract enough discriminative features during feature learning. At the same time, the high similarity of features between different small targets also exacerbates the difficulty of model classification.

[0053] It should be noted that the defects existing in the solutions in the above-mentioned prior art are the results obtained by the inventor after practice and careful research. Therefore, the discovery process of the above-mentioned problems and the solutions proposed in the embodiments of the present invention for the above-mentioned problems below should all be the contributions made by the inventor to the present invention in the process of the present invention.

[0054] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations. It should be noted that the embodiments of the present invention and the features in the embodiments can also be combined with each other without conflict.

[0055] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In the description of the present invention, the terms "first", "second", "third", "fourth", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0056] Based on the above in-depth analysis, an embodiment of the present invention provides a small target detection method and system, which adopts a two-stage detection process and combines the optimization of the threshold calculation method to provide different sample screening criteria for real labeled samples of different sizes. At the same time, the feature representation quality of small targets is improved through positive and negative sample feature learning based on feature imitation (FI), thereby achieving more accurate, precise and more effective small target detection.

[0057] like Figure 1 As shown, the small target detection method described in this embodiment includes the following steps:

[0058] Step S1, obtaining historical images and real annotated samples in the historical images as training data sets, and the training data sets include categories of real annotated samples divided by target size, including extremely small targets eS, relatively small targets rS, ordinary small targets gS and normal size targets N.

[0059] In this step, the historical image comes from the data in the standard data network; the classification of the extremely small target eS, relatively small target rS, ordinary small target gS and normal size target N is classified according to the classification method in the standard image.

[0060] Step S2, constructing a small target detection model, the model includes a convolutional neural network, an improved thick and thin pipeline network, a mapping module, a classification head, a detection head and a loss function module.

[0061] In this step, the convolutional neural network uses a backbone network to preprocess the image, and a Faster Region-based Convolutional Neural Network (Faster RCNN) model is used. The improved coarse-to-fine pipeline network (CRPN) is constructed based on the dynamic threshold method, including a coarse pipeline network and a fine pipeline network, which respectively generate and optimize candidate regions for the input image, and optimize the threshold calculation process of the coarse pipeline and the fine pipeline to achieve the improvement of CRPN.

[0062] Step S3: input the historical image into a convolutional neural network to extract a feature map of the image, wherein the feature map contains real labeled samples and category information.

[0063] Step S4, inputting the feature map into the improved CRPN, generating candidate samples through a coarse pipeline, optimizing the candidate samples through a fine pipeline, and calculating the classification score of each candidate sample.

[0064] In this step, candidate samples are generated through a rough pipeline, which includes:

[0065] In step S41, determine the categories of all real labeled samples in the feature map; for extremely small targets eS, execute step S42; for relatively small targets rS and ordinary small targets gS, execute step S43; for normal-sized targets N, execute step S44.

[0066] In this step, based on the classification of real labeled samples, different candidate sample screening methods are adopted according to the different characteristics of each sample, and anchor frames are dynamically generated in stages, so that more samples can be screened for model training while retaining the sample validity to the greatest extent, ensuring the recall rate of various target samples. In the Fine stage, more detailed and complex inferences are performed on the candidate samples, so that the obtained model can detect small targets more accurately and precisely.

[0067] Step S42, calculate the IoU values ​​of a predetermined number M of anchor boxes corresponding to each real labeled sample belonging to eS, sort the anchor boxes in descending order of IoU values, and select the first Q anchor boxes as candidate samples, with Q≤M.

[0068] In this step, extremely small targets are screened as candidate samples in a sorted manner, which is more in line with the actual situation of extremely small samples, and is distinguished from the other two types of small target samples to obtain a larger number of samples.

[0069] Step S43, calculate the IoU values ​​of a predetermined number M of anchor boxes corresponding to each real annotated sample belonging to rS and gS; construct a nonlinear function of the area of ​​each real annotated sample to calculate the IoU threshold; take the anchor boxes with IoU values ​​> IoU threshold T as candidate samples to implement the screening of samples in the rS and gS categories.

[0070] In this step, the nonlinear function of the area of ​​each real labeled sample constructed is as follows:

[0071] T=min(T max , max(αC, βC+δγ×f(wh))) (1)

[0072] In formula (1), T maxrepresents the upper limit of the IoU threshold, C represents the IoU standard value, α represents the lower limit coefficient of the IoU threshold, β represents the starting coefficient of the nonlinear function, w and h represent the width and height of the real labeled sample, f(wh) represents the nonlinear function of the product of the width and height of the real sample, γ represents the slope of the nonlinear function, and δ represents the slope coefficient.

[0073] In this step, the IoU threshold is calculated by constructing a nonlinear function of the area of ​​each real annotated sample. The constructed area nonlinear function includes a full consideration of the factors that affect the accuracy of sample selection, and reflects the factors in the calculation of the formula in the form of weights, so that the calculation of the threshold is more in line with the category characteristics of the sample. Normally, the threshold calculated using an ordinary area nonlinear function will be relatively large, resulting in a small number of samples, which is not conducive to model training. This application assigns weight values ​​to the parameters in the nonlinear function to eliminate the influence of other types of samples on the area nonlinear function, thereby obtaining more samples.

[0074] Step S44, calculate the IoU values ​​of a predetermined number M of anchor boxes corresponding to each real annotated sample belonging to N; preset an IoU threshold, and take the anchor boxes with IoU values ​​greater than the IoU threshold as candidate samples to implement the screening of samples in N categories.

[0075] Step S5, mapping the optimized candidate samples on the feature map to obtain the region of interest, pooling and aligning the original features of the region of interest, and obtaining the features of the region of interest.

[0076] Step S6, input the features of the region of interest into the classification head and the detection head to obtain the predicted categories and regression positions of all real labeled samples; at the same time, screen out positive samples and typical negative samples from the features of the region of interest to establish a positive sample teacher set and a typical negative sample teacher set.

[0077] In this step, the positive samples and typical negative samples are screened out from the features of the region of interest, specifically including the following steps:

[0078] Get the IoU value and classification score of the candidate samples corresponding to the features of the region of interest;

[0079] Define IQ_H = IoU × classification score (cls_scores), and preset the positive sample threshold T 正 , IQ_H>T 正 The features of are taken as positive samples;

[0080] Define IQ_L = (1-IoU) × classification score (cls_scores), and preset the negative sample threshold T 负 , IQ_L>T 负 The features of are taken as typical negative samples.

[0081] In the small target detection task, the interference of background noise and blurred areas is one of the main reasons for misidentification. By analyzing the visualization results, it is found that the model often misjudges areas similar to blurred backgrounds as target foregrounds during training. Such samples usually have a low IoU, but the model assigns them a higher classification score, causing these samples to be mistakenly identified as targets. This phenomenon is mainly due to the fact that the feature representation ability of the model is not sensitive enough to the difference between the background and the target, especially in small target detection, where the distinction between the background and the target is more difficult. In order to effectively solve this problem, the present invention proposes a definition of a typical negative sample, that is, by calculating the (1-IoU)*cls_scores of the sample to determine whether it is defined as a negative sample.

[0082] The IoU and classification scores are reasonably combined to screen out typical negative samples in training. By focusing on typical negative samples for training, the model can identify and ignore background noise, improving the accuracy and robustness of target detection. This method is particularly effective in detecting small targets, because small targets are often easily covered by background noise or confused with the background. By accurately defining negative samples, the model can better focus on the real target, avoid being disturbed by the background, and improve the final detection effect.

[0083] In this step, based on contrastive learning, a positive sample teacher set and a typical negative sample teacher set are collected at the same time. Based on the idea of ​​feature imitation, how to improve the feature representation quality of small targets is considered. By analyzing the visualization results, the model misidentifies a large number of blurred backgrounds as foregrounds. Such samples often have low IoU, and the model usually calculates a high classification score for such samples and misidentifies them as foregrounds. How to reduce the model's misidentification on such samples can effectively improve the model's detection effect on small targets. By collecting high-quality positive samples in training to establish a teacher set, for the newly generated target feature representation, the similarity between different feature representations is calculated through contrastive learning to provide guidance and supervision for the model to generate features; at the same time, the typical negative samples with identification errors are used as teacher features to guide feature imitation. Specifically, the high-quality samples and low-quality samples in the training set are used as guides for this process, and two dynamically updated sample feature libraries are constructed respectively. The features are mapped to their respective embedding spaces for comparison, and finally the losses are calculated by their respective loss functions and the parameters are updated. The introduction of negative sample guidance can complete the feature imitation task when high-quality samples are insufficient. By collecting high-quality positive samples in training and typical negative samples, a teacher set is established. For the newly generated target feature representation, the similarity between different feature representations is calculated through contrastive learning to provide guidance and supervision for the model to generate features. For the newly generated target feature representation, the similarity between different feature representations is calculated through contrastive learning to provide guidance and supervision for the model to generate features. The samples whose products exceed a certain value introduce the guidance of negative samples, which can help the generation of high-quality positive samples to a certain extent. In addition, the loss function of the positive and negative sample FI branch is designed based on the contrast loss. Among them, the contrast loss of the positive sample increases as the teacher sample and the student sample are more similar, while the contrast loss of the negative sample is the opposite.

[0084] Step S7, based on the predicted category, regression position, positive sample teacher set and typical negative sample teacher set, calculate the classification loss, regression loss, feature imitation (FI) positive sample comparison loss and negative sample comparison loss, calculate the comprehensive loss through the four losses, and complete the training of the small target detection model based on the comprehensive loss.

[0085] In this step, the positive sample contrast loss adopts the following loss function:

[0086]

[0087] In formula (2), ρ represents the candidate sample set, represents the current candidate sample, v p represents the current positive sample, represents the samples in the positive sample set, and τ represents the temperature coefficient for controlling contrastive learning.

[0088] The typical negative sample contrast loss uses the following loss function:

[0089]

[0090] In formula (3), ρ represents the candidate sample set, represents the current candidate sample, υ n represents the current negative sample, represents the samples in the negative sample set, and τ represents the temperature coefficient for controlling contrastive learning.

[0091] In this step, the contrast loss of positive samples and typical negative samples is calculated based on the loss function to optimize the training process. By coordinating positive sample learning and negative sample learning, the quality of sample representation in the training set is further improved, and the training is optimized and the model convergence speed is accelerated.

[0092] Step S8, inputting the current image to be detected into the trained small target detection model, and outputting all target detection results, wherein the target detection results include small target detection results.

[0093] Table 1 shows the comparison results between the small object detection method of the present invention and the most advanced method on the SODA-D dataset.

[0094] Table 1: Model results on the SODA-D benchmark set

[0095]

[0096] As shown in Table 1, compared with the existing target detection methods, the method of the present invention achieves the most advanced performance on the small target detection dataset, with an average precision (AP) of 31.4%, and all indicators are currently optimal. Specifically, compared with the baseline (Faster-RCNN), the AP is improved by 2.5%; compared with the dedicated small target detection method RFLA, the leading margins of various indicators are obvious (AP: 1.7%, AP50: 1.6%, AP75: 2.7%, APeS: 2.4%, APrS: 1.0%, APgS: 2.0%, APN 2.1%); most importantly, compared with the small target detection method CFINet with a similar structure, the AP is improved by 0.7%, AP50 is improved by 1.0%, AP75 is improved by 1.2%, APeS is improved by 0.9%, APrS is improved by 0.1%, APgS is improved by 1.0%, APN is improved by 2.1%. The experimental results prove that the method of the present invention has superiority and versatility, and has better performance and potential in the small target detection task.

[0097] like Figure 2As shown, the visualization results of the method of the present invention on the SODA-D benchmark test set, where boxes of different colors represent targets of different categories. The visualization results verify that the method of the present invention can still have excellent detection effects when facing extremely small targets.

[0098] Based on the same idea, an embodiment of the present invention further provides a small target detection system, the system comprising:

[0099] The system includes: a data acquisition module, a model construction module, a feature map extraction module, a CRPN processing module, a mapping feature map module, a classification and detection module, a positive and negative sample set construction module, a model training module and a result output module; wherein,

[0100] The data acquisition module is used to acquire historical images and real annotated samples in the historical images as training data sets, and input them into the feature map extraction module; the training data sets include categories of real annotated samples divided by target size, including extremely small targets eS, relatively small targets rS, ordinary small targets gS and normal-sized targets N; and is also used to acquire the current image to be detected, and input it into the model construction module after the model training is completed;

[0101] The model building module is used to build a small target detection model, which includes a convolutional neural network, an improved thick and thin pipeline network, a mapping module, a classification head, a detection head and a loss function module; and is also used to input the current image to be detected into the trained small target detection model after the model training is completed, and send the detection result to the result output module;

[0102] The feature map extraction module is used to input the historical image into a convolutional neural network to extract a feature map of the image, wherein the feature map contains real labeled samples and category information;

[0103] The CRPN processing module is used to input the feature map into the improved CRPN, generate candidate samples through a thick pipeline, optimize the candidate samples through a thin pipeline, and calculate the classification score of each candidate sample;

[0104] The mapping feature map module is used to map the optimized candidate samples on the feature map to obtain the region of interest, and to pool and align the original features of the region of interest to obtain the features of the region of interest;

[0105] The classification and detection module is used to map the optimized candidate samples on the feature map to obtain the region of interest, and to pool and align the original features of the region of interest to obtain the features of the region of interest;

[0106] The positive and negative sample set construction module is used to screen out positive samples and typical negative samples from the features of the region of interest, and establish a positive sample teacher set and a typical negative sample teacher set;

[0107] The model training module is used to calculate the classification loss, regression loss, positive sample contrast loss and negative sample contrast loss of feature imitation (FI) based on the predicted category, regression position, positive sample teacher set and typical negative sample teacher set, and obtain the comprehensive loss through the four loss calculations, and complete the training of the small target detection model based on the comprehensive loss;

[0108] The result output module is used to output all target detection results, including small target detection results.

[0109] In this embodiment, each module is implemented by a processor, and a memory is appropriately added when storage is required. Among them, the processor can be but is not limited to a microprocessor MPU, a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), other programmable logic devices, discrete gates, transistor logic devices, discrete hardware components, etc. The memory can include a random access memory (RAM) and can also include a non-volatile memory (NVM), such as at least one disk storage. Optionally, the memory can also be at least one storage device located away from the aforementioned processor.

[0110] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center.

[0111] It should also be noted that the small target detection system described in this embodiment corresponds to the small target detection method, and the description and limitation of the method are also applicable to the system, which will not be repeated here.

[0112] The above description is only a preferred embodiment of the present invention and an explanation of the technical principles used. It is not intended to limit the scope of the invention claimed for protection, but only represents the preferred embodiment of the present invention. Those skilled in the art should understand that the scope of the invention involved in the present invention is not limited to the technical solution formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the inventive concept. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present invention.

Claims

1. A small target detection method, characterized in that: The method comprises the following steps: Step S1, obtaining historical images and real annotated samples in the historical images as training data sets, and the training data sets include categories of real annotated samples divided by target size, including extremely small targets eS, relatively small targets rS, ordinary small targets gS, and normal-sized targets N; Step S2, constructing a small target detection model, the model includes a convolutional neural network, an improved thick and thin pipeline network CRPN, a mapping module, a classification head, a detection head and a loss function module; Step S3, inputting the historical image into a convolutional neural network to extract a feature map of the image, wherein the feature map contains real annotated samples and category information; Step S4, inputting the feature map into the improved CRPN, generating candidate samples through a coarse pipeline, optimizing the candidate samples through a fine pipeline, and calculating the classification score of each candidate sample; Step S5, mapping the optimized candidate samples on the feature map to obtain the region of interest, pooling and aligning the original features of the region of interest to obtain the features of the region of interest; Step S6, input the features of the region of interest into the classification head and the detection head to obtain the predicted categories and regression positions of all real labeled samples; at the same time, screen out positive samples and typical negative samples from the features of the region of interest to establish a positive sample teacher set and a typical negative sample teacher set; Step S7, based on the predicted category, regression position, positive sample teacher set and typical negative sample teacher set, calculate the classification loss, regression loss, positive sample contrast loss of feature imitation and negative sample contrast loss, calculate the comprehensive loss through the four losses, and complete the training of the small target detection model based on the comprehensive loss; Step S8, inputting the current image to be detected into the trained small target detection model, and outputting all target detection results, wherein the target detection results include small target detection results.

2. The small target detection method according to claim 1, characterized in that: Step S4 generates candidate samples through a rough pipeline, specifically including: Step 41, determine the categories of all real labeled samples in the feature map; for extremely small targets eS, execute step S42; for relatively small targets rS and ordinary small targets gS, execute step S43; for normal size targets N, execute step S44; Step S42, calculating the IoU values ​​of a predetermined number of M anchor boxes corresponding to each real labeled sample belonging to eS, sorting the anchor boxes in descending order of IoU values, and selecting the first Q anchor boxes as candidate samples, with Q≤M; Step S43, calculating the IoU values ​​of a predetermined number of M anchor boxes corresponding to each real annotated sample belonging to rS and gS; constructing a nonlinear function of the area of ​​each real annotated sample to calculate the IoU threshold; taking the anchor boxes with IoU values ​​greater than the IoU threshold T as candidate samples to implement the screening of samples in the rS and gS categories; Step S44, calculate the IoU values ​​of a predetermined number M of anchor boxes corresponding to each real annotated sample belonging to N; preset an IoU threshold, and take the anchor boxes with IoU values ​​greater than the IoU threshold as candidate samples to implement the screening of samples in N categories.

3. The small target detection method according to claim 2, characterized in that: The nonlinear function of the area of ​​each true labeled sample constructed is as follows: T=min(T max , max(αC, βC+δγ×f(wh))) (1) In formula (1), T max represents the upper limit of the IoU threshold, C represents the IoU standard value, α represents the lower limit coefficient of the IoU threshold, β represents the starting coefficient of the nonlinear function, w and h represent the width and height of the real labeled sample, f(wh) represents the nonlinear function of the product of the width and height of the real sample, γ represents the slope of the nonlinear function, and δ represents the slope coefficient.

4. The small target detection method according to claim 1, characterized in that: The convolutional neural network described in step S2 adopts a backbone network and a fast regional convolutional neural network Faster RCNN model.

5. The small target detection method according to claim 1, characterized in that: The improved fine pipeline network described in step S2 is constructed based on the dynamic threshold method, including a coarse pipeline network and a fine pipeline network, which respectively generate and optimize candidate regions for the input image, and improve the CRPN by optimizing the threshold calculation process of the coarse pipeline and the fine pipeline.

6. The small target detection method according to claim 1, characterized in that: Step S6 selects positive samples and typical negative samples from the features of the region of interest, and specifically includes the following steps: Get the IoU value and classification score of the candidate samples corresponding to the features of the region of interest; Define IQ_H = IoU × classification score, and preset the positive sample threshold T 正 , IQ_H>T 正 The features of are taken as positive samples; Define IQ_L = (1-IoU) × classification score, and preset the negative sample threshold T 负 , IQ_L>T 负 The features of are taken as typical negative samples.

7. The small target detection method according to claim 1, characterized in that: The positive sample comparison loss in step S7 adopts the following loss function: In formula (2), ρ represents the candidate sample set, represents the current candidate sample, υ p represents the current positive sample, represents the samples in the positive sample set, and τ represents the temperature coefficient for controlling contrastive learning.

8. The small target detection method according to claim 1, characterized in that: The typical negative sample contrast loss in step S7 adopts the following loss function: In formula (3), ρ represents the candidate sample set, represents the current candidate sample, υ n represents the current negative sample, represents the samples in the negative sample set, and τ represents the temperature coefficient for controlling contrastive learning.

9. A small target detection system, characterized in that: The system includes: a data acquisition module, a model construction module, a feature map extraction module, a CRPN processing module, a mapping feature map module, a classification and detection module, a positive and negative sample set construction module, a model training module and a result output module; wherein, The data acquisition module is used to acquire historical images and real annotated samples in the historical images as training data sets, and input them into the feature map extraction module; the training data sets include categories of real annotated samples divided by target size, including extremely small targets eS, relatively small targets rS, ordinary small targets gS and normal-sized targets N; and is also used to acquire the current image to be detected, and input it into the model construction module after the model training is completed; The model building module is used to build a small target detection model, which includes a convolutional neural network, an improved thick and thin pipeline network CRPN, a mapping module, a classification head, a detection head and a loss function module; and is also used to input the current image to be detected into the trained small target detection model after the model training is completed, and send the detection result to the result output module; The feature map extraction module is used to input the historical image into a convolutional neural network to extract a feature map of the image, wherein the feature map contains real labeled samples and category information; The CRPN processing module is used to input the feature map into the improved CRPN, generate candidate samples through a thick pipeline, optimize the candidate samples through a thin pipeline, and calculate the classification score of each candidate sample; The mapping feature map module is used to map the optimized candidate samples on the feature map to obtain the region of interest, and to pool and align the original features of the region of interest to obtain the features of the region of interest; The classification and detection module is used to map the optimized candidate samples on the feature map to obtain the region of interest, and to pool and align the original features of the region of interest to obtain the features of the region of interest; The positive and negative sample set construction module is used to screen out positive samples and typical negative samples from the features of the region of interest, and establish a positive sample teacher set and a typical negative sample teacher set; The model training module is used to calculate the classification loss, regression loss, positive sample contrast loss of feature imitation and negative sample contrast loss based on the predicted category, regression position, positive sample teacher set and typical negative sample teacher set, and obtain the comprehensive loss through the four loss calculations, and complete the training of the small target detection model based on the comprehensive loss; The result output module is used to output all target detection results, including small target detection results.

Citation Information

Patent Citations

  • Target detection method, target detection model and target detection system based on cascade detector

    CN109886286A

  • Target detection method and system based on cut area candidate network

    CN110348435A

  • Detection box determination method and device, storage medium and electronic device

    CN115830308A

  • Target detection model training method, target detection method and equipment

    CN117746008A

  • Remote sensing image directed target detection method based on smooth GIoU regression loss function

    CN117853895A