A fine-grained target detection and recognition method, device and readable storage medium

By combining the generation of multiple target samples with an adaptive grouping perception mechanism, the problems of low recognition accuracy and high computational complexity in fine-grained target detection and recognition are solved, and the recognition accuracy and efficiency are improved. It is suitable for biodiversity research and aerial vehicle recognition.

CN120510456BActive Publication Date: 2025-09-12SUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510998805.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-09-12
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Existing fine-grained target detection and recognition methods cannot effectively solve the problem of low recognition accuracy. At the same time, the complex model structure increases the computational complexity and error in the detection and recognition process, affecting biodiversity research and flight safety.

Method used

By generating multiple target samples with the same true target category, utilizing sample allocation strategy and adaptive group perception mechanism, constructing loss function for iterative training, we ensure the balanced distribution of target samples in the tail category, and introducing a small amount of parameter calculation during the training process to avoid the model being biased towards the head category.

Benefits of technology

It improves the accuracy and efficiency of fine-grained target detection and recognition, reduces computational complexity and recognition errors, and provides more accurate data support for biodiversity research and aerial vehicle identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510456B_ABST
    Figure CN120510456B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of computer vision and image processing technology, and relates to a fine-grained target detection and recognition method, device, and readable storage medium. The method generates multiple target samples for each target in an image; utilizes a fine-grained target detection and recognition model to obtain a predicted positioning box and predicted category for each target sample; utilizes a sample allocation strategy paradigm to screen the target samples to obtain an initial positive target sample set; calculates the proportion of target samples of each category in the initial positive target sample set, as well as the average number of target samples of each category in the initial positive target sample set; if the proportion of target samples in the kth category is less than a preset proportion and / or the number of target samples in the kth category is less than the average number of target samples, selects target samples from target samples in the kth category outside the initial positive target sample set to supplement the initial positive target sample set, constructs a loss function, and iteratively trains the model. This solution improves detection and recognition accuracy while reducing the computational cost of the detection and recognition process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and image processing technology, and in particular to a fine-grained target detection and recognition method, device, and computer-readable storage medium. Background Art

[0002] Fine-grained target recognition refers to further identifying the location and type of each target in an image, based on target detection. It has wide applications in fields such as biodiversity research, materials science, and aerospace. For example, in biodiversity research, fine-grained recognition of images containing multi-species information can accurately calculate the distribution characteristics and quantity information of different species in the image, providing key data for biodiversity conservation and ecosystem research. In aerospace, fine-grained recognition of images containing multiple aircraft can accurately identify the model and location of each aircraft, thereby assisting air traffic control systems in optimizing flight management and preventing collision risks.

[0003] Traditional fine-grained object detection and recognition methods iteratively train machine learning models such as object recognition neural networks to locate and classify multiple objects in an image. However, in images containing multiple objects, some common, large-scale objects often have a large number of samples, while some rare objects have sparse samples. For example, when conducting fine-grained recognition research on images containing multiple bird categories to study bird distribution characteristics, the image may contain multiple common birds such as sparrows and magpies, while the number of rare birds is extremely small. This causes the multiple objects in the image to exhibit a long-tail distribution, that is, the category containing many objects is located at the head (head category), and the category containing few objects is at the tail (tail category). Due to the unbalanced sample distribution, the model pays more attention to the characteristics of the head category targets and ignores the tail category targets, resulting in recognition errors in the model output due to the uneven sample distribution. At the same time, during the fine-grained detection and recognition process, the differences between some similar objects (such as goldfinches and rosefinches) are very subtle, which also makes it impossible for the model to accurately recognize similar objects, resulting in low fine-grained object detection and recognition accuracy.

[0004] To solve the above problems, the prior art constructs a multi-stage classifier to classify images multiple times. In the first stage, a simple classifier is used to simply identify and classify the targets in the image. For example, birds, cats, and dogs are first identified in an image containing multiple animals. In the subsequent stages, a more complex and targeted classifier is designed to further classify the targets in each major category obtained in the first stage. For example, birds are further divided into waterfowl and wading birds. Each category is then further subdivided by a subsequent classifier, such as swans and geese in the waterfowl category. This solves the problem that the model cannot accurately identify similar targets. At the same time, the gradual refinement of the recognition method can overcome the problem of low recognition accuracy caused by insufficient learning of the tail categories by the model to a certain extent. However, this method increases the amount of computation in the detection and recognition process. Moreover, as the model structure becomes more complex and the number of parameters increases, the error in the recognition results output by the model will also increase, affecting the accuracy of the recognition results. At the same time, if the model cannot detect the tail category targets in the simple recognition and classification in the first stage, it will not be able to further classify these tail category targets in the subsequent stages, resulting in the detection and recognition process completely ignoring the tail category targets, resulting in low detection and recognition accuracy.

[0005] In summary, the existing method of using multi-stage classifiers for fine-grained target detection and recognition cannot effectively solve the problem of low accuracy of fine-grained target detection and recognition results. This will lead to the inability to obtain accurate biological distribution characteristics when using fine-grained target detection and recognition to study biodiversity, and thus cannot provide an accurate data basis for biodiversity conservation and ecosystem research; at the same time, the complex model structure also increases the computational complexity and error in the detection and recognition process, reducing the accuracy and timeliness of the recognition results. When using fine-grained target detection and recognition to assist in the control of aerial vehicles, this low timeliness and low accuracy will also affect the flight safety of the aircraft. Summary of the Invention

[0006] To this end, the technical problem to be solved by the present invention is to overcome the problem that the fine-grained target detection and recognition method in the prior art is not only unable to effectively solve the problem of low accuracy of fine-grained target detection and recognition results, but also the complex model structure also increases the amount of calculation and error in the detection and recognition process.

[0007] To solve the above technical problems, the present invention provides a fine-grained target detection and recognition method, comprising:

[0008] Get the images in the training set, generate multiple target samples with the same category for each target based on the true category and true positioning box of each target in the image, and obtain the target image;

[0009] A fine-grained target detection and recognition model is used to obtain the predicted positioning box and predicted category of each target sample in the target image. The target samples are screened based on the predicted positioning box and predicted category of each target sample using the sample allocation strategy paradigm to obtain the initial positive target sample set.

[0010] Calculate the ratio of the number of target samples of each category in the initial positive target sample set to the total number of targets in the image, and obtain the ratio of target samples of each category; calculate the average number of target samples of each category in the initial positive target sample set;

[0011] If the proportion of target samples of the kth category is less than the preset proportion and / or the number of target samples of the kth category is less than the average number of target samples, then multiple target samples belonging to the kth category outside the initial positive target sample set are selected to supplement the initial positive target sample set;

[0012] The optimal positive target sample set is obtained based on the supplemented initial positive target sample set, and the negative target sample set is obtained based on the target samples outside the optimal positive target sample set, so as to calculate the value of the detection and recognition loss function and iteratively train the fine-grained target detection and recognition model until the value of the detection and recognition loss function is minimized, thereby obtaining a trained fine-grained target detection and recognition model.

[0013] Preferably, obtaining a predicted positioning box and a predicted category of each target sample in a target image using a fine-grained target detection and recognition model includes:

[0014] The target image is input into the encoder network in the fine-grained target detection and recognition model for multi-level feature extraction, and a multi-scale feature map is output;

[0015] The multi-scale feature maps are respectively input into the parallel classification feature extraction network and positioning feature extraction network in the fine-grained target detection and recognition model, and the classification feature map and positioning feature map are output;

[0016] The classification feature map and the positioning feature map are spliced ​​and input into the output layer based on the adaptive group perception mechanism in the fine-grained target detection and recognition model to output the fused feature map;

[0017] The fused feature map is input into the fully connected layer in the fine-grained target detection and recognition model, and the predicted positioning box and predicted category of each target sample in the target image are output.

[0018] Preferably, the calculation formula for the target sample ratio of each category is expressed as:

[0019] ,

[0020] in, Indicates the proportion of target samples of the kth category; Indicates the number of targets corresponding to the k-th target sample; Indicates the number of all targets in the image;

[0021] The calculation formula for the average number of target samples in each category is expressed as:

[0022] ,

[0023] in, Indicates the average number of target samples in each category; represents the number of target samples in the initial positive target sample set; Represents the number of target sample categories in the initial positive target sample set.

[0024] Preferably, selecting a plurality of target samples from target samples belonging to the kth class outside the initial positive target sample set to supplement the initial positive target sample set comprises:

[0025] Based on the number of target samples in the initial positive target sample set and the number of targets corresponding to all target samples in the initial positive target sample set, calculate the average number of target samples for each target;

[0026] Based on the difference between the preset ratio and the ratio of the k-th target sample, and the average number of target samples for each target, the number of target samples to be supplemented for the k-th target sample is calculated. ;

[0027] For each target sample belonging to the kth class outside the initial positive target sample set, calculate the intersection-and-union ratio of its predicted positioning frame and the true positioning frame of the target corresponding to the kth class target sample; select the one with the largest intersection-and-union ratio target samples are added to the initial positive target sample set.

[0028] Preferably, the calculation formula for the average number of target samples for each target is expressed as:

[0029] ,

[0030] in, represents the average number of target samples for each target; represents the number of target samples in the initial positive target sample set; Indicates the number of targets corresponding to all target samples in the initial positive target sample set;

[0031] Target quantity to be replenished The calculation formula is expressed as:

[0032] ,

[0033] in, Indicates the preset ratio; Indicates the proportion of positive target samples in the kth category.

[0034] Preferably, the target samples are screened based on the predicted positioning box and predicted category of each target sample using the sample allocation strategy paradigm, and the initial positive target sample set obtained includes:

[0035] Obtain the intersection-over-union ratio of the predicted positioning frame of each target sample and the corresponding real positioning frame of the target, as well as the predicted category confidence of each target sample;

[0036] Based on the target samples whose intersection-over-union ratio is greater than or equal to a preset intersection-over-union ratio threshold and whose predicted category confidence is greater than or equal to a preset confidence threshold, an initial positive target sample set is obtained.

[0037] Preferably, the process of constructing the detection and recognition loss function includes:

[0038] Construct a classification loss function based on the distance measurement between the target samples in the optimal positive target sample set and the negative target sample set and their classification feature maps;

[0039] Construct a positioning loss function based on the distance measurement between the predicted positioning box of each target in the optimal positive target sample set and its true positioning box;

[0040] The classification loss function and the positioning loss function are weighted and summed to obtain the detection and recognition loss function.

[0041] Preferably, the classification loss function is expressed as:

[0042] ,

[0043] in, represents the classification loss function; represents a function for calculating the distance metric between the target samples in the optimal positive target sample set and the negative target sample set and their predicted categories; represents the optimal positive target sample set; represents the negative target sample set; Represents a classification feature map;

[0044] The positioning loss function is expressed as:

[0045] ,

[0046] in, represents the positioning loss function; Represents a function for calculating the distance metric between the predicted positioning box of the target sample and the true positioning box of its corresponding target; Represents the predicted positioning box; Represents the real positioning frame; Indicates filtering;

[0047] The detection and recognition loss function is expressed as:

[0048] ,

[0049] in, Represents the detection and recognition loss function; Represents the weight of the positioning loss function.

[0050] The present invention also provides a fine-grained target detection and recognition device, comprising:

[0051] The target sample generation module is used to obtain images in the training set and generate multiple target samples of the same category for each target based on the true category and true positioning box of each target in the image to obtain the target image;

[0052] The initial sample allocation module is used to obtain the predicted positioning box and predicted category of each target sample in the target image using a fine-grained target detection and recognition model; the target samples are screened based on the predicted positioning box and predicted category of each target sample using the sample allocation strategy paradigm to obtain the initial positive target sample set;

[0053] The parameter calculation module is used to calculate the ratio of the number of target samples of each category in the initial positive target sample set to the total number of targets in the image, and obtain the ratio of target samples of each category; calculate the average number of target samples of each category in the initial positive target sample set;

[0054] An online sample tail supplement allocation module is used to select multiple target samples from target samples belonging to the kth class outside the initial positive target sample set to supplement the initial positive target sample set if the proportion of target samples of the kth class is less than a preset proportion and / or the number of target samples of the kth class is less than the average number of target samples;

[0055] The model training and acquisition module is used to obtain the optimal positive target sample set based on the supplemented initial positive target sample set, and to obtain the negative target sample set based on the target samples outside the optimal positive target sample set, so as to calculate the value of the detection and recognition loss function and iteratively train the fine-grained target detection and recognition model until the value of the detection and recognition loss function is minimized, thereby obtaining a trained fine-grained target detection and recognition model.

[0056] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned steps of fine-grained target detection and recognition are implemented.

[0057] The fine-grained target detection and recognition method provided in this application has the following beneficial effects:

[0058] 1. This application adopts a preset sample efficient supplementation strategy. Before training, multiple target samples with the same category as the real target are pre-generated around each real target in the image, and then a fine-grained target detection and recognition model is used to detect and identify the target samples to obtain the predicted category and predicted positioning box of each target sample. Specifically, in order to enable the model to more accurately identify subtle differences between similar targets, this application introduces a method based on positive and negative sample allocation, that is, the sample allocation strategy paradigm is first used to perform initial allocation based on the predicted category and predicted positioning box of the target sample to obtain an initial positive target sample set. Furthermore, in order to solve the problem of a small number of target samples in the tail category, the number ratio and average number of target samples of each category in the initial positive sample set are calculated to identify the tail category in the image, and additional target samples are supplemented from target samples outside the initial positive target sample set only for the tail category, so that the target samples of each category in the supplemented positive target sample set are evenly distributed. Finally, based on the positive and negative target samples This paper constructs a loss function based on its prediction results. On the one hand, through the distribution of positive and negative target samples, the model can learn the common features of target samples of the same category from positive target samples, and learn the boundaries of targets of different categories from negative target samples, thereby capturing the subtle differences between targets of different categories. On the other hand, by supplementing the positive target samples of the tail category, the model is no longer biased towards focusing on the target samples of the head category during the iterative update process, thereby solving the problem of low accuracy of target detection and recognition results caused by the long-tail distribution of targets. In addition, the method provided by this application only introduces a small amount of parameter calculation during the training process, that is, the computing resources are concentrated in the sample screening and allocation stages of the training process, without the need to design a complex model structure. The recognition accuracy of the model is improved without changing the model architecture. While improving the detection and recognition accuracy, the amount of calculation and the recognition error introduced by the complex model are reduced, thereby improving the recognition accuracy and efficiency of fine-grained target detection and recognition in various fields, and providing more accurate data for research in various fields.

[0059] 2. The fine-grained object detection and recognition model of this application adopts an output layer based on an adaptive grouping perception mechanism. By dividing the input feature map into multiple subgroups corresponding to different viewpoint features according to the channel dimension, and then assigning adaptive weights to each subgroup, the model can focus on different viewpoint features, thereby enabling the features of different subgroups to complement and enhance each other, thereby enabling the model to learn more representative features with limited samples, further improving the recognition accuracy of tail category targets;

[0060] 3. When using target samples outside the initial positive target sample set to supplement the tail category target samples, the intersection-over-union ratio between the predicted positioning frame of each target sample in the tail category outside the initial positive target sample set and the true positioning frame of its corresponding target is used as the selection benchmark, so as to select target samples whose predicted positioning frames are as close to the true positioning frame as possible. That is, the selected target samples have relatively more useful information related to their corresponding targets and relatively less background information unrelated to the targets, thereby providing the model with more tail category target feature information. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:

[0062] Figure 1 Flowchart of the fine-grained target detection and recognition method provided in this application;

[0063] Figure 2 This is a confusion matrix diagram of the fine-grained target detection and recognition method provided in the embodiment of the present application; wherein, Figure 2 (a) is a confusion matrix diagram for fine-grained target recognition of images containing multiple types of aircraft using the method provided by this application. Figure 2 (b) is a confusion matrix diagram of fine-grained target recognition of images containing multiple types of ships using the method provided by this application. Figure 2 (c) is a confusion matrix diagram for fine-grained target recognition of images containing multiple types of vehicles using the method provided by this application. Figure 2 (d) is a confusion matrix diagram for fine-grained target recognition of images containing multiple types of sites using the method provided by this application;

[0064] Figure 3 A schematic diagram of an ablation study of a preset ratio parameter in a fine-grained target detection and recognition method provided in an embodiment of the present application;

[0065] Figure 4 Schematic diagram of the structure of the fine-grained target detection and recognition device provided in this application. DETAILED DESCRIPTION

[0066] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.

[0067] See also Figure 1 , Figure 1 The figure shows a flow chart of the fine-grained target detection and recognition method provided by this application, which specifically includes:

[0068] S10: Obtain an image in the training set, generate multiple target samples with the same category for each target based on the true category and true positioning box of each target in the image, and obtain a target image.

[0069] Specifically, a plurality of target samples evenly distributed with the target's true positioning frame as the center are generated for each target through traditional data enhancement or sample generation methods. The true categories of these target samples are the same as the true category of the target.

[0070] Optionally, in some embodiments, the true label of each target in the image can be encoded from a serial number representation to a one-hot encoding representation, so that the one-hot encoding can be used to compare with the prediction result output by the model. For example, the hard labels of an image containing multiple birds and their corresponding one-hot encoding representations are: swan (hard label 0) - one-hot encoding is [1,0,0], wild goose (hard label 1) - one-hot encoding is [0,1,0], and redbird (hard label 2) - one-hot encoding is [0,0,1]. The specific label conversion formula is as follows:

[0071] ,

[0072] in, Indicates the true number of categories of the target.

[0073] S20: Use the fine-grained target detection and recognition model to obtain the predicted positioning box and predicted category of each target sample in the target image; use the sample allocation strategy paradigm to screen the target samples based on the predicted positioning box and predicted category of each target sample to obtain an initial positive target sample set.

[0074] S30: Calculate the ratio of the number of targets corresponding to target samples of each category in the initial positive target sample set to the number of all targets in the image, and obtain the ratio of target samples of each category; calculate the average number of target samples of each category in the initial positive target sample set.

[0075] S40: If the proportion of target samples of the kth category is less than the preset proportion and / or the number of target samples of the kth category is less than the average number of target samples, multiple target samples are selected from the target samples belonging to the kth category outside the initial positive target sample set to supplement the initial positive target sample set.

[0076] Specifically, if the proportion of target samples in the kth category is greater than or equal to the preset proportion and the number of target samples in the kth category is greater than or equal to the average number of target samples, it indicates that the kth category is not a tail category, and the initial positive target sample set contains enough target samples in the kth category, so the target samples in this category cannot be processed; if the proportion of target samples in the kth category is less than the preset proportion and / or the number of target samples in the kth category is less than the average number of target samples, it indicates that the target samples in the kth category belong to the tail category in the initial positive target sample set, so the target samples in this category need to be supplemented, so as to avoid the problem of low model recognition accuracy caused by uneven sample distribution in the initial positive target sample set.

[0077] S50: Based on the supplemented initial positive target sample set, an optimal positive target sample set is obtained, and based on the target samples outside the optimal positive target sample set, a negative target sample set is obtained, thereby calculating the value of the detection and recognition loss function and iteratively training the fine-grained target detection and recognition model until the value of the detection and recognition loss function is minimized, thereby obtaining a trained fine-grained target detection and recognition model.

[0078] An efficient strategy for supplementing preset samples is adopted. Before training, multiple target samples with the same category as the real target are pre-generated around each real target in the image. Then, a fine-grained target detection and recognition model is used to detect and recognize the target samples, and the predicted category and predicted positioning box of each target sample are obtained. Specifically, in the fine-grained detection and recognition process, the differences between some similar targets are very subtle. In order to enable the model to more accurately identify the subtle differences between similar targets, this application introduces a method based on positive and negative sample allocation, that is, the sample allocation strategy paradigm is first used to perform an initial allocation based on the predicted category and predicted positioning box of the target sample to obtain an initial positive target sample set. Furthermore, by calculating the number of target samples of each category in the initial positive sample set, the predicted category and predicted positioning box of the target sample are obtained. The method is simple and convenient to use. The method is simple and convenient to use. The method is simple and convenient to use. The method is simple and convenient to use. The method is simple and convenient to use. The method is simple and convenient to use. The method is simple and convenient to use. The method is simple and convenient to use. The method is simple and convenient to use. The method is simple and convenient to use. The method is simple and convenient to use. The method is simple and convenient to use. The method is simple and convenient to use. The method is simple and convenient to use.

[0079] Furthermore, the specific steps of obtaining the predicted positioning box and predicted category of each target sample in the target image using the fine-grained target detection and recognition model in step S20 include:

[0080] S200: Input the target image into the encoder network in the fine-grained target detection and recognition model for multi-level feature extraction, and output a multi-scale feature map.

[0081] Specifically, using the encoder network to extract multi-scale features of the target image can obtain multi-scale feature representations with different receptive fields. The feature extraction formula is expressed as:

[0082] ,

[0083] in, Represents the extracted multi-scale feature map, , Indicates the Layer-scale feature maps, Indicates the number of layers.

[0084] S201: Input the multi-scale feature map into the parallel classification feature extraction network and positioning feature extraction network in the fine-grained target detection and recognition model respectively, and output the classification feature map and the positioning feature map.

[0085] S202: The classification feature map and the positioning feature map are spliced ​​and input into the output layer based on the adaptive grouping perception mechanism in the fine-grained target detection and recognition model to output the fused feature map.

[0086] The specific implementation method of step S202 is: evenly split the input feature map into M groups along the channel dimension, perform the Head operation on each group of input feature maps, and output the corresponding view discriminant feature map; splice the view discriminant feature maps corresponding to the M groups of input feature maps to obtain a fused feature map with the same dimension as the original input feature map.

[0087] S203: Input the fused feature map into the fully connected layer in the fine-grained target detection and recognition model, and output the predicted positioning box and predicted category of each target sample in the target image.

[0088] This application adopts an output layer based on an adaptive group perception mechanism. By dividing the input feature map into multiple subgroups corresponding to different perspective features according to the channel dimension, and then assigning adaptive weights to each subgroup, the model can focus on different perspective features, thereby enabling different subgroup features to complement and enhance each other, so that the model can learn more representative features under limited samples, further improving the recognition accuracy of tail category targets.

[0089] At the same time, Table 1 shows the comparison data of the parameters (number of channel dimension split groups M, number of feature extraction layers of the encoder network, complexity, and frame rate per second) of the model including the output layer based on the adaptive group perception mechanism provided by the embodiment of the present application and the baseline model, as well as the mean average precision comparison data of the two models for fine-grained object detection and recognition on the FAIR1M-v2.0 and DOTA-v1.5 datasets:

[0090] Table 1

[0091]

[0092] It can be seen from the data in Table 1 that compared with the baseline model that does not include an output layer based on the adaptive group perception mechanism, although the present application introduces an output layer based on the adaptive group perception mechanism into the model, it does not increase the complexity of the model, nor does it have a significant impact on the frame rate of the model. At the same time, the accuracy of the model is also improved; by comparing the data with the same number of channel dimension split groups of 4, it can be seen that the model of the present application has improved in accuracy, reduced in model complexity, and has less impact on the model's frames per second.

[0093] Furthermore, in some embodiments of the present application, in step S20, the target samples are screened based on the predicted positioning boxes and predicted categories of each target sample using the sample allocation strategy paradigm, and the initial positive target sample set obtained includes:

[0094] S204: Obtain the intersection-over-union ratio of the predicted positioning frame of each target sample and the true positioning frame of its corresponding target, as well as the predicted category confidence of each target sample.

[0095] S205: Based on the target samples whose intersection-over-union ratio is greater than or equal to a preset intersection-over-union ratio threshold and whose prediction category confidence is greater than or equal to a preset confidence threshold, an initial positive target sample set is obtained.

[0096] Optionally, in some embodiments of the present application, only the distance metric between the predicted positioning frame of the target sample and the true positioning frame of its corresponding target (such as the intersection-over-union ratio, the Euclidean distance between the center points of the predicted positioning frame and the true positioning frame) can be used as the basis for allocation, or only the predicted category confidence of the target sample can be used as the basis for allocation.

[0097] Specifically, the calculation formula for the proportion of target samples of each category in step S30 is expressed as:

[0098] ,

[0099] in, Indicates the proportion of target samples of the kth category; Indicates the number of targets corresponding to the k-th target sample; Indicates the number of all targets in the image;

[0100] The calculation formula for the average number of target samples in each category is expressed as:

[0101] ,

[0102] in, Indicates the average number of target samples in each category; represents the number of target samples in the initial positive target sample set; Represents the number of target sample categories in the initial positive target sample set.

[0103] Furthermore, in step S40, among the target samples belonging to the kth class outside the initial positive target sample set, selecting multiple target samples to supplement the initial positive target sample set includes:

[0104] Based on the number of target samples in the initial positive target sample set and the number of targets corresponding to all target samples in the initial positive target sample set, the average number of target samples for each target is calculated.

[0105] Specifically, the calculation formula for the average number of target samples for each target is expressed as:

[0106] ,

[0107] in, represents the average number of target samples for each target; represents the number of target samples in the initial positive target sample set; Indicates the number of targets corresponding to all target samples in the initial positive target sample set.

[0108] Based on the difference between the preset ratio and the ratio of the k-th target sample, and the average number of target samples for each target, the number of target samples to be supplemented for the k-th target sample is calculated. .

[0109] Specifically, the target number to be replenished The calculation formula is expressed as:

[0110] ,

[0111] in, Indicates the preset ratio; Indicates the proportion of positive target samples in the kth category.

[0112] For each target sample belonging to the kth class outside the initial positive target sample set, calculate the intersection-and-union ratio of its predicted positioning frame and the true positioning frame of the target corresponding to the kth class target sample; select the one with the largest intersection-and-union ratio target samples are added to the initial positive target sample set.

[0113] In a specific example of this application, the preset ratio The value of is 0.2, for The scaling factor is in the range of approximately [1, 1.5]. It is worth noting that the tail category in this application is based on each image. Since uniformly distributed target samples are preset near each real target in step S10, when it is found that the number of tail category target samples in the initial positive target sample set is small, it can be supplemented by preset target samples. Usually, the number of preset target samples is not less than the number of target samples to be supplemented. When the number of preset target samples is less than the number of target samples to be supplemented, since the criterion for sample allocation is whether the intersection-and-union ratio of the predicted positioning frame of the target sample and the real positioning frame of its corresponding target is greater than the preset intersection-and-union ratio threshold or whether the predicted category confidence is greater than the preset confidence threshold, the number of target samples that meet the conditions during the initial allocation can be expanded by lowering the preset intersection-and-union ratio threshold or the preset confidence during the initial allocation.

[0114] Furthermore, after supplementing the initial positive target sample set to obtain the optimal positive target sample set, a classification loss function and a positioning loss function can be constructed based on the predicted category and true category of each target sample in the optimal positive target sample set and the negative target sample set, and the predicted positioning box and the true positioning box of the corresponding target, thereby constructing a detection and recognition loss function. Since the loss function is constructed based on the optimized target sample set, the model avoids the problem of excessive focus on the head category target samples and ignoring the tail category target samples during the iterative update based on the loss function.

[0115] Specifically, the construction process of the detection and recognition loss function includes:

[0116] A classification loss function is constructed based on the distance measurement between the target samples in the optimal positive target sample set and the negative target sample set and their classification feature maps.

[0117] Specifically, the classification loss function is expressed as:

[0118] ,

[0119] in, represents the classification loss function; represents a function for calculating the distance metric between the target samples in the optimal positive target sample set and the negative target sample set and their predicted categories; represents the optimal positive target sample set; represents the negative target sample set; Represents a classification feature map.

[0120] In a specific example of this application, The classic focal loss is used to calculate the distance measurement between the target samples in the optimal positive target sample set and the negative target sample set and their classification feature maps.

[0121] A positioning loss function is constructed based on the distance measurement between the predicted positioning box of each target in the optimal positive target sample set and its true positioning box.

[0122] Specifically, the positioning loss function is expressed as:

[0123] ,

[0124] in, represents the positioning loss function; Represents a function for calculating the distance metric between the predicted positioning box of the target sample and the true positioning box of its corresponding target; Represents the predicted positioning box; Represents the real positioning frame; Represents filtering, which means filtering from all predicted positioning frames in the image to obtain the predicted positioning frames corresponding to each target in the optimal positive target sample set.

[0125] In a specific example of this application, Rotated IoU Loss is used to calculate the distance metric between the predicted positioning box of the target sample and the true positioning box of its corresponding target.

[0126] The classification loss function and the positioning loss function are weighted and summed to obtain the detection and recognition loss function.

[0127] The detection and recognition loss function is expressed as:

[0128] ,

[0129] in, Represents the detection and recognition loss function; Represents the weight of the positioning loss function.

[0130] Furthermore, in order to verify the effectiveness of the fine-grained object detection and recognition method provided in the above embodiment, the embodiment of the present application also conducted a comparative test on the same data set between the above-trained fine-grained object detection and recognition model and the model in the prior art:

[0131] All experiments were conducted on a single NVIDIA RTX4090 24GB GPU, using PyTorch 1.7.0 and the MMDetection / MMRotate toolbox. Both models were trained using the SGD optimizer with momentum 0.9 and an initial learning rate of 0.0025. A 24-epoch training strategy was used on the FAIR1M-v2.0, DOTA-v1.5, and DIOR-R datasets, and a 36-epoch training strategy was used on the HRSC2016 datasets.

[0132] Parameter settings: =0.2, =1, M=4.

[0133] First, using ResNet-50 and FPN as the encoder structure, the SCS method is applied to the dataset for fine-grained rigid object recognition experiments; during model training, ATSS is selected for the sample allocation process, while following the same settings as other methods, the balance factor is set to =1.

[0134] To further compare the performance of the baseline model, we analyze the accuracy and convergence speed. In terms of accuracy, we use metrics such as precision, recall, and intersection over union (IoU) for evaluation.

[0135] in, , , , TP, FP, FN and TN represent true positive, false positive, false negative and true negative samples respectively, To predict the positioning box, is the real positioning frame.

[0136] Table 2 shows the relevant data of the comparative test, where MaxIoU, ATSS, and C2F are three different baseline methods. For each baseline method, the embodiments of the present application respectively provide the method and the recognition accuracy of different data in the dataset after using the tail sample supplementation strategy provided by the present application in the method.

[0137] Table 2

[0138]

[0139] As can be seen from Table 2, for different targets in the dataset (bridge, playground, oil tank), after adding the tail sample supplementation strategy provided by this application, the recognition accuracy of the model has been improved, which is sufficient to prove that the method provided by this application can effectively solve the problem of low accuracy of target detection and recognition results caused by uneven sample distribution.

[0140] Figure 2 The figure shows a confusion matrix diagram of the fine-grained target detection and recognition method provided by the embodiment of the present application; wherein, Figure 2 (a) is a confusion matrix diagram for fine-grained target recognition of images containing multiple types of aircraft using the method provided by this application. Figure 2 (b) is a confusion matrix diagram of fine-grained target recognition of images containing multiple types of ships using the method provided by this application. Figure 2(c) is a confusion matrix diagram for fine-grained target recognition of images containing multiple types of vehicles using the method provided by this application. Figure 2 (d) is a confusion matrix diagram of fine-grained target recognition of images containing multiple types of sites using the method provided by this application.

[0141] Specifically, the confusion matrix is ​​used to describe the distribution of correct and incorrect predictions of the model for each category. Its value range is between [0, 1]. Higher values ​​on the diagonal indicate more accurate classification of the model for that category (1 indicates 100% prediction accuracy for that category). Data outside the diagonal indicate incorrect predictions. The closer to or equal to 0, the fewer incorrect predictions the model makes. Furthermore, in this embodiment, images containing multiple types of aircraft primarily include aircraft of various models; images containing multiple types of ships primarily include passenger ships, motorboats, fishing boats, tugboats, engineering vessels, liquid cargo ships, dry cargo ships, and other ships; images containing multiple types of vehicles primarily include small cars, buses, trains, dump trucks, vans, trailers, tractors, excavators, truck tractors, and other vehicles; and images containing multiple types of venues primarily include basketball courts, tennis courts, soccer fields, baseball fields, intersections, roundabouts, and bridges.

[0142] Combine Figure 2 It can be seen from (a), (b), (c), and (d) that for images containing targets of different categories, the values ​​on the diagonal of the corresponding confusion matrix are close to 1, and the values ​​outside the diagonal are close to 0, indicating that the methods provided in this application have high accuracy in fine-grained target detection and recognition.

[0143] Figure 3 Schematic diagram of the ablation study of the preset ratio parameters in the fine-grained target detection and recognition method provided by the embodiment of the present application. The horizontal axis in the figure represents the preset ratio involved in the fine-grained target detection and recognition method provided by the present application. The value of this parameter, the left vertical axis is the mean average precision (mAP) of the fine-grained target detection and recognition method provided by this application on the DOTA-v1.5 dataset, and the right vertical axis is the mAP of the fine-grained target detection and recognition method provided by this application on the FAIR1M-v2.0 dataset.

[0144] from Figure 3 It can be seen that: with the preset ratio The value of gradually increases, and the average precision of the fine-grained target detection and recognition method on the two data sets first increases and then decreases, which shows that by changing the preset ratio The value of can optimize the recognition accuracy of fine-grained target detection and recognition methods; in addition, Figure 3 As can be seen, when the preset ratio When the value of is 0.2, the average recognition accuracy of the fine-grained target detection and recognition method provided in this application on both data sets reaches the maximum. Therefore, by comprehensively considering the comprehensive performance of this method on different data sets, 0.2 is used as the optimal value of the preset ratio in some embodiments of this application.

[0145] Based on the fine-grained target detection and recognition method provided in the above embodiment, the embodiment of the present application also provides a fine-grained target detection and recognition device, such as Figure 4 As shown, the device includes:

[0146] The target sample generation module 10 is used to obtain images in the training set, and generate multiple target samples of the same category for each target based on the true category and true positioning frame of each target in the image to obtain a target image.

[0147] The initial sample allocation module 20 is used to use the fine-grained target detection and recognition model to obtain the predicted positioning box and predicted category of each target sample in the target image; and use the sample allocation strategy paradigm to screen the target samples based on the predicted positioning box and predicted category of each target sample to obtain an initial positive target sample set.

[0148] The parameter calculation module 30 is used to calculate the ratio of the number of target samples corresponding to each category in the initial positive target sample set to the total number of targets in the image, thereby obtaining the ratio of target samples of each category; and calculate the average number of target samples of each category in the initial positive target sample set.

[0149] The online sample tail supplement allocation module 40 is used to select multiple target samples from the target samples belonging to the kth category outside the initial positive target sample set to supplement the initial positive target sample set if the proportion of the kth target sample is less than the preset proportion and / or the number of the kth target sample is less than the average number of target samples.

[0150] The model training and acquisition module 50 is used to obtain the optimal positive target sample set based on the supplemented initial positive target sample set, and to obtain the negative target sample set based on the target samples outside the optimal positive target sample set, so as to calculate the value of the detection and recognition loss function and iteratively train the fine-grained target detection and recognition model until the value of the detection and recognition loss function is minimized, thereby obtaining a trained fine-grained target detection and recognition model.

[0151] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the above-mentioned fine-grained target detection and recognition method.

[0152] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0153] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0154] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0156] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.

Claims

1. A fine-grained target detection and recognition method, characterized in that: include: Get the images in the training set, generate multiple target samples with the same category for each target based on the true category and true positioning box of each target in the image, and obtain the target image; A fine-grained target detection and recognition model is used to obtain the predicted positioning box and predicted category of each target sample in the target image. The target samples are screened based on the predicted positioning box and predicted category of each target sample using the sample allocation strategy paradigm to obtain the initial positive target sample set. Calculate the ratio of the number of target samples of each category in the initial positive target sample set to the total number of targets in the image, and obtain the ratio of target samples of each category; Calculate the average number of target samples for each category in the initial positive target sample set; If the proportion of target samples of the kth category is less than the preset proportion and / or the number of target samples of the kth category is less than the average number of target samples, then multiple target samples belonging to the kth category outside the initial positive target sample set are selected to supplement the initial positive target sample set; The optimal positive target sample set is obtained based on the supplemented initial positive target sample set, and the negative target sample set is obtained based on the target samples outside the optimal positive target sample set, thereby calculating the value of the detection and recognition loss function and iteratively training the fine-grained target detection and recognition model until the value of the detection and recognition loss function is minimized, thereby obtaining a trained fine-grained target detection and recognition model. The construction process of the detection and recognition loss function includes: Based on the distance measurement between the target samples in the optimal positive target sample set and the negative target sample set and their classification feature maps, a classification loss function is constructed: , in, represents the classification loss function; represents a function for calculating the distance metric between the target samples in the optimal positive target sample set and the negative target sample set and their predicted categories; represents the optimal positive target sample set; represents the negative target sample set; Represents a classification feature map; Based on the distance measurement between the predicted positioning box of each target in the optimal positive target sample set and its true positioning box, a positioning loss function is constructed: , in, represents the positioning loss function; Represents a function for calculating the distance metric between the predicted positioning box of the target sample and the true positioning box of its corresponding target; Represents the predicted positioning box; Represents the real positioning frame; Indicates filtering; The weighted sum of the classification loss function and the positioning loss function gives the detection and recognition loss function: , in, Represents the detection and recognition loss function; Represents the weight of the positioning loss function.

2. The fine-grained target detection and recognition method according to claim 1, characterized in that: Using the fine-grained target detection and recognition model to obtain the predicted positioning box and predicted category of each target sample in the target image includes: The target image is input into the encoder network in the fine-grained target detection and recognition model for multi-level feature extraction, and a multi-scale feature map is output; The multi-scale feature maps are respectively input into the parallel classification feature extraction network and positioning feature extraction network in the fine-grained target detection and recognition model, and the classification feature map and positioning feature map are output; The classification feature map and the positioning feature map are spliced ​​and input into the output layer based on the adaptive group perception mechanism in the fine-grained target detection and recognition model to output the fused feature map; The fused feature map is input into the fully connected layer in the fine-grained target detection and recognition model, and the predicted positioning box and predicted category of each target sample in the target image are output.

3. The fine-grained target detection and recognition method according to claim 1, characterized in that: The calculation formula for the proportion of target samples in each category is expressed as: , in, Indicates the proportion of target samples of the kth category; Indicates the number of targets corresponding to the k-th target sample; Indicates the number of all targets in the image; The calculation formula for the average number of target samples in each category is expressed as: , in, Indicates the average number of target samples in each category; represents the number of target samples in the initial positive target sample set; Represents the number of target sample categories in the initial positive target sample set.

4. The fine-grained target detection and recognition method according to claim 1, characterized in that: Among the target samples belonging to the kth class outside the initial positive target sample set, multiple target samples are selected to supplement the initial positive target sample set, including: Based on the number of target samples in the initial positive target sample set and the number of targets corresponding to all target samples in the initial positive target sample set, calculate the average number of target samples for each target; Based on the difference between the preset ratio and the ratio of the k-th target sample, and the average number of target samples for each target, the number of target samples to be supplemented for the k-th target sample is calculated. ; For each target sample belonging to the kth class outside the initial positive target sample set, calculating the intersection-over-union ratio of its predicted positioning box and the true positioning box of the target corresponding to the kth class target sample; selecting the target sample with the largest intersection-over-union ratio. The fine-grained target detection and recognition method according to claim 1 is characterized in that, among the target samples belonging to the kth class outside the initial positive target sample set, selecting multiple target samples to supplement the initial positive target sample set comprises: Based on the number of target samples in the initial positive target sample set and the number of targets corresponding to all target samples in the initial positive target sample set, calculate the average number of target samples for each target; Based on the difference between the preset ratio and the ratio of the k-th target sample, and the average number of target samples for each target, the number of target samples to be supplemented for the k-th target sample is calculated. ; For each target sample belonging to the kth class outside the initial positive target sample set, calculate the intersection-and-union ratio of its predicted positioning frame and the true positioning frame of the target corresponding to the kth class target sample; select the one with the largest intersection-and-union ratio target samples are added to the initial positive target sample set.

5. The fine-grained target detection and recognition method according to claim 4, characterized in that: The calculation formula for the average number of target samples for each target is expressed as: , in, represents the average number of target samples for each target; represents the number of target samples in the initial positive target sample set; Indicates the number of targets corresponding to all target samples in the initial positive target sample set; Target quantity to be replenished The calculation formula is expressed as: , in, Indicates the preset ratio; Indicates the proportion of positive target samples in the kth category.

6. The fine-grained target detection and recognition method according to claim 1, characterized in that: The sample allocation strategy paradigm is used to screen target samples based on the predicted positioning box and predicted category of each target sample, and the initial positive target sample set is obtained, including: Obtain the intersection-over-union ratio of the predicted positioning frame of each target sample and the corresponding real positioning frame of the target, as well as the predicted category confidence of each target sample; Based on the target samples whose intersection-over-union ratio is greater than or equal to a preset intersection-over-union ratio threshold and whose predicted category confidence is greater than or equal to a preset confidence threshold, an initial positive target sample set is obtained.

7. A fine-grained target detection and recognition device, characterized in that: include: The target sample generation module is used to obtain images in the training set and generate multiple target samples of the same category for each target based on the true category and true positioning box of each target in the image to obtain the target image; The initial sample allocation module is used to obtain the predicted positioning box and predicted category of each target sample in the target image using a fine-grained target detection and recognition model; the target samples are screened based on the predicted positioning box and predicted category of each target sample using the sample allocation strategy paradigm to obtain the initial positive target sample set; The parameter calculation module is used to calculate the ratio of the number of target samples of each category in the initial positive target sample set to the number of all targets in the image, and obtain the ratio of target samples of each category; Calculate the average number of target samples for each category in the initial positive target sample set; An online sample tail supplement allocation module is used to select multiple target samples from target samples belonging to the kth class outside the initial positive target sample set to supplement the initial positive target sample set if the proportion of target samples of the kth class is less than a preset proportion and / or the number of target samples of the kth class is less than the average number of target samples; The model training and acquisition module is used to obtain the optimal positive target sample set based on the supplemented initial positive target sample set, and obtain the negative target sample set based on the target samples outside the optimal positive target sample set, thereby calculating the value of the detection and recognition loss function and iteratively training the fine-grained target detection and recognition model until the value of the detection and recognition loss function is minimized, thereby obtaining a trained fine-grained target detection and recognition model. The construction process of the detection and recognition loss function includes: Based on the distance measurement between the target samples in the optimal positive target sample set and the negative target sample set and their classification feature maps, a classification loss function is constructed: , in, represents the classification loss function; represents a function for calculating the distance metric between the target samples in the optimal positive target sample set and the negative target sample set and their predicted categories; represents the optimal positive target sample set; represents the negative target sample set; Represents a classification feature map; Based on the distance measurement between the predicted positioning box of each target in the optimal positive target sample set and its true positioning box, a positioning loss function is constructed: , in, represents the positioning loss function; Represents a function for calculating the distance metric between the predicted positioning box of the target sample and the true positioning box of its corresponding target; Represents the predicted positioning box; Represents the real positioning frame; Indicates filtering; The weighted sum of the classification loss function and the positioning loss function gives the detection and recognition loss function: , in, Represents the detection and recognition loss function; Represents the weight of the positioning loss function.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the fine-grained target detection and recognition method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Target recognition model training method and device and electronic equipment

    CN112990432A

  • High-precision fine-grained SAR (Synthetic Aperture Radar) target detection method

    CN115661569A