Single-stage fine-grained target detection method and device based on remote sensing image

By using a single-stage detector and a convolutional neural network feature extraction network in remote sensing image detection, combined with a positive and negative sample allocation algorithm, the problems of insufficient recognition accuracy and high computational complexity in fine-grained object detection in remote sensing images are solved, and high-precision and real-time object detection are achieved.

CN119992334AActive Publication Date: 2025-05-13TIANJIN UNIV

Patent Information

Application Number
CN202510116031.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-13
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient recognition accuracy and high computational complexity in fine-grained object detection of remote sensing images, which is difficult to meet the real-time requirements.

Method used

A single-stage detector-based method is used to construct a feature extraction network through a convolutional neural network, combines backbone network, feature pyramid and object detection head network, design a positive and negative sample allocation algorithm, and use the Pytorch deep learning framework for model training.

Benefits of technology

It improves the recognition accuracy and accuracy of fine-grained targets in remote sensing images, reduces the computational complexity, and enhances the real-time detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992334A_ABST
    Figure CN119992334A_ABST
Patent Text Reader

Abstract

The invention relates to a single-stage fine-grained target detection method and device based on a remote sensing image, and the method comprises the steps: obtaining image data and label data, and carrying out the image cutting and overturning of the image data, so as to construct a training set and a test set. And constructing a feature extraction network by taking the convolutional neural network as a detection model, calling the feature extraction network to perform feature extraction on the image data, and outputting position information and category information of a fine-grained target in the image. Dividing the position information frame into positive samples and negative samples through a positive and negative sample distribution algorithm, and calling a convolutional neural network to learn distribution of the positive samples and the negative samples; and based on a Pytorch deep learning framework, compiling a network structure through a Python language, and training the convolutional neural network through the image data and the label data in the training set to obtain a target detection model. And taking the to-be-detected image data as the input of the target detection model to output the category and position information of the target object in the to-be-detected image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and image processing, and in particular to a single-stage fine-grained target detection method and device based on remote sensing images. Background Art

[0002] Computer vision is an important branch of computer science and artificial intelligence, which aims to enable computers to "see" and "understand" image and video content like humans. It simulates the human visual system to extract information from images or videos, analyze and understand them, and thus achieve various visual tasks. Object detection is a core task in the field of computer vision, and its goal is to automatically identify the location and category of specific objects in images or videos. It is an important part of image recognition technology and is widely used in security monitoring, autonomous driving, robot vision, medical image analysis, remote sensing image processing and many other fields.

[0003] Object detectors based on deep learning can be divided into single-stage detectors (such as ATSS, SSD, YOLO, etc.) and two-stage detectors (such as Faster R-CNN, Mask R-CNN, etc.). The difference between the two is the way the location of the target area is presented. Among them, the single-stage detector uses an anchor box to locate the target, and the anchor box consists of a set of predefined bounding boxes located on the feature map. The two-stage detector generates several region proposals through an additional training stage before object detection, which improves the positioning accuracy by increasing the computational cost.

[0004] At present, with the continuous progress of my country's satellite technology and the significant improvement of remote sensing image quality, the development of remote sensing image fine-grained target detection technology has been strongly promoted. Remote sensing image fine-grained target detection is applied in many important fields, including military target positioning and identification, natural environment protection, disaster detection and urban planning and construction. For example, fine-grained target detection technology can accurately identify and locate different types of aircraft, ships, military facilities, etc. in military targets, while in the civilian field it can be used for resource exploration, environmental protection and disaster monitoring. Therefore, the development of fine-grained target detection technology is an inevitable choice to meet the needs of high-precision recognition and the challenges of complex scenes, and provides strong support for the widespread application of remote sensing images in military, civilian and environmental monitoring fields.

[0005] However, since the targets in remote sensing images often have characteristics such as multi-scale, multi-angle, and complex background, and different targets in the same fine-grained category may have significant differences in appearance, while targets between different categories may have similar features, it is difficult to accurately distinguish fine-grained targets in remote sensing images. Therefore, although there are many excellent models in the existing target detection field, their accuracy in identifying fine-grained targets still needs to be further improved, and it is still easy to have false detection or inaccurate positioning in actual detection. In addition, most application scenarios have strong requirements for the real-time performance of target detection. Therefore, the detection model needs to reduce the computational complexity as much as possible while ensuring the detection accuracy. At present, most detection models for fine-grained targets use innovations such as dual-stage target detectors and superimposed attention modules, which are difficult to meet real-time requirements. Summary of the invention

[0006] Based on this, it is necessary to provide a single-stage fine-grained target detection method and device based on remote sensing images, which has high accuracy and good real-time performance in identifying fine-grained targets in order to address the above technical problems.

[0007] The present invention provides a single-stage fine-grained target detection method based on remote sensing images, the method comprising: Obtain image data and label data corresponding to each image, and perform image cropping and flipping on the image data to construct a training set and a test set; Based on the training set and the test set, a convolutional neural network is used as a detection model to construct a feature extraction network, and the feature extraction network is called to extract features from the image data, and the position information and category information of the fine-grained target in the image are output; The position information frame is divided into positive samples and negative samples by a positive and negative sample allocation algorithm, and the convolutional neural network is called to learn the allocation of the positive samples and the negative samples; Based on the Pytorch deep learning framework, the network structure is written in Python language, and the convolutional neural network is trained with the image data and label data in the training set to obtain a target detection model; Using the image data to be detected as the input of the target detection model to output the category and position information of the target object in the image data to be detected; In which, the image data is a remote sensing image, the feature extraction network includes a backbone network, a feature pyramid and a target detection head network, the backbone network adopts a residual neural network to downsample the feature map; the feature pyramid is used to perform hierarchical feature extraction on the image data and generate multi-scale features, and the backbone network and the feature pyramid are connected from bottom to top and from top to bottom; the target detection head network consists of a classification branch for identifying the category of the target object in the image and a regression branch for locating the object in the image; the positive sample is a sample containing the target object in the image data, and the negative sample is a sample containing the background in the image data.

[0008] In one embodiment, based on the training set and the test set, a convolutional neural network is used as a detection model to construct a feature extraction network, and the feature extraction network is called to extract features from the image data, and the location information and category information of fine-grained targets in the image are output, including: Calling the backbone network to downsample the image data and initialize it with a pre-trained model from ImageNet; Performing hierarchical feature extraction on the image data through the feature pyramid to generate the multi-scale features, wherein the multi-scale features are used to adapt to target object detection at different resolutions; Calling the classification branch in the target detection head network to perform category recognition on the target object in the image data based on the label data, and positioning the target object through the regression branch to output the position and category information of the target object in the image, and forming multiple anchor frames at the same time; The anchor boxes are divided into anchor boxes containing target objects and anchor boxes not containing target objects. The anchor boxes containing target objects are positive samples, and the anchor boxes not containing target objects are negative samples.

[0009] In one embodiment, dividing the position information frame into positive samples and negative samples by a positive and negative sample allocation algorithm, and calling the convolutional neural network to learn the allocation of the positive samples and negative samples includes: Get a given bounding box of a target object in an image and rotate the bounding box using a two-dimensional Gaussian distribution representation based on a rotation matrix to calculate a covariance matrix; Calculating the mean of the two-dimensional Gaussian distribution according to the coordinates of the center point of the rotated bounding box, using the generalized Jensen-Shannon divergence to measure the distance, and mapping the generalized Jensen-Shannon divergence to a bounded similarity score within the first interval; The end values ​​of the first interval are respectively a first value and a second value, and when the bounded similarity score is closer to the first value, the similarity between the corresponding anchor box and the given bounding box is lower; when the bounded similarity score is closer to the second value, the similarity between the corresponding anchor box and the given bounding box is higher, and when the bounded similarity score is equal to the second value, the corresponding anchor box is the same as the given bounding box.

[0010] In one embodiment, the dividing the position information frame into positive samples and negative samples by a positive and negative sample allocation algorithm, and calling the convolutional neural network to learn the allocation of the positive samples and negative samples, further includes: Obtain an anchor frame set consisting of multiple anchor frames and a bounded similarity score of a classification head in each anchor frame in the anchor frame set, combine position information and category information of fine-grained objects in the image, and output the highest probability that the target object in each anchor frame belongs to the target category; Based on the highest probability that the target object in each anchor frame belongs to the target category, the anchor frames in the anchor frame set are sorted, and the higher the ranking of the anchor frame in the sorting, the more consistent the anchor frame is with the true annotation frame.

[0011] In one embodiment, the dividing the position information frame into positive samples and negative samples by a positive and negative sample allocation algorithm, and calling the convolutional neural network to learn the allocation of the positive samples and negative samples, further includes: The distance between the anchor box and the target object is modeled using Gaussian distribution, and the parameters of the Gaussian distribution are estimated based on the mean and variance of the distance samples to calculate the probability density function and cumulative distribution function of the Gaussian distribution. The quality of the anchor frame is evaluated based on the probability density function and the cumulative distribution function to obtain an evaluation value, and when the evaluation value reaches a set second threshold, anchor points related to the target position and having fine-grained identification features are filtered out to obtain a positive sample set.

[0012] In one embodiment, the dividing the position information frame into positive samples and negative samples by a positive and negative sample allocation algorithm, and calling the convolutional neural network to learn the allocation of the positive samples and negative samples, further includes: The convolutional neural network is called to learn the distribution of the positive samples and negative samples through the gradient back propagation of the set loss function, and the classification loss and positioning loss in the loss function are optimized by using Focal Loss and IoU Loss respectively.

[0013] In one embodiment, the network structure is written in Python based on the Pytorch deep learning framework, and the convolutional neural network is trained by the image data and label data in the training set to obtain the target detection model, including: Using the training set as the input of the convolutional neural network, obtaining the prediction result of the convolutional neural network, and calculating the loss between the prediction result and the label data through a loss function; Back-propagating the calculated loss so that the convolutional neural network can train and learn the target features, and save the optimal model weights to obtain the target detection model; The target features include location information and category information of fine-grained targets in the image and the distribution results of positive samples and negative samples.

[0014] The present invention also provides a single-stage fine-grained target detection device based on remote sensing images, the device comprising: A data set construction module, used to obtain image data and label data corresponding to each image, and perform image cropping and flipping on the image data to construct a training set and a test set; A feature extraction module is used to construct a feature extraction network based on the training set and the test set, using a convolutional neural network as a detection model, and call the feature extraction network to extract features from the image data, and output location information and category information of fine-grained targets in the image; A sample allocation module, used to divide the position information frame into positive samples and negative samples by a positive and negative sample allocation algorithm, and call the convolutional neural network to learn the allocation of the positive samples and negative samples; A detection model training module is used to write a network structure in Python based on the Pytorch deep learning framework, and train the convolutional neural network with the image data and label data in the training set to obtain a target detection model; A target detection module, used to use the image data to be detected as the input of the target detection model to output the category and position information of the target object in the image data to be detected; In which, the image data is a remote sensing image, the feature extraction network includes a backbone network, a feature pyramid and a target detection head network, the backbone network adopts a residual neural network to downsample the feature map; the feature pyramid is used to perform hierarchical feature extraction on the image data and generate multi-scale features, and the backbone network and the feature pyramid are connected from bottom to top and from top to bottom; the target detection head network consists of a classification branch for identifying the category of the target object in the image and a regression branch for locating the object in the image; the positive sample is a sample containing the target object in the image data, and the negative sample is a sample containing the background in the image data.

[0015] The present invention also provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the single-stage fine-grained target detection method based on remote sensing images as described in any one of the above.

[0016] The present invention also provides a computer storage medium storing a computer program, wherein when the computer program is executed by a processor, the single-stage fine-grained target detection method based on remote sensing images as described in any one of the above is implemented.

[0017] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the single-stage fine-grained target detection method based on remote sensing images as described in any one of the above.

[0018] The above-mentioned single-stage fine-grained target detection method and device based on remote sensing images obtains remote sensing images and label data corresponding to each image, and performs image cropping and flipping on the remote sensing images to construct training sets and test sets. Subsequently, based on the constructed training sets and test sets, a feature extraction network is constructed with a convolutional neural network as the detection model, and the feature extraction network is called to extract features from the remote sensing images, and the location information and category information of the fine-grained targets in the images are output. After that, the location information frame is divided into positive samples and negative samples by the positive and negative sample allocation algorithm, and the convolutional neural network is called to learn the allocation of positive samples and negative samples. Based on the Pytorch deep learning framework, the network structure is written in Python language, and the convolutional neural network is trained by the remote sensing images and label data in the training set to obtain the target detection model. Finally, the image data to be detected is used as the input of the target detection model to output the category and location information of the target object in the image data to be detected. This method improves the model's ability to learn fine-grained target feature information by designing a positive and negative sample allocation algorithm, thereby improving the precision and accuracy of target detection, and effectively reducing the single-stage target detector's category detection errors for fine-grained targets. The algorithm process is relatively simple and has good real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 One of the flow charts of the single-stage fine-grained target detection method based on remote sensing images provided by the present invention; Figure 2A schematic diagram of a target detection architecture of a single-stage fine-grained target detection method based on remote sensing images in a specific embodiment provided by the present invention; Figure 3 A schematic diagram of a sample allocation process of a single-stage fine-grained target detection method based on remote sensing images in a specific embodiment provided by the present invention; Figure 4 The second flowchart of the single-stage fine-grained target detection method based on remote sensing images provided by the present invention; Figure 5 The third flowchart of the single-stage fine-grained target detection method based on remote sensing images provided by the present invention; Figure 6 The fourth flowchart of the single-stage fine-grained target detection method based on remote sensing images provided by the present invention; Figure 7 The fifth flowchart of the single-stage fine-grained target detection method based on remote sensing images provided by the present invention; Figure 8 The sixth flowchart of the single-stage fine-grained target detection method based on remote sensing images provided by the present invention; Fig. 9 A schematic diagram of the structure of a single-stage fine-grained target detection device based on remote sensing images provided by the present invention; Fig.10 This is a diagram of the internal structure of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0022] Combine the following Figures 1 to 10 The present invention describes a single-stage fine-grained target detection method and device based on remote sensing images.

[0023] like Figure 1 As shown, in one embodiment, a single-stage fine-grained target detection method based on remote sensing images includes the following steps: Step S110, acquiring image data and label data corresponding to each image, and performing image cropping and flipping on the image data to construct a training set and a test set.

[0024] Among them, the image data is a remote sensing image.

[0025] Specifically, the workstation server obtains remote sensing image data and corresponding label data in each image, and performs image cropping and flipping on the remote sensing images to construct a training set and a test set containing remote sensing images and their label data.

[0026] Combination Figure 2 As shown, in a specific embodiment, the single-stage fine-grained target detection method based on remote sensing images provided by the present invention first reads remote sensing image data by a computer. In the process of the computer reading data, two fine-grained remote sensing data sets are used, including the VEDAI data set and the HRSC2016 data set. Among them, the VEDAI data set is a remote sensing vehicle fine-grained target detection data set containing RGB images and infrared aerial images. In this example, only RGB images are used, and the resolution of each image is 1024×1024. Categories with less than 50 instances are excluded, and the remaining 8 target categories are: cars, pickups, campers, trucks, tractors, ships, vans and others. The HRSC2016 data set is a remote sensing ship fine-grained data set, which contains a total of 1061 remote sensing images with resolutions ranging from 300×300 to 1500×900. The HRSC2016 data set provides three types of labels. This example uses the most fine-grained labels to meet the requirements of fine-grained target detection tasks. According to the label, the objects are classified into the following 19 types: submarine (Sub.), Nimitz (Nim.), Enterprise (Ent.), Arleigh Burke (Arl.), White Bay Island (Whi.), Perry (Per.), San Antonio (San.), Ticonderoga (Tic.), Admiral (Adm.), Austin (Aus.), Tarawa (Tar.), container (Con.), command ship A (Com.A), car carrier A (Car.A), container ship A (Con.A), medical ship (Med.), car carrier B (Car.B), Midway (Mid) and Root (Inv.). During the training process, the training set is input into the neural network and a certain amount of data augmentation is performed on it.

[0027] Step S120, based on the training set and the test set, using the convolutional neural network as the detection model, constructing a feature extraction network, and calling the feature extraction network to extract features from the image data, and outputting the location information and category information of the fine-grained targets in the image.

[0028] The feature extraction network includes a backbone network, a feature pyramid, and a target detection head network. The backbone network uses a residual neural network to downsample the feature map. The feature pyramid is used to extract hierarchical features from image data and generate multi-scale features. The backbone network and the feature pyramid are connected from bottom to top and from top to bottom. The target detection head network consists of a classification branch for identifying the category of the target object in the image and a regression branch for locating the object in the image.

[0029] Specifically, based on the training set and test set constructed in step S110, the workstation server uses a convolutional neural network as a detection model to construct a feature extraction network, and calls the constructed feature extraction network to extract features from the remote sensing image, and outputs the location information and category information of fine-grained targets in the remote sensing image through the target detection head network.

[0030] Combination Figure 2 As shown, in a specific embodiment, the single-stage fine-grained target detection method based on remote sensing images provided by the present invention, the target detection structure diagram proposed in this example, mainly includes two parts: feature extraction and data processing. Among them, feature extraction is composed of a backbone network feature pyramid network and a target detection head network. The backbone network adopts a residual neural network ResNet-50, which is used to downsample the feature map and is initialized using a pre-trained model from ImageNet. The feature pyramid network is responsible for hierarchical feature extraction and generates multi-scale features to adapt to object detection at different resolutions. The target detection head network is the last part of feature extraction and consists of two branches: a classification branch for class recognition and a regression branch for object positioning.

[0031] Step S130, dividing the position information frame into positive samples and negative samples through a positive and negative sample allocation algorithm, and calling a convolutional neural network to learn the allocation of positive samples and negative samples.

[0032] Among them, the positive sample is the sample containing the target object in the image data, and the negative sample is the sample containing the background in the image data.

[0033] Specifically, the workstation server divides the position information frame outputted in step S120 into positive samples containing the target object in the image and negative samples containing the background in the image through a positive and negative sample allocation algorithm.

[0034] Combination Figure 2 and Figure 3As shown, in a specific embodiment, the single-stage fine-grained target detection method based on remote sensing images provided by the present invention, after passing through the feature extraction network, the output of the target detection head network is the position and category information of the target in the image, and a series of anchor frames are formed. Since these anchor frames may contain objects or may not contain objects, positive samples (samples containing target objects) are used to allow the model to learn how to identify the features and positions of target objects, while negative samples (samples that do not contain target objects) are used to allow the model to learn to distinguish the background. Therefore, it is necessary to assign positive and negative samples to the anchor frames so that the model can better understand which areas are target objects and which areas are backgrounds, thereby improving the accuracy and robustness of detection. In the prior art, label assignment is performed by emphasizing the positional relationship between the anchor frame and the target object, for example, the intersection over union (IoU) of the anchor frame and the true annotation frame is used to measure the quality of anchor frame detection. If the intersection over union is greater than the set threshold, the anchor frame is considered to be a positive sample, otherwise it is a negative sample. However, in the task of fine-grained object detection, the discriminative features that distinguish objects usually exist in the local position of the object. If we still only use the anchor box position information for quality assessment, the anchor box will find it difficult to capture the subtle differences between objects of different classes. A large number of low-quality (including non-highly discriminative features of fine-grained objects) may still be identified as valid positive samples, thereby increasing the possibility of false detection.

[0035] In this embodiment, the label assignment process is optimized by developing a comprehensive score to evaluate the anchor boxes, which is affected not only by the target location information but also by the classification prediction of the model. Given the key role played by local high-discriminative features in improving the accuracy of category prediction in fine-grained recognition tasks, combining the prediction scores helps to obtain high-quality (covering the local high-discriminative features of fine-grained targets) positive samples. The scoring mechanism of this example prioritizes high-quality positive anchor boxes over low-quality positive anchor boxes, and reduces the possibility of false detection by utilizing the discriminative features present in these anchor boxes.

[0036] The process of obtaining the positive sample set is mainly divided into four steps. The first step is to measure the distance from the anchor box to the target as the position quality of the anchor box, and use the quality score for screening; the second step is to use the prediction score of the anchor box for ranking; the third step is to calculate the fine-grained score based on sorting, and use the score to re-screen; the fourth step is to use the positive sample sets obtained in the first and third steps and take the union to obtain the positive sample set, and the rest are used as negative samples, thereby completing the allocation of positive and negative samples.

[0037] In this embodiment, the target ground truth box in the remote sensing image is annotated using a rotated bounding box. Based on this configuration, we use the generalized Jensen-Shannon divergence (GJSD) as a metric to evaluate the distance between the target ground truth box and the anchor box. There are two main reasons for using the generalized Jensen-Shannon divergence: First, for rotated targets, the intersection-over-union metric is very sensitive to angle changes, while the divergence-based metric provides a more robust evaluation. Second, unlike the Kullback-Leibler Divergence (KLD), the generalized Jensen-Shannon divergence can still be used even when the distributions do not completely overlap.

[0038] In this embodiment, based on a given bounding box , first according to the rotation matrix Using a 2D Gaussian distribution Represents the bounding box and calculates the covariance matrix , whose expression is:

[0039] In the formula, is the rotation angle of the bounding box, is the height of the bounding box, is the rotation speed of the bounding box, are the 2D coordinate positions of the bounding box respectively.

[0040] According to the center point coordinates of the bounding box (rotated bounding box) ( ), we can get the mean of the two-dimensional Gaussian distribution , whose expression is: .

[0041] In addition, since GJSD itself is a divergence measure rather than a similarity measure, in order to use the generalized Jensen-Shannon divergence for distance measurement, it is necessary to transform GJSD to a bounded similarity score in the range of (0,1], where the mid-range value 1 represents exactly the same distribution, and values ​​close to 0 represent the maximum dissimilarity, and the bounded similarity score is equal to 1 plus the inverse of the GJSD value.

[0042] For two distributions and , their GJSD is defined as:

[0043] Where KLD represents the Courbet-Léibler divergence, is the average (mixed) distribution of P and Q.

[0044]

[0045]

[0046] Based on the above reasoning, the i-th anchor box can be calculated and the jth target object The distance between , whose expression is:

[0047] This method can be used to filter the anchor boxes related to the target location. represents the number of anchor boxes, Indicates the number of targets.

[0048] Since the concept of fine-grained targets is related to the detailed classification of objects, it is crucial to evaluate the quality of anchor boxes using the prediction score of each anchor box in the fine-grained target detection task. In addition, since the prediction score of the anchor point is continuously updated during the training phase as the object detector is optimized, it is better to use a ranking function instead of directly relying on a fixed threshold to evaluate the quality score of the positive anchor point, so as to reduce the potential problem caused by too many or insufficient positive samples due to the use of inappropriate score thresholds.

[0049] Given a set of anchor boxes , each anchor box Will get the prediction score of the classification head, expressed as ,in Anchor frame The corresponding features, Indicates the target category. The highest probability among them is:

[0050] Anchor Box The category is classified as Target, belonging to ,Right now The position of the target is used to filter the anchor boxes related to the target position and obtain the anchor box set :

[0051] and The order of the anchor boxes is as follows:

[0052] in, Represents the anchor box By predicting the score The ranking obtained, anchor box The ranking is included in Within, expressed as . The ranking subset is A part of, each element in the subset corresponds to an anchor box exist Ranking in . The higher the value, the stronger the anchor box The higher the degree of consistency with the target real frame in terms of positioning accuracy. , low-quality anchor boxes will get lower ranking numbers during ranking.

[0053] In this embodiment, anchor boxes with fine-grained discriminative feature regions can obtain higher prediction scores than anchor boxes lacking fine-grained discriminative feature regions. Therefore, we use a sorting operation to organize anchor boxes according to the prediction scores, and the sorting results indicate the extent to which the anchor boxes contain fine-grained discriminative features.

[0054] and They are used to provide fine-grained scores from the perspective of classification and regression, respectively, and their relationship is:

[0055] and Have a common normalization term , indicating that the scores for classification and regression are based on Since the ranking order shows a continuous upward trend, when the prediction score of the anchor point is low, both the classification and regression scores will decrease. This proves that and Strong correlation with fine-grained discriminative features. Therefore, is considered as a fine-grained score in the label assignment process.

[0056] Combined with the above evaluation methods, in order to perform statistical analysis, the distance between the anchor box and the target object is first modeled using Gaussian distribution, and the parameters of the Gaussian distribution are estimated using the sample mean and variance. After the parameters are estimated, the probability density function of the Gaussian distribution is cumulatively distributed. In addition, to evaluate the significance of the optimal threshold, a suitable standard value can be given to determine the optimal threshold. In this example, the standard value can be 0.9.

[0057] In this embodiment, the positive sample set should also include anchor points with fine-grained discriminative features and the top k positive sample anchor frames, which together constitute the positive sample set. In deep learning, network parameters are back-propagated through the gradient of the set loss function to complete the learning of target features. In the loss function of the model in this example, the classification loss is optimized using FocalLoss to solve the imbalance problem in the object detection training stage, and the positioning loss is optimized using IoU Loss.

[0058] Step S140, based on the Pytorch deep learning framework, write the network structure in Python language, and train the convolutional neural network with the image data and label data in the training set to obtain the target detection model.

[0059] Specifically, the workstation server is based on the Pytorch deep learning framework, uses Python to write the network structure, and trains the convolutional neural network with image data and label data in the training set to obtain the target detection model.

[0060] In a specific embodiment, the single-stage fine-grained target detection method based on remote sensing images provided by the present invention, in the process of training a convolutional neural network, after constructing a detection model, the image training set is input into the model under the Pytorch deep learning framework to obtain the predicted value of the neural network. After the predicted value and the loss of the true label of the training data are obtained through the loss function, the loss is reversely propagated by gradient, thereby achieving the effect of target feature learning by the network. When the training is stable, the optimal model weights are saved.

[0061] Step S150: using the image data to be detected as input of the target detection model to output the category and position information of the target object in the image data to be detected.

[0062] Specifically, the workstation server receives the remote sensing image to be detected, calls the trained target detection model to identify and process it, and then outputs the category and location information of the target object in the remote sensing image to be detected, and draws it in the image.

[0063] In a specific embodiment, the single-stage fine-grained target detection method based on remote sensing images provided by the present invention inputs the image to be detected into a trained neural network, obtains the category and location information of the target, and draws it in the image.

[0064] In this embodiment, two fine-grained remote sensing datasets are used, including the VEDAI dataset and the HRSC2016 dataset. The results obtained from the verification on the two datasets show that the detection accuracy of the target detection algorithm designed in this example is improved to a certain extent compared with the comparison algorithm, thereby proving the effectiveness of this example for fine-grained target detection. The results on the two datasets are respectively shown in Table 1 and Table 2, where Table 1 is the comparison results on the remote sensing vehicle fine-grained target detection dataset VEDAI, and Table 2 is the comparison results on the remote sensing ship fine-grained target detection dataset HRSC2016 dataset: Table 1 Table 2 As shown in the table, the model designed in this example can effectively detect fine-grained targets in images with high positioning accuracy.

[0065] The above-mentioned single-stage fine-grained target detection method based on remote sensing images obtains remote sensing images and label data corresponding to each image, and performs image cropping and flipping on the remote sensing images to construct training sets and test sets. Subsequently, based on the constructed training sets and test sets, a feature extraction network is constructed with a convolutional neural network as the detection model, and the feature extraction network is called to extract features from the remote sensing images, and the location information and category information of the fine-grained targets in the images are output. After that, the location information frame is divided into positive samples and negative samples by the positive and negative sample allocation algorithm, and the convolutional neural network is called to learn the allocation of positive samples and negative samples. Based on the Pytorch deep learning framework, the network structure is written in Python language, and the convolutional neural network is trained by the remote sensing images and label data in the training set to obtain the target detection model. Finally, the image data to be detected is used as the input of the target detection model to output the category and location information of the target object in the image data to be detected. This method improves the model's ability to learn fine-grained target feature information by designing a positive and negative sample allocation algorithm, thereby improving the precision and accuracy of target detection, and effectively reducing the single-stage target detector's category detection errors for fine-grained targets. The algorithm process is relatively simple and has good real-time performance.

[0066] like Figure 4 As shown, in one embodiment, the single-stage fine-grained target detection method based on remote sensing images provided by the present invention is based on a training set and a test set, uses a convolutional neural network as a detection model, constructs a feature extraction network, and calls the feature extraction network to extract features from image data, and outputs the location information and category information of the fine-grained target in the image, specifically including the following steps: Step S121, calling the backbone network to downsample the image data, and initializing it through a pre-trained model from ImageNet.

[0067] Step S122 , performing hierarchical feature extraction on the image data through a feature pyramid to generate multi-scale features, and the multi-scale features are used to adapt to target object detection at different resolutions.

[0068] Step S123, calling the classification branch in the target detection head network to identify the category of the target object in the image data based on the label data, and locating the target object through the regression branch to output the position and category information of the target object in the image, and forming multiple anchor frames at the same time.

[0069] Among them, the anchor boxes are divided into anchor boxes containing target objects and anchor boxes not containing target objects. The anchor boxes containing target objects are positive samples, and the anchor boxes not containing target objects are negative samples.

[0070] like Figure 5 As shown, in one embodiment, the single-stage fine-grained target detection method based on remote sensing images provided by the present invention divides the location information frame into positive samples and negative samples through a positive and negative sample allocation algorithm, and calls a convolutional neural network to learn the allocation of positive samples and negative samples, specifically including the following steps: Step S131 , obtaining a given bounding box of a target object in an image, and rotating the bounding box using a two-dimensional Gaussian distribution representation based on a rotation matrix to calculate a covariance matrix.

[0071] Step S132, calculating the mean of the two-dimensional Gaussian distribution according to the coordinates of the center point of the rotated bounding box, using the generalized Jensen-Shannon divergence to measure the distance, and mapping the generalized Jensen-Shannon divergence to a bounded similarity score within the first interval.

[0072] The end values ​​of the first interval are the first value and the second value, respectively. When the bounded similarity score is closer to the first value, the corresponding anchor box has a lower similarity with the given bounding box. When the bounded similarity score is closer to the second value, the corresponding anchor box has a higher similarity with the given bounding box, and when the bounded similarity score is equal to the second value, the corresponding anchor box is the same as the given bounding box.

[0073] like Figure 6 As shown, in one embodiment, the single-stage fine-grained target detection method based on remote sensing images provided by the present invention divides the location information frame into positive samples and negative samples through a positive and negative sample allocation algorithm, and calls a convolutional neural network to learn the allocation of positive samples and negative samples, and specifically includes the following steps: Step S133, obtaining an anchor frame set consisting of multiple anchor frames and a bounded similarity score of a classification head in each anchor frame in the anchor frame set, combining the position information and category information of fine-grained objects in the image, and outputting the highest probability that the target object in each anchor frame belongs to the target category.

[0074] Step S134, based on the highest probability that the target object in each anchor frame belongs to the target category, the anchor frames in the anchor frame set are sorted, and the higher the ranking of the anchor frame in the sorting, the more consistent the anchor frame is with the true annotation frame.

[0075] like Figure 7 As shown, in one embodiment, the single-stage fine-grained target detection method based on remote sensing images provided by the present invention divides the location information frame into positive samples and negative samples through a positive and negative sample allocation algorithm, and calls a convolutional neural network to learn the allocation of positive samples and negative samples, and specifically includes the following steps: Step S135 , using Gaussian distribution to model the distance between the anchor box and the target object, and estimating the parameters of the Gaussian distribution according to the mean and variance of the distance samples to calculate the probability density function and cumulative distribution function of the Gaussian distribution.

[0076] Step S136, evaluating the quality of the anchor frame based on the probability density function and the cumulative distribution function to obtain an evaluation value, and when the evaluation value reaches a set second threshold, filtering out anchor points related to the target position and having fine-grained identification features to obtain a positive sample set.

[0077] Step S137, calling the convolutional neural network to learn the distribution of positive samples and negative samples through the gradient back propagation of the set loss function, and using Focal Loss and IoU Loss to optimize the classification loss and positioning loss in the loss function respectively.

[0078] like Figure 8 As shown, in one embodiment, the single-stage fine-grained target detection method based on remote sensing images provided by the present invention is based on the Pytorch deep learning framework, the network structure is written in Python language, and the convolutional neural network is trained by the image data and label data in the training set to obtain the target detection model, which specifically includes the following steps: Step S141, using the training set as the input of the convolutional neural network to obtain the prediction result of the convolutional neural network, and calculating the loss between the prediction result and the label data through the loss function.

[0079] Step S142, back-propagating the calculated loss so that the convolutional neural network can train and learn the target features, and save the optimal model weights to obtain the target detection model.

[0080] Among them, the target features include the location information and category information of fine-grained targets in the image and the distribution results of positive samples and negative samples.

[0081] The single-stage fine-grained target detection device based on remote sensing images provided by the present invention is described below. The single-stage fine-grained target detection device based on remote sensing images described below and the single-stage fine-grained target detection method based on remote sensing images described above can be referenced to each other.

[0082] like Fig. 9 As shown, in one embodiment, a single-stage fine-grained target detection device based on remote sensing images includes a data set construction module 910, a feature extraction module 920, a sample allocation module 930, a detection model training module 940 and a target detection module 950.

[0083] The data set construction module 910 is used to obtain image data and label data corresponding to each image, and perform image cropping and flipping on the image data to construct a training set and a test set.

[0084] The feature extraction module 920 is used to construct a feature extraction network based on the training set and the test set, using the convolutional neural network as the detection model, and call the feature extraction network to extract features from the image data, and output the location information and category information of fine-grained targets in the image.

[0085] The sample allocation module 930 is used to divide the position information frame into positive samples and negative samples through a positive and negative sample allocation algorithm, and call a convolutional neural network to learn the allocation of positive samples and negative samples.

[0086] The detection model training module 940 is used to write the network structure in Python based on the Pytorch deep learning framework, and train the convolutional neural network with the image data and label data in the training set to obtain the target detection model.

[0087] The target detection module 950 is used to use the image data to be detected as the input of the target detection model to output the category and position information of the target object in the image data to be detected.

[0088] The image data is a remote sensing image, and the feature extraction network includes a backbone network, a feature pyramid, and a target detection head network. The backbone network uses a residual neural network to downsample the feature map. The feature pyramid is used to extract hierarchical features from the image data and generate multi-scale features, and the backbone network and the feature pyramid are connected from bottom to top and from top to bottom. The target detection head network consists of a classification branch for identifying the category of the target object in the image and a regression branch for locating the object in the image. The positive sample is a sample containing the target object in the image data, and the negative sample is a sample containing the background in the image data.

[0089] In this embodiment, the single-stage fine-grained target detection device based on remote sensing images provided by the present invention, the feature extraction module 920 is specifically used for: The backbone network is called to downsample the image data and initialized with a pre-trained model from ImageNet.

[0090] Hierarchical feature extraction is performed on image data through a feature pyramid to generate multi-scale features, which are used to adapt to target object detection at different resolutions.

[0091] The classification branch in the target detection head network is called to identify the category of the target object in the image data based on the label data, and the target object is located through the regression branch to output the position and category information of the target object in the image and form multiple anchor boxes at the same time.

[0092] Among them, the anchor boxes are divided into anchor boxes containing target objects and anchor boxes not containing target objects. The anchor boxes containing target objects are positive samples, and the anchor boxes not containing target objects are negative samples.

[0093] In this embodiment, the single-stage fine-grained target detection device based on remote sensing images provided by the present invention, the sample allocation module 930 is specifically used for: Get a given bounding box of a target object in an image and rotate the bounding box using a 2D Gaussian distribution representation based on a rotation matrix to calculate the covariance matrix.

[0094] The mean of the two-dimensional Gaussian distribution is calculated according to the coordinates of the center point of the rotated bounding box, and the generalized Jensen-Shannon divergence is used for distance measurement, and the generalized Jensen-Shannon divergence is mapped to a bounded similarity score within a first interval.

[0095] The end values ​​of the first interval are the first value and the second value, respectively. When the bounded similarity score is closer to the first value, the corresponding anchor box has a lower similarity with the given bounding box. When the bounded similarity score is closer to the second value, the corresponding anchor box has a higher similarity with the given bounding box, and when the bounded similarity score is equal to the second value, the corresponding anchor box is the same as the given bounding box.

[0096] In this embodiment, in the single-stage fine-grained target detection device based on remote sensing images provided by the present invention, the sample allocation module 930 is further used to: Obtain an anchor box set consisting of multiple anchor boxes and the bounded similarity score of the classification head in each anchor box in the anchor box set, combine the location information and category information of the fine-grained target in the image, and output the highest probability that the target object in each anchor box belongs to the target category.

[0097] Based on the highest probability that the target object in each anchor frame belongs to the target category, the anchor frames in the anchor frame set are sorted, and the higher the ranking of the anchor frame in the sorting, the more consistent the anchor frame is with the true annotation frame.

[0098] In this embodiment, in the single-stage fine-grained target detection device based on remote sensing images provided by the present invention, the sample allocation module 930 is further used to: The distance between the anchor box and the target object is modeled using Gaussian distribution, and the parameters of the Gaussian distribution are estimated based on the mean and variance of the distance samples to calculate the probability density function and cumulative distribution function of the Gaussian distribution.

[0099] The quality of the anchor frame is evaluated based on the probability density function and the cumulative distribution function to obtain an evaluation value, and when the evaluation value reaches a set second threshold, the anchor points related to the target position and having fine-grained identification features are filtered out to obtain a positive sample set.

[0100] In this embodiment, in the single-stage fine-grained target detection device based on remote sensing images provided by the present invention, the sample allocation module 930 is further used to: The convolutional neural network is called to learn the distribution of positive and negative samples through the gradient back propagation of the set loss function, and Focal Loss and IoU Loss are used to optimize the classification loss and positioning loss in the loss function respectively.

[0101] In this embodiment, the single-stage fine-grained target detection device based on remote sensing images provided by the present invention, the detection model training module 940 is specifically used for: The training set is used as the input of the convolutional neural network to obtain the prediction results of the convolutional neural network, and the loss between the prediction results and the label data is calculated through the loss function.

[0102] The calculated loss is back-propagated so that the convolutional neural network can train and learn the target features and save the optimal model weights to obtain the target detection model.

[0103] Among them, the target features include the location information and category information of fine-grained targets in the image and the distribution results of positive samples and negative samples.

[0104] Fig.10 An example of a physical structure diagram of an electronic device is shown. The electronic device may be a smart terminal, and its internal structure diagram may be as follows: Fig.10As shown. The electronic device includes a processor, an internal memory and a network interface connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a single-stage fine-grained target detection method based on remote sensing images is implemented, and the method includes: Obtain image data and label data corresponding to each image, and perform image cropping and flipping on the image data to construct a training set and a test set; Based on the training set and the test set, a convolutional neural network is used as the detection model to construct a feature extraction network, which is then called to extract features from the image data and output the location information and category information of fine-grained targets in the image. The location information frame is divided into positive samples and negative samples through the positive and negative sample allocation algorithm, and the convolutional neural network is called to learn the allocation of positive samples and negative samples; Based on the Pytorch deep learning framework, the network structure is written in Python, and the convolutional neural network is trained with the image data and label data in the training set to obtain the target detection model; The image data to be detected is used as the input of the target detection model to output the category and location information of the target object in the image data to be detected; Among them, the image data is a remote sensing image, the feature extraction network includes a backbone network, a feature pyramid and a target detection head network, the backbone network adopts a residual neural network to downsample the feature map; the feature pyramid is used to perform hierarchical feature extraction on the image data and generate multi-scale features, and the backbone network and the feature pyramid are connected from bottom to top and from top to bottom; the target detection head network consists of a classification branch for identifying the category of the target object in the image and a regression branch for locating the object in the image; the positive sample is a sample containing the target object in the image data, and the negative sample is a sample containing the background in the image data.

[0105] Those skilled in the art will understand that Fig.10 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present invention, and does not constitute a limitation on the electronic device to which the scheme of the present invention is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0106] On the other hand, the present invention also provides a computer storage medium storing a computer program, which implements the above-mentioned single-stage fine-grained target detection method based on remote sensing images when executed by a processor.

[0107] In another aspect, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the above-mentioned single-stage fine-grained target detection method based on remote sensing images is implemented.

[0108] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory.

[0109] By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0110] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0111] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A single-stage fine-grained target detection method based on remote sensing images, characterized in that: The method comprises: Obtain image data and label data corresponding to each image, and perform image cropping and flipping on the image data to construct a training set and a test set; Based on the training set and the test set, a convolutional neural network is used as a detection model to construct a feature extraction network, and the feature extraction network is called to extract features from the image data, and the position information and category information of the fine-grained target in the image are output; The position information frame is divided into positive samples and negative samples by a positive and negative sample allocation algorithm, and the convolutional neural network is called to learn the allocation of the positive samples and the negative samples; Based on the Pytorch deep learning framework, the network structure is written in Python language, and the convolutional neural network is trained with the image data and label data in the training set to obtain a target detection model; Using the image data to be detected as the input of the target detection model to output the category and position information of the target object in the image data to be detected; In which, the image data is a remote sensing image, the feature extraction network includes a backbone network, a feature pyramid and a target detection head network, the backbone network adopts a residual neural network to downsample the feature map; the feature pyramid is used to perform hierarchical feature extraction on the image data and generate multi-scale features, and the backbone network and the feature pyramid are connected from bottom to top and from top to bottom; the target detection head network consists of a classification branch for identifying the category of the target object in the image and a regression branch for locating the object in the image; the positive sample is a sample containing the target object in the image data, and the negative sample is a sample containing the background in the image data.

2. The single-stage fine-grained target detection method based on remote sensing images according to claim 1 is characterized in that: Based on the training set and the test set, a convolutional neural network is used as a detection model to construct a feature extraction network, and the feature extraction network is called to extract features from the image data, and the location information and category information of fine-grained targets in the image are output, including: Calling the backbone network to downsample the image data and initialize it with a pre-trained model from ImageNet; Performing hierarchical feature extraction on the image data through the feature pyramid to generate the multi-scale features, wherein the multi-scale features are used to adapt to target object detection at different resolutions; Calling the classification branch in the target detection head network to perform category recognition on the target object in the image data based on the label data, and positioning the target object through the regression branch to output the position and category information of the target object in the image, and forming multiple anchor frames at the same time; The anchor boxes are divided into anchor boxes containing target objects and anchor boxes not containing target objects. The anchor boxes containing target objects are positive samples, and the anchor boxes not containing target objects are negative samples.

3. The single-stage fine-grained target detection method based on remote sensing images according to claim 2 is characterized in that: The dividing the position information frame into positive samples and negative samples by a positive and negative sample allocation algorithm, and calling the convolutional neural network to learn the allocation of the positive samples and the negative samples, includes: Get a given bounding box of a target object in an image and rotate the bounding box using a two-dimensional Gaussian distribution representation based on a rotation matrix to calculate a covariance matrix; Calculating the mean of the two-dimensional Gaussian distribution according to the coordinates of the center point of the rotated bounding box, using the generalized Jensen-Shannon divergence to measure the distance, and mapping the generalized Jensen-Shannon divergence to a bounded similarity score within the first interval; The end values ​​of the first interval are respectively a first value and a second value, and when the bounded similarity score is closer to the first value, the similarity between the corresponding anchor box and the given bounding box is lower; when the bounded similarity score is closer to the second value, the similarity between the corresponding anchor box and the given bounding box is higher, and when the bounded similarity score is equal to the second value, the corresponding anchor box is the same as the given bounding box.

4. The single-stage fine-grained target detection method based on remote sensing images according to claim 3 is characterized in that: The method of dividing the position information frame into positive samples and negative samples by a positive and negative sample allocation algorithm, and calling the convolutional neural network to learn the allocation of the positive samples and negative samples, further includes: Obtain an anchor frame set consisting of multiple anchor frames and a bounded similarity score of a classification head in each anchor frame in the anchor frame set, combine position information and category information of fine-grained objects in the image, and output the highest probability that the target object in each anchor frame belongs to the target category; Based on the highest probability that the target object in each anchor frame belongs to the target category, the anchor frames in the anchor frame set are sorted, and the higher the ranking of the anchor frame in the sorting, the more consistent the anchor frame is with the true annotation frame.

5. The single-stage fine-grained target detection method based on remote sensing images according to claim 4 is characterized in that: The method of dividing the position information frame into positive samples and negative samples by a positive and negative sample allocation algorithm, and calling the convolutional neural network to learn the allocation of the positive samples and negative samples, further includes: The distance between the anchor box and the target object is modeled using Gaussian distribution, and the parameters of the Gaussian distribution are estimated based on the mean and variance of the distance samples to calculate the probability density function and cumulative distribution function of the Gaussian distribution. The quality of the anchor frame is evaluated based on the probability density function and the cumulative distribution function to obtain an evaluation value, and when the evaluation value reaches a set second threshold, anchor points related to the target position and having fine-grained identification features are filtered out to obtain a positive sample set.

6. The single-stage fine-grained target detection method based on remote sensing images according to claim 5 is characterized in that: The method of dividing the position information frame into positive samples and negative samples by a positive and negative sample allocation algorithm, and calling the convolutional neural network to learn the allocation of the positive samples and negative samples, further includes: The convolutional neural network is called to learn the distribution of the positive samples and negative samples through the gradient back propagation of the set loss function, and the classification loss and positioning loss in the loss function are optimized by using Focal Loss and IoU Loss respectively.

7. The single-stage fine-grained target detection method based on remote sensing images according to claim 6 is characterized in that: The network structure is written in Python based on the Pytorch deep learning framework, and the convolutional neural network is trained using the image data and label data in the training set to obtain the target detection model, including: Using the training set as the input of the convolutional neural network, obtaining the prediction result of the convolutional neural network, and calculating the loss between the prediction result and the label data through a loss function; Back-propagating the calculated loss so that the convolutional neural network can train and learn the target features, and save the optimal model weights to obtain the target detection model; The target features include location information and category information of fine-grained targets in the image and the distribution results of positive samples and negative samples.

8. A single-stage fine-grained target detection device based on remote sensing images, characterized in that: The device comprises: A data set construction module, used to obtain image data and label data corresponding to each image, and perform image cropping and flipping on the image data to construct a training set and a test set; A feature extraction module is used to construct a feature extraction network based on the training set and the test set, using a convolutional neural network as a detection model, and call the feature extraction network to extract features from the image data, and output location information and category information of fine-grained targets in the image; A sample allocation module, used to divide the position information frame into positive samples and negative samples by a positive and negative sample allocation algorithm, and call the convolutional neural network to learn the allocation of the positive samples and negative samples; A detection model training module is used to write a network structure in Python based on the Pytorch deep learning framework, and train the convolutional neural network with the image data and label data in the training set to obtain a target detection model; A target detection module, used to use the image data to be detected as the input of the target detection model to output the category and position information of the target object in the image data to be detected; In which, the image data is a remote sensing image, the feature extraction network includes a backbone network, a feature pyramid and a target detection head network, the backbone network adopts a residual neural network to downsample the feature map; the feature pyramid is used to perform hierarchical feature extraction on the image data and generate multi-scale features, and the backbone network and the feature pyramid are connected from bottom to top and from top to bottom; the target detection head network consists of a classification branch for identifying the category of the target object in the image and a regression branch for locating the object in the image; the positive sample is a sample containing the target object in the image data, and the negative sample is a sample containing the background in the image data.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Remote sensing image target detection method based on new frame regression loss function

    CN111091105A

  • Rotating frame remote sensing target detection method based on lightweight deep neural network

    CN114005045A

  • SAR ship target rotation detection method and system

    CN116310837A

  • Remote sensing image fine-grained target detection method and system based on feature balance

    CN118570447A

  • Improved rotating small target detection method

    CN118823315A

Cited By

  • Bidirectional guide remote sensing image fine-grained target detection method

    CN121962956A