Remote Sensing Image Target Detection Method and System Based on Consistency Relationship Reasoning

By adopting a method based on consistency relationship inference in remote sensing image object detection, the spatial and semantic consistency of object distribution is learned, and the problem of degradation of detection performance under the condition of constrained visual features is solved, achieving higher detection accuracy and robustness.

CN119863617BActive Publication Date: 2025-06-10NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510352280.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-10
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing remote sensing image object detection methods have significantly reduced their detection performance under conditions such as small target size, blurred image or severe occlusion, and failed to effectively utilize context-related information between objects.

Method used

Using a method based on consistency relationship inference, by constructing a consistency relationship inference network, learning the spatial and semantic consistency of object distribution, and constructing distribution consistency characteristics, thereby reconstructing the higher-order semantic information of the obstructed object and realizing object detection.

Benefits of technology

The performance of the object detector under limited visual characteristics is significantly improved, and the accuracy and robustness of object detection in complex environments are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119863617B_ABST
    Figure CN119863617B_ABST
Patent Text Reader

Abstract

Aiming at the problem that the detection accuracy of existing remote sensing image target detection methods decreases in the target detection task under the condition of limited visual features, a remote sensing image target detection method and system based on consistency relationship reasoning are proposed, which relates to the technical field of target detection. The method includes: obtaining a remote sensing image data set and preprocessing the remote sensing images in the data set; constructing a target detection network based on consistency relationship reasoning; training the constructed target detection network based on the preprocessed remote sensing images; evaluating the detection accuracy of the trained target detection network. If the preset accuracy is not reached, return to the previous step to retrain until the preset accuracy is reached, and output the target detection network; realizing the target detection of the remote sensing image based on the output target detection network. The method proposed by the present invention uses visual relationships to improve the performance of target detection and can perform reliable reasoning without relying on visual features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and particularly to a remote sensing image object detection method and system based on consistency relationship reasoning. Background Art

[0002] The remote sensing object detection task aims to accurately locate the position of an object of interest by a rotated bounding box and determine its category. In recent years, the remote sensing image object detection task has become one of the current research hotspots and plays an important role in many research fields. Currently, existing remote sensing image object detection algorithms have made significant breakthroughs in aspects such as rotation-invariant feature extraction, rotated bounding box representation, and the continuity of angle regression, effectively improving the detection performance of the detector for objects in any direction, with large aspect ratios, and densely distributed. However, most of these existing methods focus on locating and classifying objects of interest based on regional features, ignoring the context-related information between objects. Under conditions where visual features are limited, such as small object sizes, blurred images, or severe occlusions, the detection performance drops sharply.

[0003] Introducing context relationship reasoning can significantly improve the performance of object detectors. According to whether external prior knowledge is introduced, existing relationship reasoning-based methods are mainly divided into two categories. Among them, the first category of methods integrates external prior knowledge by defining a knowledge graph within the network. However, they only consider regional features and cannot achieve reliable relationship reasoning when visual features are limited. The second category of methods aggregates spatial or semantic priors through context information fusion operations and constructs context relationships between objects through a graph structure to improve the object localization ability of the detector. Since this method constructs multiple local graphs for information propagation, its performance is greatly affected by the way of initializing the relationship graph.

[0004] In summary, existing remote sensing image object detection methods have obvious deficiencies when dealing with object detection tasks under limited visual feature conditions, and it is necessary to further study relationship reasoning-based object detection methods to improve the accuracy and robustness of object detection in complex environments. Summary of the Invention

[0005] Aiming at the problem that the performance of existing remote sensing image object detection algorithms drops when dealing with object detection tasks with limited visual features such as small object sizes, blurred images, or severe occlusions, the present invention proposes a remote sensing image object detection method based on consistency relationship reasoning under limited visual feature conditions, using visual relationships to improve the performance of object detection.

[0006] The technical solution adopted by the present invention is as follows: A remote sensing image object detection method based on consistency relationship reasoning under limited visual feature conditions, and the specific steps of this method are as follows:

[0007] Step 1: Obtain a remote sensing image dataset and preprocess the remote sensing images in the dataset;

[0008] Step 2: Construct an object detection network based on consistency relation reasoning;

[0009] Step 3: Based on the preprocessed remote sensing images, train the object detection network constructed in Step 2;

[0010] Step 4: Evaluate the detection accuracy of the trained object detection network. If the preset accuracy is not reached, return to Step 3 to retrain until the preset accuracy is achieved, and output the object detection network;

[0011] Step 5: Implement object detection of remote sensing images based on the object detection network output in Step 4.

[0012] Furthermore, the specific process of Step 1 includes:

[0013] Step 1.1: Select the remote sensing image datasets DOTA and DIOR-R;

[0014] Step 1.2: For different detection tasks, preprocess the selected datasets respectively;

[0015] Among them, different detection tasks include remote sensing image object detection tasks and remote sensing image occluded object reasoning tasks;

[0016] For the remote sensing image object detection task, no preprocessing is performed on the DOTA and DIOR-R datasets;

[0017] For the remote sensing image occluded object reasoning task, select 25% of the total number of objects in each image in the DOTA and DIOR-R datasets, and divide the selected objects into m × m image patches. Randomly select 75% of the image patches and apply random cloud and fog occlusion, and denote the datasets after applying random cloud and fog occlusion as Occluded DOTA and Occluded DIOR-R respectively.

[0018] Furthermore, the object detection network includes a region proposal network, a relation proposal network, and a consistency reasoning network, where:

[0019] The region proposal network is used to use a trained two-stage remote sensing image object detector as a candidate box extractor to detect unoccluded objects and locate the regions of occluded objects;

[0020] The relation proposal network is used to construct coarse relation pairs based on the detection results of the region proposal network, and select the most likely interacting coarse relation pairs as relation proposals;

[0021] A consistency relationship inference network is used to construct distribution consistency features by learning the spatial and semantic consistency of object distributions based on relationship suggestions, so as to reconstruct the high-order semantic information of occluded objects and achieve object detection.

[0022] Further, the working process of the object detection network described in step 2 includes the following steps:

[0023] Step 2.1: The region proposal network uses a remote sensing image object detector to extract candidate boxes of unoccluded objects v and candidate boxes of occluded objects , and based on o the network extracts the region features of unoccluded objects according to RoI Align and the multi-scale features output by the remote sensing image object detector , and predicts the class labels of unoccluded objects P , where: v ; ; ; ;

[0024] ;

[0025] ;

[0026] In the formula, ( x , y ) is the center point coordinate of the candidate box, ( w , h ) are the length and width of the candidate box respectively, is the angle information of the candidate box, is the i th unoccluded object, is the i th occluded object, is the number of unoccluded objects, is the number of occluded objects, is the i th class of the unoccluded object;

[0027] Step 2.2: For each occluded object , based on the nearest neighbor principle, select n unoccluded objects to construct a coarse relationship pair, then use the relationship suggestion network to obtain the interaction score of the interaction presence feature, and then select n coarse relationship pairs with the highest k interaction scores as relationship suggestions;

[0028] Step 2.3: Use the consistency relation inference network to strengthen the relationship pair region features in the relationship suggestions. Then, based on the strengthened relationship pair region features and relationship suggestions, construct consistency predicate features. Use the consistency predicate features for spatial and semantic interaction reasoning and output distribution consistency features. Finally, perform occluded object detection based on the distribution consistency features and output the detection results.

[0029] Further, the specific steps of Step 2.2 include:

[0030] Step 2.2.1: Select n unoccluded objects that are closest to the occluded object in terms of spatial distance to construct rough relationship pairs. Each rough relationship pair contains relationship pair region features , spatial features and category vectors , where D v is the dimension of the relationship pair features, N p is the number of relationship pairs, D s , D l are the dimensions of the spatial features and category vectors respectively;

[0031] Step 2.2.2: Based on the constructed rough relationship pairs, obtain the spatial encoding that the relationship pairs depend on. Take as the position encoding, and use the interaction of and to obtain the interactive presence feature :

[0032] ,

[0033] ;

[0034] In the formula, represents the relationship encoder, respectively represent the unoccluded object, occluded object, and the intersection and union regions of the two objects in the rough relationship pair, i and j respectively represent the numbers of the unoccluded object and the occluded object; Concat represents the connection operation;

[0035] Then, input into a multi-layer perceptron network to obtain 's interactive score, and then use Sigmoid activation function to normalize the obtained interactive score to the interval [0, 1]:

[0036] ;

[0037] In the formula, represents the normalized interaction score;

[0038] Step 2.2.3: Select k coarse relation pairs with the highest interaction scores as relation suggestions.

[0039] Furthermore, the specific steps of Step 2.3 include:

[0040] Step 2.3.1: Construct the self-attention mechanism of all relation pair regions in each relation suggestion and the spatial prior knowledge, and strengthen the regional features of the relation suggestion through this self-attention mechanism:

[0041] ;

[0042] ;

[0043] ;

[0044] In the formula, N p is the number of relation pairs in the relation suggestion; are the query value, key value, and value respectively; represents the spatial prior knowledge; D q , D k , D v represent the number of channels of the query value, key value, and value respectively; represents the Kronecker product;

[0045] Step 2.3.2: Calculate the regional feature v j of each strengthened relation suggestion by aggregating all attention-weighted values :

[0046] ;

[0047] In the formula, is the attention weight of the value v , and each attention mechanism is obtained by normalizing through the following Softmax function:

[0048] ;

[0049] Wherein, is the attention weight, The calculation formula of

[0050] ;

[0051] Wherein, is the learnable embedding matrix, is a two-layer MLP network for encoding spatial features;

[0052] Step 2.3.3: Obtain the position encoding of the relationship proposal region through the following formula :

[0053] ;

[0054] Wherein, and are respectively the lengths and widths of the candidate boxes of the unoccluded object v and the occluded object o , and are respectively the angle information of the unoccluded object v and the occluded object o ;

[0055] Step 2.3.4: Based on the semantic prior knowledge, the position encoding of the relationship proposal region, and the enhanced regional features of the relationship proposal , construct the consistency predicate feature :

[0056] ;

[0057] Wherein, ε is a small constant;

[0058] Step 2.3.5: Construct the interaction between the consistency predicate feature and the spatial and semantic priors of the relationship proposal through the cross-attention mechanism, so as to reconstruct the high-order semantic features of the occluded object. The cross-attention mechanism is:

[0059] ;

[0060] ;

[0061] ;

[0062] Wherein, represents the spatial transformation position encoding between two targets in the relationship proposal; represents the class embedding of the corresponding unoccluded target, obtained by sampling ;

[0063] Next, the interaction with the corresponding query value is realized through the spatial and semantic features of each key, and spatial and semantic interaction reasoning is performed to output the attention weights. :

[0064] ;

[0065] Among them, represents the query value in the i th cross-attention mechanism; represents the key value in the j th cross-attention mechanism;

[0066] Step 2.3.6: Aggregate all attention weights The weighted values to obtain the distribution consistency feature, and through a two-layer MLP network and combined with Sigmoid the activation function to obtain the category of the occluded target, and then combined with the center point coordinates, length, width, and angle information of the candidate box of the occluding object, the detection result of the occluding object is output.

[0067] The second aspect of the present invention provides a remote sensing image target detection system based on consistency relationship reasoning. This system is implemented by using the remote sensing image target detection method based on consistency relationship reasoning. This system includes:

[0068] A dataset acquisition module, which is used to acquire a remote sensing image dataset and preprocess the remote sensing images in the dataset;

[0069] A detection network construction module, which is used to construct a target detection network based on consistency relationship reasoning;

[0070] A network training module, which is used to train the constructed target detection network based on the preprocessed remote sensing images;

[0071] An accuracy evaluation module, which is used to evaluate the detection accuracy of the trained target detection network. If the preset accuracy is not reached, the network training module is used to retrain until the preset accuracy is reached, and the target detection network is output;

[0072] A target detection module, which is used to implement the target detection of remote sensing images based on the target detection network output by the accuracy evaluation module.

[0073] Preferably, the detection network construction module includes:

[0074] A region proposal network establishment unit, which is used to establish a region proposal network. This region proposal network uses a trained two-stage remote sensing image target detector as a candidate box extractor to detect unoccluded objects and locate the regions of occluded objects;

[0075] A relationship suggestion network building unit is used to build a relationship suggestion network. The relationship suggestion network constructs rough relationship pairs based on the detection results of the region proposal network and selects the rough relationship pairs most likely to have interactions as relationship suggestions.

[0076] A consistency relationship reasoning network building unit is used to build a consistency relationship reasoning network. The consistency relationship reasoning network constructs distribution consistency features based on the learned spatial and semantic consistency of object distributions in relationship suggestions, thereby reconstructing the high-order semantic information of occluded objects and achieving object detection.

[0077] Therefore, compared with the existing technologies, the present invention has the following beneficial effects:

[0078] First, the present invention converts the visual relationship learning task into a scene graph completion task, proposes a visual relationship learning framework based on consistency reasoning to learn the spatial and semantic consistency of object distributions, and then endows the object detection algorithm with visual relationship reasoning ability.

[0079] Second, the visual relationship learning method based on consistency reasoning proposed by the present invention encodes spatial dependencies through a self-attention mechanism to refine relationship features and selects the relationship pairs with the highest interaction degree for further reasoning.

[0080] Third, the query-key interaction mechanism proposed by the present invention learns the distribution consistency features of objects by integrating spatial and semantic prior knowledge and reconstructs the semantic features of occluded objects. Description of the Drawings

[0081] Figure 1 It is a structural diagram of a remote sensing image object detection network based on consistency relationship reasoning.

[0082] Figure 2 It is a flowchart for building a consistency relationship reasoning network.

[0083] Figure 3 It is a framework diagram of a mask-reconstruction-based relationship learning.

[0084] Figure 4 It is an example diagram of object detection results using the baseline method and the object detection algorithm based on consistency reasoning on the DOTA dataset and the DIOR-R dataset. Detailed Embodiments

[0085] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0086] Step 1: Obtain a remote sensing image dataset, and preprocess the remote sensing images in the dataset for different detection tasks;

[0087] Step 2: Construct an object detection network based on consistency relationship reasoning, which includes a relationship proposal network, a region proposal network, and a consistency reasoning network, as shown in the appended Figure 1 figure;

[0088] The main working process of this object detection network is as follows: First, use a trained two-stage remote sensing image object detector as the region proposal network to extract candidate boxes, detect unoccluded objects and locate the regions of occluded objects; Then, the relationship proposal network constructs rough relationship pairs based on the detection results of the detector, and selects the rough relationship pairs that are most likely to have interactions as relationship proposals for subsequent reasoning;

[0089] Finally, the consistency relationship reasoning network constructs distribution consistency features by learning the spatial and semantic consistency of object distributions based on the relationship proposals, reconstructs the high-order semantic information of occluded objects, and finally realizes object detection;

[0090] Step 3: Combine the appended Figure 3 figure to train the object detection network based on consistency relationship reasoning constructed in Step 2; Input the preprocessed remote sensing images in Step 1 into the object detection network constructed in Step 2 for network training, calculate and record the loss function, and stop training until the preset number of epochs is reached;

[0091] Step 4: Evaluate the detection accuracy of the trained object detection network. If the preset detection accuracy is not reached, return to Step 3 to re-train the network until the preset detection accuracy is met, and output the object detection network;

[0092] Step 5: Perform object detection on remote sensing images based on the object detection network output in Step 4 and output the detection results.

[0093] Furthermore, the specific operation steps of Step 1 include:

[0094] Step 1.1: Select the remote sensing image datasets DOTA and DIOR-R;

[0095] DOTA is a challenging remote sensing image oriented target detection dataset, containing 2,806 satellite optical images with sizes ranging from 800×800 to 4,000×4,000 pixels, and a total of 188,282 annotated target instances; these instances are divided into 15 common categories, including airplane (plane, PL), baseball diamond (BD), bridge (BR), ship (SH), harbor (HA), roundabout (RA), storage tank (ST), tennis court (TC), small vehicle (SV), large vehicle (LV), basketball court (BC), ground track field (GTF), soccer ball field (SBF), swimming pool (SP), helicopter (HC);

[0096] DIOR-R is a comprehensive dataset for large-scale remote sensing image detection, containing 23,463 images and a total of 192,518 instances. The dataset contains 20 common target categories, including airplane (airplane, APL), airport (airport, APO), baseball field (BF), basketball court (BC), bridge (BR), chimney (CH), expressway service area (ESA), expressway toll station (ETS), dam (DAM), golf field (GF), ground track field (GTF), harbor (HA), overpass (OP), ship (SH), stadium (STA), storage tank (ST), tennis court (TC), train station (TS), vehicle (VE), windmill (WM);

[0097] Step 1.2: Preprocess DOTA and DIOR-R respectively for different detection tasks;

[0098] 1) Remote sensing image object detection task: The rotated object detection task is a conventional object detection task. Using the conventional settings of the DOTA dataset and the DIOR-R dataset, there is no need to perform special processing on the objects in the image, that is, there is no need to preprocess the remote sensing image dataset for this task.

[0099] 2) Remote sensing image occluded object reasoning task: For the remote sensing images in the DOTA and DIOR-R datasets, 25% of the total number of objects in each remote sensing image is selected, and the selected objects are segmented into m × m image patches. 75% of the image patches are randomly selected and randomly occluded with clouds. To avoid excessive computational complexity, m ranges from [8, 16], and the datasets after applying random cloud occlusion are denoted as Occluded DOTA and Occluded DIOR-R respectively.

[0100] Furthermore, the object detection steps of the object detection network based on consistency relationship reasoning constructed in Step 2 include:

[0101] Step 2.1: Use a two-stage object detector in the existing technology as the remote sensing image object detector, and use the trained remote sensing image object detector as the candidate box extractor to extract the candidate boxes of unoccluded objects v and the candidate boxes of occluded objects and occluded objects o where, are the center point coordinates, length and width, and angle information of the candidate box respectively, is the i th unoccluded object, is the i th occluded object, N v is the number of unoccluded objects, N o is the number of occluded objects;

[0102] RoI Align The network extracts the regional features B v of unoccluded objects according to P and the multi-scale features v detected and output by the remote sensing image object detector, F v and according to F v ​​​Predict the class label of the unoccluded object , where is the class of the i th unoccluded object;

[0103] Step 2.2: Construct the rough relation pairs;

[0104] For each occluded object, select the n unoccluded objects with the closest spatial distance to the occluded object according to the nearest neighbor principle to construct rough relation pairs. Each rough relation pair contains the relation pair region feature , the spatial feature and the class vector , where is the dimension of the relation pair feature, is the number of relation pairs, are the dimensions of the spatial feature and the class vector respectively; the relation pair region feature is the visual feature of the region where the rough relation pair is extracted by the RoI Align network; the spatial feature is the position encoding of the geometric attributes in the rough relation pair; the class vector is the high-dimensional embedding of the class label of the unoccluded object;

[0105] Step 2.3: Based on the constructed rough relation pairs, obtain the interactive presence feature and calculate the interaction score;

[0106] First, according to the constructed rough relation pairs, construct the spatial encoding that depends on the relation pairs:

[0107] ;

[0108] where respectively represent the unoccluded object and the occluded object in the rough relation pair, I , U respectively represent the intersection and union regions of the unoccluded object and the occluded object; Concat represents the concatenation operation;

[0109] Second, use the relation pair region feature as the key, query, and value of the self-attention mechanism in the relation encoder, as the position encoding, and use to obtain the interactive presence feature through the interaction with the position encoding:

[0110] (1);

[0111] where represents the relation encoder, respectively represent the unoccluded object, the occluded object, and the intersection and union regions of the two objects in the coarse relation pair;

[0112] Finally, input into a two-layer MLP network to obtain the interactive presence feature of the interactive score. Then, through Sigmoid the activation function, normalize the interactive score to the interval [0, 1]. This process can be expressed as: (2);

[0113] where represents the normalized interactive score;

[0114] Step 2.4: Select k coarse relation pairs with the highest interactive scores as relation suggestions for subsequent relation reasoning;

[0115] Step 2.5: Refer to Appendix Figure 2 to construct a consistency relation reasoning network. Use this network to construct the relation pair region features k of each relation pair in the relation suggestions and the self-attention mechanism of the spatial prior knowledge to strengthen the relation pair region features . The self-attention mechanism is:

[0116] (3);

[0117] In the formula, N p is the number of relation pairs in the relation suggestion; are the query value, key value, and value respectively; represents the spatial prior knowledge; D q , D k , D v represent the channel numbers of the query value, key value, and value respectively; represents the Kronecker product;

[0118] Calculate the region feature v j of each strengthened relation suggestion by aggregating all attention-weighted values :

[0119] (4);

[0120] where Wv is the attention weight for the value, and each attention mechanism v obtains it through the following α ij function normalization: Softmax

[0121] (5);

[0122] wherein, is the attention weight, and can be obtained by the following formula:

[0123] (6);

[0124] In the formula, is a learnable embedding matrix, and is a two-layer MLP network for encoding spatial features;

[0125] Step 2.6: Construct the position encoding of the semantic prior knowledge and the relationship proposal region;

[0126] The semantic prior knowledge is the high-dimensional vector encoding of the class labels of all objects of interest , and is the position encoding of the relationship proposal region;

[0127] Based on the position information of the unoccluded object and the occluded object, construct the position encoding of the relationship proposal region ;

[0128] (7);

[0129] wherein, is the candidate box of the unoccluded object v , is the candidate box of the occluded object o , , , and are the position information of the unoccluded object and the occluded object, and is the length and width of the candidate box of the unoccluded object v , is the length and width of the candidate box of the occluded object o , and are the angle information of the unoccluded object and the occluded object respectively;

[0130] Step 2.7: Based on the semantic prior knowledge , the position encoding of the relationship proposal region ​And the region features of the enhanced relationship suggestions , construct the consistency predicate features :

[0131] (8);

[0132] Among them, ε is a small constant added to the vector to avoid taking the logarithm to zero;

[0133] Step 2.8: Based on the predicate features containing visual, spatial, and semantic features obtained in Step 2.7 , through the cross-attention mechanism represented by formula (9), construct the interaction of the spatial and semantic priors between the consistency predicate features and the relationship suggestions (filtered coarse relationship pairs), thereby reconstructing the high-order semantic features of the occluded object; the cross-attention mechanism is: (9);

[0134] In the formula, represents the spatial transformation position encoding between two targets in the relationship suggestion region, which is used to interact with the pairwise position encoding in the predicate features; is the category embedding corresponding to the unoccluded target, obtained by sampling ;

[0135] Realize the interaction with the corresponding query value through the spatial and semantic features of each key to perform spatial and semantic interaction reasoning for reasoning, and output the attention weight :

[0136] (10);

[0137] Among them, represents the query value in the i th cross-attention mechanism; represents the key value in the j th cross-attention mechanism;

[0138] Step 2.9: Aggregate all attention weights The weighted values to obtain the distribution consistency features, and then through a two-layer MLP network and combined with Sigmoid the activation function to obtain the category of the occluded object, and combine the position information of the occluded object obtained in Step 2.1 (i.e., the center point coordinates, length, width, and angle information of the candidate box), and output the detection result of the occluded object;

[0139] Further, in step 4, for the object detection network based on consistency relation reasoning after parameter tuning, a remote sensing image object detection task is constructed through the test sets of the DOTA dataset and the DIOR-R dataset to evaluate the detection accuracy; a remote sensing image occluded object reasoning task is constructed through the test sets of the Occluded DOTA dataset and the Occluded DIOR-R dataset to evaluate the reasoning accuracy under the condition of limited target visual features.

[0140] Example:

[0141] To further illustrate the effect of the remote sensing image object detection method based on consistency relation reasoning (Consistency Reasoning Transformer, hereinafter referred to as the CRTr method) proposed in the present invention, a remote sensing image object detection task is constructed on the DOTA dataset (Table 1, Table 3) and the DIOR-R dataset (Table 2) to evaluate the detection accuracy; a remote sensing image occluded object reasoning task is constructed on the Occluded DOTA dataset (Table 4) and the Occluded DIOR-R dataset (Table 5) to evaluate the reasoning accuracy under the condition of limited target visual features;

[0142] 1. Remote sensing image object detection task:

[0143] The baseline models used in the experiment are the Faster RCNN-O and Gliding Vertex algorithms.

[0144] 1) DOTA: Table 1 shows the performance comparison of the proposed CRTr method with existing rotated object detection algorithms on the DOTA dataset. When trained for 12 Epochs, the CRTr method with ResNet50 as the feature extraction network reached 76.81% mAP, exceeding many current advanced methods, such as S 2 A-Net, Gliding Vertex, KFIoU, and Oriented RCNN methods. In addition, without using additional training techniques, the CRTr method with ResNet101 as the feature extraction network obtained 77.44% mAP. The performance of this method exceeded all the listed advanced rotated object detection methods and reached the state-of-the-art level. Figure 4The object detection results using the baseline method (i.e., Faster RCNN-O) and the object detection method based on consistency reasoning proposed in the present invention on the DOTA dataset and the DIOR-R dataset are shown. In the figure, the white dashed boxes are the object detection results before the implementation of consistency reasoning, and the white solid boxes are the object detection results after the implementation of consistency reasoning. From the results in the figure, it can be seen that the proposed method improves the object detection performance through consistency relationship reasoning and has a higher detection accuracy compared to the baseline method. Therefore, the quantitative and qualitative results prove the superiority of the proposed method.

[0145] Table 1 Comparison of detection performance with the state-of-the-art rotated object detection methods on the DOTA dataset: ;

[0146] 2) DIOR-R: The performance comparison with existing methods on the DIOR-R dataset is shown in Table 2. On the DIOR-R dataset, the CRTr method based on ResNet50 achieved 64.02% mAP, which is better than other two-stage and one-stage detection methods. The experimental results show that the proposed method improved 4.48% mAP compared to the baseline method Faster RCNN-O, and is respectively better than the RetinaNet-O method, the FCOS-O method, and the Gliding Vertex method by 6.47% mAP, 4.22% mAP, and 3.98% mAP. It is worth noting that the proposed method achieved the best performance in 9 object categories. Especially in those more challenging categories, such as airports, dams, and bridges, there is a huge improvement compared to existing methods. The detection results on the DIOR-R dataset further demonstrate the generalization and stability of the proposed CRTr method.

[0147] Table 2 Comparison of detection performance with the state-of-the-art rotated object detection methods on the DIOR-R dataset: ;

[0148] 3) Comparison with other inference-based detectors:

[0149] Table 3 compares the CRTr method with existing inference-based detectors to further verify the superiority of the proposed method. For CRTr based on ResNet101, the mAP reaches 75.90%, showing better detection performance than other detectors. DRGN endows the object detection model with the ability of inference by propagating visual and spatial embeddings. Compared with DRGN, the CRTr method achieves a 5.20% performance improvement (75.90% vs 70.70%). Even when compared with the state-of-the-art PCI method and CTM method, the proposed method based on ResNet101 improves the mAP by 1.20% and 1.69% respectively. In addition, the classic inference-based method Reasoning RCNN was reproduced on MMRotate. Under the same settings, the CRTr method improves the mAP by 0.88% in detection performance compared to Reasoning RCNN (75.90% vs 74.82%). In conventional rotated object detection tasks, the proposed method uses easily detectable objects to infer difficult-to-detect objects and has better detection accuracy than other detectors for some easily confused categories, such as small vehicles, large vehicles, basketball courts, and helicopters. In addition, Figure 4 Visualization results of object detection before and after implementing consistency reasoning are shown. As shown in the figure, in the face of challenges such as blurred image features, shadow occlusion, and small targets, the method proposed in the present invention can obtain more consistent detection results through visual relationship reasoning. The experimental results prove the powerful reasoning ability of the CRTr method.

[0150] Table 3 Comparison of detection performance with other inference-based methods on the DOTA dataset: ;

[0151] 2. Remote sensing image occluded target reasoning task:

[0152] The baseline models used in the experiments are Faster RCNN-O and Gliding Vertex algorithms.

[0153] 1) Occluded DOTA: Table 4 gives the comparison of the detection performance of the CRTr method and the current state-of-the-art methods on Occluded DOTA. The data in parentheses indicate that the performance of existing methods degrades when dealing with occluded objects.

[0154] Table 4 Comparison of detection performance with advanced rotated object detection methods on the Occluded DOTA dataset:

[0155] ;

[0156] The results show that when dealing with the occlusions in the experimental setup, all methods experienced a performance drop of approximately 9%. For example, the performance of Faster RCNN-O dropped by 9.10%, that of RoI Transformer dropped by 10.59%, the performance of R 3 Det dropped by 8.72%, the performance of S 2 A-Net dropped by 8.93%, and the performance of CFL dropped by 9.86%. Even the currently relatively advanced two-stage method OrientedR-CNN only achieved a mAP of 73.25%, with a performance drop of 9.77%. However, the baseline method applying CRTr (using ResNet50 as the feature extraction network) achieved a very large gain (79.76% vs 73.25%). When using ResNet101 as the feature extraction network, the proposed algorithm even reached a detection performance of 80.52%, achieving a 6.61% performance improvement compared to the ReDet algorithm using a powerful feature extraction network (ReR50-ReFPN). The experiments show that the proposed method can infer and detect occluded objects under the condition of limited visual features, and the proposed method can be integrated into existing object detection algorithms as an easy-to-use component.

[0157] 2) Occluded DIOR-R:

[0158] The Occluded DIOR-R dataset contains rich categories and semantic relationships. As shown in Table 5, on the Occluded DIOR-R dataset, the proposed method obtained 61.54% and 63.14% mAP when using ResNet50 and ResNet101 respectively, outperforming other two-stage and one-stage methods. The superior results on the Occluded DIOR-R dataset demonstrate the reliability and superiority of the proposed method.

[0159] Table 5 Comparison of detection performance with advanced rotated object detection methods on the Occluded DIOR-R dataset:

[0160] ;

[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A remote sensing image target detection method based on consistency relationship reasoning, characterized in that: The following steps are involved: Step 1: Obtain a remote sensing image dataset and preprocess the remote sensing images in the dataset; Step 2: Build an object detection network based on consistency relationship reasoning; Step 3: Based on the preprocessed remote sensing image, train the target detection network constructed in step 2; Step 4: Evaluate the detection accuracy of the trained target detection network. If the preset accuracy is not reached, return to step 3 and retrain until the preset accuracy is reached, and output the target detection network. Step 5: Implement target detection of remote sensing images based on the target detection network output in step 4; The object detection network includes a region proposal network, a relationship proposal network and a consistency reasoning network, wherein: Region Proposal Network, which uses a trained two-stage remote sensing image object detector as a candidate box extractor to detect unoccluded objects and locate the regions of occluded objects; The relation proposal network is used to construct coarse relation pairs based on the detection results of the region proposal network, and select the coarse relation pairs with the greatest possibility of interaction as relation proposals; The consistency relation reasoning network is used to build distribution consistency features based on the spatial and semantic consistency of object distribution learning based on relation proposals, thereby reconstructing the high-order semantic information of occluded objects and achieving target detection; The workflow of the target detection network described in step 2 includes the following steps: Step 2.1: The region proposal network uses the remote sensing image object detector to extract unobstructed objects v Candidate box and occluded objects o Candidate box , and based on RoI Align Network Based and multi-scale features output by remote sensing image object detectors P Extract unobstructed objects v Regional characteristics , and according to Predicting the class labels of unobstructed objects ,in: , ; In the formula, ( x , y ) is the center point coordinate of the candidate box, ( w , h ) are the length and width of the candidate box respectively, is the angle information of the candidate box, For the i An unobstructed object, For the i occluded object, is the number of unobstructed objects, is the number of occluded objects, For the i The category of unobstructed objects; Step 2.2: For each occluded object , based on the nearest neighbor principle n Unobstructed objects Construct a coarse relationship pair, then use the relationship suggestion network to obtain the interactivity score of the interactive existence feature, and then n Select from the coarse relation pairs k The coarse relationship pairs with the highest interaction scores are taken as relationship suggestions; Step 2.3: Use the consistent relationship reasoning network to strengthen the relationship pair region features in the relationship proposal, and then build the consistent predicate features based on the strengthened relationship pair region features and relationship proposals. Use the consistent predicate features to perform spatial and semantic interactive reasoning, and output distribution consistency features. Finally, perform occluded object detection based on the distribution consistency features and output the detection results.

2. The remote sensing image target detection method based on consistency relationship reasoning as claimed in claim 1, characterized in that: The specific process of step 1 includes: Step 1.1: Select remote sensing image datasets DOTA and DIOR-R; Step 1.2: Preprocess the selected data sets for different detection tasks; Among them, different detection tasks include remote sensing image target detection tasks and remote sensing image occluded target reasoning tasks; For the remote sensing image object detection task, no preprocessing is performed on the DOTA and DIOR-R datasets; For the remote sensing image occluded target reasoning task, 25% of the total number of objects in each image in the DOTA and DIOR-R datasets are selected and the selected targets are divided into m × m , randomly select 75% of the image blocks and apply random cloud occlusion, and record the data sets after applying random cloud occlusion as Occluded DOTA and Occluded DIOR-R respectively.

3. The method for remote sensing image target detection based on consistency relationship reasoning as claimed in claim 2, characterized in that: The specific steps of step 2.2 include: Step 2.2.1: Select n and the occluded object The closest unobstructed object in space Construct coarse relationship pairs, each of which contains relationship pair regional features , spatial characteristics and the category vector ,in, D v is the dimension of the relationship pair feature, N p The number of relationship pairs, D s , D l are the dimensions of spatial features and category vectors respectively; Step 2.2.2: Based on the constructed coarse relation pairs, obtain the spatial encoding of relation pair dependencies ,Will As position code, use and Interaction obtains interactive existence characteristics : , ; In the formula, represents the relational encoder, They represent the unobstructed objects, obstructed objects, and the intersection and union areas of two objects in the rough relationship pair, respectively. i and j Respectively represent the numbers of unobstructed objects and obstructed objects; Concat Indicates a connection operation; Next, Input into a multi-layer perceptron network and get The interactivity score is then used Sigmoid The activation function normalizes the obtained interactivity score to the interval [0, 1]: ; In the formula, represents the normalized interactivity score; Step 2.2.3: Select from the coarse relationship pairs constructed in step 2.2.1 k Interactivity score The highest coarse relationship pairs are used as relationship suggestions.

4. The method for remote sensing image target detection based on consistency relationship reasoning as claimed in claim 3, characterized in that: The specific steps of step 2.3 include: Step 2.3.1: Construct regional features for all relationship pairs in each relationship proposal The self-attention mechanism with spatial prior knowledge strengthens the regional features of relationship suggestions through this self-attention mechanism: ; ; ; In the formula, N p is the number of relation pairs in the relation proposal; They are query value, key value and value respectively; Representing spatial prior knowledge; D q , D k , D v The number of channels representing query value, key value and value respectively; represents the Kronecker product; Step 2.3.2: Aggregate all attention weighted values v j To calculate the regional features of each enhanced relationship proposal : ; In the formula, For value v The attention weight of each attention mechanism Through the following Softmax The function is normalized to obtain: ; In the formula, is the attention weight, The calculation formula is: ; In the formula, is the learnable embedding matrix, It is a two-layer MLP network used to encode spatial features; Step 2.3.3: Get the position encoding of the relation proposal region by the following formula : ; in, , Unobstructed objects v and occluding objects o The length and width of the candidate box, and Unobstructed objects v and occluding objects o Angle information of Step 2.3.4: Position encoding of relation proposal regions and enhanced regional features of relation proposals based on semantic prior knowledge , construct consistent predicate features : ; in, ε is a small constant; semantic prior knowledge E lay ( L v ) is the category label of all objects of interest L v High-dimensional vector encoding of ; Step 2.3.5: Reconstruct the high-order semantic features of the occluded object by constructing the interaction between the consistent predicate features and the spatial and semantic priors of the relation proposal through the cross-attention mechanism. The cross-attention mechanism is: ; ; ; In the formula, Represents the spatial transformation position encoding between two objects in the relation proposal; Represents the category embedding corresponding to the unoccluded target, through sampling get; Next, the spatial and semantic features of each key are used to interact with the corresponding query value, perform spatial and semantic interaction reasoning, and output the attention weight : ; in, Indicates i The query value in the criss-cross attention mechanism; Indicates j The key value in the cross-attention mechanism; Step 2.3.6: Aggregate all attention weights The weighted values ​​are used to obtain the distribution consistency feature, which is then combined with a two-layer MLP network. Sigmoid The activation function obtains the category of the occluded target, and then combines the center point coordinates, length, width, and angle information of the candidate box of the occluded object to output the detection result of the occluded object.

5. A remote sensing image target detection system based on consistency relationship reasoning, characterized in that: The method according to any one of claims 1 to 4 is implemented, the system comprising: The data set acquisition module is used to acquire the remote sensing image data set and preprocess the remote sensing images in the data set; Detection network building module, used to build an object detection network based on consistency relationship reasoning; The network training module is used to train the constructed target detection network based on the preprocessed remote sensing images; The accuracy evaluation module is used to evaluate the detection accuracy of the trained target detection network. If the preset accuracy is reached, the network training module is used to retrain until the preset accuracy is reached, and the target detection network is output; The target detection module is used to realize target detection of remote sensing images based on the target detection network output by the accuracy assessment module.

6. The remote sensing image target detection system based on consistency relationship reasoning as claimed in claim 5, characterized in that: The detection network building module includes: A region proposal network building unit is used to build a region proposal network, which uses a trained two-stage remote sensing image object detector as a candidate box extractor to detect unobstructed objects and locate the region of occluded objects; A relationship proposal network building unit, used to build a relationship proposal network, which builds coarse relationship pairs based on the detection results of the region proposal network and selects the coarse relationship pairs that are most likely to interact as relationship suggestions; The consistent relation reasoning network building unit is used to build a consistent relation reasoning network. The consistent relation reasoning network builds distribution consistency features based on the spatial and semantic consistency of object distribution learned by relation proposals, thereby reconstructing high-order semantic information of occluded objects and realizing target detection.

Citation Information

Patent Citations

  • Scene graph generation method based on self-supervised pre-training

    CN112989927A

  • Scene graph generation method based on super relation learning network

    CN113065587A