Remote sensing image aircraft detection and identification method combined with geometric prior progressive instance enhancement

By combining geometric priors and gradual example enhancement methods, the problem of aircraft target detection and recognition in remote sensing images is solved, and the accurate detection and recognition of aircraft targets is achieved, and the detection accuracy is improved.

CN120107831AActive Publication Date: 2025-06-06HARBIN INST OF TECH

Patent Information

Application Number
CN202510196594.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-06
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

In remote sensing images, the special cross-shaped structure and low duty cycle of the aircraft target make it difficult to accurately determine its position in the image, the categories are easily confused, and the prior art is difficult to effectively detect and identify.

Method used

Using a method combining geometric priori progressive instance enhancement, the geometric semantic information of the aircraft target is fully explored through the progressive class-related dual-branch and instance-guided enhancement module, adaptively generate high-quality predicted point sets, and explicitly enhance the authenticable features of the target.

Benefits of technology

Accurate detection and identification of aircraft targets is achieved, positioning and type recognition capabilities are improved, the impact of background interference is reduced, and detection accuracy is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107831A_ABST
    Figure CN120107831A_ABST
Patent Text Reader

Abstract

A remote sensing image aircraft detection and identification method combined with geometric prior progressive instance enhancement comprises the following steps: step 1, a main branch gradually generates a representation point set through an initial stage and a refining stage to realize position determination and type identification; 2, the progressive class correlation double branches serve as auxiliary branches to be parallel to the main branch, the implicit guidance network learns more robust appearance and semantic feature embedding, and a high-quality prediction point set is generated in a self-adaptive mode; 3, an instance guide enhancement module fully utilizes sufficient class-related semantic information in the affiliated branches, and explicitly enhances identifiable features of targets in the main branches; and 4, an optimization target in the training process mainly comprises three parts of classification, positioning and instance segmentation, and the feature representation and generalization ability of the model is improved through multi-task learning. According to the method, the special cross-shaped geometric structure of the target is fully considered, the multi-task learning thought and the interactive attention mechanism are combined, and the method can be used for achieving accurate aircraft detection and recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of target detection and recognition, and relates to an aircraft target detection and recognition method for optical remote sensing images, and specifically to a remote sensing image aircraft detection and recognition method combined with geometric prior progressive instance enhancement. Background Art

[0002] In the context of countries around the world paying more and more attention to air superiority, in order to ensure that air superiority is not affected, the detection and identification of high-strategic-value targets represented by aircraft is currently a hot topic in the field of remote sensing image processing, and its development is of great significance to both military and civil applications. In the military field, accurate detection of various types of military aircraft can be used to evaluate and analyze enemy military dynamics, provide support for battlefield defense warning, and also support military deployment in air superiority areas, providing strong protection for national security; in the civilian field, detection and identification of aircraft targets based on remote sensing images can provide strong technical support for air traffic control. By analyzing the number and location of parked aircraft at the airport and historical information, flights, inbound and outbound routes and parking locations can be scientifically arranged.

[0003] Unlike natural images where targets are extracted through horizontal boxes, remote sensing images usually rely on rotating boxes to accurately locate targets due to their special bird's-eye view. Rotating target detection can determine the target orientation while determining the target position and category. It can be mainly divided into two categories: anchor box detection and key point detection. The anchor box detection-based method inherits the processing flow in RCNN and YOLO. This method outputs the orientation box by adding additional angle regression to the horizontal anchor box, and relies on special feature alignment operators to alleviate the mismatch between the horizontal candidate box and the rotated target. However, due to the inherent periodicity of the angle orientation, this type of modified angle regression often brings problems of loss discontinuity and regression inconsistency.

[0004] In comparison, the key point detection-based method abandons the manually set anchor point box and uses the prediction of key points such as diagonal points and center to achieve target detection. Among them, Oriented Repooints uses point sets to represent the bounding box, which has the potential to achieve target detection with different orientations, shapes, and postures. The steps of predicting point sets in this type of method are mainly divided into two parts. First, the point set is initially generated through ordinary convolution, and then the target geometric structure information is further mined through deformable convolution to generate a refined point set. After the point set is generated, the rotation box is generated through the transformation function. However, this type of method lacks effective information guidance in the point set generation process, and only realizes the network parameter update through the reverse gradient propagation of the bounding box generated based on the point set envelope and the loss of the true value. The geometric structure of the aircraft target in the remote sensing bird's-eye view scene is more special than that of the ship and vehicle targets, and the overall structure is a cross-shaped special structure. In the absence of effective information supervision, it is more difficult for the network to generate a predicted point set that represents the shape of the aircraft target. In addition, the low occupancy of the aircraft will introduce more background interference, affecting the representation of the key features of the target, further increasing the probability of generating abnormal prediction points. In fact, as a type of rigid body target, airplanes have obvious geometric features, and can be identified by humans through the cross-shaped structure composed of the fuselage and wings. Therefore, introducing additional supervision information containing prior knowledge of airplanes can guide the point set to more fully represent the geometric and semantic key features, and achieve more accurate detection and recognition of airplane targets. Summary of the invention

[0005] Aiming at the problem that the special structure of aircraft in remote sensing scenes and the low duty cycle make it difficult to accurately determine the position and the category is easily confused, this paper proposes a remote sensing image aircraft detection and recognition method combined with geometric prior progressive instance enhancement. The proposed method fully considers the special cross-shaped geometric structure of the target, combines the multi-task learning idea and the interactive attention mechanism, and constructs a progressive class-related dual branch and instance-guided enhancement module to fully mine the geometric semantic information of the aircraft target, implicitly guide and explicitly enhance the representation of the effective features of the target, so as to achieve accurate aircraft detection and recognition.

[0006] The objective of the present invention is achieved through the following technical solutions:

[0007] A method for remote sensing image aircraft detection and recognition combined with geometric prior progressive instance enhancement, the method is:

[0008] Step 1: Based on the backbone network and feature pyramid to extract multi-scale features of the image, the main branch gradually generates a representation point set and extracts key features of the target through the initial stage and refinement stage, and realizes position determination and type recognition in the positioning branch and classification branch respectively.

[0009] Step 2: The progressive class-related dual branches are parallel to the main branch as subsidiary branches. The coarse instance branch and the fine instance branch provide additional supervision information containing the geometric prior of the aircraft target, implicitly guiding the network to learn more robust appearance and semantic feature embedding and adaptively generate high-quality prediction point sets.

[0010] Step 3: The instance-guided enhancement module makes full use of the sufficient class-related semantic information in the subsidiary branches, realizes the flow of instance-level information through the interactive attention mechanism, and explicitly enhances the identifiable features of the targets in the main branch in the initial stage and refinement stage.

[0011] Step 4: The optimization objectives during training include classification, localization, and instance segmentation. Multi-task learning is used to simultaneously learn multiple related tasks in the same model to improve the model's feature representation and generalization capabilities. The three losses come from the classification branch and localization branch in step 1 and the progressive class-related dual branch in step 2.

[0012] Compared with the prior art, the present invention has the following advantages:

[0013] (1) The present invention constructs an anchor-free aircraft target detection and recognition network based on point set representation, which fully mines the inherent geometric semantic information of rigid body targets by using progressive class-related dual branches and instance-guided enhancement modules, and adaptively integrates class-related semantic information into positioning and classification branches, thereby fully improving the positioning and type recognition capabilities of aircraft targets.

[0014] (2) The present invention designs a progressive class-correlated dual branch to make full use of the special cross-shaped geometric structure prior of the aircraft target and implicitly guide the generation of high-quality prediction point sets. Based on the idea of ​​multi-task learning, this branch uses a coarse foreground mask and a fine cross-shaped mask to gradually guide the network to learn more robust appearance and semantic feature embedding, thereby suppressing the generation of abnormal prediction points in redundant background areas.

[0015] (3) The present invention proposes an instance-guided enhancement module to make full use of the efficient feature embedding in the subsidiary branch and explicitly enhance the identifiable features of the aircraft target. This module adaptively fuses the rich class-related semantic information in the mask branch with the feature maps in the localization and classification branches through the interactive attention mechanism, thus realizing the flow of instance-level information. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 The overall framework of the method proposed in the present invention;

[0017] Figure 2 Example of ground truth masks used for the progressive class-dependent dual branch;

[0018] Figure 3-Figure 6 This is an example of aircraft target detection results. DETAILED DESCRIPTION

[0019] The technical solution of the present invention will be clearly and completely described below in conjunction with the drawings and embodiments. Obviously, the described embodiments are only part of the embodiments of the invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0020] The present invention provides a remote sensing image aircraft detection and recognition method combining geometric prior progressive instance enhancement, based on the anchor-free frame point set representation mode and aircraft geometric structure prior, through progressive class-related dual-branch implicit guidance and instance-guided enhancement module to explicitly enhance the effective geometric and semantic feature representation of the target, to achieve accurate aircraft detection and recognition. Figure 1 As shown, the method comprises the following steps:

[0021] Step 1: Based on the backbone network and feature pyramid to extract multi-scale features of the image, the main branch gradually generates a representation point set and extracts key features of the target through the initial stage and refinement stage, and realizes position determination and type recognition in the positioning branch and classification branch respectively.

[0022] Step 1-1: The backbone network extracts features of different scales in the input image, and then fuses high-level semantic features and low-level detail features through the feature pyramid. The feature pyramid outputs multi-level features with a downsampling number of 3-7, denoted as F n (n=3,4,5,6,7).

[0023] Step 1-2: In order to adapt to the special geometric structure and variable scale and orientation of aircraft targets in remote sensing scenes, the main branch of the prediction head uses the representation point set for prediction, including regressing the target position through the positioning branch and determining the target category through the classification branch.

[0024] Step 1-2-1: After the positioning branch is attached to the feature pyramid, the initialization and refinement stages are used to gradually generate the offset of the representation point set, and then the target bounding box is generated through the conversion function. The model based on the representation point set prediction can adapt to target detection of different scales, orientations, and shapes, and the target is represented by the point set R adaptively generated at the feature grid point:

[0025]

[0026] In the formula, n is the number of representation points, p i is the i-th point in it. In this representation mode, the relative positions of different points in the point set imply the size, orientation, geometry and morphology information of the target.

[0027] Step 1-2-1-1: In the initial stage, the features generated by the feature pyramid in step 1-1 are firstly obtained through l 3×3 cascade convolution layers (l defaults to 3) to obtain the initial positioning features. Then, a 1×1 convolution is performed to generate a learnable initial point set offset at each feature grid point. If (x, y) is set to represent the initial position of the grid point in the feature map, then the initialization point R 1 Denoted as:

[0028]

[0029] Step 1-2-1-2: In the refinement stage, the initial point set offset δ in the feature map is extracted by deformable convolution 1 The corresponding refined positioning features

[0030]

[0031] In the formula, w i Represents the contribution of each point to the feature, which is a learnable parameter in the deformable convolution. Further generate the adjusted point set offset Then the representation point set of the final output of the positioning branch is recorded as:

[0032]

[0033] Step 1-2-1-3: In order to facilitate the supervised training of the network under the existing true value and realize the prediction of the target rotation box, it is necessary to transform the point set into the rotation box OB surrounding the target through the conversion function. The conversion process is recorded as:

[0034] OB=G(R 2 )

[0035] Where OB is the rotation box transformed by the point set, including the center position, size and rotation angle of the target. G is the conversion function. The differentiable function ConvexHull is used to optimize the learning process of the adaptive points during training. MinAeraRect is used for post-processing during inference to predict the rotation box of the target.

[0036] Step 1-2-2: In the classification branch, the features generated by the feature pyramid in step 1-1 are first passed through l 3×3 cascaded convolutional layers (l defaults to 3) to obtain the initial classification features. Refined classification features in the feature map are extracted based on the deformable convolution and the bias predicted in the localization branch. Finally, Softmax is used to confirm the target category.

[0037] Step 2: The progressive class-related dual branches are parallel to the main branch as subsidiary branches. The coarse instance branch and the fine instance branch provide additional supervision information containing the geometric prior of the aircraft target, implicitly guiding the network to learn more robust appearance and semantic feature embedding and adaptively generate high-quality prediction point sets.

[0038] Step 2-1: The structures of the coarse instance branch and the fine instance branch are concise and similar to ensure that the performance is improved without adding too much additional computational burden. The difference lies in the different supervision information during the training process. H, W, and C represent the height, width, and number of channels of the feature map, respectively. First, a class-related reference feature map is obtained through l 3×3 convolutions (l defaults to 3)

[0039] Step 2-2: For the class-related reference feature map obtained in step 2-1, directly generate the instance segmentation prediction result through a 1×1 convolution N is the number of target categories, and N+1 represents the background category.

[0040] Step 2-3: During the training process, the instance segmentation prediction results R in the two branches n The loss is calculated with the mask truth value, and the back propagation optimizes the parameter update, implicitly guiding the representation point set to gradually move along the background-coarse foreground-fine body path. At the same time, under the supervision of the class-related mask truth value, R n The instance-level information is implicit in R, and the instance-guided enhancement module in step 3 can be used to n It is then fused with the feature map in the main branch to achieve explicit enhancement of the target features.

[0041] Step 2-3-1: Coarse mask truth value M in coarse instance branch corase The true value of the rotating box is directly converted, the value in the background area is 0, and the foreground area is filled with the corresponding target category. This branch is used to assist in focusing on the target foreground area, corresponding to the initial stage in the positioning branch.

[0042] Step 2-3-2: Considering the cross-shaped structure of the aircraft target, it is difficult to mine its intrinsic geometric features only through rectangular instance segmentation, so the fine instance branch uses the cross-shaped fine mask truth M refine This type of truth value is more consistent with the actual contour of the aircraft target, which can reduce the impact of redundant background on the characterization of key features of the target and further guide the adaptive point set to mine the inherent geometric structure characteristics of the target.

[0043] Step 2-3-2-1: The aircraft target is in the shape of a cross consisting of a fuselage and wings, which can be simulated by two mutually perpendicular ellipses. Let the target true value be the center position (x c ,yc ), size (w, h) and orientation angle θ, then let the two ellipses used to represent the cross mask of the aircraft target be along the length and width of the rotation frame, respectively, to simulate the fuselage and wing. The long sides of the two ellipses are the target true values ​​h and w, and the short sides are w / τ and h / τ, respectively. Then the equations of the two ellipses are written as:

[0044]

[0045]

[0046] Where τ is an adjustment factor used to change the thickness of the long side relative to the short side. When τ is 1, the true value degenerates into an elliptical mask.

[0047] Step 2-3-2-2: Fill the inside of the two elliptical areas with the target category cls and the rest of the area with 0, then a cross-shaped mask truth value can be generated for supervision of the fine instance branch. Cross-shaped mask truth value M refine Denoted as:

[0048]

[0049] Step 3: The instance-guided enhancement module makes full use of the sufficient class-related semantic information in the subsidiary branches, realizes the flow of instance-level information through the interactive attention mechanism, and explicitly enhances the identifiable features of the targets in the main branch in the initial stage and refinement stage.

[0050] Step 3-1: In order to achieve progressive enhancement of high-quality point set generation, the feature maps of the classification branch and the localization branch in the initial stage (corresponding to step 1-2-1-1) and step 1-2-2 ), enhanced by the reference features in the coarse instance branch; and the feature maps of the classification branch and the localization branch in the refinement stage (corresponding to step 1-2-1-2 and step 1-2-2 ), which is enhanced by reference features in the fine instance branch.

[0051] Step 3-2: Let the features to be enhanced in the main branch be The reference features in the subsidiary branches used for enhancement are Then the enhanced features are obtained by outputting the feature-guided enhancement module

[0052] Step 3-2-1: The feature-guided enhancement module fully utilizes the contextual semantic information through the interactive attention mechanism, where the query Q and value V are derived from F ori , used to query the key value and determine the feature blocks to be aggregated and enhanced, and is defined as Q = F ori and V = Fori W v .W v is the embedding matrix implemented by 1×1 convolution. The key K is derived from R and is defined as K=R, where sufficient instance information is matched by the query to be used for F ori Explicit enhancement of .

[0053] Step 3-2-2: Process the key K through k×k grouped convolution to mine the static instance information K between local neighborhoods within the k×k grid s . K s It reflects the static contextual representation among the local neighborhood of R.

[0054] Step 3-2-3: Represent K based on the concatenated static context s and query Q, generate attention weights writing:

[0055] A=[K s ,Q]W 1 W 2

[0056] In the formula, C h is the number of heads in the multi-head attention, W 1 and W 2 is an alternating 1×1 convolution, W 1 Add the ReLU activation function. The features at all spatial positions in A are a k×k×C h dimensional vector. Instead of directly performing dot products on isolated query-key pairs, this attention s Under the guidance of , we further explore the relationship between each query and key.

[0057] Step 3-2-4: Based on the obtained attention weight, the features of value V in the k×k grid are aggregated and enhanced to obtain the feature K after attention enhancement enh ,writing:

[0058]

[0059] In the formula, is the local matrix dot product operator in the k×k grid. enh The dynamic interaction relationship between reference features and input features is mined, which is called dynamic instance enhancement of input features.

[0060] Step 3-2-5: Use the channel attention fusion structure to integrate the static instance information K s and the feature K dynamically enhanced by instance information enh Perform adaptive fusion to obtain the final enhanced feature F fuseChannel attention can adaptively learn the importance of different channel features and achieve feature aggregation through weight redistribution.

[0061] Step 3-2-5-1: First, s and K enh Perform element-by-element summation to obtain F sum , and then perform global pooling in the spatial dimension to obtain channel-related global features The calculation method of the k-th channel feature of G is:

[0062]

[0063] In the formula, F sum (i,j,k) is F sum The value at the i-th row, j-th column, k-th channel.

[0064] Step 3-2-5-2: Further explore the interdependence between channels through the fully connected layer to obtain a more compact feature descriptor writing:

[0065] G c =F fc (G) = ReLU(B(W 3 G))

[0066] In the formula, F fc is a fully connected layer, consisting of the embedding matrix It consists of batch normalization B and nonlinear activation function ReLU. r is the channel scaling ratio, which is used to reduce the dimension and improve the computational efficiency.

[0067] Step 3-2-5-3: Based on feature descriptor G c And Softmax operator, calculate the channel soft attention weight w 1 and w 2 ,writing:

[0068]

[0069] In the formula, A(k) and B(k) are and The kth row of w 1 (k) and w 2 (k) are w 1 and w 2 The kth element of .

[0070] Step 3-2-5-4: The attention weights generated by 3-2-5-3 are K s and K enh The final enhanced feature F is obtained by weighting fuse ,writing:

[0071] F fuse =w 1 ·K s +w 2 ·K enh

[0072] Step 4: The optimization objectives during training include classification, localization, and instance segmentation. Multi-task learning is used to simultaneously learn multiple related tasks in the same model to improve the model's feature representation and generalization capabilities. The three losses come from the classification branch and localization branch in step 1 and the progressive class-related dual branch in step 2.

[0073] The overall loss function during training is written as:

[0074] L=L cls +L loc +λL mask

[0075] In the formula, λ is used to adjust the loss weight (default is 0.05), L cls , L loc and L mask They are classification loss, localization loss and instance segmentation loss respectively.

[0076] Step 4-1: Classification loss L cls Derived from classification loss, it is directly calculated by focal loss.

[0077] Step 4-2: Locate the loss L loc Derived from the localization branch, including the initial stage loss L loc1 and the refinement stage loss L loc2 The positioning loss in both stages is based on the intersection-over-union loss L IoU and the spatial constraint loss L s.c Composition. IoU The bounding box is generated by the intersection-over-union constraint, written as:

[0078]

[0079] Where N p is the number of positive samples, F GIoU is the GIoU function, and Represent the bounding box and category assigned the true value respectively, A directed polygon obtained by transforming a point set.

[0080] L s.c It is used to constrain the representation point to be generated within the bounding box. o and rc are points outside the true value box and the true value center point respectively, then the penalty function ρ ij and L s.c writing:

[0081]

[0082] Where N o and N a They are the number of points outside the true value box in each point set and the number of positive sample point set samples assigned to each target.

[0083] Step 4-3: Instance segmentation loss L mask Derived from the progressive class-related dual branch, the coarse mask loss L mask1 and fine mask loss L mask2 The mask loss is calculated by the pixel-level Softmax cross entropy function.

[0084] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A method for remote sensing image aircraft detection and recognition combined with geometric prior progressive instance enhancement, characterized in that: The method comprises the following steps: Step 1: Based on the backbone network and feature pyramid to extract multi-scale features of the image, the main branch gradually generates a representation point set and extracts key features of the target through the initial stage and refinement stage, and realizes position determination and type recognition in the localization branch and classification branch respectively; Step 2: The progressive class-related dual branches are parallel to the main branch as subsidiary branches. The coarse instance branch and the fine instance branch provide additional supervision information containing the geometric prior of the aircraft target, implicitly guiding the network to learn more robust appearance and semantic feature embedding and adaptively generate high-quality prediction point sets. Step 3: The instance-guided enhancement module makes full use of the sufficient class-related semantic information in the subsidiary branches, realizes the flow of instance-level information through the interactive attention mechanism, and explicitly enhances the discriminable features of the targets in the main branch in the initial stage and refinement stage; Step 4: The optimization objectives in the training process include classification, localization, and instance segmentation. Multi-task learning is used to simultaneously learn multiple related tasks in the same model to improve the model's feature representation and generalization capabilities. The three losses come from the classification branch in step 1, the localization branch, and the progressive class-related dual branch in step 2.

2. The method for remote sensing image aircraft detection and recognition combined with geometric prior progressive example enhancement according to claim 1, characterized in that: The specific steps of step 1 include: Step 1-1: The backbone network extracts features of different scales in the input image, and then fuses high-level semantic features and low-level detail features through the feature pyramid. The feature pyramid outputs multi-level features with a downsampling number of 3-7, denoted as F n (n=3,4,5,6,7); Step 1-2: In order to adapt to the special geometric structure and variable scale and orientation of aircraft targets in remote sensing scenes, the main branch of the prediction head uses the representation point set for prediction, including regressing the target position through the positioning branch and determining the target category through the classification branch.

3. The method for remote sensing image aircraft detection and recognition combined with geometric prior progressive example enhancement according to claim 2, characterized in that: The specific steps of steps 1-2 are: Step 1-2-1: After the positioning branch is attached to the feature pyramid, the initialization and refinement stages are used to gradually generate the offset of the characterization point set, and then the target bounding box is generated through the conversion function. The model based on the characterization point set prediction can adapt to target detection of different scales, orientations, and shapes. The target is represented by the point set R adaptively generated at the feature grid point: In the formula, n is the number of representation points, p i is the i-th point in the set. In this representation mode, the relative positions of different points in the set imply the size, orientation, geometry and morphology of the target. Step 1-2-2: In the classification branch, the features generated by the feature pyramid in step 1-1 are first obtained through l 3×3 cascaded convolutional layers to obtain the initial classification features. Refined classification features in the feature map are extracted based on the deformable convolution and the bias predicted in the localization branch. Finally, Softmax is used to confirm the target category.

4. The method for remote sensing image aircraft detection and recognition based on geometric prior progressive example enhancement according to claim 3 is characterized in that The specific steps of step 1-2-1 are: Step 1-2-1-1: In the initial stage, the features generated by the feature pyramid in step 1-1 are first obtained by passing through l 3×3 cascade convolution layers to obtain the initial positioning features. Then, a 1×1 convolution is performed to generate a learnable initial point set offset at each feature grid point. If (x, y) is set to represent the initial position of the grid point in the feature map, the initialization representation point R1 is recorded as: Step 1-2-1-2: In the refinement stage, the refined positioning features corresponding to the initial point set offset δ1 in the feature map are extracted through deformable convolution In the formula, w i Represents the contribution of each point to the feature, which is a learnable parameter in the deformable convolution and further generates the adjusted point set offset Then the representation point set of the final output of the positioning branch is recorded as: Step 1-2-1-3: In order to facilitate the supervised training of the network under the existing true value and realize the prediction of the target rotation box, it is necessary to transform the point set into the rotation box OB surrounding the target through the conversion function. The conversion process is recorded as: OB=G(R2) Where OB is the rotation box transformed by the point set, including the center position, size and rotation angle of the target; G is the conversion function. The differentiable function ConvexHull is used to optimize the learning process of the adaptive points during training. MinAeraRect is used for post-processing during inference to predict the rotation box of the target.

5. The method for remote sensing image aircraft detection and recognition combined with geometric prior progressive example enhancement according to claim 2, characterized in that: The specific steps of step 2 are: Step 2-1: For the nth level feature of the feature pyramid output H, W, and C represent the height, width, and number of channels of the feature map, respectively. First, a class-related reference feature map is obtained through l 3×3 convolutions. Step 2-2: For the class-related reference feature map obtained in step 2-1, directly generate the instance segmentation prediction result through a 1×1 convolution N is the number of target categories, and N+1 represents the background category. Step 2-3: During the training process, the instance segmentation prediction results R in the two branches n The loss is calculated with the mask truth value, and the back propagation optimization parameter update is performed to implicitly guide the representation point set to gradually move along the background-coarse foreground-fine body path. At the same time, under the supervision of the class-related mask truth value, R n The instance level information is implicit in the 6. The method for remote sensing image aircraft detection and recognition combined with geometric prior progressive example enhancement according to claim 5, characterized in that The specific steps of step 2-3 are: Step 2-3-1: Coarse mask truth value M in coarse instance branch corase Directly transform the true value of the rotating box, the value in the background area is 0, and the foreground area is filled with the corresponding target category. This branch is used to assist in focusing on the target foreground area, corresponding to the initial stage in the positioning branch; Step 2-3-2: The fine instance branch uses the cross-shaped fine mask truth M refine .

7. The method for remote sensing image aircraft detection and recognition combined with geometric prior progressive example enhancement according to claim 6, characterized in that The specific steps of step 2-3-2 are: Step 2-3-2-1: The aircraft target is in the shape of a cross consisting of a fuselage and wings, which can be simulated by two mutually perpendicular ellipses. Let the target true value be the center position (x c ,y c ), size (w, h) and orientation angle θ, then let the two ellipses used to represent the cross mask of the aircraft target be along the length and width of the rotation frame, respectively, to simulate the fuselage and wing, the long sides of the two ellipses are the target true values ​​h and w, and the short sides are w / τ and h / τ, respectively, then the two ellipse equations are written as: Where τ is the adjustment factor, which is used to change the thickness of the long side relative to the short side. When τ is 1, the true value degenerates into an elliptical mask. Step 2-3-2-2: Fill the inside of the two elliptical areas with the target category cls and the rest of the area with 0, then a cross-shaped mask truth value can be generated for supervision of the fine instance branch. The cross-shaped mask truth value M refine Denoted as:

8. The method for remote sensing image aircraft detection and recognition combined with geometric prior progressive example enhancement according to claim 5 is characterized in that The specific steps of step 3 are: Step 3-1: In order to achieve progressive enhancement for the generation of high-quality point sets, the feature maps of the classification branch and the localization branch in the initial stage are enhanced by the reference features in the coarse instance branch; The feature maps of the classification branch and the localization branch in the refinement stage are enhanced by the reference features in the fine instance branch; Step 3-2: Let the features to be enhanced in the main branch be The reference features in the subsidiary branches used for enhancement are Then the enhanced features are obtained by outputting the feature-guided enhancement module 9. The method for remote sensing image aircraft detection and recognition combined with geometric prior progressive example enhancement according to claim 8, characterized in that: The specific steps of step 3-2 are: Step 3-2-1: The feature-guided enhancement module fully utilizes the contextual semantic information through the interactive attention mechanism, where the query Q and value V are derived from F ori , used to query the key value and determine the feature blocks to be aggregated and enhanced, and is defined as Q = F ori and V = F ori W v , W v is the embedding matrix implemented by 1×1 convolution, with the key K derived from R and defined as K=R, where sufficient instance information will be matched by the query for F ori Explicit enhancement of Step 3-2-2: Process the key K through k×k grouped convolution to mine the static instance information K between local neighborhoods within the k×k grid s , K s Reflects the static context representation between the local neighborhood of R; Step 3-2-3: Represent K based on the concatenated static context s and query Q, generate attention weights writing: A [K s ,Q]W1W2 In the formula, C h is the number of heads in the multi-head attention, W1 and W2 are alternating 1×1 convolutions, ReLU activation function is added to W1, and the features at all spatial positions in A are a k×k×C h dimensional vector, which is different from directly performing dot products on isolated query-key pairs. This attention is in the static context K s Under the guidance of , we can further explore the relationship between each query and key; Step 3-2-4: Based on the obtained attention weight, the features of value V in the k×k grid are aggregated and enhanced to obtain the feature K after attention enhancement enh ,writing: In the formula, is the local matrix dot multiplication operator in the k×k grid, K enh The dynamic interaction between reference features and input features is mined, which is called dynamic instance enhancement of input features; Step 3-2-5: Use the channel attention fusion structure to integrate the static instance information K s and the feature K dynamically enhanced by instance information enh Perform adaptive fusion to obtain the final enhanced feature F fuse ,Channel attention can adaptively learn the importance of different channel features and achieve feature aggregation through weight redistribution.

10. The method for remote sensing image aircraft detection and recognition combined with geometric prior progressive example enhancement according to claim 9, characterized in that: The specific steps of step 3-2-5 are: Step 3-2-5-1: First, s and K enh Perform element-by-element summation to obtain F sum , and then perform global pooling in the spatial dimension to obtain channel-related global features The calculation method of the k-th channel feature of G is: In the formula, F sum (i,j,k) is F sum The value at row i, column j, channel k; Step 3-2-5-2: Further explore the interdependence between channels through the fully connected layer to obtain a more compact feature descriptor writing: G c =F fc (G)=ReLU(B(W3G)) In the formula, F fc is a fully connected layer, consisting of the embedding matrix It consists of batch normalization B and nonlinear activation function ReLU, where r is the channel scaling ratio, which is used to reduce dimensions and improve computational efficiency; Step 3-2-5-3: Based on feature descriptor G c And Softmax operator, calculate the channel soft attention weights w1 and w2, written as: In the formula, A(k) and B(k) are and The k-th row of , w1(k) and w2(k) are the k-th elements of w1 and w2 respectively; Step 3-2-5-4: The attention weights generated by 3-2-5-3 are K s and K enh The final enhanced feature F is obtained by weighting fuse ,writing: F fuse =w1·K s +w2·K enh 。 11. The method for remote sensing image aircraft detection and recognition combined with geometric prior progressive example enhancement according to claim 1, characterized in that: The specific steps of step 4 are: The overall loss function during training is written as: L=L cls +L loc +λL mask In the formula, λ is used to adjust the loss weight, L cls , L loc and L mask They are classification loss, localization loss and instance segmentation loss respectively; Step 4-1: Classification loss L cls Derived from classification loss, directly calculated by focal loss; Step 4-2: Locate the loss L loc Derived from the localization branch, including the initial stage loss L loc1 and the refinement stage loss L loc2 The positioning loss of the two parts and the two stages is still based on the intersection-over-union loss L IoU and the spatial constraint loss L s.c Composition, L IoU The bounding box is generated by the intersection-over-union constraint, written as: Where N p is the number of positive samples, F GIoU is the GIoU function, and Represent the bounding box and category assigned the true value respectively, To represent the directed polygon obtained by the transformation function of the point set; L s.c It is used to constrain the representation point to be generated within the bounding box. If r o and r c are points outside the true value box and the true value center point respectively, then the penalty function ρ ij and L s.c writing: Where N o and N a The number of points outside the true value box in each point set and the number of positive sample point set samples assigned to each target; Step 4-3: Instance segmentation loss L mask Derived from the progressive class-related dual branch, the coarse mask loss L mask1 and fine mask loss L mask2 The composition corresponds to the initial stage and fine-tuning stage of the positioning process, and the mask loss is calculated by the pixel-level Softmax cross entropy function.

Citation Information

Patent Citations

  • A convolutional neural network-based remote sensing image target detection method

    CN109800629A

  • Heterogeneous face recognition method based on instance-level spatial perception guidance

    CN114581975A

  • Remote sensing rotating target detection method based on semantic and geometric feature enhancement

    CN118366017A

  • Anchor-frame-free remote sensing image rotating target detection method under attention mechanism

    CN118379617A

Cited By

  • Remote sensing image segmentation method and system based on convolution additive interactive attention and multi-level feature fusion

    CN121937722A