Aircraft Target Detection Method Based on Cross-Attention Fusion Feature Pyramid Network

By adopting a cross-attention fusion feature pyramid network in aircraft target detection, the problem that traditional technology is difficult to identify small or long-distance aircraft targets is solved, and high-precision and robust aircraft target detection are achieved.

CN119313878BActive Publication Date: 2025-05-27耕宇牧星(北京)空间科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411427729.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-05-27
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

Traditional aircraft target detection technology is difficult to accurately identify small or long-distance aircraft targets when processing large-scale remote sensing data, and is susceptible to background noise.

Method used

The aircraft target detection method based on the cross attention fusion feature pyramid network is adopted. By constructing a cross attention fusion module and a cross attention fusion feature pyramid network, the characteristics of different scales are integrated, information flow and fusion strategies are optimized, and the ability to capture key features of the target is enhanced.

Benefits of technology

It significantly improves the accuracy and robustness of aircraft target detection in remote sensing images, and can more comprehensively extract and characterize data characteristics, achieving high-precision aircraft target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119313878B_ABST
    Figure CN119313878B_ABST
Patent Text Reader

Abstract

The present invention discloses an aircraft target detection method based on a cross-attention fusion feature pyramid network, belonging to the technical field of remote sensing image processing. The method includes: constructing a cross-attention fusion module; building a cross-attention fusion feature pyramid network based on the cross-attention fusion module and its corresponding aircraft target detection model; then training the aircraft target detection model for aircraft target detection tasks; and finally using the trained aircraft target detection model to perform aircraft target detection on the remote sensing image to be measured. By effectively utilizing the advantages of the attention mechanism in capturing global semantic information associations, the present invention enriches the semantic information of the fusion features and improves the multi-scale feature expression ability. This characteristic enables the model to more comprehensively extract and characterize data features in the aircraft target detection task of remote sensing images, better understand the image content, and thus achieve high-precision aircraft target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to an aircraft target detection method based on a cross-attention fusion feature pyramid network. Background Art

[0002] In modern aviation monitoring and management systems, the accurate detection of aircraft targets is crucial for ensuring air traffic safety, implementing border defense, and air traffic management. Traditional aircraft detection technologies mainly rely on manual interpretation of high-resolution remote sensing images and radar signal recognition. These technologies face significant challenges in terms of high labor costs, low processing efficiency, and limited accuracy. Especially when dealing with large-scale remote sensing data, traditional methods are often difficult to accurately identify small or distant aircraft targets due to resolution limitations and visual complexity. In recent years, although object detection technologies based on deep learning have achieved remarkable achievements in fields such as autonomous driving and medical image analysis, these technologies still face certain limitations in the application of aircraft detection in remote sensing images. In particular, traditional convolutional neural networks (CNNs) and feature fusion methods, although able to extract rich features from remote sensing images, still have room for improvement in dealing with the scale diversity of aircraft and complex backgrounds in images. These methods often ignore the interaction between features at different scales, resulting in inaccurate detection of small or heavily occluded aircraft targets and susceptibility to background noise interference in practical applications.

[0003] Therefore, there is an urgent need to design a suitable technology to fully integrate the features between different scales, optimize the information flow and fusion strategy between features, explore the dependence relationship between features at different scales, strengthen the ability to capture key features of the target, and effectively improve the recognition accuracy and robustness of the detection model for aircraft. Summary of the Invention

[0004] In view of this, the present invention provides an aircraft target detection method and system based on a cross-attention fusion feature pyramid network that can solve the above technical problems, which is beneficial to improving the accuracy and robustness of aircraft target detection in remote sensing images.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] In a first aspect, an embodiment of the present invention provides an aircraft target detection method based on a cross-attention fusion feature pyramid network, and the method includes the following steps:

[0007] S1. Construct a cross-attention fusion module;

[0008] S2. Build a cross-attention fusion feature pyramid network in combination with the cross-attention fusion module;

[0009] S3. Build an aircraft target detection model based on the cross-attention fusion feature pyramid network;

[0010] S4. Train the aircraft target detection model; use the trained aircraft target detection model to detect aircraft targets in the remote sensing image to be measured.

[0011] Further, in the step S1, the constructed cross-attention fusion module includes: a layer normalization module, a multi-head attention module, and a multi-layer perceptron module; the cross-attention fusion module takes feature maps of adjacent scales as inputs, performs attention calculation on the feature maps of adjacent scales, and obtains the feature dependence relationship between different scales.

[0012] Further, in the multi-head attention module, each attention head independently performs attention calculation, and multiple attention heads process in parallel.

[0013] Further, in the step S2, in the constructed cross-attention fusion feature pyramid network, taking multi-scale feature maps as inputs, perform dimensional adjustment on the feature maps of different scales through convolution operations, the adjusted feature maps perform feature fusion between adjacent feature maps through the cross-attention fusion module, and then obtain multi-scale fusion feature maps through one convolutional layer.

[0014] Further, in the step S3, the constructed aircraft target detection model includes, connected in sequence: an image preprocessing network, a feature extraction network, a cross-attention fusion feature pyramid network, and a detection head, where:

[0015] The image preprocessing network performs image preprocessing operations on the remote sensing image; the preprocessed image data is sent into the feature extraction network for multi-scale feature extraction; the obtained multi-scale features are sent into the cross-attention fusion feature pyramid network to perform cross-attention feature fusion calculation to obtain multi-scale fusion feature maps; the multi-scale fusion feature maps are input into the detection head to locate and classify aircraft targets.

[0016] Further, the image preprocessing operations of the image preprocessing network include: horizontal flipping, noise injection, image sharpening, and blurring;

[0017] The detection head includes a regression network and a classification network. Based on predefined anchor boxes, the detection head uses the regression network to perform regression adjustment on the anchor boxes in each hierarchical feature map to finely determine the position of the aircraft detection box; at the same time, the classification network performs classification operations on all feature points in the feature map to evaluate the class probability of each detection box.

[0018] Further, in the step S4, use the cross-entropy loss function and the mean squared error loss function to train the model for the aircraft target detection task in remote sensing images.

[0019] In a second aspect, an aircraft target detection system based on a cross-attention fusion feature pyramid network is further provided in an embodiment of the present invention. By applying the above-mentioned aircraft target detection method based on a cross-attention fusion feature pyramid network, accurate aircraft target detection task recognition is achieved. The system includes:

[0020] A target detection model establishment module, configured to construct a cross-attention fusion module, build a cross-attention fusion feature pyramid network in combination with the cross-attention fusion module, and build an aircraft target detection model based on the cross-attention fusion feature pyramid network;

[0021] A target detection model training module, configured to train the aircraft target detection model for aircraft target detection tasks;

[0022] A target detection model application module, configured to use the trained aircraft target detection model to perform aircraft target detection on a remote sensing image to be measured.

[0023] Compared with the prior art, the present invention has at least the following beneficial effects:

[0024] 1. The present invention provides an aircraft target detection method based on a cross-attention fusion feature pyramid network, in which a cross-attention fusion module is built, and a cross-attention fusion feature pyramid network and its corresponding aircraft target detection model are constructed in combination with the cross-attention fusion module. Using this aircraft target detection model for aircraft target detection in remote sensing images facilitates the realization of fully automated aircraft target detection without manual intervention, and improves the accuracy and robustness of aircraft target detection in remote sensing images.

[0025] 2. The aircraft target detection model based on the cross-attention fusion feature pyramid network proposed by the present invention effectively utilizes the advantage of the attention mechanism in capturing global semantic information associations, enriches the semantic information of the fusion features, and enhances the multi-scale feature expression ability. This characteristic enables the model to more comprehensively extract and represent data features in the aircraft target detection task of remote sensing images, better understand the image content, and thus achieve high-precision aircraft target detection.

[0026] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification and the drawings.

[0027] The technical solutions of the present invention will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0029] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention.

[0030] Figure 1 It is a schematic flow chart of an aircraft target detection method based on a cross-attention fusion feature pyramid network provided by an embodiment of the present invention.

[0031] Figure 2 It is a schematic diagram of a multi-head attention structure provided by an embodiment of the present invention.

[0032] Figure 3 It is a schematic diagram of a cross-attention fusion module provided by an embodiment of the present invention.

[0033] Figure 4 It is a schematic diagram of a cross-attention fusion feature pyramid network provided by an embodiment of the present invention.

[0034] Figure 5 It is a schematic diagram of an aircraft target detection model based on a cross-attention fusion feature pyramid network provided by an embodiment of the present invention. Detailed implementation manners

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.

[0036] In the description of the present invention, it should be noted that in some processes described in the specification and accompanying drawings of the present application, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. In addition, various serial numbers, etc. are only for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0037] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0038] See Figure 1 As shown, the present invention provides an aircraft target detection method based on a cross-attention fusion feature pyramid network, and this method mainly includes the following steps:

[0039] S1. Construct a cross-attention fusion module;

[0040] S2. Build a cross-attention fusion feature pyramid network by combining the cross-attention fusion module;

[0041] S3. Build an aircraft target detection model based on the cross-attention fusion feature pyramid network;

[0042] S4. Train the aircraft target detection model; use the trained aircraft target detection model to detect aircraft targets in the remote sensing image to be measured.

[0043] The following details the specific implementation manners of the method of the present invention:

[0044] In the embodiment of the present invention, an efficient cross-attention fusion module is designed, a cross-attention fusion feature pyramid network and a corresponding aircraft target detection model are built. By fully fusing the features between different scales, optimizing the information flow and fusion strategy between features, and exploring the dependency relationship between different-scale features, the ability to capture key features of the target is strengthened; it can enable the model to accurately identify and locate aircraft targets in complex backgrounds, significantly improving the detection accuracy and robustness, so as to facilitate the identification of accurate aircraft target detection tasks. The specific implementation manners of the present invention are as follows:

[0045] I. Build a cross-attention fusion module:

[0046] In the embodiment of the present invention, the cross-attention fusion module (CrossAttention FusionModule, CAFM) calculates the attention of feature maps at adjacent scales to obtain the feature dependency relationship between different scales. Specifically, the cross-attention fusion module takes the feature maps at adjacent scales as inputs, and the two feature maps at different scales have been upsampled or downsampled before being input into the cross-attention fusion module to make the sizes of the two inputs unified. The specific implementation strategies of upsampling and downsampling will be described in detail in step II.

[0047] First, the two feature maps are transformed into embedding vectors that can be calculated by the attention mechanism through a deformation operation. Let the input feature maps be F a and F b , then the deformation operation can be expressed as:

[0048]

[0049] Among them, ReshapeA is a reshaping operation that transforms a three-dimensional feature map into an embedding vector. and are the transformed embedding vectors respectively. Then, through a layer normalization operation, namely and the input vector X of the multi-head attention module is obtained Ea and X Eb .

[0050] Combined with Figure 2 and Figure 3 As shown, in the multi-head attention mechanism, each "head" independently performs attention calculation, and multiple heads can process in parallel. In this object detection model, 8 attention heads are set. Through this parallel computing method, the model can extract input features from multiple perspectives and integrate this diverse perspective information to generate a richer and more comprehensive attention feature representation.

[0051] Before X Ea and X Eb are fed into the multi-head attention, they will each be copied twice, that is, three identical Xs Ea and X Eb are obtained respectively, and one exchange is performed, that is, two X Ea vectors and one X Eb vector will be fed into the same multi-head attention module for cross-attention calculation. At the same time, two X Eb vectors and one X Ea vector will be fed into another multi-head attention module for cross-attention calculation. In this embodiment, each attention head has its own independent three fully connected layers, which map the input three vectors to query (Query, Q), key (Key, K), and value (Value, V), where the query comes from the exchanged input vector. The specific operation can be expressed by the following formula:

[0052]

[0053] Among them, and respectively represent the independent learnable weight matrices in the j-th attention head. In the attention mechanism, first calculate the dot product of the query vector and the key vector and perform scaling to obtain the attention scores. Subsequently, normalize these scores and multiply the normalized attention scores by the corresponding value vectors to complete the subsequent calculation process. Both multi-head attentions perform the above calculation operations. The attention calculation process can be formally expressed as follows:

[0054]

[0055] Among them, and represent the outputs of each head of the two multi-head attentions. d k represents the dimension of the key vector, which is used for the scaling operation to prevent overly large numerical values before applying the softmax function, thereby ensuring the stability of numerical calculations. The softmax function is a normalized exponential function used to convert the attention scores into a probability distribution, reflecting the degree of attention of each output element to each element in the input features. According to the obtained probability distribution, the value vectors are weighted and summed to form the final output of the attention mechanism. This process can be understood as extracting and combining important features from the input features according to the attention weights. Finally, the outputs of all attention heads are concatenated and processed through a linear transformation layer to integrate the information extracted by each head. Both multi-head attentions perform the above calculation operations, and the formal representation is as follows:

[0056]

[0057] where W O is a learnable weight matrix, representing the processing of the linear transformation layer, and Concat represents the concatenation operation.

[0058] Finally, the outputs of the multi-head attention are successively subjected to residual connection and layer normalization operations, then passed through a multi-layer perceptron (MLP), and finally another residual connection operation and layer normalization operation are performed. The entire process can be formally expressed as follows:

[0059]

[0060] AttentionOut a = LayerNorm(Output a + MLP(Output a ))

[0061] Another cross-attention branch performs the corresponding same operations to obtain AttentionOut b , and then AttentionOut a and AttentionOut b are concatenated, and the concatenated embedding vector is converted back to the three-dimensional feature map form through a deformation operation. The specific operation can be expressed by the following formula:

[0062] CAO = ReshapeB(Concat(AttentionOut a , AttentionOut b ))

[0063] Among them, CAO (Cross Attention Output) is the final output of cross attention, and ReshapeB is the reverse deformation operation of ReshapeA. The complete cross-attention fusion calculation can be expressed by the following formula:

[0064] CAO = CAFM(F a , F b )

[0065] Among them, CAFM represents the calculation of the complete cross-attention fusion module.

[0066] II. Building a Cross-Attention Fusion Feature Pyramid Network:

[0067] In the embodiment of the present invention, the Cross Attention Fusion Feature Pyramid Network (CAF-FPN) takes multi-scale feature maps as input. As Figure 4 shown, let the input multi-scale feature maps be C i , in this object detection task, there are a total of four feature maps with different scales, which are respectively denoted as c 1 , C 2 , C 3 and C 4 . Among them, the scale of C 3 is twice that of C 4 , the scale of C 2 is twice that of C 3 , and the scale of C 1 is twice that of C 2 . These four initial feature maps need to be preliminarily fused. Specifically, the dimensions of the feature maps with different scales are adjusted through 1×1 convolution operations to obtain the feature maps F 1 , F 2 , F 3 and F 4 . This operation can be formally expressed as follows:

[0068] F i = Conv 1×1 (C i )

[0069] Furthermore, the obtained feature maps F 1 , F 2 , F 3 , F 4 will perform feature fusion between adjacent feature maps through the cross-attention fusion module. First, F 3 is downsampled so that its scale is the same as that of F 4 , and then F 4 is combined with the downsampled F3 They are input into the cross-attention fusion module for calculation, and then the output of the attention calculation passes through a convolutional layer to obtain the fused feature map P 4 , and the specific operation can be expressed by the following formula:

[0070] CAO 4 = CAFM(F 4 , Downsampling(F 3 ))

[0071] P 4 = Conv 3×3 (CAO 4 )

[0072] Among them, Downsampling represents the downsampling operation, and CAO 4 represents the obtained intermediate fused feature. The scale of CAO 4 is the same as that of F 4 . Similarly, perform an upsampling operation on CAO 4 to make its scale the same as that of F 3 . Then, input the upsampled CAO 4 and F 3 into the cross-attention fusion module for calculation. In addition, the obtained cross-attention fusion output feature is further subjected to a cross-attention fusion calculation with the downsampled F 2 . The specific operation can be expressed in the following form:

[0073]

[0074] P 3 = Conv 3×3 (CAO 3 )

[0075] Among them, Upsampling represents the upsampling operation, and CAO 3 represent the intermediate fused features. Therefore, P 3 passes through the cross-attention fusion module and simultaneously contains the dependency relationships with the features of the previous layer and the next layer. The calculation method of the fused feature P 2 is corresponding to that of P 3 , which is shown by the following formula:

[0076]

[0077] P 2 = Conv 3×3 (CAO 2 )

[0078] To obtain P1 , CAO 2 After upsampling, it is combined with F 1 Perform cross-attention fusion calculation, and then pass through a convolutional layer to obtain P 1 , that is, as shown in the following formula:

[0079] CAO 1 = CAFM(F 1 , Downsampling(CAO 2 ))

[0080] P 1 = Conv 3×3 (CAO 1 )

[0081] Finally, the obtained multi-scale fusion feature map includes P 1 , P 2 , P 3 , P 4 . The calculation of the entire cross-attention fusion pyramid network can be summarized as:

[0082] P 1 , P 2 , P 3 , P 4 = CAFFPN(C 1 , C 2 , C 3 , C 4 )

[0083] Among them, CAFFPN represents the cross-attention fusion pyramid network.

[0084] III. Build a complete aircraft target detection model:

[0085] As Figure 5 shown, the complete aircraft target detection model mainly consists of image preprocessing, feature extraction network, cross-attention fusion feature pyramid network, and detection head.

[0086] First of all, the remote sensing image needs to go through the image preprocessing process. The image preprocessing includes a series of image enhancement operations, such as horizontal flipping, noise injection, image sharpening and blurring, etc. Then, the preprocessed image is sent into the multi-scale feature extraction network for the extraction of multi-scale features, which can be specifically expressed as the following formula:

[0087] C 1 , C 2 , C 3 , C 4 = BackboneNet(I)

[0088] Among them, I is the remotely sensed image after image preprocessing, and BackboneNet is a multi-scale feature extraction network. A common network can be selected, such as the ResNet network.

[0089] Subsequently, the obtained multi-scale feature maps C 1 、C 2 、C 3 、C 4 will be used as the input of the cross-attention fusion feature pyramid network, and cross-attention feature fusion calculation will be performed therein to obtain the fused multi-scale feature maps P 1 、P 2 、P 3 、P 4 。

[0090] Finally, the detection head receives the fused multi-scale feature maps and performs precise localization and classification of the aircraft target. Specifically, based on the predefined anchor boxes, the detection head uses the regression network to perform regression adjustment on the anchor boxes in each hierarchical feature map to precisely determine the position of the aircraft detection box. At the same time, the classification network performs classification operations on all feature points in the feature map to evaluate the class probability of each detection box. To optimize the detection results, first, the detection boxes with low confidence are filtered out through the confidence threshold, and then the non-maximum suppression (NMS) technique is applied to eliminate redundant detection boxes. Combining these two steps, the final precise localization and classification results are output, thus effectively realizing the target detection task. The specific operation can be expressed by the following formula:

[0091] y = Score(ClsNet(P 1 ,P 2 ,P 3 ,P 4 ))

[0092] bbox = NMS(RegNet(P 1 ,P 2 ,P 3 ,P 4 ))

[0093] Among them, RegNet represents the regression network, and ClsNet represents the classification network, both of which are composed of multiple convolutional networks. The classification network contains a Sigmoid activation function at its last layer, which is used to convert the output into a probability value between 0 and 1. Score represents the confidence filtering mechanism, which is used to discard the detection results with confidence lower than the set threshold; NMS is responsible for filtering overlapping prediction boxes to retain the optimal detection results. Let the detection head be DetHead, then the complete prediction calculation process can be formally expressed as follows:

[0094] y, bbox = DetHead(P 1 , P 2 , P 3 , P 4 )

[0095] Among them, y and bbox correspond one by one. bbox represents the detected bounding box obtained by prediction, while y represents the predicted probability of the category to which the bounding box belongs.

[0096] IV. Training and testing the model with remote sensing image aircraft data:

[0097] In the embodiment of the present invention, after constructing an aircraft target detection model based on a cross-attention fusion feature pyramid network, the model uses a cross-entropy loss function (Cross-entropy Loss, L CE ) and a mean squared error loss function (Mean Squared Error, L MSE ) for training. Finally, the total loss function of the model is the sum of these two loss functions, and the specific expression is:

[0098] L CE = L CE (y, y label ) + L MSE (bbox, bbox label )

[0099] Among them, bbox label and y label represent the true label values of the bounding box and classification respectively. When the value of the loss function in the model training process tends to be stable and no longer significantly decreases, it indicates that the model has converged and reached a stable state. At this time, the training process ends, and an aircraft target detection model based on a cross-attention fusion feature pyramid network is obtained.

[0100] Subsequently, use the trained aircraft target detection model based on a cross-attention fusion feature pyramid network to conduct a target detection experiment on the test image. This process can be formally expressed as follows:

[0101]

[0102] Among them, (Cross-Attention-Fusion-Feature-Pyramidbased DetectionModel) represents the aircraft target detection model based on a cross-attention fusion feature pyramid network obtained through training; I Test represents the remote sensing test image to be detected; and y test and bbox testThey respectively represent the class probabilities and bounding boxes of the aircraft target detection results generated during the test.

[0103] From the description of the above embodiments, those skilled in the art can learn that the present invention provides an aircraft target detection method based on a cross-attention fusion feature pyramid network. In this method, a cross-attention fusion module is constructed, and a cross-attention fusion feature pyramid network and its corresponding aircraft target detection model are built in combination with the cross-attention fusion module. Using this aircraft target detection model for remote sensing image aircraft target detection, without manual intervention, accurate remote sensing image aircraft target detection can be achieved. The aircraft target detection model based on the cross-attention fusion feature pyramid network proposed by the method of the present invention effectively utilizes the advantage of the attention mechanism in capturing global semantic information associations, enriches the semantic information of the fusion features, and improves the multi-scale feature expression ability. This characteristic enables the model to more comprehensively extract and characterize data features in the remote sensing image aircraft target detection task, better understand the image content, and thus achieve high-precision aircraft target detection. The method of the present invention provides a more reliable and efficient solution for the field of aircraft target detection, and provides a new perspective and possibility for the development of aircraft target detection technology.

[0104] Furthermore, the present invention also provides an aircraft target detection system based on a cross-attention fusion feature pyramid network, which is applied to the aircraft target detection method based on a cross-attention fusion feature pyramid network in the above embodiments to improve the performance of remote sensing image target detection and facilitate the identification of accurate aircraft target detection tasks. The system includes:

[0105] A target detection model establishment module, configured to construct a cross-attention fusion module, build a cross-attention fusion feature pyramid network in combination with the cross-attention fusion module, and build an aircraft target detection model based on the cross-attention fusion feature pyramid network;

[0106] A target detection model training module, configured to train the aircraft target detection model for aircraft target detection tasks;

[0107] A target detection model application module, configured to use the trained aircraft target detection model to perform aircraft target detection on the remote sensing image to be measured.

[0108] For the system provided by the embodiments of the present invention, its implementation principle and the technical effects produced are the same as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the system embodiments, reference may be made to the corresponding content in the foregoing method embodiments, and details will not be repeated here.

[0109] In addition, an embodiment of the present invention further provides a storage medium, on which one or more programs readable by a computing device are stored. The one or more programs include instructions that, when executed by the computing device, cause the computing device to execute a method for aircraft target detection based on a cross-attention fusion feature pyramid network in the above embodiment.

[0110] In an embodiment of the present invention, the storage medium may be, for example, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of the storage medium include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded device, and any suitable combination of the above.

[0111] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.

[0112] It should be noted that the word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The present invention can be implemented by means of hardware including several different components and by means of a suitably programmed computer.

[0113] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0114] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for detecting aircraft targets based on a cross-attention fusion feature pyramid network, characterized in that: The method comprises the following steps: S1, build cross attention fusion module; S2. Build a cross-attention fusion feature pyramid network in combination with the cross-attention fusion module; S3. Building an aircraft target detection model based on the cross-attention fusion feature pyramid network; S4, training the aircraft target detection model; using the trained aircraft target detection model to perform aircraft target detection on the remote sensing image to be tested; In the step S1, the constructed cross-attention fusion module includes: a layer normalization module, a multi-head attention module and a multi-layer perceptron module; the cross-attention fusion module takes feature maps of adjacent scales as input, performs attention calculation on feature maps of adjacent scales, and obtains feature dependencies between different scales; the feature map is converted into an embedding vector that can be calculated by the attention mechanism through a deformation operation, and the input vector of the multi-head attention module is obtained through a layer normalization operation; before the input vector is sent to the multi-head attention module, it will be copied twice and exchanged once; In the step S2, in the constructed cross-attention fusion feature pyramid network, a multi-scale feature map is used as input, and the dimensions of feature maps of different scales are adjusted through convolution operations. The adjusted feature maps are subjected to feature fusion between adjacent feature maps through a cross-attention fusion module, and then a multi-scale fusion feature map is obtained through a convolution layer. The specific process includes: Assume that the input multi-scale feature map is a feature map of four different scales, and the feature maps after dimension adjustment are F1, F2, F3 and F4 respectively; where: F3 is downsampled to make its scale consistent with F4, and F4 and the downsampled F3 are input into the cross attention fusion module to calculate the intermediate fusion feature CAO4, and then CAO4 is passed through a convolution layer to obtain the fusion feature map P4; CAO4 is upsampled to make its scale consistent with F3. The upsampled CAO4 and F3 are input into the cross attention fusion module for calculation. The obtained cross attention fusion output feature is cross-attentive fused with the downsampled F2 to obtain the intermediate fusion feature CAO3. CAO3 is then passed through a convolutional layer to obtain the fusion feature map P3. The calculation method of the intermediate fusion feature CAO2 and the fusion feature map P2 is consistent with the calculation method of CAO3 and P3; After upsampling CAO2, cross-attention fusion calculation is performed with F1 to obtain the intermediate fusion feature CAO1, and then CAO1 is passed through a convolution layer to obtain the fusion feature map P1.

2. The method for detecting aircraft targets based on a cross-attention fusion feature pyramid network according to claim 1, characterized in that: In the multi-head attention module, each attention head performs attention calculation independently, and multiple attention heads are processed in parallel.

3. The method for detecting aircraft targets based on a cross-attention fusion feature pyramid network according to claim 1, characterized in that: In step S3, the aircraft target detection model constructed includes: an image preprocessing network, a feature extraction network, a cross-attention fusion feature pyramid network and a detection head connected in sequence, wherein: The image preprocessing network performs image preprocessing operations on remote sensing images; the preprocessed image data is sent to the feature extraction network for multi-scale feature extraction; the obtained multi-scale features are sent to the cross-attention fusion feature pyramid network to perform cross-attention feature fusion calculations to obtain a multi-scale fusion feature map; the multi-scale fusion feature map is input into the detection head to locate and classify aircraft targets.

4. The method for detecting aircraft targets based on a cross-attention fusion feature pyramid network according to claim 3 is characterized in that: The image preprocessing operations of the image preprocessing network include: horizontal flipping, noise injection, image sharpening and blurring; The detection head includes a regression network and a classification network. The detection head uses the regression network to regress the anchor frames in the feature maps of each level based on the predefined anchor frames to accurately determine the position of the aircraft detection frame; at the same time, the classification network performs classification operations on all feature points in the feature map to evaluate the category probability of each detection frame.

5. The method for detecting aircraft targets based on a cross-attention fusion feature pyramid network according to claim 1, characterized in that: In step S4, the model is trained for the remote sensing image aircraft target detection task using a cross entropy loss function and a mean square error loss function.

6. An aircraft target detection system based on a cross-attention fusion feature pyramid network, characterized in that: The method for detecting an aircraft target based on a cross-attention fusion feature pyramid network as described in any one of claims 1 to 5 is applied to realize accurate aircraft target detection task recognition, and the system comprises: A target detection model building module is used to build a cross-attention fusion module, build a cross-attention fusion feature pyramid network in combination with the cross-attention fusion module, and build an aircraft target detection model based on the cross-attention fusion feature pyramid network; A target detection model training module, used for training the aircraft target detection model for an aircraft target detection task; The target detection model application module is used to use the trained aircraft target detection model to perform aircraft target detection on the remote sensing images to be tested.