Target contour detection method based on multi-view cross attention mechanism

By introducing a multi-view cross-attention mechanism in the remote sensing detection method, combining the multi-scale feature pyramid network and spatial transformation network, the fitting difficulties of scale changes, aspect ratio and angle distribution characteristics in remote sensing target profile detection are solved, and the detection recall rate is significantly improved.

CN120219944APending Publication Date: 2025-06-27BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510171461.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In optical remote sensing scenarios, existing remote sensing detection methods are difficult to accurately detect remote sensing target profiles with large scale variation ranges, extreme aspect ratios and arbitrary angle distributions.

Method used

The target contour detection method based on the multi-view cross attention mechanism is adopted, features are extracted through the skeleton network, multi-scale semantic feature fusion is achieved using the multi-scale feature pyramid network, and multi-view feature maps are generated through the spatial transformation network, combined with the cross attention mechanism to enhance the feature map, and finally target detection and positioning is performed with an anchor-free frame detection framework.

Benefits of technology

The recall rate of remote sensing target profile detection is significantly improved, and the target profile with multiple scales, extreme aspect ratios and arbitrary angle distribution in large field of view optical remote sensing scenarios can be effectively handled.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219944A_ABST
    Figure CN120219944A_ABST
Patent Text Reader

Abstract

The invention provides a target contour detection method based on a multi-view cross attention mechanism, and belongs to the technical field of optical remote sensing image target detection, and the method specifically comprises the steps: inputting a remote sensing image, extracting features through a skeleton network, and achieving the multi-scale semantic feature fusion through a multi-scale feature pyramid network; processing the fused multi-scale semantic feature map by using a spatial transformation network to generate a feature map with a plurality of different view angles; taking the fused multi-scale semantic feature map as a query Query and a value Key, taking the generated feature map with a plurality of different view angles as a key Value, and generating a view angle-view angle similarity matrix from the query Query and the key Value by using a cross attention mechanism; enhancing the feature maps of the plurality of different visual angles by using a similarity matrix; and on the basis of the generated feature map, predicting the central point, the width, the height and the angle of the target by using a detection framework without an anchor frame so as to realize detection and positioning of the remote sensing target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting the contour of a remote sensing target, in particular to a method for detecting the contour of a target based on a multi-view cross-attention mechanism, and belongs to the technical field of optical remote sensing image target detection. Background Art

[0002] With the development of remote sensing technology, optical payloads exhibit observation characteristics such as multi-platform, multi-resolution, large swath width, and high revisit rate, and gradually form a natural remote sensing big data. At the same time, driven by the remote sensing big data, target detection algorithms are widely applied to the intelligent processing of optical remote sensing, and they can play an important role in fields such as national territorial sovereignty maintenance, urban planning and construction, intelligent transportation / trade control, ecological environment protection, emergency search and rescue, and military reconnaissance. Currently, the common optical remote sensing target detectors mainly include two types: anchor box detectors and anchor-free detectors. Among them, the anchor box detector can match the preset prior anchor boxes with the true bounding boxes in advance to prepare positive and negative samples for the training of the detector, and achieve the positioning and classification capabilities of optical remote sensing targets; the anchor-free detector directly locates the target position by learning the key points of the remote sensing target, and uses a pixel-level prediction method to regress the width and height of the target corresponding to the key point positions, without the need to introduce additional artificial prior information.

[0003] With the continuous improvement of the spatial resolution of optical remote sensing images, the traditional horizontal box and rotated box target modeling methods can no longer provide sufficient accurate positioning granularity. Therefore, target contour detection has become an important way for refined target positioning. However, in high-resolution optical remote sensing images with a large field of view coverage, there are characteristics such as a large scale change range, extreme aspect ratios, and arbitrary angle distributions in the distribution of multi-category typical ground object targets. For the anchor box detector, there will be a large number of "scale gaps" in the fitting of the set discrete multi-scale / multi-aspect ratio initial anchor boxes during the preprocessing process, resulting in difficulty for the detection network to learn the feature representations of targets at corresponding scales, causing serious deficiencies in the performance of remote sensing target detection. Similarly, in the anchor-free detection framework, the receptive fields constructed by general feature extraction methods cannot effectively perceive remote sensing targets with a large scale transformation range. At the same time, the convolutional neural network lacks the ability to perceive and model the angle information of targets, and it is also difficult for the anchor-free detector to achieve good target contour detection effects for remote sensing targets with complex scales, angles, and aspect ratio distributions in complex optical remote sensing scenes. In summary, neither the anchor box detector nor the anchor-free detector can accurately and effectively provide good target contour description accuracy for remote sensing ground object targets with complex shape distributions, restricting the development of remote sensing target contour detection algorithms. Summary of the Invention

[0004] The object of the present invention is to address the deficiencies and defects existing in the prior art, and to propose an object contour detection method based on a multi-view cross-attention mechanism to solve the problem that the existing remote sensing detection method has low detection accuracy for remote sensing target contours with large-scale variation ranges, extreme aspect ratios, and arbitrary angular distributions in optical remote sensing scenarios.

[0005] The method of the present invention is realized through the following technical solutions.

[0006] An object contour detection method based on a multi-view cross-attention mechanism, comprising the following steps:

[0007] Step 1: Input a remote sensing image, extract features through a backbone network, and then use a multi-scale feature pyramid network to achieve multi-scale semantic feature fusion;

[0008] Step 2: Use a spatial transformation network to process the multi-scale semantic feature map fused in Step 1 to generate feature maps with several different views;

[0009] Step 3: Use the multi-scale semantic feature map fused in Step 1 as the query Query and the value Key, use the feature maps with several different views generated in Step 2 as the key Value, and use the cross-attention mechanism to generate a view-view similarity matrix from the query Query and the key Value;

[0010] Step 4: Use the similarity matrix in Step 3 to enhance the feature maps with several different views;

[0011] Step 5: Based on the feature map generated in Step 4, use an anchor-free detection framework to predict the center point, width, height, and angle of the object, so as to realize the detection and positioning of the remote sensing target.

[0012] Optionally, the specific process of Step 1 of the present invention is: cut the input remote sensing image into multiple image slices, use the backbone network to perform basic feature extraction on each image slice to obtain several feature maps with different resolutions, sort the multi-scale feature maps according to the resolution and perform adjacent layer scale feature fusion to generate a multi-scale semantic feature map F i 。

[0013] Optionally, the backbone network selected in the present invention is a ResNet-101 network.

[0014] Optionally, the specific process of Step 2 of the present invention is:

[0015] Step 2.1, adopt a spatial transformation network f loc , and generate an affine transformation matrix according to the input multi-scale semantic feature map F i

[0016] Step 2.2, use the affine matrix as the parameter of the spatial affine transformation to control the multi-scale semantic feature map F i to complete the corresponding feature perspective transformation;

[0017] Step 2.3, for each multi-scale feature map, repeat Step 2.1 and Step 2.2 multiple times to generate feature maps with several different perspectives.

[0018] Optionally, the spatial transformation network f described in the present invention loc is composed of two convolutional layers with a convolutional kernel size of 3×3 and a stride of 2, a global average pooling layer, and two fully connected layers.

[0019] Optionally, the specific process of Step 3 in the present invention is as follows:

[0020] Step 3.1, adjust the dimensions of the multi-scale semantic feature map fused in Step 1 and the feature maps with several different perspectives generated in Step 2;

[0021] Step 3.2, the multi-scale semantic feature map generates a key query and a value through the corresponding multi-layer perceptron, and the multi-perspective feature map generates a query through the corresponding multi-layer perceptron;

[0022] Step 3.3, perform a dot product operation on the key query and the query, and then obtain the perspective-perspective similarity matrix W through a soft maximum calculation i .

[0023] Optionally, the key query, value, and query in the present invention are expressed as follows:

[0024]

[0025] where, Concat(V i 0 ,V i 1 ,...,V i j ) represents feature concatenation of the multi-perspective feature map V i j ; Repeat(F i ,M v ) represents repeating and stacking the initial feature F i M v times to complete the feature stacking operation, f q and f k,v represent the functions for generating q i ,k i and v iMulti-layer perceptron.

[0026] Optionally, the specific process of step 4 of the present invention is as follows:

[0027] Step 4.1: Multiply the value Key in step 3 with the perspective-perspective similarity matrix W i matrix to obtain a processing result P i ;

[0028] Step 4.2: Adjust the feature dimension of P i and then perform average pooling on the first dimension of the feature to obtain the finally integrated result Q i ; Use the residual structure to add and fuse Q i with the multi-scale semantic feature map F i ;

[0029] Step 4.3: For each multi-scale semantic feature map F i , repeat the above steps 4.1 and 4.2 to enhance the initial feature map and obtain the enhanced multi-scale feature map.

[0030] Optionally, the specific process of step 5 of the present invention is as follows:

[0031] Step 5.1: Use the multi-scale feature map constructed in step 4 and, according to the constructed center point and contour point prediction network layer, predict the positioning information of the remote sensing target contour;

[0032] Step 5.2: Combine the predicted center point and contour point and output the target contour.

[0033] Optionally, the present invention further calculates the loss function of the above prediction network layer using the predicted center point and contour point for supervised training.

[0034] Beneficial effects

[0035] First of all, the present invention selects a feature pyramid network to fuse multi-scale semantic feature information of optical remote sensing images containing multi-category typical remote sensing targets, providing guarantee for the subsequent generation of multi-view feature maps. Then, a spatial transformation network is used to generate feature maps from multiple output features of the feature pyramid for several perspectives, which is used to simulate the changes in the scale, aspect ratio, and direction characteristics of remote sensing targets. Then, cross-attention is used to interact the initial feature map with the transformed multi-view feature maps in the form of query-key-value pairs, and through the coupling discrimination relationship learning between different perspective feature maps, coupled discrimination information is generated to support the subsequent detection and recognition of remote sensing targets. Finally, the contour detection and positioning of optical remote sensing targets are carried out by taking the center point-based anchor-free detection method as a benchmark. This method can effectively solve the problem of difficult contour fitting of targets with multi-scale, extreme aspect ratio, and arbitrary angle distribution characteristics in large field-of-view optical remote sensing scenarios, can significantly improve the recall rate of the remote sensing target contour detection task, and has good practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0037] Figure 1 It is the overall flowchart of this method.

[0038] Figure 2 It is the schematic diagram of the positioning network structure in this method.

[0039] Figure 3 It is the schematic diagram of the multi-view cross-attention mechanism in this method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The embodiments of the present invention will be described in detail below with reference to the drawings.

[0041] It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other; and, based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.

[0042] It should be noted that the following description pertains to various aspects of embodiments within the scope of the appended claims. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of the aspects set forth herein can be used to implement a device and / or practice a method. Additionally, this device can be implemented and this method can be practiced using other structures and / or functionality in addition to one or more of the aspects set forth herein.

[0043] Embodiment

[0044] An embodiment of this application is a target contour detection method based on a multi-perspective cross-attention mechanism, as Figure 1 shown, which specifically includes the following steps:

[0045] Step 1: Input the original optical remote sensing image. After extracting features through a deep backbone network, use a feature pyramid network to achieve multi-scale semantic feature fusion and output 3 multi-scale semantic fusion feature maps.

[0046] In the specific implementation of this step, the following process can be adopted:

[0047] Step 1.1: Scan the original optical remote sensing image, perform cutting processing on the input high-resolution remote sensing image, and obtain L image slices with a resolution of H×W.

[0048] Step 1.2: Use the deep backbone network to perform basic feature extraction on each image slice S n , n∈[1, L] to obtain multi-scale feature maps with resolutions of H / 4×W / 4, H / 8×W / 8, H / 16×W / 16, and H / 32×W / 32 output by 4 different network blocks in the deep backbone network structure.

[0049] In this embodiment, the mentioned deep backbone network selects the ResNet-101 network, as Figure 1 shown. This deep backbone network can output 4 feature maps with different resolutions, namely C1, C2, C3, and C4.

[0050] Step 1.3: Arrange the above 4 groups of feature maps with different resolutions from low resolution to high resolution, and perform adjacent layer scale feature fusion to obtain three fused multi-scale feature maps F0, F1, and F2, whose resolutions are arranged from large to small.

[0051] The following is an example for illustration.

[0052] Sort C1, C2, C3, and C4 in ascending order of resolution. Through linear interpolation operation, upsample the low-resolution feature layer to make its resolution consistent with that of its adjacent higher-resolution feature layer. Then, perform a feature addition operation on the upsampled feature layer and the higher-resolution feature layer. Input the result of the addition into a convolutional layer with a kernel size of 1×1 to complete the fusion of adjacent-scale features. And so on, complete the fusion of the subsequent two higher-resolution features, thereby outputting three multi-scale feature layers, namely where B represents the batch size; N represents the number of channels of the feature map; H i and W i represent the height and width of the i-th feature map respectively, that is, H i = H / 2 i+1 , W i = W / 2 i+1 .

[0053] Step 2: Input the multi-scale semantic feature map fused in Step 1 into the spatial transformation network for processing to generate feature maps with several different perspectives.

[0054] In the specific implementation of this step, the following process can be adopted:

[0055] Step 2.1: As Figure 2 shown, in order to generate feature maps with different perspectives, first, an affine transformation matrix needs to be generated. For this purpose, a learnable localization network f loc (i.e., the spatial transformation network) is used to adaptively generate the affine transformation matrix This localization network consists of two convolutional layers with a kernel size of 3×3 and a stride of 2, a global average pooling layer, and two fully connected layers. Its specific composition structure is as Figure 3 shown. By inputting the obtained in Step 1.3 into the above-mentioned learnable localization network, a 2×3 affine matrix is generated, as shown in formula (1):

[0056]

[0057] where and represent the transformation coefficients of rotation and scaling respectively, represents the transformation coefficient of translation.

[0058] Step 2.2: Use the affine matrix as the parameter of the spatial affine transformation to control the initial feature F i to complete the corresponding feature perspective transformation. The specific operation is as shown in formula (2) below.

[0059]

[0060] Among them, G = (x t , y t ) represents the coordinate point position of the output feature map after affine transformation; (x s , y s ) represents the coordinate point position of the initial perspective feature map; represents the coordinate point corresponding to the original initial perspective feature map according to each coordinate point position of the output feature map according to the affine transformation matrix .

[0061] Then, the values at the corresponding coordinate points of the original feature map are assigned to the coordinate points corresponding to the output feature after affine transformation by means of interpolation sampling.

[0062] Step 2.3: Perform the operations in Step 2.1 and Step 2.2 on each feature map of the multi-scale feature map by using two localization networks respectively, aiming to obtain feature maps with two different perspectives, denoted as M v which is set to 2 in this step.

[0063] Step 3: Use the multi-scale semantic feature map fused in Step 1 as the query (Query) and value (Key), and use the feature maps with several different perspectives generated in Step 2 as the key (Value), and use the cross-attention mechanism to generate a similarity matrix between the query (Query) and the key (Value).

[0064] After generating the visual multi-perspective feature map V i j , use the cross-attention mechanism to interact the initial feature F i and the multi-perspective feature map V i j , aiming to establish a global dependency relationship between different perspective feature maps and model the different shapes, orientations, and scales of the same object under different perspectives.

[0065] In the specific implementation of this step, the following process can be adopted:

[0066] Step 3.1: Considering that the initial feature and the multi-perspective feature map are both feature layers generated by a convolutional neural network, and their feature dimensions are inconsistent with the sequence dimensions required by the cross-attention mechanism in the vision transformer. Therefore, first adjust the feature maps generated by the above convolutional neural network from 4D feature B×N×H i ×W i to 3D feature B×N×H iW i 。

[0067] Step 3.2: In the attention mechanism of the vision transformer network structure, the query, key, and value features input into the attention mechanism are all generated from the same input features. In contrast, in the cross-attention mechanism of this method, the key and value are both generated from the initial features through the corresponding multi-layer perceptron, while the query is generated from the multi-view feature map through the corresponding multi-layer perceptron. The calculation formulas are shown in the following formulas (3) and (4):

[0068]

[0069] where Concat(V i 0 , V i 1 ,..., V i j ) represents concatenating the features of the multi-view feature map V i j ; Repeat(F i , M v ) represents repeating and stacking the initial feature F i M v times to complete the feature stacking operation, aiming to keep the dimensions aligned with the features to be concatenated above. In addition, f q and f k,v represent the multi-layer perceptrons used to generate q i , k i and v i .

[0070] Step 3.3: After obtaining q i , k i and v i , first perform a dot product operation on q i and k i , and then perform a softmax calculation to obtain the view-view similarity matrix W i . The specific formula is shown in the following formula (5):

[0071]

[0072] Then, as shown in the following formula (6), based on the view-view similarity matrix W i ,

[0073] Step 4: Perform a matrix multiplication operation on the similarity matrix in Step 3 and the value (Key) in Step 3, and use the constructed similarity matrix to enhance the initial feature map to realize the learning of the coupling discriminant relationship between different perspectives.

[0074] In the specific implementation of this step, the following process can be adopted:

[0075] Step 4.1, multiply the value (Key) v in Step 3 i by the perspective-perspective similarity matrix W i to obtain the processing result P i :

[0076]

[0077] Step 4.2: In order to keep the feature dimensions of the processing result P i and the initial feature F i consistent, first adjust the feature dimension of P i to and then perform average pooling on the first dimension of the feature. The specific formula is shown in Equation (7) below.

[0078]

[0079] where Q i represents the finally integrated result. Then, introduce a residual structure to add and fuse Q i and the initial feature F i to avoid the negative impact of low-quality affine changes on model training in the initial stage of detector training.

[0080] Step 4.3: Complete the above process on feature layers of different scales to construct a multi-scale regression layer.

[0081] Step 5: Based on the feature map generated in Step 4, use an anchor-free detection framework to predict the target center point and target contour points to realize the contour detection of remote sensing targets.

[0082] In the specific implementation of this step, the following process can be adopted:

[0083] Step 5.1: Use the multi-scale feature map constructed in Step 4 to construct a center point and contour point prediction network layer to predict the positioning information of the remote sensing target contour.

[0084] The following is an example for illustration.

[0085] Both of the above two prediction network layers are composed of two convolutional structures: the first convolutional structure adopts a unified structure, that is, a 3×3 convolutional layer with independent parameters, with a stride of 1, padding of 1, and an output channel of 64. The second convolutional structure adopts a 1×1 convolutional layer with independent parameters, with a stride of 1, padding of 0, and the output channel varies according to different prediction information. The specific number of output channels is as follows: in the center point prediction network layer, the output channel is the same as the number of categories included in the training data; in the contour point prediction network layer, the number of output channels is the number N of sampling points of the target contour.

[0086] Step 5.2: After completing the prediction of the center point and contour point information in Step 5.1, combine the center point and contour point to output the target contour.

[0087] Step 5.3: Use the center point and contour point loss functions to supervise the training of the network.

[0088] The following is an example.

[0089] The loss function of the center point is shown in the following formula (8):

[0090]

[0091] where C represents the category, w and h represent the width and height of the output feature map; The predicted value at the position (x, y) on the channel of the c-th category; Y xyc = 1 The true value at the position (x, y) on the channel of the c-th category.

[0092] In addition, the loss function of the target contour point adopts the Smooth-L1 norm loss function, which is specifically shown in the following formula (9):

[0093]

[0094] where x represents the difference between the true value and the predicted value, that is, a p -a g , where a represents the coordinate value of the sampled target contour point.

[0095] Without departing from the principle of the present invention, several improvements can also be made, which should also be regarded as belonging to the protection scope of the present invention.

[0096] In summary, the above is only a preferred embodiment of the present invention, and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A target contour detection method based on multi-view cross-attention mechanism, characterized in that: The following steps are involved: Step 1: Input the remote sensing image, extract features through the skeleton network, and then use the multi-scale feature pyramid network to realize multi-scale semantic feature fusion; Step 2: Use the spatial transformation network to process the multi-scale semantic feature map fused in step 1 to generate feature maps with several different perspectives; Step 3: Use the multi-scale semantic feature map fused in step 1 as the query Query and value Key, and use the feature map with several different perspectives generated in step 2 as the key Value. Use the cross attention mechanism to generate the perspective-perspective similarity matrix from the query Query and the key Value. Step 4: using the similarity matrix in step 3 to enhance the feature maps of the several different perspectives; Step 5: Based on the feature map generated in step 4, the center point, width, height and angle of the target are predicted using a detection framework without an anchor box, thereby achieving detection and positioning of the remote sensing target.

2. According to claim 1, the target contour detection method based on multi-view cross attention mechanism is characterized in that: The specific process of step 1 is as follows: the input remote sensing image is cut into multiple image slices, basic features are extracted from each image slice using a skeleton network, several feature maps with different resolutions are obtained, multi-scale feature maps are sorted according to resolution and adjacent scale features are fused to generate a multi-scale semantic feature map F i .

3. According to claim 2, the target contour detection method based on multi-view cross attention mechanism is characterized in that: The skeleton network uses the ResNet-101 network.

4. According to claim 2, the target contour detection method based on multi-view cross attention mechanism is characterized in that: The specific process of step 2 is: Step 2.1, use the spatial transformation network f loc , according to the input multi-scale semantic feature map F i Generate an affine transformation matrix Step 2.2, transform the affine matrix As the parameters of the spatial affine transformation to control the multi-scale semantic feature map F i Complete the corresponding feature perspective transformation; Step 2.3: For each multi-scale feature map, repeat steps 2.1 and 2.2 multiple times to generate feature maps with several different perspectives.

5. According to claim 4, the object contour detection method based on multi-view cross attention mechanism is characterized in that: The spatial transformation network f loc It consists of two convolutional layers with a kernel size of 3×3 and a stride of 2, a global average pooling layer, and two fully connected layers.

6. The target contour detection method based on multi-view cross attention mechanism according to claim 1 is characterized in that: The specific process of step 3 is as follows: Step 3.1, adjust the dimensions of the multi-scale semantic feature map fused in step 1 and the feature map with several different perspectives generated in step 2; Step 3.2, the multi-scale semantic feature map is passed through the corresponding multi-layer perceptron to generate the key query and value value, and the multi-view feature map is passed through the corresponding multi-layer perceptron to generate the query query; Step 3.3, perform a dot multiplication operation on the key query and the query query and then perform a soft maximum calculation to obtain the view-view similarity matrix W i .

7. The target contour detection method based on multi-view cross attention mechanism according to claim 6 is characterized in that: The key query, value value and query query are expressed as follows: Among them, Concat(V i 0 ,V i 1 ,...,V i j ) represents the multi-view feature map V i j Perform feature splicing; Repeat(F i ,M v ) represents the initial feature F i Repeat stacking M v times to complete the feature stacking operation, f q and f k,v Indicates the use to generate q i ,k i and v i Multi-layer perceptron.

8. The target contour detection method based on multi-view cross attention mechanism according to claim 1 or 6, characterized in that: The specific process of step 4 is as follows: Step 4.1: Compare the value Key in step 3 with the view-view similarity matrix W i Matrix multiplication to obtain the processing result P i ; Step 4.2, for P i The feature dimension is adjusted, and then average pooling is performed on the first dimension of the feature to obtain the final integrated result Q i ; Using the residual structure, Q i and multi-scale semantic feature map F i Perform additive fusion; Step 4.3: for each multi-scale semantic feature map F i , using the above process, the initial feature map is enhanced to obtain the enhanced multi-scale feature map.

9. The target contour detection method based on multi-view cross attention mechanism according to claim 1 is characterized in that: The specific process of step 5 is as follows: Step 5.1, using the multi-scale feature map constructed in step 4, according to the constructed center point and contour point prediction network layer, predict the positioning information of the remote sensing target contour; Step 5.2, combine the predicted center point and contour point to output the target contour.

10. The target contour detection method based on multi-view cross attention mechanism according to claim 1, characterized in that: The predicted center points and contour points are further used to calculate the prediction network layer loss function for supervised training.