Optical Remote Sensing Small-Scale Target Detection Method Based on Adaptive Local Context Embedding

Through the adaptive local context embedding method and anchor-free frame detection technology, the detection accuracy and accuracy of small-scale targets in optical remote sensing images are improved, the shortcomings of target detection in complex environments are solved, and efficient and accurate target positioning is achieved.

CN115410089BActive Publication Date: 2025-07-18BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210754167.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-28
Publication Date
2025-07-18
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

The existing anchor-free frame method still has room for improvement in the detection accuracy and accuracy of small-scale targets in optical remote sensing scenarios, especially in complex environments and densely distributed targets.

Method used

Adaptive local context embedding method is adopted, combined with multi-scale optimization of hourglass feature extraction network and channel attention algorithm, through corner point pooling and center point pooling operations, the target position is corrected using the cross entropy loss function and push-pull loss function to achieve anchor-free box detection.

Benefits of technology

It significantly improves the detection effect of small-scale targets in large field of view and high-resolution optical remote sensing images, reduces false alarm rate, and improves positioning accuracy, adapts to detection performance under complex conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115410089B_ABST
    Figure CN115410089B_ABST
Patent Text Reader

Abstract

The present invention discloses an optical remote sensing small-scale target detection method based on adaptive local context embedding. First, a multi-scale optimized hourglass feature extraction network is used to extract features of small-scale targets in remote sensing images. Second, combining an adaptive local context embedding algorithm and a channel attention algorithm, the features extracted in the first step are deeply optimized. Third, using the feature map generated in the second step, through corner pooling and center point pooling operations, the positions of the upper left corner point, the lower right corner point, and the center point of the target are obtained. Then, the cross-entropy loss function and the push-pull loss function are used to correct the coordinates of the corner points and the center point, and finally the target position is determined to achieve anchor-free detection of the corner points and center points of the entire image. The present invention can solve the problems of detection accuracy and detection precision for small-sized targets in large-field-of-view high-resolution optical remote sensing images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image target detection, and particularly relates to an optical remote sensing small-scale target detection method based on adaptive local context embedding. Background Art

[0002] Target detection, as a basic problem in tasks such as remote sensing image interpretation and analysis, is the basis for algorithms such as image segmentation, image description, target tracking, and scene understanding. Target detection uses computer vision algorithms to search for the presence of targets of interest in an image and determine the positions where the targets appear.

[0003] Since the number and size of targets in an image are uncertain, and accurate positioning of the targets is often required, target detection algorithms are relatively complex algorithms in computer vision. In the early stage, target detection was usually based on traditional image processing methods, using a large number of manually designed features to describe a small number of specific categories of targets, and then using traditional machine learning methods to complete target detection. Since convolutional neural networks have been widely used, due to their powerful feature extraction and deep learning capabilities, the performance and stability of general-category target detection have been significantly improved.

[0004] Currently, in target detection, for methods of generating target candidate regions, they mainly include the anchor box method and the anchor-free method. Most traditional target detection algorithms mainly use the anchor box method. By using a set of rectangular candidate boxes with different scale and ratio combinations to generate candidate regions of the target, the convolutional neural network judges whether the region within the box contains a target. If it contains a target, its category attribution is determined, and the candidate box is regressed to a more accurate position. However, such methods have obvious disadvantages such as low operation efficiency and data redundancy. When the target scale is small, or when the target scales in the image vary greatly, it is difficult to determine the scale of the anchor box, seriously affecting the target detection performance.

[0005] Different from the anchor box method, the anchor-free method does not have an explicit candidate region generation process. Instead, it models the target by means of key points and feature lines, and uses encoding and decoding to complete the regression of the target border. It can avoid the limitation of the anchor box on the target size matching degree, improve the phenomenon of unbalanced positive and negative samples, reduce the introduction of hyperparameters, reduce complexity, and can improve the detection performance of small-scale targets in remote sensing images. However, the existing anchor-free methods still have some limitations. Due to the lack of a large number of artificially added anchor boxes, some anchor-free methods based on key point detection have relatively high requirements for the richness of semantic information in the feature map; in addition, at present, the anchor-free method is mainly applied in natural scenes. Aiming at the problems that the target scale in the optical remote sensing scene is relatively smaller and difficult to identify, there is still a large room for improvement in the detection accuracy and accuracy of the anchor-free method. Summary of the Invention

[0006] In view of this, the present invention provides an optical remote sensing small-scale target detection method based on adaptive local context embedding, which can address the deficiencies and defects of the existing technologies and solve the problems of detection accuracy and precision for small-sized targets in large field-of-view high-resolution optical remote sensing images. Among them, the small-sized targets generally refer to target images with pixel numbers greater than 8*8 pixels and less than 20*20 pixels.

[0007] The technical solution of the present invention is implemented as follows:

[0008] An optical remote sensing small-scale target detection method based on adaptive local context embedding includes the following steps:

[0009] Step 1: Use a multi-scale optimized hourglass feature extraction network to extract features of small-scale targets in the remote sensing image;

[0010] Step 2: Combine an algorithm based on adaptive local context embedding and a channel attention algorithm to deeply optimize the features extracted in Step 1;

[0011] Among them, for the adaptive local context embedding algorithm, by using multi-scale convolutional kernels, the convolutional kernel size with the best discriminative power is adaptively selected according to different inputs to capture context information; for the channel attention algorithm, by screening the aggregated information of each channel in the feature map, the channel information helpful for target detection is retained, redundant channels are suppressed, and the network detection performance is improved; then, the two are connected in series before and after, and by simulating the characteristics of human vision, a feature map with concentrated attention is generated;

[0012] Step 3: Use the feature map generated in Step 2, through corner pooling and center point pooling operations, to obtain the positions of the upper left corner point, lower right corner point, and center point of the target; then, use the cross-entropy loss function and push-pull loss function to correct the corner and center point coordinates, and finally determine the target position to achieve anchor-free detection of "corner-center point" for the entire image.

[0013] Further, Step 1 is specifically: by means of skip layer connection and multi-scale aggregation methods, feature fusion is performed on feature maps with different depths in the hourglass network to obtain the feature extraction result.

[0014] Further, in Step 1, the multi-scale optimized hourglass feature extraction network mainly includes three parts: extracting deep image features, skip connection feature fusion, and multi-scale feature aggregation.

[0015] Further, first, an optical remote sensing image is input. After image preprocessing, multi-scale extraction of the deep features of the image is performed. In the neural network, since shallow information is helpful for small-scale target detection, skip connections are used to fuse the shallow information and the deep information to obtain feature maps of different scales. Finally, the feature maps of different scales after feature fusion are multi-scale aggregated, and finally, a feature map with both shallow and deep feature information is output.

[0016] Further, the adaptive local context embedding algorithm in step two mainly includes two parts: capturing context information with multi-scale convolution and adaptively selecting the optimal size of the convolution kernel. First, the feature map of step 1 is input, and the context information is captured by using multi-scale convolution kernels. The input feature map is multi-scale convolved by selecting convolution kernels with scales of 3*3, 5*5, 7*7, …, (2i+1)*(2i+1) to capture context information of different scales. Secondly, the results after convolution are superimposed, and the superimposed results are screened by an attention mechanism to obtain a set of feature descriptors with multi-scale information. The different-scale convolutions are multiplied by the feature descriptors respectively, and the appropriate spatial scale information is selected. The feature maps after being screened by the attention mechanism are added up to obtain the output result of the adaptive local context embedding algorithm.

[0017] Beneficial effects:

[0018] The method of the present invention has the following advantages compared with the prior art:

[0019] 1. The method of the present invention can significantly improve the detection effect of small-scale targets in the scene of large-field high-resolution optical remote sensing images. Especially when facing complex conditions such as a more complex environment and dense target distribution, good results can also be obtained. On the basis of improving the detection rate, the positioning accuracy is improved, and the false alarm rate is greatly reduced.

[0020] 2. By reducing the depth of the feature extraction network in step one, the method of the present invention balances the relationship between the detection speed and the detection accuracy of the network, and has good practical application value.

[0021] 3. By adopting the adaptive local context embedding algorithm in step two, the method of the present invention uses a multi-scale adaptive feature extraction method to improve the ability to obtain context information of small-scale targets and improve the positioning accuracy; at the same time, combined with the channel attention algorithm, redundant feature information is suppressed and eliminated, and the false alarm rate is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a flowchart of the method of the present invention.

[0023] Figure 2 is a schematic diagram of the multi-scale optimized hourglass feature extraction network of the present invention.

[0024] Figure 3 This is a schematic diagram of the adaptive local context embedding algorithm of the present invention.

[0025] Figure 4 This is a schematic diagram of the channel attention module of the present invention.

[0026] Figure 5 This is a schematic diagram of the "corner point - center point" anchor - free detection of the present invention. Detailed implementation manners

[0027] The following is a detailed description of the present invention by way of examples in conjunction with the accompanying drawings.

[0028] The present invention provides an optical remote - sensing small - scale target detection method based on adaptive local context embedding, as Figure 1 shown, including the following steps:

[0029] Step 1: Use a multi - scale optimized hourglass feature extraction network to extract features from the optical remote - sensing image.

[0030] The present invention gives a specific implementation method, as Figure 2 shown, including the following steps:

[0031] Step 1.1: Extract the deep features of the image.

[0032] First, input the optical remote - sensing image S A into the multi - scale optimized hourglass feature extraction network (composed of residual modules and pooling layers). After pre - processing through a convolutional network with a convolution kernel of 7×7 and a stride of 2 and 1 residual network, through 4 convolutional and down - sampling pooling operations, deep - feature information is extracted, and the feature maps after down - sampling in each layer are respectively denoted as

[0033] Then, perform 4 same convolutional and up - sampling operations on the feature map The up - sampled feature maps are respectively denoted as

[0034] The following is an example for illustration.

[0035] First, input the optical remote - sensing image S with a size of 640×640 A into the multi - scale optimized hourglass feature extraction network for pre - processing. In this embodiment, the network includes a convolutional network with a convolution kernel of 7×7 and a stride of 2 and 1 residual network, reducing the image size to 160×160 and increasing the number of channels to 256.

[0036] Then, the preprocessed image is input into the multi-scale optimized hourglass feature extraction network. After 4 groups of residual and max pooling operations, the image is denoted as The size changes from the input 160×160 to {80×80, 40×40, 20×20, 10×10} in sequence, and the number of channels is {256, 384, 384, 384} in sequence. After another 4 groups of residual and upsampling operations that are completely symmetric to the downsampling process, the feature maps from the inner layer to the outer layer are denoted as The size changes from 10×10 to {20×20, 40×40, 80×80, 160×160} in sequence, and the number of channels is {384, 384, 384, 256} in sequence.

[0037] Step 1.2: Skip connection feature fusion.

[0038] Since some information is lost during the pooling operation, in order to improve the hourglass network's ability to obtain more comprehensive feature information, skip connection layers are used to retain the information lost during the pooling process. The feature map before downsampling is superimposed on the feature map after upsampling as shown in Equation 1:

[0039]

[0040] In Equation 1, represents the feature map after the i-th upsampling; Γ represents the convolutional operation in the skip connection part, including a residual module with a 3×3 convolutional kernel and a stride of 1.

[0041] Step 1.3: Multi-scale feature aggregation.

[0042] Before upsampling in the traditional hourglass network, a convolutional operation is required first, but this will lead to an increase in the number of network layers and is not conducive to retaining the feature information of the shallow layer.

[0043] Therefore, in order to effectively improve the information richness of the feature map output by the backbone network, in the present invention, each layer of the feature map in the upsampling part of the hourglass network is directly superimposed, and the result after superimposition is multiplied by the input feature map to obtain the final output, as shown in Equation 2:

[0044]

[0045] Among them, ξ i (i = 2, 3, 4) represents the upsampling and 1×1 convolutional operations on the feature map ; S out represents the feature map output by the multi-scale optimized hourglass feature extraction network.

[0046] Step 2: Use the algorithm combining the adaptive local context embedding algorithm and the channel attention algorithm to deeply optimize the feature map extracted in Step 1.

[0047] Among them, the adaptive local context embedding algorithm captures context information by using multi-scale convolutional kernels and adaptively selecting the convolutional kernel size with the best discriminative power according to different inputs. As Figure 3 shown, it includes the following steps:

[0048] Step 2.1: Perform multi-scale convolutional kernels to capture context information.

[0049] Among them, the number of convolutional kernels is represented by i, i = 1, 2, 3,..., n; the size of the convolutional kernel is (2i + 1) × (2i + 1);

[0050] Taking i = 3 as an example, three scales of convolutional kernels of 3×3, 5×5, and 7×7 are selected to perform multi-scale convolutional operations on the feature map S out respectively, and the results are represented by S 3×3 , S 5×5 , S 7×7 respectively.

[0051] The multi-scale convolutional operation is shown in Equation 3:

[0052] S (2i+1)×(2i+1) = Conv (2i+1)×(2i+1) (S out ) (3)

[0053] Among them, Conv (2i+1)×(2i+1) represents the convolutional operation with a convolutional kernel of (2i + 1) × (2i + 1).

[0054] Step 2.2: Adaptively select the optimal size of the convolutional kernel.

[0055] First, the results S 3×3 , S 5×5 , S 7×7 ∈R C×H×W after convolution are superimposed to obtain S Add ∈R C×H×W , where R represents the real number field, C represents the number of channels, H represents the image height, W represents the image width, and S Add represents the superimposed result.

[0056] In this embodiment, S Add = S 3×3 + S 5×5 + S 7×7 .

[0057] Then, use average pooling and max pooling operations to aggregate S AddSpatial information is obtained to get the result S c1 , S c2 ∈R C ×1×1 :

[0058]

[0059] Among them, AvgPool H×W and MaxPool H×W respectively represent average pooling and max pooling operations on the spatial direction (H×W) of the image.

[0060] After that, the obtained results S c1 , S c2 are respectively input into the fully connected layer to get z1, z2 ∈ R 3C×1×1 :

[0061] z i = Fc(S ci ) (5)

[0062] Fc(S ci ) = Conv1(δ(BN(Conv2(S ci )))) (6)

[0063] Among them, z i (i = 1, 2) respectively represent the results after S c1 , S c2 are input into the fully connected layer; S ci (i = 1, 2) represents the feature map after average pooling (i = 1) / max pooling (i = 2) operation; Fc represents the fully connected layer, Conv1 is a convolution operation with a convolution kernel of 1×1, the number of input channels is C / r, and the number of output channels is C; Conv2 is a convolution operation with a convolution kernel of 1×1, the number of input channels is C, and the number of output channels is C / r; BN represents batch normalization (Batch Normalization); δ represents the activation function ReLU.

[0064] Then, the results z1, z2 output from the fully connected layer Fc are superimposed and the softmax operation is performed to obtain a feature descriptor η ∈ R 3C×1×1 , and the feature descriptor η is decomposed into η = [η1, η2, η3] (η i ∈ R C×1×1 , i = 1, 2, 3). The feature descriptor η = [η1, η2, η3] is multiplied with S 3×3 , S 5×5 , S 7×7 respectively to get S3' ×3 , S5' ×5 , S7'×7 Finally, perform scale screening and summation on S3', ×3 , S5', ×5 , S7', ×7 to obtain the final output S SA-out .

[0065] S ( ' 2i+1)×(2i+1) = η 2i+1 × S (2i+1)×(2i+1) (7)

[0066]

[0067] In Equations 7 and 8, the value of i is a non-zero natural number and is determined by the number of multi-scale convolution kernels in Step 2.1.

[0068] Step 2.3: Channel attention filters out valid information.

[0069] Channel attention filters the aggregated information of each channel in the feature map, retains the channel information helpful for target detection, suppresses redundant channels, and improves the network detection performance.

[0070] As Figure 4 shown, first extract the channel information. Perform max pooling and average pooling operations on the feature map S SA-out ∈ R C×H×W to aggregate the spatial information and generate c max , c avg ∈ R C×1×1 two different spatial feature descriptors respectively. c max , c avg represent the results of average pooling and max pooling operations on the spatial directions H×W of the image respectively, and contain feature channel information, as shown in Equation 9:[[]]

[0071]

[0072] Then, input c max , c avg into the fully connected layer respectively, and add the outputs of the fully connected layer. The output vector of the final result channel attention is represented by C out , as shown in Equations 10 and 11:[[]]

[0073] C out = σ(Fc(MaxPool(S SA-out )) + Fc(AvgPool(S SA-out ))) (10)

[0074] Fc(S Max(Avg) ) = Conv1(δ(BN(Conv2(SMax(Avg) )))) (11)

[0075] Among them, Fc represents the fully connected layer; σ represents the Sigmoid function; S Max(Avg) represents the feature map after average pooling / maximum pooling operation; Conv1 is a convolution operation with a convolution kernel of 1×1, the number of input channels C / r, and the number of output channels C; Conv2 is a convolution operation with a convolution kernel of 1×1, the number of input channels C, and the number of output channels C / r; BN represents Batch Normalization; δ represents the activation function ReLU.

[0076] Finally, multiply the result after addition by the input to obtain the feature map Sc, as shown in Equation 12, to achieve channel screening:

[0077] S c = C out × S SA-out (12)

[0078] Step 3: Use the feature map S c generated in Step 2, through corner pooling and center point pooling operations, to obtain the positions of the upper left corner point, lower right corner point, and center point of the target. Use the cross-entropy loss function and push-pull loss function to correct the corner and center point coordinates, and finally determine the target position to achieve anchor-free detection of "corner-center point" for the entire image, and finally obtain the small-scale target detection result with high precision and low false alarm rate.

[0079] As Figure 5 shown, Step 3 includes the following steps:

[0080] Step 3.1: Use the feature map S c generated in Step 2, and successively through center point pooling and cascaded corner pooling operations, to obtain the positions of the upper left corner point, lower right corner point, and center point of the target, and obtain the k bounding boxes with the highest scores from the corner points.

[0081] Among them, the methods of center point pooling and cascaded corner pooling are as follows:

[0082] In center point pooling, for the feature map S c with a size of H×W, find the maximum value in the horizontal and vertical directions, and add the maximum values to find the central key point of the target.

[0083] In cascaded corner pooling, convolve the feature map S c with a convolution kernel of 3×3 to obtain S c_conv3 . Perform left pooling on S c_conv3 to obtain S lp . Pooling result S lp and convolution result Sc_conv3 Add them together to obtain Take Through a convolution with a 3×3 convolution kernel and vertical pooling, the final result is obtained

[0084] As shown in Equations 13 and 14:

[0085]

[0086]

[0087] Among them, Conv3 represents the convolution with a convolution kernel of 3; LP and TP represent left pooling and vertical pooling respectively.

[0088] For left pooling and vertical pooling, assume the input feature map is f t , with a size of H×W. Take f t Perform a max pooling operation on all the feature vectors within the range of (i,j) and (i,H) to obtain the feature vector t ij ; Then, take the feature map f l Perform a max pooling operation on all the feature vectors within the range of (i,j) and (W,j), and obtain the feature vector l ij . Finally, add l ij and t ij together. As shown in Equations 15 and 16:

[0089]

[0090]

[0091] Among them, f t , f l are the feature maps input to the pooling layer; and are the feature vectors of f t and f l at (i,j) respectively, l ij and t ij represent the feature vectors of f l after left max pooling and f t after vertical max pooling at (i,j) respectively. represents the feature vector of f t at (H,j), represents the feature vector of f l at (i,W). t (i+1)j represents the feature vector of f t after max pooling at (i+1,j), l i(j+1) represents the feature vector of f lThe feature vector after max pooling at (i, j+1).

[0092] Step 3.2: According to the scores, select the top k central key points with the highest scores and map the k central key points back to the input image.

[0093] Step 3.3: In the k bounding boxes obtained from the corner points in Step 3.1, define a central region and check whether the central key points obtained in Step 3.2 exist in the central region. On the premise of ensuring the same labels for both, retain the corresponding bounding boxes. If there are no central key points in the central region of the bounding box, discard the bounding box.

[0094] Among them, the method for selecting the central region of the bounding box is as follows:

[0095] In the central region of the bounding box, it is determined by the formula:

[0096]

[0097] Among them, (tl x , tl y ), (br x , br y ) are the coordinates of the upper left corner point and the lower right corner point of the bounding box; (ctl x , ctl y ), (cbr x , cbr y ) are the upper left corner point and the lower right corner point of the central region. n is an odd number used to determine the size of the central region of the bounding box.

[0098] Step 3.4: Use the loss function to correct the coordinates of the corner points and the center points, and finally determine the target position to achieve anchor-free detection of "corner point - center point" for the entire image.

[0099] Among them, the loss function includes three parts: "corner point and central key point loss function", "push - pull loss function", and "offset loss function". The complete loss function expression L push is:

[0100]

[0101] Among them, is the loss function of the corner point and the center point; represents the minimum value of the distances between two corner points of the same target, represents the maximum value between two corner points of different targets; is the loss function for predicting the corner point and the center point. In Equation 18, α, β, γ represent the "pull - loss function", "push - loss function", and "offset loss function" respectively, and their values can be set to 0.1, 0.1, and 1 respectively.

[0102] Corner and center key point loss function \(L\) det is as follows:

[0103]

[0104] where \(p\) cij represents the value at the \((i, j)\) position in the \(c\)-th channel of the feature map output after predicting the corner points and the center point; \(y\) cij represents the ground truth value at the corresponding position; \(N\) represents the number of targets.

[0105] "Push - Pull" loss function \(L\) pull and \(L\) push are as follows:

[0106]

[0107]

[0108] The relationship between two "corner points" belonging to the same target calculated by the "Push - Pull" loss function. Among them, and are the "embeddings" of the upper - left and lower - right corners of the \(k\)-th target respectively, \(e\) k is and 's mean value, and \(\Delta\) takes 1. \(e\) j represents the mean value of the "embeddings" of the upper - left and lower - right corners of the \(j\)-th target and .

[0109] The offset loss function is as follows:

[0110]

[0111]

[0112]

[0113] where \(o\) k represents the offset of the predicted corner point mapped from the original image to the feature map, \(L\) off represents the offset loss function, \(o'\) k represents the offset of the actual corner point of the target mapped from the original image to the feature map, \(x\) represents \(o\) k - \(o'\) k in this example, and SmoothL1Loss represents the smooth L1 loss function. \(x\) k , \(y\) k are the corner point coordinate values of the \(k\)-th target, and \(m\) is the down - sampling factor.

[0114] In summary, the above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An optical remote sensing small-scale target detection method based on adaptive local context embedding, characterized in that It includes the following steps: Step 1: Use a multi-scale optimized hourglass feature extraction network to extract features of small-scale targets in remote sensing images; Step 2: Combine the adaptive local context embedding algorithm and the channel attention algorithm to deeply optimize the features extracted in Step 1; Among them, the adaptive local context embedding algorithm captures context information by using multi-scale convolutional kernels and adaptively selecting the convolutional kernel size with the best discriminative power according to different inputs; the channel attention algorithm screens the aggregated information of each channel in the feature map, retains the channel information helpful for target detection, suppresses redundant channels, and improves the network detection performance; then, the two are connected in series before and after, and a feature map with concentrated attention is generated by simulating the characteristics of human vision; Step 3: Use the feature map generated in Step 2, and through corner pooling and center point pooling operations, obtain the positions of the upper left corner point, lower right corner point, and center point of the target; then, use the cross-entropy loss function and the push-pull loss function to correct the corner and center point coordinates, and finally determine the target position to achieve anchor-free detection of the corner-center points of the entire image; In Step 1, the multi-scale optimized hourglass feature extraction network includes three parts: extracting deep image features, skip connection feature fusion, and multi-scale feature aggregation; The adaptive local context embedding algorithm in Step 2 is divided into two parts: capturing context information with multi-scale convolution and adaptively selecting the optimal size of the convolutional kernel; first, input the feature map of Step 1, use multi-scale convolutional kernels to capture context information, and perform multi-scale convolution on the input feature map through convolutional kernels with scales of 3*3, 5*5, 7*7,..., (2i+1)*(2i+1) to capture context information at different scales; secondly, superimpose the convolution results and perform attention mechanism screening on the superimposed results to obtain a set of feature descriptors with multi-scale information; Multiply the multi-scale convolutions with the feature descriptors respectively, select the appropriate spatial scale information, and add and sum the feature maps after being screened by the attention mechanism to obtain the output result of the adaptive local context embedding algorithm.

2. The optical remote sensing small-scale target detection method with adaptive local context embedding according to claim 1, wherein Specifically, Step 1 is: by means of skip layer connection and multi-scale aggregation, perform feature fusion on the feature maps at different depths in the hourglass network to obtain the feature extraction result.

3. The optical remote sensing small-scale target detection method with adaptive local context embedding according to claim 1, characterized in that, First, input the optical remote sensing image, and after image preprocessing, perform multi-scale extraction of deep image features; in the neural network, since shallow information helps in small-scale target detection, the shallow information and deep information are fused by skip connection to obtain feature maps at different scales; finally, the feature maps at different scales after feature fusion are subjected to multi-scale aggregation, and finally a feature map with both shallow and deep feature information is output.