A Low-Altitude Remote Sensing Image Target Recognition Method Based on a Neighborhood Self-Attention Spatiotemporal Model

A target recognition method for low-altitude remote sensing images was constructed by using a neighborhood self-attention spatiotemporal model. This method solves the problems of target frame loss and deformation caused by the rapid movement of UAVs, and improves the recognition accuracy and stability of low-altitude remote sensing images, especially for small targets and targets in complex backgrounds.

CN121811204BActive Publication Date: 2026-05-26CHENGDU UNIV OF INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU UNIV OF INFORMATION TECH
Filing Date
2026-03-09
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing low-altitude remote sensing image target recognition methods cannot effectively handle the problems of target frame loss, deformation and distortion caused by the rapid movement of UAVs, and the recognition accuracy is insufficient in high dynamic scenes, especially the recognition effect is poor for small targets and complex backgrounds.

Method used

A target recognition method for low-altitude remote sensing images based on a neighborhood self-attention spatiotemporal model is adopted. The encoder module extracts and fuses spatial and spatiotemporal features, the decoder module fuses multi-scale features, and the target recognition module generates anchor points and performs recognition. The model is trained using historical time-series data, and the recognition effect is enhanced by combining a multi-anchor point selection mechanism and clustering technology.

Benefits of technology

It improves the stability and accuracy of recognition under dynamic conditions, enhances the detection rate of small and irregular targets, reduces missed detections, and improves boundary clarity and semantic discriminability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811204B_ABST
    Figure CN121811204B_ABST
Patent Text Reader

Abstract

This invention discloses a method for target recognition in low-altitude remote sensing images based on a neighborhood self-attention spatiotemporal model, belonging to the field of low-altitude remote sensing image target recognition technology. The method includes the following steps: constructing a low-altitude remote sensing image target recognition model based on a neighborhood self-attention spatiotemporal model, the low-altitude remote sensing image target recognition model including an encoder module, a decoder module, and a target recognition module connected in sequence; training the low-altitude remote sensing image target recognition model using historical time-series data of the low-altitude remote sensing image to obtain a trained low-altitude remote sensing image target recognition model; and obtaining the low-altitude remote sensing image target recognition result based on real-time time-series data of the low-altitude remote sensing image and the trained low-altitude remote sensing image target recognition model. This invention effectively solves the problems of frame loss, distortion, boundary blurring, and missed detection of small targets caused by high dynamic range and low contrast in low-altitude remote sensing images, improving the accuracy and robustness of target recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of low-altitude remote sensing image target recognition technology, and specifically to a low-altitude remote sensing image target recognition method based on a neighborhood self-attention spatiotemporal model. Background Technology

[0002] With the nation's strategic focus on the low-altitude economy, industries related to low-altitude drone applications are growing rapidly. Leveraging the advantages of drones such as low cost, high speed, and strong maneuverability, they can accomplish various tasks under diverse conditions and complex scenarios, and are currently widely used in security surveillance, road traffic control, express delivery, environmental monitoring, military operations, and many other fields.

[0003] Because low-altitude remote sensing images are typically acquired using drones, the size of the scene is related to battery life and flight speed. Therefore, the camera's viewpoint changes rapidly between frames during data acquisition. Highly dynamic scenes can lead to dropped frames, distortion, or targets appearing at different scales, structures, and shapes. Existing target recognition methods cannot effectively handle these issues, resulting in misclassification and blurred boundaries. Furthermore, when the shooting distance is too far, targets appear as a small percentage of pixel area and often form low contrast with complex backgrounds, sometimes lacking obvious color, shape, or texture features. Therefore, existing methods, relying solely on a limited number of visual features, are prone to missed detections, impacting recognition accuracy.

[0004] In conclusion, how to achieve accurate target detection and recognition using low-altitude remote sensing images under high dynamic range and low contrast is a research direction with great practical significance and an urgent problem to be solved. Summary of the Invention

[0005] To address the aforementioned shortcomings in existing technologies, this invention provides a low-altitude remote sensing image target recognition method based on a neighborhood self-attention spatiotemporal model.

[0006] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:

[0007] A target recognition method for low-altitude remote sensing images based on a neighborhood self-attention spatiotemporal model includes the following steps:

[0008] A low-altitude remote sensing image target recognition model is constructed based on a neighborhood self-attention spatiotemporal model. The low-altitude remote sensing image target recognition model includes an encoder module, a decoder module, and a target recognition module connected in sequence. The encoder module is used to extract and fuse spatial and spatiotemporal features from the low-altitude remote sensing image to obtain multi-scale features. The decoder module is used to fuse multi-scale features to obtain multi-scale fused features. The target recognition module is used to generate anchor points by clustering based on the multi-scale fused features, and obtain the target recognition result of the low-altitude remote sensing image based on the anchor points.

[0009] The target recognition model for low-altitude remote sensing images is trained using historical time-series data of low-altitude remote sensing images to obtain the trained target recognition model for low-altitude remote sensing images.

[0010] Based on real-time time-series data of low-altitude remote sensing images and a trained low-altitude remote sensing image target recognition model, the target recognition results of low-altitude remote sensing images are obtained.

[0011] Furthermore, the encoder module includes a first-scale spatiotemporal state neighborhood adaptive unit, a second-scale spatiotemporal state neighborhood adaptive unit, a third-scale spatiotemporal state neighborhood adaptive unit, and a fourth-scale spatiotemporal state neighborhood adaptive unit connected in sequence; each scale spatiotemporal state neighborhood adaptive unit includes a convolutional layer and a multi-scale spatiotemporal state neighborhood adaptive module connected in sequence.

[0012] Furthermore, the multi-scale spatiotemporal state neighborhood adaptation module includes parallel spatial domain branches and spatiotemporal domain branches, as well as a dual-gated fusion unit connected to the outputs of both the spatial domain branches and the spatiotemporal domain branches. The spatial domain branch is used to calculate spatial neighborhood attention and adaptively extract local details and boundary features in low-altitude remote sensing images through spatial neighborhood attention to obtain the final spatial features. The spatiotemporal domain branch is used to calculate temporal neighborhood attention and adaptively extract temporal dependencies between multiple frames in low-altitude remote sensing images through temporal neighborhood attention to obtain the final spatiotemporal features. The dual-gated fusion unit is used to dynamically adjust the weights of the final spatial features and the final spatiotemporal features according to regional characteristics and adaptively fuse them to obtain multi-scale features.

[0013] Furthermore, the expression for calculating spatial neighborhood attention is:

[0014] ,

[0015] ,

[0016]

[0017] in: The first in low-altitude remote sensing image The feature has a spatial neighborhood size of The attention below, The softmax activation function is used. The first in low-altitude remote sensing image Feature query projection and its Neighborhood attention prediction vectors among the nearest neighbors, The number of feature channels in a low-altitude remote sensing image. The first in low-altitude remote sensing image Features The set of value vectors within the n nearest neighbors, where each row vector is the nth row vector. Features The nearest neighbor projection value, The first in low-altitude remote sensing image The query vector of the feature. for The key vector of the first pixel, The first in low-altitude remote sensing image The feature is offset relative to the position of its first nearest neighbor. for The key vector of the second pixel, The first in low-altitude remote sensing image The feature is offset relative to its second nearest neighbor. for With the A key vector of 1 pixel, The first in low-altitude remote sensing image Features and their first The relative position offset of the nearest neighbor, Let the value vector be the spatial neighborhood. The first in low-altitude remote sensing image The first nearest neighbor of the feature The first in low-altitude remote sensing image The second nearest neighbor of the feature The first in low-altitude remote sensing image The first feature The nearest neighbor, This is the transpose symbol.

[0018] Furthermore, the expression for calculating temporal neighborhood attention is:

[0019] ,

[0020] ,

[0021]

[0022] in: The first in low-altitude remote sensing image The spatial feature of the first Each spatiotemporal sub-feature has a size of [value] in its spatiotemporal neighborhood. The attention below, The softmax activation function is used. The first in low-altitude remote sensing image The spatial feature of the first Attention weight vectors for each spatiotemporal sub-feature. The number of feature channels in a low-altitude remote sensing image. The first in low-altitude remote sensing image The spatial feature of the first The spatiotemporal neighborhood value matrix of each spatiotemporal sub-feature, each row of which corresponds to A value vector of spatiotemporal features in a neighborhood The first in low-altitude remote sensing image The spatial feature of the first A query vector for each spatiotemporal sub-feature. for The key vector of the first spatiotemporal sub-feature, It is the transpose symbol. The first in low-altitude remote sensing image The spatial feature of the first The relative position offset of each spatiotemporal sub-feature to its first nearest neighbor, for The key vector of the second spatiotemporal sub-feature, The first in low-altitude remote sensing image The spatial feature of the first The relative position offset of each spatiotemporal sub-feature to its second nearest neighbor, for With the Key vectors of spatiotemporal sub-features The first in low-altitude remote sensing image The spatial feature of the first The spatiotemporal sub-feature and its first The relative position offset of the nearest neighbor, The first in low-altitude remote sensing image The spatial feature of the first The neighborhood index of each spatiotemporal sub-feature and the first spatiotemporal sub-feature. The first in low-altitude remote sensing image The first feature The neighborhood index of the first spatiotemporal feature and the second spatiotemporal sub-feature. The first in low-altitude remote sensing image The first feature The first spatiotemporal feature and the second Neighborhood index of spatiotemporal sub-features.

[0023] Furthermore, the calculation expression for the dual-gated fusion unit is as follows:

[0024]

[0025] in: For multi-scale features, As the first gating weight, For the final spatial characteristics, The symbol for element-wise multiplication. As the second gating weight, This represents the final spatiotemporal characteristics.

[0026] Furthermore, the decoder module includes a multi-branch feature input layer, a first stitching layer, an element-gated attention fusion submodule, and a second stitching layer; the input of the multi-branch feature input layer is connected to the output of the encoder module, the output of the multi-branch feature input layer is simultaneously connected to the input of the first stitching layer and the input of the second stitching layer, the output of the first stitching layer is connected to the input of the element-gated attention fusion submodule, the output of the element-gated attention fusion submodule is connected to the input of the second stitching layer, and the output of the second stitching layer is connected to the target recognition module.

[0027] Furthermore, the element-gated attention fusion submodule includes three parallel gating units, and the calculation expression for each gating unit is:

[0028] ,

[0029]

[0030] in: For the first The intermediate features of a gated unit For the first The input of each gated unit, The first learning parameter for the gating unit. This is the second learning parameter for the gating unit. This is the third learning parameter for the gating unit. It is the ReLU activation function. For the Sigmoid function, The symbol for element-wise multiplication. This is the output of the gating unit.

[0031] The beneficial effects of this invention are as follows:

[0032] (1) This invention constructs a low-altitude remote sensing image target recognition model based on a neighborhood self-attention spatiotemporal model. The low-altitude remote sensing image target recognition model includes an encoder module, a decoder module, and a target recognition module connected in sequence. The encoder module can extract and fuse spatial and spatiotemporal features from the low-altitude remote sensing image to obtain multi-scale features. The decoder module can fuse multi-scale features to obtain multi-scale fused features. The target recognition module can cluster and generate anchor points based on the multi-scale fused features, and obtain the target recognition result of the low-altitude remote sensing image based on the anchor points. The entire low-altitude remote sensing image target recognition model can adaptively enhance key frames and local detail features, effectively alleviating the target frame loss, deformation, and distortion problems caused by the rapid movement of UAVs, and making the recognition result more stable and reliable under dynamic conditions.

[0033] (2) The decoder module constructed in this invention strengthens the clarity of boundaries and semantic discriminability by adopting a step-by-step information propagation and gating weighting mechanism, especially improving the detection rate and positioning accuracy of targets with small pixel proportions and similar backgrounds;

[0034] (3) The target recognition module constructed in this invention adopts a cluster-based multi-anchor selection mechanism, which gets rid of the constraints of traditional preset anchor frames, can adapt to the actual distribution of targets, significantly reduces the missed detection of small targets, and enhances the ability to recognize irregular and occluded targets. Attached Figure Description

[0035] Figure 1 This is a schematic diagram of the process for low-altitude remote sensing image target recognition based on a neighborhood self-attention spatiotemporal model. Detailed Implementation

[0036] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0037] like Figure 1 As shown, the low-altitude remote sensing image target recognition method based on the neighborhood self-attention spatiotemporal model includes steps S1-S3, as detailed below:

[0038] S1. A low-altitude remote sensing image target recognition model is constructed based on a neighborhood self-attention spatiotemporal model. The low-altitude remote sensing image target recognition model includes an encoder module, a decoder module, and a target recognition module connected in sequence. The encoder module is used to extract and fuse spatial and spatiotemporal features from the low-altitude remote sensing image to obtain multi-scale features. The decoder module is used to fuse multi-scale features to obtain multi-scale fused features. The target recognition module is used to generate anchor points by clustering based on the multi-scale fused features and obtain the target recognition result of the low-altitude remote sensing image based on the anchor points.

[0039] In an optional embodiment of the present invention, the encoder module includes a first-scale spatiotemporal state neighborhood adaptive unit, a second-scale spatiotemporal state neighborhood adaptive unit, a third-scale spatiotemporal state neighborhood adaptive unit, and a fourth-scale spatiotemporal state neighborhood adaptive unit connected in sequence; each scale spatiotemporal state neighborhood adaptive unit includes a convolutional layer and a multi-scale spatiotemporal state neighborhood adaptive module connected in sequence.

[0040] The multi-scale spatiotemporal state neighborhood adaptation module includes parallel spatial domain branches and spatiotemporal domain branches, as well as a dual-gated fusion unit connected to the outputs of both the spatial domain branches and the spatiotemporal domain branches. The spatial domain branch is used to calculate spatial neighborhood attention and adaptively extract local details and boundary features in low-altitude remote sensing images through spatial neighborhood attention to obtain the final spatial features. The spatiotemporal domain branch is used to calculate temporal neighborhood attention and adaptively extract temporal dependencies between multiple frames in low-altitude remote sensing images through temporal neighborhood attention to obtain the final spatiotemporal features. The dual-gated fusion unit is used to dynamically adjust the weights of the final spatial features and the final spatiotemporal features according to regional characteristics and adaptively fuse them to obtain multi-scale features.

[0041] The expression for calculating spatial neighborhood attention in this invention is as follows:

[0042] ,

[0043] ,

[0044]

[0045] in: The first in low-altitude remote sensing image The feature has a spatial neighborhood size of The attention below, The softmax activation function is used. The first in low-altitude remote sensing image Feature query projection and its Neighborhood attention prediction vectors among the nearest neighbors, The number of feature channels in a low-altitude remote sensing image. The first in low-altitude remote sensing image Features The set of value vectors within the n nearest neighbors, where each row vector is the nth row vector. Features The nearest neighbor projection value, The first in low-altitude remote sensing image The query vector of the feature. for The key vector of the first pixel, The first in low-altitude remote sensing image The feature is offset relative to the position of its first nearest neighbor. for The key vector of the second pixel, The first in low-altitude remote sensing image The feature is offset relative to its second nearest neighbor. for With the A key vector of 1 pixel, The first in low-altitude remote sensing image Features and their first The relative position offset of the nearest neighbor, Let the value vector be the spatial neighborhood. The first in low-altitude remote sensing image The first nearest neighbor of the feature The first in low-altitude remote sensing image The second nearest neighbor of the feature The first in low-altitude remote sensing image The first feature The nearest neighbor, This is the transpose symbol.

[0046] Specifically, the spatial neighborhood attention in the spatial domain branch is responsible for capturing the spatial dependencies between any two locations in the feature map. This not only fully utilizes the local smoothness of space but also enhances the perception of local detailed features. Finally, this invention uses spatial neighborhood attention to... Spatial features are obtained by weighting the inner pixels. .

[0047] Then, this invention inputs spatial features into the SS2D layer (2D-Selective-Scan) to perform spatial dimension selective scanning, establishing long-distance spatial dependencies with linear complexity. Spatial dimension selective scanning traverses each pixel along the spatial dimension, and each sequence element corresponds to the d×1 spectral feature of a spatial pixel. The calculation expression for the SS2D layer in the spatial domain branch is:

[0048] ,

[0049] ,

[0050]

[0051] Where z∈{1,2,3,4} represents four different scanning directions. For the scan unfolding operation, For scan merging operations, Spatial features are spatially embedded features processed by a selectively scanned spatial state sequence model. This is a selective scanning spatial state sequence model. For spatial embedding features, This represents the final spatial characteristics.

[0052] The expression for calculating temporal neighborhood attention in this invention is as follows:

[0053] ,

[0054] ,

[0055]

[0056] in: The first in low-altitude remote sensing image The spatial feature of the first Each spatiotemporal sub-feature has a size of [value] in its spatiotemporal neighborhood. The attention below, The softmax activation function is used. The first in low-altitude remote sensing image The spatial feature of the first Attention weight vectors for each spatiotemporal sub-feature. The number of feature channels in a low-altitude remote sensing image. The first in low-altitude remote sensing image The spatial feature of the first The spatiotemporal neighborhood value matrix of each spatiotemporal sub-feature, each row of which corresponds to A value vector of spatiotemporal features in a neighborhood The first in low-altitude remote sensing image The spatial feature of the first A query vector for each spatiotemporal sub-feature. for The key vector of the first spatiotemporal sub-feature, It is the transpose symbol. The first in low-altitude remote sensing image The spatial feature of the first The relative position offset of each spatiotemporal sub-feature to its first nearest neighbor, for The key vector of the second spatiotemporal sub-feature, The first in low-altitude remote sensing image The spatial feature of the first The relative position offset of each spatiotemporal sub-feature to its second nearest neighbor, for With the Key vectors of spatiotemporal sub-features The first in low-altitude remote sensing image The spatial feature of the first The spatiotemporal sub-feature and its first The relative position offset of the nearest neighbor, The first in low-altitude remote sensing image The spatial feature of the first The neighborhood index of each spatiotemporal sub-feature and the first spatiotemporal sub-feature. The first in low-altitude remote sensing image The first feature The neighborhood index of the first spatiotemporal feature and the second spatiotemporal sub-feature. The first in low-altitude remote sensing image The first feature The first spatiotemporal feature and the second Neighborhood index of spatiotemporal sub-features.

[0057] Specifically, the temporal-temporal domain branch's temporal neighborhood attention focuses on relevant features within the temporal domain, fully utilizing the smooth changes in local temporal features, adaptively weighting key features, and enhancing the ability to perceive global context and capture long-distance spatial dependencies, thereby highlighting key features and reducing information redundancy. Finally, this invention uses temporal neighborhood attention to focus on neighborhood... Spatiotemporal features are obtained by weighting the inner pixels. .

[0058] Then, this invention inputs the spatiotemporal features into the SS2D layer (2D-Selective-Scan) to perform a spatiotemporal sequence dimension selective scan, establishing temporal context dependencies with linear complexity. The temporal dimension selective scan involves traversing the features of each frame along the time series. Each sequence element corresponds to a w×h spatial feature under a temporal frame, where w is the width of the input low-altitude remote sensing image and h is the height of the input low-altitude remote sensing image. The calculation expression for the SS2D layer in the spatiotemporal domain branch is:

[0059] ,

[0060] ,

[0061]

[0062] Where z∈{1,2,3,4} represents four different scanning directions. For the scan unfolding operation, For scan merging operations, The spatiotemporal features are spatiotemporal features after selective scanning of spatial state sequence model for spatiotemporal embedding features. This is a selective scanning spatial state sequence model. For spatiotemporal embedding features, This represents the final spatiotemporal characteristics.

[0063] The calculation expression for the dual-gated fusion unit is:

[0064]

[0065] in: For multi-scale features, As the first gating weight, For the final spatial characteristics, The symbol for element-wise multiplication. As the second gating weight, This represents the final spatiotemporal characteristics.

[0066] The decoder module includes a multi-branch feature input layer, a first stitching layer, an element-gated attention fusion submodule, and a second stitching layer. The input of the multi-branch feature input layer is connected to the output of the encoder module, and the output of the multi-branch feature input layer is simultaneously connected to the input of the first stitching layer and the input of the second stitching layer. The output of the first stitching layer is connected to the input of the element-gated attention fusion submodule, the output of the element-gated attention fusion submodule is connected to the input of the second stitching layer, and the output of the second stitching layer is connected to the target recognition module.

[0067] The element-gated attention fusion submodule consists of three parallel gating units, each of which includes a convolutional layer, two fully connected layers, and a sigmoid function for output normalization.

[0068] Specifically, the multi-branch feature input layer includes a first feature input layer, a second feature input layer, and a third feature input layer. The first and second feature input layers each include a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a downsampling layer connected in sequence. The third feature input layer includes a convolutional layer, a batch normalization layer, and a ReLU activation function layer connected in sequence.

[0069] The element-gating attention fusion submodule includes three parallel gating units. The calculation expression for each gating unit is:

[0070] ,

[0071]

[0072] in: For the first The intermediate features of a gated unit For the first The input of each gated unit, The first learning parameter for the gating unit. This is the second learning parameter for the gating unit. This is the third learning parameter for the gating unit. It is the ReLU activation function. For the Sigmoid function, The symbol for element-wise multiplication. This is the output of the gating unit.

[0073] This invention uses multiple gating units to balance information at different scales, optimizes feature combination by dynamically adjusting the contribution of each scale, adjusts the weight of each feature element by element, enhances global feature representation, and suppresses noise and irrelevant features. Then, this invention uses SS2D to model the fused global context features, learning a more discriminative global context feature F4, whose expression is:

[0074] ,

[0075] ,

[0076]

[0077] Where z∈{1,2,3,4} represents four different scanning directions. For the scan unfolding operation, For scan merging operations, For contextual features, These are the context features processed by a selective scan spatial state order model. This is a selective scanning spatial state sequence model. For more discriminative global context features.

[0078] This invention utilizes a second splicing layer to fuse F4 with the outputs F1 of the first feature input layer, F2 of the second feature input layer, and F3 of the third feature input layer to obtain multi-scale fused features. .

[0079] Specifically, for the target recognition module:

[0080] Conventional deep learning-based object detection algorithms all require pre-defined anchor boxes for object classification and location regression. Feature extraction becomes difficult when detecting small targets or occluded objects, and they cannot achieve fine-grained object detection at different scales. Furthermore, the anchor boxes contain reference size and scale parameters, introducing numerous hyperparameters, increasing computational complexity, and making recognition accuracy susceptible to hyperparameter influence. To avoid the difficulty of pre-defined anchor box configuration, reduce the number of hyperparameters used, and adapt to object detection tasks with different sizes and aspect ratios, this invention designs a clustering-based multi-anchor selection mechanism in the object recognition module. This mechanism explicitly pushes target sample features to multiple anchor points to obtain a better potential representation. This invention first uses K-means clustering for each class of samples, using the cluster centers as anchor points, and clusters them by minimizing the following errors. There are several clusters, and their expressions are:

[0081] ,

[0082]

[0083] in: For L2 distance, The centroid of the cluster, The number of clusters, The number of images in the cluster. This is a multi-scale fusion feature.

[0084] The present invention then measures each sample in this class with its nearest anchor point (centroid). The distance between anchor points is used as the anchor point, and the furthest distance is selected. Then, the present invention generates multiple anchor boxes (not exceeding the furthest distance) around the anchor point, forming a cluster structure containing target objects of the same category. This is more advantageous for detection tasks that need to handle irregular targets and targets with large aspect ratio differences. The number of anchor points is affected by the number of target types, and multiple anchor points can fully characterize their distribution characteristics. Finally, cross-entropy loss and smoothed L1 loss are used to predict the category probability and location of each anchor box, ultimately obtaining the target location and category. Thus, the present invention avoids the problem in images containing small targets where the small target region is too small, resulting in a limited number of matching anchor points and thus reducing the probability of small targets being detected.

[0085] S2. Train the low-altitude remote sensing image target recognition model using historical time-series data of low-altitude remote sensing images to obtain the trained low-altitude remote sensing image target recognition model.

[0086] In an optional embodiment of the present invention, the present invention utilizes historical time-series data of collected and labeled low-altitude remote sensing images. First, the data is preprocessed by standardization and registration. Then, the preprocessed image sequence is input into the low-altitude remote sensing image target recognition model. The encoder extracts and fuses spatiotemporal features, and the decoder fuses multi-scale contextual information. The target recognition module outputs the prediction result based on anchor points generated by clustering. During training, a multi-task loss function composed of cross-entropy loss and smoothing L1 loss is used as the optimization objective. The backpropagation algorithm is used to iteratively update all parameters of the model until the loss converges, thereby obtaining the trained low-altitude remote sensing image target recognition model.

[0087] S3. Based on real-time time-series data of low-altitude remote sensing images and the trained low-altitude remote sensing image target recognition model, obtain the target recognition results of low-altitude remote sensing images.

[0088] In an optional embodiment of the present invention, the real-time time series data of low-altitude remote sensing images are input into a trained low-altitude remote sensing image target recognition model to obtain low-altitude remote sensing image target recognition results.

[0089] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.

Claims

1. A method for target recognition in low-altitude remote sensing images based on a neighborhood self-attention spatiotemporal model, characterized in that, Includes the following steps: A low-altitude remote sensing image target recognition model is constructed based on a neighborhood self-attention spatiotemporal model. The low-altitude remote sensing image target recognition model includes an encoder module, a decoder module, and a target recognition module connected in sequence. The encoder module is used to extract and fuse spatial and spatiotemporal features from low-altitude remote sensing images to obtain multi-scale features. The decoder module is used to fuse multi-scale features to obtain multi-scale fused features; The target recognition module is used to generate anchor points by clustering based on multi-scale fusion features, and to obtain target recognition results of low-altitude remote sensing images based on the anchor points; The encoder module includes a first-scale spatiotemporal state neighborhood adaptive unit, a second-scale spatiotemporal state neighborhood adaptive unit, a third-scale spatiotemporal state neighborhood adaptive unit, and a fourth-scale spatiotemporal state neighborhood adaptive unit connected in sequence; each scale spatiotemporal state neighborhood adaptive unit includes a convolutional layer and a multi-scale spatiotemporal state neighborhood adaptive module connected in sequence. The multi-scale spatiotemporal state neighborhood adaptation module includes parallel spatial domain branches and spatiotemporal domain branches, as well as a dual-gated fusion unit that is simultaneously connected to the output of the spatial domain branch and the output of the spatiotemporal domain branch. The spatial domain branch is used to calculate spatial neighborhood attention and to adaptively extract local details and boundary features in low-altitude remote sensing images through spatial neighborhood attention to obtain the final spatial features. The spatiotemporal branch is used to compute temporal neighborhood attention, and the temporal dependencies between multiple frames in low-altitude remote sensing images are adaptively extracted through temporal neighborhood attention to obtain the final spatiotemporal features; The dual-gated fusion unit is used to dynamically adjust the weights of the final spatial features and the final spatiotemporal features according to the regional characteristics, and adaptively fuse them to obtain multi-scale features. The target recognition model for low-altitude remote sensing images is trained using historical time-series data of low-altitude remote sensing images to obtain the trained target recognition model for low-altitude remote sensing images. Based on real-time time-series data of low-altitude remote sensing images and a trained low-altitude remote sensing image target recognition model, the target recognition results of low-altitude remote sensing images are obtained.

2. The low-altitude remote sensing image target recognition method based on a neighborhood self-attention spatiotemporal model according to claim 1, characterized in that, The expression for calculating spatial neighborhood attention is: , , in: The first in the low-altitude remote sensing image The feature has a spatial neighborhood size of The attention below, The softmax activation function is used. The first in the low-altitude remote sensing image Feature query projection and its Neighborhood attention prediction vectors among the nearest neighbors, The number of feature channels in a low-altitude remote sensing image. The first in low-altitude remote sensing image Features The set of value vectors within the n nearest neighbors, where each row vector is the nth row vector. Features The nearest neighbor projection value, The first in the low-altitude remote sensing image The query vector of the feature. for The key vector of the first pixel, The first in low-altitude remote sensing image The feature is offset relative to the position of its first nearest neighbor. for The key vector of the second pixel, The first in the low-altitude remote sensing image The feature is offset relative to its second nearest neighbor. for With the A key vector of 1 pixel, The first in low-altitude remote sensing image Features and their first The relative position offset of the nearest neighbor, Let the value vector be the spatial neighborhood. The first in low-altitude remote sensing image The first nearest neighbor of the feature The first in the low-altitude remote sensing image The second nearest neighbor of the feature The first in the low-altitude remote sensing image The first feature The nearest neighbor, This is the transpose symbol.

3. The low-altitude remote sensing image target recognition method based on a neighborhood self-attention spatiotemporal model according to claim 1, characterized in that, The expression for calculating temporal neighborhood attention is: , , in: The first in the low-altitude remote sensing image The spatial feature of the first Each spatiotemporal sub-feature has a size of [value] in its spatiotemporal neighborhood. The attention below, The softmax activation function is used. The first in the low-altitude remote sensing image The spatial feature of the first Attention weight vectors for each spatiotemporal sub-feature. The number of feature channels in a low-altitude remote sensing image. The first in the low-altitude remote sensing image The spatial feature of the first The spatiotemporal neighborhood value matrix of each spatiotemporal sub-feature, each row of which corresponds to A value vector of spatiotemporal features in a neighborhood The first in the low-altitude remote sensing image The spatial feature of the first A query vector for each spatiotemporal sub-feature. for The key vector of the first spatiotemporal sub-feature, It is the transpose symbol. The first in the low-altitude remote sensing image The spatial feature of the first The relative position offset of each spatiotemporal sub-feature to its first nearest neighbor, for The key vector of the second spatiotemporal sub-feature, The first in the low-altitude remote sensing image The spatial feature of the first The relative position offset of each spatiotemporal sub-feature to its second nearest neighbor, for With the Key vectors of spatiotemporal sub-features The first in the low-altitude remote sensing image The spatial feature of the first The spatiotemporal sub-feature and its first The relative position offset of the nearest neighbor, The first in the low-altitude remote sensing image The spatial feature of the first The neighborhood index of each spatiotemporal sub-feature and the first spatiotemporal sub-feature. The first in the low-altitude remote sensing image The first feature The neighborhood index of the first spatiotemporal feature and the second spatiotemporal sub-feature. The first in the low-altitude remote sensing image The first feature The first spatiotemporal feature and the second Neighborhood index of spatiotemporal sub-features.

4. The low-altitude remote sensing image target recognition method based on a neighborhood self-attention spatiotemporal model according to claim 1, characterized in that, The calculation expression for the dual-gated fusion unit is: in: For multi-scale features, As the first gating weight, For the final spatial characteristics, The symbol for element-wise multiplication. As the second gating weight, This represents the final spatiotemporal characteristics.

5. The low-altitude remote sensing image target recognition method based on a neighborhood self-attention spatiotemporal model according to claim 1, characterized in that, The decoder module includes a multi-branch feature input layer, a first stitching layer, an element-gated attention fusion submodule, and a second stitching layer. The input of the multi-branch feature input layer is connected to the output of the encoder module, and the output of the multi-branch feature input layer is simultaneously connected to the input of the first stitching layer and the input of the second stitching layer. The output of the first stitching layer is connected to the input of the element-gated attention fusion submodule, the output of the element-gated attention fusion submodule is connected to the input of the second stitching layer, and the output of the second stitching layer is connected to the target recognition module.

6. The low-altitude remote sensing image target recognition method based on a neighborhood self-attention spatiotemporal model according to claim 5, characterized in that, The element-gating attention fusion submodule includes three parallel gating units. The calculation expression for each gating unit is: , in: For the first The intermediate features of a gated unit For the first The input of each gated unit, The first learning parameter for the gating unit. This is the second learning parameter for the gating unit. This is the third learning parameter for the gating unit. It is the ReLU activation function. For the Sigmoid function, The symbol for element-wise multiplication. This is the output of the gating unit.