A Robust Recognition Method for Small-Sample Remote Sensing Targets with Cross-Scene Multi-Domain Fusion

Through improved domain mapping and meta-learning strategies, improved mask autoencoder and multi-domain heterogeneous image registration technology, combined with cross-modal attention mechanism and lightweight object detection model, problems such as small sample learning, domain offset and multi-modal data fusion in remote sensing image processing are solved, and high-precision and robust object detection are achieved.

CN118918476BActive Publication Date: 2025-06-03BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410997142.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2025-06-03
Estimated Expiration
2044-07-24

AI Technical Summary

Technical Problem

In an adversarial environment, remote sensing image processing and object detection and recognition face many challenges, including small sample learning, domain offset, low multimodal data fusion accuracy, insufficient feature extraction and high detection model complexity, resulting in insufficient accuracy and robustness of object detection.

Method used

Improved domain mapping strategy and meta-learning strategy are used to generalize small sample learning across scenes, combined with improved mask autoencoder for characterization learning, feature extraction and fusion is performed through multi-domain heterogeneous image registration technology and cross-modal attention mechanism, and a lightweight object detection model is built to improve real-time and computational efficiency.

Benefits of technology

It significantly improves the accuracy and robustness of object detection, reduces dependence on large amounts of sample data, enhances the adaptability and stability of the model in different scenarios, and meets the requirements of real-time detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118918476B_ABST
    Figure CN118918476B_ABST
Patent Text Reader

Abstract

The present invention provides a robust recognition method for small-sample remote sensing targets with cross-scenario multi-domain fusion, belonging to the technical field of target recognition. This method fuses cross-scenario multi-domain multi-modal remote sensing images, and uses meta-learning strategies and masked autoencoders for representation learning, significantly improving the efficient learning and accurate generalization ability of the recognition model under limited labeled data and complex and changeable environments; through improved multi-modal image registration technology and improved cross-modal multi-head self-attention mechanism, combined with improved multi-level feature fusion method, the registration accuracy and feature expression ability are improved; based on a lightweight real-time target recognition model of an improved deep neural network, combined with the anchor box mechanism and non-maximum suppression algorithm, the real-time and accuracy of target detection are realized. The present invention innovates cross-scenario generalization small-sample learning and representation learning methods, innovates multi-domain heterogeneous image registration, fusion and lightweight real-time target recognition technologies, and significantly improves cross-scenario target recognition accuracy and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target recognition, and particularly to a robust recognition method for small-sample remote sensing targets with cross-scene multi-domain fusion. By fusing cross-scene multi-domain multi-modal remote sensing image data, the accuracy and robustness of target detection and recognition are improved. Background Art

[0002] Remote sensing image processing and target detection technologies play an important role in fields such as environmental monitoring, disaster assessment, urban planning, etc. By acquiring and analyzing remote sensing images, precise monitoring and recognition of surface objects and phenomena can be achieved, thus providing crucial decision-making support. The application of these technologies can not only improve the efficiency of resource management and environmental protection, but also enhance the reconnaissance and monitoring capabilities in the field of security.

[0003] However, remote sensing image processing and target detection and recognition in adversarial environments face many challenges and problems. In adversarial environments, advanced combat equipment often has a high degree of concealment and technological novelty, which makes it extremely difficult to obtain sufficient target data through conventional means before the war. Coupled with the rapidly changing war environment, it is very difficult to accumulate enough sample data for model training in a short time. The lack of sample data directly limits the effective perception and recognition capabilities of new or rare targets. Due to the short intelligence collection window period, it is difficult to widely collect key information data such as the enemy's new weaponry and tactical layouts, resulting in a very limited amount of available data. Moreover, the successfully obtained target information data often focuses on specific scenarios or conditions, which is quite different from the diverse environments in actual combat. The target recognition model constructed based on these limited and scenario-restricted data often exhibits poor generalization performance when facing new environments or new targets, and is unable to accurately recognize or adapt to the characteristic changes of targets in different environments.

[0004] When traditional intelligent algorithms handle such small-sample learning problems, their feature representation capabilities are limited, and it is difficult to extract sufficient and representative features from a small number of samples. In addition, their generalization capabilities are insufficient. Even if they perform well on the training data, they cannot accurately predict or classify unseen samples. Especially in the case of cross-scenes, the training set and the test set come from different domains and have different data distributions, while in traditional intelligent algorithms, the training set and the test set come from the same domain. The resulting domain shift problem will have a negative impact on the performance of small-sample learning, seriously hindering the rapid and accurate recognition of targets in complex and changing battlefield environments, and thus affecting the timeliness and accuracy of decision-making.

[0005] In addition, it is difficult for single-modal data to comprehensively reflect the target features in complex scenarios, thus it is difficult to ensure high-precision target detection. Remote sensing images come from a wide range of sources, including multiple modalities such as visible light, infrared, synthetic aperture radar (SAR), etc. The data of each modality has different imaging principles and characteristics, and can provide information in different aspects. For example, visible light images can provide rich color and texture information, infrared images can acquire data at night or under adverse weather conditions, and SAR images can penetrate clouds and vegetation to provide surface structure information. In remote sensing image processing, images of different modalities each have unique advantages, but also have their own limitations: Visible light images can provide rich color and texture information, have strong detail representation ability, are intuitive and easy to understand, but are sensitive to lighting conditions. At night or under adverse weather, the imaging quality drops significantly, and affected by occlusion, they cannot penetrate occlusions such as clouds, smoke or vegetation, which limits their application in complex environments. Infrared images can acquire data at night or under adverse weather conditions, are sensitive to temperature differences, and are suitable for detecting heat sources, but have low resolution. Usually, the resolution of infrared images is lower than that of visible light images, with poor detail representation ability, and at the same time, the lack of color information leads to poor visual intuitiveness of the images. Affected by heat interference, heat sources and temperature changes will affect the imaging effect, and there is a possibility of false alarms. Synthetic aperture radar (SAR) images can penetrate clouds and vegetation to provide surface structure information, are not affected by lighting conditions, and are suitable for all-weather imaging, but have high noise. Speckle noise often exists in SAR images, affecting image quality and detail representation, being difficult to interpret, with poor image intuitiveness, and it is difficult to directly extract useful information from the images, requiring professional processing and analysis. Geometric distortion, due to different imaging principles, SAR images may have geometric distortion and need geometric correction. By combining the advantages of these modalities and overcoming their respective disadvantages, it is possible to detect and identify targets in complex scenarios more comprehensively and accurately, and improve the overall performance of remote sensing image processing.

[0006] The registration of multimodal data is a prerequisite for data fusion. However, due to geometric distortions, resolution differences, and changes in imaging conditions between different modal images, traditional registration methods (such as gray-scale or feature-point-based methods) are difficult to achieve high-precision registration in complex scenarios. This directly affects the accuracy of subsequent feature extraction and fusion. Second, there is insufficient feature extraction. Feature extraction is a key step in object detection. Existing methods mostly rely on single-modal feature extraction or simple multimodal feature fusion, and fail to fully exploit and utilize the complementary information of multimodal data. For example, traditional handcrafted feature extraction methods (such as SIFT, SURF) are not robust enough in complex backgrounds and between different modalities, while some deep learning methods, although able to extract relatively rich features, often fail to effectively combine the advantages of each modality when dealing with multimodal data. Third, the complexity of the detection model is high. To improve the accuracy of object detection, some existing methods have designed very complex detection models. These models usually contain a large number of parameters and computational amounts, and it is difficult to meet the real-time and computational resource limitations in practical applications. For example, although the detection method based on deep neural networks has high accuracy, the computational complexity during training and inference is relatively high, which limits its promotion in real-time applications. Finally, the multimodal data fusion method is insufficient: In the existing technology, although there are some multimodal data fusion methods, when the existing methods deal with the fusion of different modal features, there are often problems of information loss and redundancy, and due to the use of the stitching weighted fusion strategy, the complementary nature of different modal data cannot be fully exploited, resulting in insufficient expressive ability of the fused features.

[0007] To address the above problems, the present invention proposes a cross-scenario generalization few-shot learning based on improved domain mapping and a lightweight real-time object recognition model based on a deep neural network, which has the following innovation points and advantages:

[0008] It comprehensively applies the meta-learning strategy and representation learning based on an improved masked autoencoder to improve the accuracy and generalization ability of remote sensing object recognition. By adopting the meta-learning strategy to construct a K-way-N-shot learning scenario, only a small number of labeled samples are required to drive the rapid iterative upgrade of the object recognition model. When encountering a new task, it can quickly adapt and accurately execute the classification task, significantly reducing the dependence on a large amount of sample data.

[0009] Pre-train using a masked autoencoder without relying on labeled data for self-supervised learning of image features. The masked autoencoder randomly occludes some regions of the input image and uses the efficient encoder of Vision Transformer (ViT) to generate low-dimensional latent variables containing global context and local details. The encoder part is used as a feature extractor to effectively decouple domain features and target features, thereby establishing a common feature space between the source domain and the target domain, which helps to narrow the data distribution difference.

[0010] Adopt the multi - domain heterogeneous image registration technology based on feature matching. Extract significant boundary features and feature points by a gray - level - based method. Dynamically adjust the Scale - Invariant Feature Transform (SIFT) threshold, and calculate the weight value to strengthen important feature points, ensuring high - precision registration in complex scenarios. Secondly, fuse the registration results based on feature points and boundary features, use the incremental optimization method to improve the registration accuracy, and design new weights according to the accuracy of object detection, iteratively optimizing the registration algorithm to improve its adaptability in different scenarios.

[0011] Use the Deep Residual Network (ResNet - 101) to extract features from the registered visible light, infrared, and SAR images, and enhance the high - level and low - level features through the cross - modal attention mechanism. Adopt a multi - level feature fusion strategy, splice and reduce the dimension of the high - level and low - level features to form the output feature map of the encoder part, fully retaining important information.

[0012] Construct a lightweight object detection model. By reducing the number of network layers and optimizing the network structure, use lightweight modules (such as depth - separable convolution) to reduce the number of parameters and computational complexity to ensure the efficient operation of the model in resource - constrained environments. Use the anchor box mechanism to delimit the bounding boxes, preset multi - scale and multi - ratio anchor boxes on each grid cell of the input image, and adjust the position and size of the anchor boxes through model learning to achieve the precise positioning of the target object. Take the maximum IoU value as the loss function. During the training process, optimize the model parameters through backpropagation to improve the prediction accuracy, and perform post - processing through the Non - Maximum Suppression (NMS) algorithm. Sort the detection results by confidence, starting from the bounding box with the highest confidence, compare the IoU values of other bounding boxes one by one, remove redundant detections, and retain the optimal detection results to ensure the accuracy and effectiveness of the final detection results.

[0013] Through these innovative points, the present invention can solve problems such as low registration accuracy of multi - modal remote - sensing images, insufficient feature extraction, high complexity of the detection model, and poor multi - modal data fusion effect, thereby achieving high - precision and high - efficiency object detection in complex scenarios and significantly improving the reliability and practicality of remote - sensing image processing. Summary of the Invention

[0014] In view of this, the present invention provides a robust recognition method for cross - scene multi - domain fusion of small - sample remote - sensing targets, reducing or improving the problem of domain shift caused by large differences in feature data distributions between the source - domain dataset and the target - domain dataset in small - sample learning in the prior art, and solving problems such as low registration accuracy of multi - modal data fusion, insufficient feature extraction, and high complexity of the detection model, thereby improving the accuracy and robustness of object detection.

[0015] The present invention provides a cross-scenario generalization few-shot learning based on improved domain mapping, which is characterized by adopting a meta-learning strategy to construct a K-way-N-shot learning framework, quickly iteratively updating the model through a small number of labeled samples, and enhancing the adaptability of the target recognition model to new tasks. Based on a certain degree of similarity between the training set categories (old categories) and the test categories (new categories), a model that performs well on the training set can also be applicable to the new categories. Meta-learning constructs the training set into a series of few-shot tasks with the same settings as the test tasks, and then trains the model to perform well on a large number of few-shot tasks to ensure that it can also quickly classify when adapting to new tasks. During the learning process, a few-shot classifier and a domain classifier are used. The method includes the following steps:

[0016] Pre-training stage: Use an improved masked autoencoder to train on the source domain dataset, and perform self-supervised learning of image features. Take the encoder module of the trained masked autoencoder as the feature extractor, and the feature extractor decouples the image features of the source domain and the target domain to obtain domain-invariant features and domain features. Map the features of the target domain to the source domain feature space, thereby helping to narrow the difference in data distribution between the source domain and the target domain.

[0017] Meta-learning stage: In order to weaken the impact of domain shift on the performance of few-shot learning and reduce the distribution difference, the core idea is to map the features of the target domain to the source domain, and guide the learning of the model by adding a very small number of labeled data belonging to the target domain to the training dataset, that is, use a very small amount of labeled data in the target domain as auxiliary data to improve the ability of cross-domain few-shot learning to cope with domain shift. Set a very small number of samples in the target domain as the auxiliary dataset through the dataset mixer, and fuse it with the dataset in the source domain to form a mixed dataset, and re-partition it into a source domain support set, an auxiliary support set, and a mixed query set, effectively promoting the feature fusion and migration between the source domain and the target domain, and alleviating the domain shift phenomenon. Send the source domain support set, the auxiliary support set, and the mixed query set into the feature extractor to obtain their respective domain-invariant features and domain features. Denote the domain-invariant features and domain features by H1 and H2 respectively, then the domain-invariant features and domain features of the source domain support set are respectively expressed as and The domain-invariant features and domain features of the auxiliary support set are respectively expressed as and The domain-invariant features and domain features of the mixed query set are respectively expressed as and Send the domain-invariant features into the few-shot classifier g fsl for classification and compare with the label to obtain the loss function L FSL ; Send the domain-invariant features and domain features into the domain classifier g dom for domain classification and compare with the domain label to obtain the loss function and Overall loss Optimize L using the Adam gradient descent optimizer to minimize L.

[0018] The described few-shot classifier is characterized in that With Input the few-shot classifier to obtain the classification loss With Input the few-shot classifier to obtain the classification loss Since the mixed query set is a mixture of the source domain query set and the auxiliary query set with a ratio of λ, the loss function of the few-shot classifier is

[0019] The described domain classifier is characterized in that for domain-related features, the label of the source domain is set to 1 and the label of the target domain is set to 0, and for domain-unrelated features, the domain label vector is set to y1 = [0.5, 0.5]; For domain-unrelated features, input the domain classifier g dom Classify in and compare with the label y1, and calculate the loss using the Kullback-Leibler divergence (KL) For domain-related features, input and compare with their respective domain labels, and calculate the loss using cross-entropy (CE)

[0020] The present invention provides a representation learning based on an improved masked autoencoder, which is characterized in that based on a random masking mechanism, randomly partial regions of the input image are masked, and the original complete image is predicted through a reconstruction operation process to achieve effective encoding of image data; in the masked autoencoder, the encoder is composed of a combination of VisionTransformer (ViT) modules, specifically ViT-L / 16, which is stacked by 24 layers of ViT modules, and each image patch is mapped to a 1024-dimensional feature vector; the decoder can adopt a lightweight Transformer module, which is stacked by 8 layers of ViT modules, and each image patch is mapped to a 512-dimensional feature vector, and the computational amount of the decoder is only 10% of that of the encoder, improving the training and execution efficiency; since it is necessary to infer complete information from the incomplete input, the masked autoencoder learns how to capture global context information and local details, thereby generating a more robust and generalizable feature representation, enabling the masked autoencoder to also have a good performance under cross-scene conditions.

[0021] Among them, ViT is a model based on the Transformer architecture. By modifying the encoder part of the traditional Transformer to adapt to computer vision tasks, it marks the successful migration of the Transformer architecture from the natural language processing (NLP) field to the computer vision field. It mainly includes the following key modules:

[0022] Input processing: First, assume that the dimensions of the input image are H×W×C, representing height, width, and number of channels respectively, and it is divided into N fixed-size patches. The dimension of each patch block is P×P×C, then N = HW / P 2 . Flatten each patch block into a one-dimensional vector with a dimension of P 2 C, and then map it to a vector space with a dimension of D through a linear layer to obtain the embedding vector. This process helps the ViT model learn a higher-level feature representation.

[0023] Position encoding: Since the Transformer architecture itself does not contain position information, position encoding vectors need to be added to the embedding vectors corresponding to each patch to inform the model of the position of each patch in the original image. These position encoding vectors are added to the result of the patch's embedding vector and used as the input to the Transformer encoder together.

[0024] Transformer encoder: After the input sequence undergoes layer normalization, it passes through the multi-head self-attention layer to concurrently focus on different positions of the input sequence, and calculates the Query, Key, and Value matrices to enable the ViT model to understand the dependencies between different patches. Then it passes through the feed-forward neural network layer, including two linear layers and a non-linear activation function (such as the ReLU function), for further transforming and enriching the feature representation. Residual connections are used after both the self-attention layer and the feed-forward neural network layer to alleviate the problems of gradient disappearance and gradient explosion, promote feature extraction, and accelerate the training process.

[0025] The present invention provides an improved multi-domain heterogeneous image registration technology based on feature matching, characterized in that this method performs registration by extracting significant boundary features and feature points based on a gray-scale method to ensure high-precision registration in complex scenarios. This method includes the following steps:

[0026] Multi-modal Remote Sensing Image Registration Based on Local Boundary Features: According to the inlier maximization and redundant point control method, feature points are extracted by extracting significant boundary features and using a gray-scale-based method. The scale-invariant feature transform (SIFT) threshold is dynamically adjusted to extract feature points, and weight values are calculated to strengthen important feature points, ensuring high-precision registration in complex scenarios. The method for dynamically adjusting the SIFT threshold is carried out through the following formula:

[0027]

[0028] where α is the adjustment coefficient, and the intensity of the feature point i is the intensity of the i-th feature point, and n is the total number of feature points.

[0029] Based on an improved method for extracting significant boundary feature points, multi-modal remote sensing image registration of local boundary features is realized. The specific process includes:

[0030] After inputting the reference image R and the image S to be registered, where the reference image R is the standard image for reference, and the image S to be registered is the image to be aligned with the reference image. By extracting significant boundary features and using a gray-scale-based method to extract feature points, the scale-invariant feature transform (SIFT) threshold is dynamically adjusted, feature points are extracted and weight values are calculated to strengthen important feature points, ensuring high-precision registration in complex scenarios. The specific implementation process includes: for the general case of the input image, that is, the image I of either the reference image R or the image S to be registered, calculate its set of significant boundary feature points F = {f i | i = 1, 2, …, n}, and the weight value of the feature point where, g i is the gray value of the feature point f i , and α and β are adjustment parameters;

[0031] The improved incremental optimization method detects boundaries and boundary feature points, calculates local boundary features and boundary descriptors, and performs feature matching. The specific implementation process includes: at each iteration step t, calculate the loss L total = L match + L boundary , and update the parameters based on the gradient of the loss function

[0032] Based on the incremental optimization method, fuse the registration results based on feature points and boundary features, evaluate and compare the accuracy metrics of the respective registration results. Design new weights according to the target recognition accuracy, and iteratively optimize the registration algorithm to improve its adaptability in different scenarios. The specific implementation process includes: at each incremental optimization step t, update the feature point weight where TA represents the target recognition accuracy, and CA represents the current recognition accuracy;

[0033] Based on an improved image registration and exchange method, the registered image is transformed to ensure the accuracy and robustness of object detection. The calculation method of the image transformation matrix is as follows:

[0034] Based on the object detection evaluation index analysis method, the detection results are analyzed to evaluate the effects of registration and detection, and the mean Average Precision (mAP) is obtained to ensure the applicability and reliability of the system in complex scenarios. The calculation of the mean Average Precision is as follows:

[0035] Based on the incremental optimization method, the registration results of feature points and object boundary features are fused, and the accuracy indexes of their respective registration results are evaluated and compared. A new weight is designed according to the object recognition accuracy rate, and the registration algorithm is iteratively optimized to improve the adaptability in different scenarios. Specifically, the evaluation of the registration result is carried out through the following formula:

[0036] RA represents the registration accuracy, CRP represents the number of correctly registered points, and TRP represents the total number of registered points.

[0037] Design a new weight w and iteratively optimize the registration algorithm to improve the adaptability in different scenarios. The calculation of the new weight design is through the following formula:

[0038] At each incremental optimization step t, update the feature point weight where TA represents the object recognition accuracy rate, and CA represents the current recognition accuracy rate;

[0039] Based on the improved multi-level feature fusion method, introduce a cross-modal attention mechanism and adaptive weight adjustment to fuse features at different levels to improve the detection effect. The general multi-level feature fusion method is where α i represents the fusion weight of the i-th layer feature, which is optimized through training. The specific implementation process includes the following steps: extract multi-level features from the reference image and the image to be registered, perform cross-modal attention enhancement on the extracted features, calculate the attention weight, assign adaptive weights to features at different levels, and dynamically adjust the weight parameters through training. The improved fusion formula is where A i is the attention weight matrix of the i-th layer feature, and α i is the fusion weight dynamically optimized through training. The specific implementation process includes: extract multi-level features {f i |1, 2, …, n} from the reference image R and the image to be registered S, calculate the attention weight A i = Softmax(Sim(f i , fi ), where dynamically optimize the fusion weight α through training i , so that the fused features can better adapt to different scenarios and task requirements. Through the above improvement method, more robust and efficient feature fusion can be achieved under different scenarios, improving the object detection accuracy and robustness of multimodal remote sensing images;

[0040] Based on the backbone network feature extractor of the deep residual network (ResNet-101), extract features from the registered visible light, infrared, and SAR images. Enhance the high-level and low-level features through the cross-modal attention mechanism to improve the feature expression ability. Multilevel feature fusion splices and reduces the dimension of the high-level and low-level features to form the output feature map of the encoder part to retain important information. The enhancement effect of the cross-modal attention mechanism is quantified by the following formula: F ResNet-101 = {f 1 , f 2 , …, f n}, where F ResNet-101 represents the set of multi-level features extracted by ResNet-101;

[0041] Among them, the attention weight is calculated by the following formula:

[0042] AW = Softmax(SM)

[0043] The calculation formula for the enhanced features is:

[0044] EF = OF × AW

[0045] EF represents the enhanced features (Enhanced Features, EF), OF represents the original features (Original Features, OF), AW represents the attention weights (Attention Weights, AW), and SM represents the similarity matrix (Similarity Matrix, SM). The similarity is calculated by methods such as cosine similarity. The relevant formulas are as follows:

[0046] Calculation formula for the dot product of vectors:

[0047] Calculation formula for the norm of a vector:

[0048] Calculation formula for cosine similarity:

[0049] Through these technologies, the present invention can solve problems such as low registration accuracy of multimodal remote sensing images, insufficient feature extraction, high complexity of detection models, and poor multimodal data fusion effect, thereby achieving high-precision and high-efficiency target detection in complex scenarios and significantly improving the reliability and practicality of remote sensing image processing.

[0050] An improved cross-modal multi-head self-attention mechanism is introduced. By introducing a cross-modal attention mechanism among the features of multiple modalities including visible light, infrared, and Synthetic Aperture Radar (SAR) images, multi-head segmentation and parallel processing of features of different modalities are performed to capture more diverse feature relationships and semantic information, ultimately enhancing the correlation between features and improving the accuracy of target detection. The attention weights are obtained by calculating the similarity through the Softmax function, and the relevant formula is: A attention = softmax(S)·V. Where S represents the similarity matrix, which is obtained by calculating the cosine similarity, and V represents the value matrix. The calculation formula of the cosine similarity is: Where, Q i and K j represent the query vector and the key vector respectively, · represents the dot product operation, and ‖·‖ represents the norm of the vector Where, H represents the number of heads, W O is the output weight matrix;

[0051] The present invention also provides an improved method for designing and constructing a lightweight real-time target recognition model based on a deep neural network. It is characterized in that a lightweight target detection model is constructed based on a deep neural network, and an improved anchor box mechanism is used for bounding box delimitation. The Intersection over Union (IoU) value is maximized as the loss function to optimize the deep learning model to ensure the accuracy and real-time performance of detection. The Non-Maximum Suppression (NMS) algorithm is used to post-process the detection results, retaining the best detection results and removing redundant results. The method specifically includes the following steps:

[0052] First, the delimitation of the bounding box is realized based on the improved anchor box mechanism. The improved anchor box mechanism is characterized in that the position and size of the anchor box are adjusted through model learning to achieve precise positioning of the target object. The specific process includes:

[0053] By presetting multi-scale and multi-ratio anchor boxes on each grid unit of the input image to use similar anchor boxes to match the target to be predicted. The method for setting the anchor box size is set, L is the loss function, L conf represents the confidence loss, L loc represents the position loss, L cls represents the classification loss, λ 1 and λ2 If \(L\) is the weight coefficient, then \(L = L\) conf +\(\lambda\) 1 L loc +\(\lambda\) 2 L cls ; Then define it as \(A\) to represent the set of all preset anchor boxes. Let the grid size be \(S\times S\), and the number of preset anchor boxes on each grid cell be \(k\). Then the anchor boxes on each grid cell \((i, j)\) can be expressed as:

[0054]

[0055] Among them, and are the center coordinates of the bounding box, \(w\) k and \(h\) k are the width and height of the bounding box.

[0056] The adjustment amount \(\Delta=(\Delta x,\Delta y,\Delta w,\Delta h)\) of the predicted bounding box is obtained through model learning, and the position and size of the bounding box are adjusted according to the adjustment amount:

[0057]

[0058]

[0059]

[0060]

[0061] Generate multiple preset bounding boxes at different scales and ratios. Let the scale set be \(\{s\) 1 ,s 2 ,\(\cdots\),s n \}\), and the ratio set be \(\{r\) 1 ,r 2 ,\(\cdots\),r m \}\), then the width and height of the anchor box are respectively:

[0062]

[0063] To reduce the number of network layers of the model and optimize the structure of the network, a lightweight deep neural network model is constructed. The depthwise separable convolution lightweight module is used to reduce the number of parameters and the computational complexity, so that the model can also run efficiently in resource-constrained environments. The depthwise separable convolution decomposes the standard convolution into depthwise convolution and pointwise convolution to reduce the computational complexity. Let the size of the input feature map be \(H\times W\times C\), the size of the convolution kernel be \(K\times K\), and the number of output channels be \(C'\).

[0064] The calculation complexity formula of the standard convolution is:

[0065] Standard Convolution FLOPs = H × W × C × K 2 × C′

[0066] The relevant formula for the computational complexity of depthwise separable convolution is as follows:

[0067] Depthwise Convolution FLOPs = H × W × C × K 2

[0068] Pointwise Convolution FLOPs = H × W × C × C′

[0069] Total FLOPs = H × W × C × K 2 + H × W × C × C′

[0070] Then, based on the IoU optimization and non - maximum suppression (NMS) post - processing method, redundant detections are removed, and the optimal detection results are retained to ensure the accuracy and effectiveness of the final detection results. The specific process includes:

[0071] Taking the maximization of the IoU value as the loss function, during the training process, the goal is to maximize the IoU value between the predicted bounding box and the ground - truth bounding box. The model parameters are optimized through backpropagation to improve the prediction accuracy. The specific loss function can be expressed as:

[0072] IoU loss = 1 - IoU(B p , B t )

[0073] Among them, B p and B t represent the predicted bounding box and the ground - truth bounding box respectively. The calculation formula for IoU is:

[0074]

[0075] The non - maximum suppression (NMS) algorithm is used for post - processing. This algorithm can retain the best object detection results by removing the bounding boxes with high overlap degrees. All predicted bounding boxes are sorted by confidence. Starting from the bounding box with the highest confidence, the IoU values between it and other bounding boxes are compared one by one. The specific steps are as follows:

[0076] 1. Sort the predicted bounding boxes from high to low confidence.

[0077] 2. Select the bounding box with the highest confidence from the sorted list as the current bounding box.

[0078] 3. Calculate the IoU value between the current bounding box and the remaining bounding boxes. If the IoU value is greater than the set threshold (such as 0.5), then remove these bounding boxes.

[0079] 4. Repeat steps 2 and 3 until all bounding boxes are processed.

[0080] The formula of the NMS algorithm is as follows:

[0081] NMS(B, scores, IoU th ) = {Bi | IoU(Bi, Bj) < IoU threshold}

[0082] In this formula, B represents the set of predicted bounding boxes, scores represents the corresponding confidence scores, and IoU th represents the set threshold.

[0083] Based on the dynamic setting of the improved threshold and weight, a dynamic adjustment mechanism based on incremental optimization is implemented. The relevant parameters will be updated in real time according to the performance metrics of the current model during each training cycle. The specific process includes:

[0084] During the model training process, the thresholds and weights of various parameters are dynamically adjusted through cross-validation to ensure the best detection effect. Different from traditional methods, the specific implementation process includes: θ * = argmin θ E[L(θ)], where θ * represents the optimization objective, θ represents the model parameters, and L(θ) represents the loss function.

[0085] Carry out model training and optimization, data preparation and preprocessing: Collect a multi-modal remote sensing image dataset, and perform annotation and preprocessing (such as image enhancement and denoising) to ensure data quality. Model training: Use the preprocessed dataset for training, and gradually optimize the model performance through batch training and hyperparameter tuning. Performance evaluation and adjustment: Evaluate the detection performance of the model on the test set, including metrics such as precision, recall, and F1-score, and adjust the model structure and parameters according to the evaluation results to ensure that the model reaches the best performance.

[0086] Actual application and deployment, deploy the optimized model into the target detection system, and optimize the model running efficiency in combination with hardware devices (such as GPUs or embedded devices) to ensure real-time detection capabilities. Real-time detection: Input the real-time obtained remote sensing image, and the model performs target detection and outputs the bounding box and class information. Result verification and correction: Manually verify and make necessary corrections to the detection results to ensure the reliability and accuracy of the detection results.

[0087] The multi-modal remote sensing image registration method realizes more robust and accurate image registration through the extraction of significant boundary feature points and the construction of local boundary descriptors. The incremental optimization method can design new weights according to the target detection accuracy, significantly improving the adaptability and stability of the registration method in different scenarios.

[0088] The feature extraction and fusion method further enhances the accuracy and robustness of target detection through a cross-modal attention mechanism and a multi-level feature fusion strategy. The target detection model reduces the computational complexity through lightweight design and optimization, meeting the real-time requirements in practical applications.

[0089] On the other hand, the present invention provides a system for implementing the above method, characterized in that the system adopts the above steps and methods, improves the accuracy and robustness of target detection by fusing multi-modal remote sensing image data, solves the problems existing in the prior art, and provides an effective solution for remote sensing image processing and target detection.

[0090] The beneficial effects of the present invention are at least:

[0091] The present invention provides a robust recognition method for small-sample remote sensing targets with cross-scenario multi-domain fusion, belonging to the fields of remote sensing image processing and target detection and recognition. This method is based on improved domain mapping for cross-scenario generalization of small-sample learning. By mapping the features of the target domain to the source domain feature space, it effectively reduces the data distribution difference between the source domain and the target domain, improves the robustness and generalization ability of the target recognition model in different scenarios, especially in the small-sample learning scenario, significantly enhances the recognition ability of the target recognition model for new tasks and unseen categories, and adopts a meta-learning strategy, enabling the target recognition model to quickly adjust and update using a small number of labeled samples. Based on the representation learning of an improved masked autoencoder, using self-supervised learning methods, and adopting ViT as the core module, it improves the learning efficiency of image features by using data parallel processing, strengthens the feature extractor's understanding of global context and capture of local details, and generates more robust and generalized feature representations. For the multi-modal remote sensing image registration based on local boundary features, by extracting the significant boundary features of the image pair and combining with the gray-based method to extract feature points, high-precision registration in complex scenarios is achieved; based on dynamically adjusting the SIFT threshold to extract feature points and calculating weight values to strengthen important feature points, ensuring the robustness and accuracy of registration; using a deep residual network (ResNet-101) backbone network feature extractor to extract features from the registered visible light, infrared, and SAR images, and enhancing the high-level and low-level features through a cross-modal attention mechanism to further optimize the feature expression ability. Based on the multi-level feature fusion strategy provided by the present invention, it can retain important information, improve the accuracy and robustness of target detection; by incrementally optimizing the method to fuse the registration results based on feature points and boundary features, and designing a new weight iteration optimization registration algorithm, it significantly improves the adaptability and stability of the registration method in different scenarios; constructing a lightweight target detection model based on a deep neural network, using the anchor box mechanism to delimit the bounding box, and using the maximum IoU value as the loss function to ensure the accuracy and real-time performance of detection. Using the non-maximum suppression algorithm to post-process the detection results, effectively retaining the best detection results and removing redundant results. Through the innovative methods and systems provided by the present invention, it can significantly improve the registration accuracy, feature extraction and fusion effect of multi-modal remote sensing images, as well as the accuracy and real-time performance of target detection, not only meeting the real-time requirements in practical applications, but also effectively solving the problems of low registration accuracy, insufficient feature extraction, and high complexity of detection models in the prior art, providing an effective solution for remote sensing image processing and target detection.

[0092] The additional advantages, objects, and features of the present invention will be partially described below, and will become partially apparent to those of ordinary skill in the art after studying the following text, or can be learned from the practice of the present invention. The objects and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification and the drawings.

[0093] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present invention are not limited to those specifically described above, and the above and other purposes that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention. In the drawings:

[0095] Figure 1 It is a schematic diagram of the steps of a cross-scenario generalization few-shot learning based on improved domain mapping in an embodiment of the present invention.

[0096] Figure 2 It is a schematic diagram of the process of a cross-scenario generalization few-shot learning based on improved domain mapping in an embodiment of the present invention.

[0097] Figure 3 It is a schematic diagram of the steps of an object detection method for multi-domain heterogeneous data fusion enhancement in an embodiment of the present invention.

[0098] Figure 4 It is a schematic diagram of the process of an object detection method for multi-domain heterogeneous data fusion enhancement in an embodiment of the present invention.

[0099] Figure 5 It is a schematic diagram of the process of a multi-domain heterogeneous image registration technology based on feature matching in an embodiment of the present invention.

[0100] Figure 6 It is a framework diagram of a multi-modal remote sensing image semantic segmentation scheme in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0101] To make the purposes, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.

[0102] Here, it should also be noted that in order to avoid obscuring the present invention with unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, while other details less related to the present invention are omitted.

[0103] It should be emphasized that the term "comprising / including" when used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0104] Here, it should also be noted that, unless otherwise specified, the term "connection" in this article can not only refer to direct connection, but also indirect connection with intermediate substances.

[0105] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0106] As Figure 1 shown, the cross-scenario generalization few-shot learning based on improved domain mapping and the representation learning based on improved masked autoencoder of the present invention include the following steps S101 to S105:

[0107] Step S101: Use the source domain dataset as samples to train an improved masked autoencoder. Randomly mask some regions of the input image, input the masked image into the encoder part of the autoencoder to obtain the feature representation of the image, and then restore it to the original complete image through a lightweight decoder. After training, use the encoder part as a feature extractor.

[0108] Step S102: Set a small number of target domain samples as the auxiliary dataset, mix them with the source domain dataset through a dataset mixer to obtain a mixed dataset, and re-partition the mixed dataset into a source domain support set, an auxiliary support set, and a mixed query set. The source domain support set maintains the representative examples of the source domain and helps the target recognition model maintain and consolidate the learning results in the source domain. The auxiliary support set contains a small number of samples from the target domain and serves as auxiliary information to guide the model to gradually adapt to the feature distribution of the target domain, reducing the domain shift problem. The mixed query set is composed of a mixture of source domain and target domain datasets and is used to test the generalization ability of the model when facing unknown or cross-domain samples.

[0109] Step S103: Input the image data of the source domain support set, the auxiliary support set, and the mixed query set into the feature extractor to obtain domain-agnostic features and domain features. Domain-agnostic features refer to features that are not affected by a specific domain and are commonly present in various domains, such as some basic visual patterns or the general contours of objects; domain features refer to features related to a specific domain, and these features reflect the specificities of different scenarios or environments.

[0110] Step S104: Send the domain-agnostic features into a few-shot classifier to perform a classification task, compare them with the corresponding class labels to obtain the target classification loss function. Input the domain-agnostic features of the three sets into a domain classifier for domain classification and compare them with the domain labels to evaluate and obtain the loss function of domain classification. The domain features of the three sets are also processed in the same way to obtain the loss value of domain classification. Add up all the loss functions to obtain the overall loss function.

[0111] Step S105: Use the Adam optimization algorithm to perform an efficient gradient descent optimization process on the overall loss function, thereby iteratively optimizing the parameters of the target recognition model to ensure the continuous improvement of the model performance and the gradual convergence of the overall error.

[0112] As Figure 2 shown, it is a schematic flow chart of cross-scenario generalization few-shot learning based on improved domain mapping in a cross-scenario multi-domain fusion few-shot remote sensing target robust recognition method.

[0113] In step S101, first, a masked autoencoder is used to learn the feature representation of the source domain dataset. A random masking strategy is adopted to partially mask the image, enabling the autoencoder to learn the image feature representation in the absence of information. The encoder then encodes the processed image to generate a feature vector containing information such as local texture, context relationship, and semantic information. The lightweight decoder receives these feature vectors and attempts to reconstruct the original image. This process not only ensures the effectiveness of the features but also enhances the generalization ability of the model. After training, the encoder part is selected as the feature extractor. This process belongs to the pre-training stage.

[0114] In step S102, to address the challenges of cross-domain learning, a small number of samples from the target domain are introduced as an auxiliary dataset and fused with the source domain dataset through a dataset mixer to create a mixed dataset that contains both source domain characteristics and target domain features. This mixed dataset is then reorganized into a source domain support set, an auxiliary support set, and a mixed query set, with each part carrying a specific training mission: the source domain support set ensures the model's performance in the familiar domain, the auxiliary support set guides the model to gradually explore and adapt to the target domain features, and the mixed query set tests the model's generalization ability under unknown or cross-domain conditions.

[0115] In step S103, the feature extractor is used to deeply analyze the images in the source domain support set, the auxiliary support set, and the mixed query set to extract domain-invariant features and domain features. Domain-invariant features are the common features that span multiple domains and form the general basis for the recognition task; while domain features specifically correspond to the unique properties of a specific domain and help the model identify and distinguish the feature distributions of different scenarios.

[0116] In step S104, an overall loss function is designed to ensure the synchronous optimization of the model in both the classification task and the domain discrimination task. The domain-invariant features are fed into a few-shot classifier for target class prediction and compared with the actual labels to calculate the classification loss. At the same time, the domain-invariant features and domain features are fed into the domain classifier, and the domain classification losses are calculated through relative entropy and cross-entropy respectively. The sum of these loss functions forms the overall optimization objective of the model, reflecting the comprehensive pursuit of the model in recognition accuracy and domain adaptability.

[0117] In step S105, in order to efficiently adjust the model parameters, the Adam optimization algorithm is adopted. This is a method of adaptive learning rate, which can dynamically adjust the learning rate based on historical gradient information, so as to more accurately guide the model to approach the global minimum loss point during the gradient descent process. Through continuous iterative optimization, the performance of the model is continuously improved, while ensuring the gradual reduction of the overall error, and finally achieving high accuracy and strong robustness in cross-scenario small-sample target recognition.

[0118] As Figure 3 shown, the multi-domain heterogeneous image registration technology based on feature matching and the lightweight real-time target recognition model based on deep neural network of the present invention include the following steps S301 to S307:

[0119] Step S301: Adopt a multi-modal remote sensing image registration method based on local boundary features. By extracting the significant boundary features of the image pair and combining with the gray-based method to extract feature points for registration. In order to improve the registration accuracy, dynamically adjust the Scale-Invariant Feature Transform (SIFT) threshold to extract feature points, and calculate the weight value to strengthen the important feature points. For the multi-modal remote sensing image registration based on boundary features, by extracting significant boundary feature points and constructing local boundary descriptors for feature matching, high-precision registration in complex scenes is ensured.

[0120] Step S302: Fusion the registration results based on feature points and boundary features through an incremental optimization method, and evaluate and compare the accuracy indexes of their respective registration results. Design new weights according to the target detection accuracy, and iteratively optimize the registration algorithm to improve the adaptability in different scenarios.

[0121] Step S303: Use the backbone network feature extractor of the deep residual network (ResNet-101) to extract features from the registered visible light, infrared and SAR images. Feature enhancement of high-level and low-level features is performed through a cross-modal attention mechanism to improve the feature expression ability. Concatenate and reduce the dimension of the high-level and low-level features to form the output feature map of the encoder part, and achieve multi-level feature fusion to retain important information.

[0122] Step S304: A lightweight object detection model was constructed based on a deep neural network, ensuring the accuracy and real-time performance of detection. In this model, the anchor box mechanism was used to delimit the bounding boxes. By predefined a set of anchor boxes with different scales and aspect ratios, and adjusting the positions and sizes of these anchor boxes during the training process, the accurate enclosure of the target object was achieved. The model generates multiple anchor boxes on each grid cell of the input image and predicts the offset and class probability of each anchor box to achieve the delimitation of the bounding box for object detection. Then, the maximum Intersection over Union (IoU) value was adopted as the loss function to optimize the deep learning model, improving the accuracy of the target position. IoU is used to measure the overlap degree between the predicted bounding box and the ground truth bounding box. In the model training, by maximizing the IoU value as the loss function, the performance of the model was effectively improved. Finally, the non-maximum suppression (NMS) algorithm was used to post-process the detection results, removing redundant detection results and retaining the best detection results, further improving the accuracy and real-time performance of object detection.

[0123] Step S305: In the object detection method, a deep residual network (ResNet-101) backbone network feature extractor was used to extract features from the registered visible light, infrared, and SAR images. In this step, images of different modalities were input into the ResNet-101 network, and the high-level and low-level features of the images were gradually extracted through the hierarchical structure of the network.

[0124] Step S306: In object detection, to improve the feature expression ability, a cross-modal attention mechanism was introduced to enhance the high-level and low-level features extracted by ResNet-101. This attention mechanism can adaptively adjust the weights between features of different modalities, so that more important features receive more attention, thereby improving the performance of object detection.

[0125] Step S307: Based on the cross-modal attention mechanism, a multi-level feature fusion strategy was adopted in object detection to fuse the high-level and low-level features. Specifically, for the feature maps of each modality, they were concatenated with the feature maps of other modalities, and the dimension of the features was reduced through a dimensionality reduction operation, finally forming the output feature map of the encoder part. This multi-level feature fusion strategy can retain important information and improve the accuracy and robustness of object detection.

[0126] As Figure 4 shown, it is a flowchart of the multi-domain heterogeneous image registration technology based on feature matching and the lightweight real-time object recognition model based on a deep neural network in a cross-scene multi-domain fusion small sample remote sensing target robust recognition method.

[0127] In step S301, first, significant boundary features are extracted from the image by using an edge detection algorithm (such as Canny). The edge detection algorithm can effectively identify the boundaries and contours in the image, thereby extracting significant boundary feature points. These boundary feature points usually have high stability and reliability, providing strong support for subsequent image registration.

[0128] Next, feature points are extracted based on the gray level method. By analyzing the image through the gray level co-occurrence matrix (GLCM), feature points reflecting the texture features of the image can be extracted. The gray level co-occurrence matrix is a statistical method for describing the co-occurrence information of image gray levels, which can capture the spatial relationship of pixel gray levels in the image, thereby extracting representative feature points.

[0129] To further improve the registration accuracy, the threshold of the scale-invariant feature transform (SIFT) is dynamically adjusted to extract feature points. According to the different features of the image, the threshold of the SIFT algorithm needs to be dynamically adjusted in order to extract more stable and representative feature points. The SIFT algorithm extracts and matches feature points by finding key points in the image and describing these key points.

[0130] After extracting the feature points, the importance weight value of each feature point is calculated to strengthen the significant feature points. The importance weight value of the feature points is calculated by various methods, such as comprehensively evaluating factors such as the repeatability and stability of the feature points. By assigning higher weight values to significant feature points, the role of these feature points can be more prominent in the registration process, improving the registration accuracy.

[0131] Finally, local boundary descriptors are constructed using the significant boundary feature points and feature matching is performed to achieve image registration. The local boundary descriptor is a descriptor for describing the local boundary features of the image, which can effectively characterize the local structure and shape information of the image. By using these local boundary descriptors for feature matching, high-precision image registration can be achieved in complex scenes. The entire registration process ensures high-precision registration in complex scenes through the extraction of significant boundary feature points and the construction of local boundary descriptors.

[0132] In step S302, first, the registration results based on feature points and boundary features are fused. In the process of multi-modal remote sensing image registration, the registration results of different feature points and boundary features have their own advantages and disadvantages. By fusing these two registration results, the information of both can be comprehensively utilized to improve the overall registration accuracy. Specifically, the registration result based on feature points is integrated with the registration result based on boundary features to generate a comprehensive registration result.

[0133] After that, evaluate the accuracy metrics of each registration result. During the fusion process, it is necessary to evaluate the accuracy of different registration results to select the best registration result. Commonly used evaluation metrics include mean square error, the number and distribution of matching points, etc. By comprehensively analyzing these metrics, the performance of each registration method in different scenarios can be determined, so as to select the optimal registration result.

[0134] Design new weights according to the accuracy of object detection. To further optimize the registration result, it is necessary to design a new weight allocation strategy according to the accuracy of object detection. Specifically, by analyzing the accuracy of object detection, it can be determined which feature points and boundary features contribute more to object detection. According to this information, adjust the weights of different feature points and boundary features in the registration result, so that the final registration result can maximize the accuracy of object detection.

[0135] Finally, iteratively optimize the registration algorithm. The incremental optimization method optimizes the registration result through continuous iteration. In each iteration process, adjust and optimize the registration algorithm according to the new weights designed in the previous step. Through repeated iteration, gradually improve the adaptability and stability of the registration algorithm in different scenarios. The registration result after each iteration will be further used for evaluation and weight adjustment until the expected accuracy and robustness goals are achieved. The whole process ensures that the registration algorithm can maintain high-efficiency and reliable performance in complex and changeable remote sensing image scenarios.

[0136] In step S303, first fuse the registered visible light, infrared, and SAR images. Through the previous registration steps, the corresponding relationship of different modality images in the same coordinate system has been obtained. Next, it is necessary to fuse these registered image data. Specifically, stack the visible light image, infrared image, and SAR image in a spatially aligned manner so that their respective pixel points are spatially aligned. During the fusion process, methods such as weighted average and maximum value selection can be used to ensure that the image information of different modalities can be effectively integrated, thereby generating a fused image containing rich information.

[0137] After that, construct the initial fused image data. The fused image data contains rich information from three different modalities: visible light, infrared, and SAR, with higher signal-to-noise ratio and more comprehensive feature performance. The goal of this step is to generate an initial fused image, providing a high-quality data basis for subsequent feature extraction and object detection. When constructing the initial fused image data, various image processing techniques can be used, such as image filtering, denoising, and enhancement, to further improve the quality and clarity of the fused image. Through these preprocessing steps, it can be ensured that the fused image has a high visual effect and information fidelity, thus laying a solid foundation for subsequent deep feature extraction and object detection.

[0138] In step S304, a lightweight object detection model based on a deep neural network was first designed. This model was optimized for the real-time requirements in object detection tasks by reducing the number of network layers, optimizing the network structure and the number of parameters to reduce the computational complexity of the model. Such a lightweight design ensures efficient operation in resource-constrained environments and also has remarkable performance in object detection.

[0139] Next, the anchor box mechanism was used to delimit the bounding boxes. The anchor box mechanism is one of the commonly used techniques in object detection models. It predefines a set of anchor boxes with different scales and aspect ratios, and enables the model to learn to adjust the positions and sizes of these anchor boxes during the training process to accurately enclose the target objects. Specifically, on each grid cell of the input image, the model generates multiple anchor boxes and predicts the offsets and class probabilities of each anchor box, thereby delimiting the bounding boxes for object detection. This mechanism can significantly improve the model's detection ability for target objects of different scales and shapes.

[0140] Then, the maximum IoU value was adopted as the loss function. IoU is an index to measure the overlapping degree between the predicted bounding box and the ground truth bounding box. During the model training process, by adopting the maximum IoU value as the loss function, the model can be effectively optimized to make it more accurately predict the position of the target object. Specifically, the loss function calculates the IoU value between the predicted box and the ground truth box, and adjusts the model parameters through backpropagation to increase the IoU value, thereby improving the accuracy of object detection.

[0141] Finally, the non-maximum suppression (NMS) algorithm was used to post-process the detection results. Since the object detection model may predict multiple overlapping bounding boxes for the same target object, in order to remove these redundant detection results, the NMS algorithm was introduced. During the NMS process, first, the confidences of all predicted bounding boxes are sorted, and then starting from the bounding box with the highest confidence, the IoU values of other bounding boxes are compared one by one. If the IoU value exceeds the set threshold, these bounding boxes are considered redundant and removed. In this way, the NMS algorithm can retain the optimal detection results, remove the redundant results, and thus improve the accuracy and efficiency of object detection.

[0142] In step S305, images of different modalities were first input into the ResNet-101 network. To make full use of the information in multi-modal remote sensing images, the registered visible light, infrared, and SAR images were input into the pre-trained ResNet-101 network. ResNet-101 is a deep residual network with 101 layers. By introducing residual connections, the problem of gradient disappearance during the training of deep networks can be effectively alleviated, ensuring the stability and effectiveness of feature extraction.

[0143] Then, through the ResNet-101 network, high-level and low-level features of the image are extracted. In the ResNet-101 network, the first few layers are mainly responsible for extracting low-level features of the image, which usually contain local information such as edges and textures; while the last few layers extract high-level features of the image, which can capture global information such as the shape and category of objects. By extracting multi-level features of the image layer by layer, it is ensured that image information of different modalities can be fully mined and expressed, providing rich feature representations for subsequent feature fusion and object detection.

[0144] In step S306, first, an improved cross-modal multi-head self-attention module is introduced to adaptively adjust feature weights. During the multi-modal feature fusion process, the contributions of image features of different modalities to object detection may vary. Therefore, an improved cross-modal multi-head self-attention module is introduced. By calculating the importance weights of each modality feature, these weights are adaptively adjusted so that more important features receive more attention during the fusion process. This adaptive weight adjustment mechanism can effectively improve the effect of multi-modal feature fusion.

[0145] Then, through the attention mechanism, the expression ability of high-level and low-level features is enhanced. The improved cross-modal multi-head self-attention module can not only adjust the weights between different modality features but also enhance different-level features of the same modality. In the specific implementation process, the attention mechanism calculates the attention weights of the feature map, amplifies the feature responses in the high-weight regions, thereby enhancing the expression ability of high-level and low-level features. The feature map enhanced by the attention mechanism can more accurately represent the target objects in the image, providing more precise feature support for the final object detection.

[0146] In step S307, first, a multi-level feature fusion strategy is adopted to fuse high-level and low-level features. To make full use of the information of different-level features, a multi-level feature fusion strategy is adopted, that is, the high-level and low-level features extracted from the deep residual network (such as ResNet-101) are fused. High-level features usually contain more semantic information, while low-level features contain more detailed information. By fusing these two types of features, various characteristics of the target object can be better captured, improving the effect of object detection.

[0147] Next, the feature maps of different modalities are concatenated, and the feature dimension is reduced through a dimensionality reduction operation. During the fusion process, the feature maps of different modalities (such as the feature maps of visible light, infrared, and SAR images) are concatenated together to form a multi-modal feature map. To reduce the computational complexity and avoid too high a feature map dimension, a dimensionality reduction operation is used to process the concatenated feature map, retaining the main information while reducing redundant features. This step ensures the compactness and efficiency of the fused features.

[0148] such asFigure 5 As shown, it is a multi-domain heterogeneous image registration technology based on feature matching. As shown in the figure, this technology is executed in an image registration system and includes the following steps:

[0149] Step S501: Multi-modal remote sensing image registration based on local boundary features. Input the reference image R and the image S to be registered. By extracting significant boundary features and feature points using a gray-scale-based method, dynamically adjust the Scale-Invariant Feature Transform (SIFT) threshold, extract feature points, and calculate weight values to strengthen important feature points, ensuring high-precision registration in complex scenarios.

[0150]

[0151] Among them, w i is the weight of the feature point p i , and score(p i ) is the score of the feature point.

[0152] Step S502: Multi-domain heterogeneous image registration based on feature points. Use the extracted SIFT feature points for feature matching, and adopt weight calculation and feature matching strategies to achieve high-precision image registration.

[0153]

[0154] Among them, dist(p i .p j ) is the distance between the feature points p i and p j , and ∈ is the distance threshold.

[0155] Step S503: Multi-domain heterogeneous image registration based on boundary features. By detecting boundaries and boundary feature points, calculate local boundary features and boundary descriptors, and perform feature matching.

[0156]

[0157] Among them, desc(p i ) is the boundary descriptor of the feature point p i , and are the gradients of the image in the x and y directions respectively.

[0158] Step S504: Incremental optimization registration. Integrate the incremental optimization method with the registration results based on feature points and boundary features, evaluate and compare the accuracy metrics of their respective registration results. Design new weights according to the target detection accuracy, and iteratively optimize the registration algorithm to improve adaptability in different scenarios.

[0159]

[0160] Among them, loss is the loss function, and w i is the weight of the feature point, is the distance between the feature point p i and its corresponding point thereof.

[0161] Step S505: Image registration and transformation. Transform the registered image to ensure the accuracy and robustness of object detection.

[0162]

[0163] Among them, T is the registration transformation matrix.

[0164] Step S506: Analysis of object detection evaluation metrics. Analyze the detection results to evaluate the effects of registration and detection, and ensure the applicability and reliability of the system in complex scenarios.

[0165]

[0166] Among them, IoU is the Intersection over Union.

[0167] The figure shows the specific implementation process of each step, including inputting the reference image R and the image S to be registered, extracting SIFT feature points, calculating weights and performing feature matching, adopting incremental optimization for registration, finally realizing the registration and transformation of the image, as well as the evaluation and analysis of the object detection results.

[0168] The present invention also provides a scheme framework for semantic segmentation of multimodal remote sensing images. As Figure 6 shown, this technology is executed in an image segmentation system, including the following steps:

[0169] Step S501: Input the reference image and the image to be registered, including visible light images, infrared images, and SAR images. For the input visible light, infrared, and SAR images, extract low-level and high-level features through the backbone network. Let the input image be I, and the low-level and high-level features extracted by the backbone network are respectively represented as F low and F high

[0170] F low = Backbone low (I)

[0171] F high = Backbone high (I)

[0172] Step S502: Process the low-level features through the cross-modal attention module to achieve feature concatenation and obtain the low-level concatenated features. Let the concatenation result of the low-level features be A low-dw

[0173] A low-dw = Attention low (F low,OPT , F low,SAR )

[0174] Step S503: Process the high-level features through the cross-modal attention module to achieve feature concatenation and obtain the high-level concatenated features. A high-ms

[0175] A high-ms = Attention high (F high,OPT , F high,SAR )

[0176] Step S504: Perform comprehensive concatenation on the low-level and high-level concatenated features to generate the final concatenated feature A low-high :

[0177] A low-high = Concatenate(A low-dw , A high-ms )

[0178] Step S505: Input the concatenated features into the regression predictor based on the deep neural network for object detection and semantic segmentation, and output the prediction result P

[0179] P = RegressionNetwork(A low-high )

[0180] Step S506: Post-process and optimize the detection results. Use the non-maximum suppression (NMS) algorithm to remove redundant detection boxes, set the confidence threshold to t, retain the best detection results, and output the final object detection results. The processing formula of NMS is:

[0181]

[0182] The figure shows the specific implementation process of each step, including inputting the reference image and the image to be registered, feature extraction, cross-modal attention module processing, feature concatenation, regression prediction based on the deep neural network, and post-processing optimization process. Finally, the accurate semantic segmentation of multi-modal remote sensing images is achieved

[0183] Then, the output feature maps of the encoder part are formed, retaining important information. The feature maps after multi-level fusion and dimensionality reduction are used as the output feature maps of the encoder part. These feature maps integrate feature information of different modalities and different levels, can better represent various information in the image, and provide a solid foundation for subsequent object detection.

[0184] Finally, based on the fused feature maps, object detection and recognition are carried out to improve the detection accuracy and robustness. Using an optimized object detection model (such as a lightweight deep neural network model), the fused feature maps are input into the detection model for the recognition and localization of target objects. Through the above multi-level feature fusion strategy, the accuracy and robustness of object detection can be effectively improved, ensuring excellent detection performance even in complex scenarios. The non-maximum suppression (NMS) algorithm further removes redundant results in the post-processing of detection results to ensure the accuracy of detection results.

[0185] In summary, the present invention provides a method for object detection and image registration based on multi-modal remote sensing images, belonging to the fields of image processing and remote sensing. This method is based on edge detection algorithms (such as Canny), gray-level co-occurrence matrix (GLCM), and dynamically adjusted scale-invariant feature transform (SIFT). By extracting significant boundary features, texture features, and stable feature points, accurate image registration is achieved; and by calculating the importance weights of feature points and constructing local boundary descriptors, the role of significant feature points is strengthened to ensure high-precision image registration. This method also improves the accuracy and robustness of registration by fusing the registration results based on feature points and boundary features, evaluating the accuracy metrics of registration results, and designing new weights. Based on the registered multi-modal images, a lightweight deep neural network object detection model is designed, and using the anchor box mechanism, the maximum IoU value loss function, and the non-maximum suppression algorithm, efficient and accurate object detection is achieved. Finally, by introducing an improved cross-modal multi-head self-attention module, a multi-level feature fusion strategy, and dimensionality reduction operations, the feature information of different modalities and different levels is effectively fused, improving the accuracy and robustness of object detection. The entire method ensures efficient and accurate object detection and image registration in complex and variable remote sensing image scenarios, providing an effective solution for remote sensing image processing.

[0186] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.

[0187] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.

[0188] In the present invention, the features described and / or illustrated for one embodiment can be used in the same or a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0189] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A robust recognition method for small sample remote sensing targets with cross-scene multi-domain fusion, characterized in that: The method comprises the following steps: Step 1: Based on the cross-scenario generalized small-shot learning of improved domain mapping, a meta-learning strategy is used to build a K-way-N-shot learning scenario, and a small number of labeled samples are used to drive the iterative upgrade of the model; Step 2: Based on the representation learning of the improved masked autoencoder, a random mask mechanism is used to extract high-level semantic feature representations from large-scale image data without relying on labeled data. Step 3: Based on the improved feature matching multi-domain heterogeneous image registration technology, the significant boundary features of the multimodal remote sensing images are extracted, and the images are registered with the feature points extracted based on the grayscale method. The visible light, infrared and SAR images after image registration and exchange are subjected to feature extraction again, and the improved incremental optimization method is combined with the improved anchor frame mechanism to delineate the boundary frame. Using an improved method for extracting significant boundary feature points, multimodal remote sensing image registration based on local boundary features, after inputting a reference image R and an image to be registered S, the reference image R is used as a standard image for reference, and the image to be registered S is used as an image aligned with the reference image; feature points are extracted by extracting significant boundary features and a grayscale-based method, the scale-invariant feature transformation threshold is dynamically adjusted, feature points are extracted, and weight values ​​are calculated. The specific implementation process includes: for the general case of the input image, that is, an image I of one of the reference image R or the image to be registered S, its significant boundary feature point set F = {f i |i=1,2,…,n}, weight value of feature point where g i is the feature point f i The gray value of , ɑ and β are adjustment parameters; The image registration and exchange method transforms the registered image, and the image transformation matrix is ​​calculated as follows: Using the improved incremental optimization method, by detecting the boundary and boundary feature points, calculating the local boundary features and boundary descriptors, and performing feature matching, the specific implementation process includes: at each iteration step t, calculating the loss L of the current registration result total =L match +L boundary , update the parameters based on the gradient of the loss function The improved incremental optimization method is used to fuse the registration results based on feature points and boundary features, evaluate and compare the accuracy indicators of each registration result, and then design new weights according to the target recognition accuracy, and iteratively optimize the registration algorithm. The specific implementation process includes: in each incremental optimization step t, update the feature point weight Among them, TA represents the target recognition accuracy, and CA represents the current recognition accuracy; Step 4: Based on the lightweight real-time target recognition model of the improved deep neural network, the ResNet-101 backbone network feature extractor is used to extract features of the registered multimodal remote sensing images; and target detection and recognition are performed through the improved cross-modal self-attention mechanism and multi-level feature fusion method.

2. The cross-scene multi-domain fusion small sample remote sensing target robust recognition method according to claim 1 is characterized in that: The cross-scenario generalized small sample learning based on improved domain mapping includes the following steps: A very small number of samples in the target domain are set as auxiliary datasets through the dataset mixer, and merged with the dataset in the source domain, and re-divided into source domain support set, auxiliary support set and mixed query set; In the pre-training stage, the masked autoencoder is used for training on the dataset of the source domain to learn image features in a self-supervised manner. The encoder part of the trained masked autoencoder is used as a feature extractor. The feature extractor decouples the image features of the source domain and the target domain to obtain domain-independent features and domain features, and maps the features of the target domain to the source domain, thereby helping to narrow the gap in data distribution between the source domain and the target domain, and converting the data of the source domain and the target domain into a common feature space; the learning of the model is guided by obtaining a very small amount of labeled data belonging to the target domain and adding it to the training dataset, that is, using the very small amount of labeled data of the target domain as auxiliary data to improve the ability of cross-domain small sample learning to cope with domain shift, and enable the meta-learning-based object recognition model to have good generalization ability on the target domain dataset; In the meta-learning stage, the source domain support set, auxiliary support set, and mixed query set are sent to the feature extractor to obtain their respective domain-independent features and domain features. H1 and H2 are used to represent the domain-independent features and domain features, respectively. The domain-independent features and domain features of the source domain support set are expressed as and The domain-independent features and domain features of the auxiliary support set are expressed as and The domain-independent features and domain features of the mixed query set are expressed as and Feed the domain-independent features into the small-shot classifier g fsl Classify and compare with the label to get the loss function L FSL ; Send domain-independent features and domain features to the domain classifier g dom Perform domain classification in and compare with domain labels to obtain the loss function and Overall loss Use the Adam gradient descent optimizer to optimize L to minimize L; Will and Input small sample classifier and get classification loss and Input small sample classifier and get classification loss Since the mixed query set is a mixture of the source domain query set and the auxiliary query set with a ratio of λ, the loss function of the small sample classifier is For domain-dependent features, the source domain label is set to 1 and the target domain label is set to 0. For domain-independent features, the domain label vector is set to y1 = [0.5, 0.5]; is a domain-independent feature, and is input into the domain classifier g dom It is classified and compared with the label y1, and the loss is calculated using the relative entropy KL divergence The input is compared with the respective domain labels, and the loss is calculated using cross entropy 3. The cross-scene multi-domain fusion small sample remote sensing target robust recognition method according to claim 1 is characterized in that: The representation learning based on the improved masked autoencoder includes the following steps: The random masking mechanism is used to mask random areas of the high-dimensional input image, and the reconstruction operation of the autoencoder is used to predict the area hidden by the mask. In the masked autoencoder, the encoder is composed of a combination of ViT modules, specifically ViT-L / 16; the decoder uses a lightweight Transformer module; The ViT module transforms the encoder part of the traditional Transformer and uses the attention mechanism to parallelize the training of the mask autoencoder, speed up the training and inference of the mask autoencoder, and handle problems with long-term dependencies. The image is divided into blocks of the same size, and then these image blocks are linearly mapped into embedded vectors through a fully connected layer. Since the attention mechanism lacks position information representation, the vector converted from the image block is coupled with the vector representing the position information as the input of ViT. With the help of position encoding, the model can capture more refined position features. The lightweight Transformer module has a depth of 8, that is, it is composed of 8 layers of ViT modules stacked together, and a width of 512, that is, each image block is processed into a 512-dimensional feature vector; the computational complexity of the decoder composed of it is only 10% of that of the encoder.

4. The cross-scene multi-domain fusion small sample remote sensing target robust recognition method according to claim 1 is characterized in that: The lightweight real-time target recognition model based on the improved deep neural network includes the following steps: By using the ResNet-101 backbone network feature extractor and an improved cross-modal multi-head self-attention mechanism and an improved multi-level feature fusion method, the accuracy and robustness of target detection and recognition are enhanced, and the computational complexity is reduced through lightweight design and optimization methods combined with improved threshold and weight dynamic settings; The ResNet-101 backbone network feature extractor extracts multi-layer features of the image through the ResNet-101 backbone network feature extractor; The improved cross-modal multi-head self-attention mechanism introduces a cross-modal attention mechanism between multi-modal features including visible light, infrared and synthetic aperture radar images, performs multi-head segmentation and parallel processing on features of different modalities, and the attention weight is obtained by calculating the similarity through the Softmax function. The specific implementation process includes: A attention =softmax(S)·V, where S represents the similarity matrix, which is calculated by cosine similarity, and V represents the value matrix. The calculation formula of cosine similarity is: Where Q i and K j denote the query vector and key vector respectively, · denotes the dot product operation, ‖·‖ denotes the norm of the vector, Where H represents the number of heads, W O is the output weight matrix; The improved multi-level feature fusion method introduces a cross-modal attention mechanism and adaptive weight adjustment to fuse features at different levels; the improved fusion formula is: Among them A i represents the attention weight matrix of the i-th layer feature, α i represents the fusion weight dynamically optimized through training. The specific implementation process of the improved fusion method includes: extracting multi-level features {f i |1,2,…,n}, calculate the attention weight A i =Softmax(Sim(f i ,f i )),in Dynamically optimize fusion weights ɑ through training i ; The lightweight design and optimization method reduces the computational complexity of the model through depthwise separable convolution.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Image foreground segmentation method combining subject determination and edge precision

    CN113592893A

  • General target detection method for adaptive attention guidance mechanism

    WO2021139069A1