Remote sensing image surface anomaly detection method and device based on image-text cooperative processing
Through a method based on graphic and text collaborative processing, combined with a multi-scale spatial adapter and a cross-band alignment module, the feature extraction and fusion of multi-spectral remote sensing images is solved, and the problem of insufficient surface anomaly detection accuracy and generalization in the prior art is achieved, and more efficient surface anomaly recognition is achieved.
Patent Information
- Application Number
- CN202510679108.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-26
AI Technical Summary
The existing remote sensing image surface anomaly detection methods have shortcomings in terms of accuracy and generalization, making it difficult to achieve cross-scene and cross-species surface anomaly detection. In addition, traditional methods rely on limited band combinations, resulting in missed and mis-checked surface anomaly.
Using a method based on graphic and text collaborative processing, the text features corresponding to the labels of the surface anomaly type in various places are obtained, and RGB and SWIR images of multi-spectral remote sensing images are combined, and feature extraction and spectral adjustment are used for multi-spectral spatial adapter and cross-band alignment module. Finally, feature fusion is performed through dynamic spectral weighted fusion devices to realize the identification of surface anomaly types.
It significantly improves the accuracy and generalization of surface anomaly detection, can more effectively extract the surface anomaly features in multi-spectral remote sensing images, reduce missed detection and missed detection, and is suitable for detection of various types of surface anomaly.
Smart Images

Figure CN120217263A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of remote sensing image processing, and particularly to a method and device for detecting surface anomalies of remote sensing images based on collaborative text and image processing. Background Art
[0002] Earth Surface Anomaly Detection (ESAD) in remote sensing images plays an important role in fields such as emergency response, disaster assessment, and environmental monitoring. In recent years, with global climate change and rapid economic growth, various surface anomalies occur frequently due to natural or human factors, and their rapid evolution characteristics pose a threat to public safety and sustainable development, which also puts forward higher requirements for surface anomaly detection technology.
[0003] There are various types of surface anomalies, and their remote sensing response characteristics are different. The key to surface anomaly detection lies in extracting effective anomaly features, including spectral features, spatial features, etc. For example, there is a quantitative relationship between surface anomalies such as earthquakes, landslides, and debris flows and the response characteristics of visible light (spectrum, geometry, and texture) and radar (backscattering and texture); vegetation surface anomalies caused by forest and grassland fires are mainly related to the radiation temperature of thermal infrared. Traditional methods use corresponding prior knowledge to design targeted artificial features, such as various remote sensing indices, and then combine manual interpretation to achieve surface anomaly detection. Especially with the development of sensor technology, the detection range and spectral information of high-resolution satellite remote sensing have been continuously broadened and enriched, and the cost of large-scale surface anomaly detection has been continuously reduced. However, such methods based on prior knowledge have a long development cycle, and due to the complex diversity of data distribution, it is difficult to be extended to surface anomaly events in different time and space ranges.
[0004] The rise of deep learning technology has changed this situation, and data-driven deep learning models have gradually replaced expert-driven feature engineering. In the field of remote sensing image processing, deep learning methods mainly include Convolutional Neural Network (CNN) and Vision Transformer (ViT) models: The CNN automatically extracts local features in images, such as edges and textures, through convolutional layers, and reduces the feature dimensions through pooling layers to improve computational efficiency; The ViT model can capture long-range dependencies between any positions in the input image sequence through the self-attention mechanism, thus enhancing the understanding of global context. By deeply mining the high-level semantic information of remote sensing images, deep learning methods have surpassed the vast majority of artificial feature-based methods in terms of accuracy and transferability, but it is still difficult to achieve cross-scene and cross-category surface anomaly detection. In addition, deep learning models usually require a large amount of training data to obtain good generalization ability, which is a challenge for remote sensing images of surface anomalies with low accessibility and high annotation costs.
[0005] Existing surface anomaly detection methods often rely on limited band combinations, such as only using visible light remote sensing images, resulting in insufficient utilization of spectral information and causing missed and false detections of surface anomalies. To more effectively obtain surface anomaly features, many improved methods add additional spectral band data as auxiliary branches, such as near-infrared images and Synthetic Aperture Radar (SAR) images, and then achieve surface anomaly detection through feature fusion. However, although this design can capture key anomaly features to a certain extent, it has significant limitations: First, the additional spectral band data increases the acquisition and processing costs, especially for heterologous SAR images that require registration; Second, the spectral response characteristics of different surface anomalies are often different, and existing improved methods are often limited to a certain type of surface anomaly and lack a general paradigm applicable to most surface anomalies. In addition, modal semantic conflicts limit the representation ability of deep learning models for anomaly features, and the linear fusion method cannot mine deep band associations, resulting in ineffective fusion of multi-modal features.
[0006] In recent years, due to their powerful generalization ability, large pre-trained models have been widely applied in the field of computer vision. Taking the Contrastive Language-Image Pre-training (CLIP) model as an example, it is a powerful classification model that realizes cross-modal semantic understanding and association by aligning the semantic spaces of images and texts. However, when directly applying the CLIP model to the task of surface anomaly detection, two major generalization bottlenecks will be faced: First, surface anomaly events themselves have the characteristics of low occurrence frequency and uneven distribution. Their remote sensing images are also affected by cloud cover, etc., and the cost of sample annotation is high, resulting in a scarcity of high-quality training samples and the model being prone to overfitting. Second, existing large pre-trained models, such as the CLIP model, mainly use natural image datasets to train weights, but there are significant differences between remote sensing images and natural images, including spectral distribution differences and spatial scale differences. How to establish a cross-domain generalization mechanism for the spectral and spatial characteristics of remote sensing images, and use a small number of samples to learn highly generalized cross-scene surface anomaly features is an open challenge that is constantly being explored.
[0007] Existing surface anomaly detection methods still have many deficiencies in the large-scale practical application stage, especially in open detection. Traditional methods are mostly closed-set detection, that is, certain anomaly categories need to be predefined during the training stage. For new types that have not been seen before, they can only be retrained on the existing basis. However, in real-world scenarios, the connotation and extension of surface anomalies are constantly deepening and expanding, and various surface anomalies are intertwined and superimposed on each other, showing complex and diverse evolution trends. Therefore, from the perspectives of the effectiveness of surface anomaly emergency response and the training cost of deep models, there is a serious disconnection between the limited types of the closed-set hypothesis and the infinite demands of actual detection. In addition, due to the intra-class variability and inter-class similarity of remote sensing ground objects, the manifestation forms of surface anomalies in remote sensing images change with the scene, and existing models lack the ability to model the essential features of surface anomalies, resulting in poor performance in the detection task of unknown anomalies.
[0008] In summary, the current methods for detecting surface anomalies in remote sensing images have relatively low detection accuracy and generalization ability for surface anomalies. Summary of the Invention
[0009] Based on this, in view of the above technical problems, it is necessary to provide a method, device, electronic device, and storage medium for detecting surface anomalies in remote sensing images based on text-image collaborative processing, which can improve the detection accuracy and generalization ability of surface anomalies.
[0010] The present application provides a method for detecting surface anomalies in remote sensing images based on text-image collaborative processing, including the following steps: Step 1: Obtain the text features corresponding to each surface anomaly type label. The text features corresponding to the surface anomaly type labels are pre-converted into corresponding natural sentences according to the surface anomaly type labels and a custom text prompt template, and the natural sentences are subjected to feature extraction by a text encoder to obtain text features represented by vectors. Step 2: Perform band extraction and fusion on the multi-spectral remote sensing image to be detected to obtain the RGB image and the SWIR image of the multi-spectral remote sensing image. Step 3: Input the RGB image and the SWIR image into an image encoder module for feature extraction to obtain initial RGB features and initial SWIR features. Step 4: Input the initial RGB features and the initial SWIR features into a multi-scale spatial adapter module for spatial adjustment to obtain multi-scale enhanced spatially adapted RGB features and multi-scale enhanced spatially adapted SWIR features. Step 5: Input the multi-scale enhanced spatially adapted RGB features and the multi-scale enhanced spatially adapted SWIR features into a cross-band alignment module for spectral adjustment to obtain RGB image features and SWIR image features. Step 6: Input the RGB image features and the SWIR image features into a dynamic spectral weighting fuser for feature fusion to generate fused image features. Step 7: Perform cosine similarity analysis on the fused image features and the text features corresponding to each surface anomaly type label to determine the surface anomaly type of the multi-spectral remote sensing image.
[0011] In one embodiment, the multi-scale spatial adapter module includes two multi-scale spatial adapters with the same network structure. The step of inputting the initial RGB features and the initial SWIR features into a multi-scale spatial adapter module for spatial adjustment to obtain multi-scale enhanced spatially adapted RGB features and multi-scale enhanced spatially adapted SWIR features includes: Input the initial RGB features into one of the multi-scale spatial adapters of the multi-scale spatial adapter module for spatial adjustment to obtain multi-scale enhanced spatially adapted RGB features. Input the initial SWIR features into the other multi-scale spatial adapter of the multi-scale spatial adapter module for spatial adjustment to obtain multi-scale enhanced spatially adapted SWIR features.
[0012] In one embodiment, the processing process of the multi-scale spatial adapter is as follows: Separate the initial features input into the multi-scale spatial adapter into a CLS sequence and an image patch sequence. The sequence of image blocks is reconstructed into a 2D feature map that preserves spatial information: The 2D feature map is input into a convolutional enhancement unit for multi-scale convolutional processing to obtain multi-scale fused features. The convolutional enhancement unit is a four-branch multiple convolutional structure; The multi-scale fused features are input into a global feature extraction unit for feature extraction to obtain global features; The global features and the CLS sequence are added together to obtain an enhanced CLS sequence; The multi-scale fused features are restored to a sequence format and recombined with the enhanced CLS sequence to obtain multi-scale enhanced spatially adapted features.
[0013] In one embodiment, the cross-band alignment module includes two linear layers with the same network structure and two cross-band alignment units with the same network structure; The process of inputting the multi-scale enhanced spatially adapted RGB features and the multi-scale enhanced spatially adapted SWIR features into the cross-band alignment module for spectral adjustment to obtain RGB image features and SWIR image features includes: Input the multi-scale enhanced spatially adapted RGB features into one of the linear layers of the cross-band alignment module for processing to obtain RGB query features, RGB key features, and RGB numerical features; Input the multi-scale enhanced spatially adapted SWIR features into the other linear layer of the cross-band alignment module for processing to obtain SWIR query features, SWIR key features, and SWIR numerical features; Input the SWIR query features, the RGB key features, and the RGB numerical features into one of the cross-band alignment units of the cross-band alignment module for spectral adjustment to obtain RGB image features; Input the RGB query features, the SWIR key features, and the SWIR numerical features into the other cross-band alignment unit of the cross-band alignment module for spectral adjustment to obtain SWIR image features.
[0014] In one embodiment, the process of inputting the SWIR query features, the RGB key features, and the RGB numerical features into one of the cross-band alignment units of the cross-band alignment module for spectral adjustment to obtain RGB image features includes: The SWIR query features, the RGB key features, and the RGB numerical features undergo similarity measurement and weighted summation to obtain a multi-spectral attention weighted RGB feature with attention concentrated on the similar features of the RGB image and the SWIR image. The expression is: ; where, Represents the multi - spectral attention - weighted RGB feature, Represents the Softmax activation function, Is the attention weight coefficient of one of the cross - band alignment units, Represents the RGB numerical feature, where R represents the real - number space, N represents the length of the feature, and C represents the number of channels of the feature; Among them, the spectral information divergence is used for the similarity measurement of the SWIR query feature And the RGB key feature To obtain the attention weight coefficient of one of the cross - band alignment units , and the expression is: ; Among them, Refers to the similarity measurement method, Is the natural exponential function, Represents the spectral information divergence matrix between the SWIR query feature and the RGB key feature. The spectral information divergence of the th row and the th column of the spectral information divergence matrix between the SWIR query feature and the RGB key feature The expression is: ; Among them, Represents the th sequence after probability normalization, , Is the logarithmic function with base 10, Represents the th sequence after probability normalization, , and the superscript Represents the transpose; After subtracting the matrix of the RGB numerical feature from the multi - spectral attention - weighted RGB feature and performing linear mapping, it is connected in residual with the SWIR query feature as a residual structure to obtain the RGB difference feature; After the RGB difference feature is processed by layer normalization and a multi - layer perceptron, it is connected in residual with the RGB difference feature again to obtain the RGB image feature; Inputting the RGB query feature, the SWIR key feature, and the SWIR numerical feature into another cross - band alignment unit of the cross - band alignment module for spectral adjustment to obtain the SWIR image feature, including: The RGB query feature, the SWIR key feature, and the SWIR numerical feature are subjected to similarity measurement and weighted summation to obtain a multi-spectral attention weighted SWIR feature with attention concentrated on the similar features of the SWIR image and the RGB image. The expression is as follows: ; Wherein, represents the multi-spectral attention weighted SWIR feature, is the attention weight coefficient of another cross-band alignment unit, represents the SWIR numerical feature; Among them, the spectral information divergence is used for the similarity measurement of the RGB query feature and the SWIR key feature to obtain the attention weight coefficient of another cross-band alignment unit. The expression is as follows: ; Wherein, represents the spectral information divergence matrix between the RGB query feature and the SWIR key feature. The spectral information divergence in the th column of the row of the spectral information divergence matrix between the RGB query feature and the SWIR key feature is calculated as follows: ; Wherein, represents the th sequence after probability normalization, , represents the th sequence after probability normalization, ; After subtracting the matrix of the SWIR numerical feature from the multi-spectral attention weighted SWIR feature and performing linear mapping, it is connected with the RGB query feature as a residual structure to obtain the SWIR difference feature; The SWIR difference feature is processed by layer normalization and a multi-layer perceptron, and then connected with the SWIR difference feature by residual connection to obtain the SWIR image feature.
[0015] In one embodiment, the step of inputting the RGB image feature and the SWIR image feature into a dynamic spectral weighted fusion device for feature fusion to generate a fused image feature includes: The RGB image features are input into the dynamic spectral weighted fusion device, and after layer normalization operation, feature separation is performed to obtain the CLS sequence of the RGB image features and the image patch sequence of the RGB image features; The SWIR image features are input into the dynamic spectral weighted fusion device, and after layer normalization operation, feature separation is performed to obtain the CLS sequence of the SWIR image features and the image patch sequence of the SWIR image features; The image patch sequence of the RGB image features and the image patch sequence of the SWIR image features are reconstructed into the 2D feature map of the RGB image features and the 2D feature map of the SWIR image features; After the 2D feature map of the RGB image features and the 2D feature map of the SWIR image features are concatenated along the channel dimension, they are input into a gating mechanism to generate an element-wise gating signal; According to the element-wise gating signal, multi-modal feature adaptive fusion is performed on the 2D feature map of the RGB image features and the 2D feature map of the SWIR image features to obtain multi-modal fusion features, and the expression is: ; where, represents reshaping the array shape, represents the element-wise gating signal, H and W respectively represent the length and width of the 2D feature map, represents a matrix with all elements being 1, represents element-wise multiplication; represents the multi-modal fusion features, represents the image patch sequence of the RGB image features, represents the image patch sequence of the SWIR image features; After performing global average pooling on the multi-modal fusion features, they are concatenated and flattened with the CLS sequence of the RGB image features and the CLS sequence of the SWIR image features to obtain the flattened image features; Channel re-weighting operation with a three-branch parallel structure is performed on the flattened image features to obtain the fused image features, and the expression of the channel re-weighting operation with a three-branch parallel structure is: ; where, represents the Softmax activation function, represents the 1D convolution operation, represents the fused image features, represents the flattened image features, represents the fully connected layer.
[0016] In one embodiment, the cosine similarity analysis of the fused image features and the text features corresponding to each surface anomaly type label to determine the surface anomaly type of the multispectral remote sensing image includes: Map the fused image features and the text features corresponding to each surface anomaly type label to a shared semantic space with the same dimension, and calculate the cosine similarity between the fused image features and the text features corresponding to each surface anomaly type; Process the cosine similarity between the fused image features and the text features corresponding to each surface anomaly type through the Softmax activation function to obtain the probability that the fused image features belong to each surface anomaly type label, and use the surface anomaly type label with the highest probability as the surface anomaly type of the multispectral remote sensing image.
[0017] A remote sensing image surface anomaly detection device based on text-image collaborative processing includes: A text feature acquisition module, configured to acquire the text features corresponding to each surface anomaly type label, where the text features corresponding to the surface anomaly type label are pre-converted into corresponding natural sentences according to the surface anomaly type label and a custom text prompt template, and the natural sentences are subjected to feature extraction by a text encoder to obtain text features represented by vectors; A band extraction module, configured to perform band extraction and fusion on the multispectral remote sensing image to be detected to obtain the RGB image and the SWIR image of the multispectral remote sensing image; An image encoder module, configured to perform feature extraction on the RGB image and the SWIR image to obtain initial RGB features and initial SWIR features; A multi-scale spatial adapter module, configured to perform spatial adjustment on the initial RGB features and the initial SWIR features to obtain multi-scale enhanced spatially adapted RGB features and multi-scale enhanced spatially adapted SWIR features; A cross-band alignment module, configured to perform spectral adjustment on the multi-scale enhanced spatially adapted RGB features and the multi-scale enhanced spatially adapted SWIR features to obtain RGB image features and SWIR image features; A dynamic spectral weighted fusion device, configured to perform feature fusion on the RGB image features and the SWIR image features to generate fused image features; A similarity analysis module, configured to perform cosine similarity analysis on the fused image features and the text features corresponding to each surface anomaly type label to determine the surface anomaly type of the multispectral remote sensing image.
[0018] A computer device includes a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the processor executes the computer program, the steps of the remote sensing image surface anomaly detection method based on graphic and text collaborative processing are implemented.
[0019] A computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the remote sensing image surface anomaly detection method based on graphic and text collaborative processing are implemented.
[0020] The above-mentioned remote sensing image surface anomaly detection method based on image-text collaborative processing obtains the text features corresponding to each surface anomaly type label, and the text features corresponding to the surface anomaly type label are converted into corresponding natural sentences in advance according to the surface anomaly type label and the customized text prompt template, and the natural sentences are subjected to feature extraction through the text encoder to obtain the text features represented by vectors; the multispectral remote sensing image to be detected is subjected to band extraction and fusion to obtain the RGB image and SWIR image of the multispectral remote sensing image; the RGB image and the SWIR image are input into the image encoder module for feature extraction to obtain the initialized RGB features and the initialized SWIR features; the initialized RGB features are extracted from the image encoder module to obtain the initialized RGB features and the initialized SWIR features; the initialized RGB features are extracted from the image encoder module to obtain the initialized RGB features and the initialized SWIR features. The initialized SWIR features are input into the multi-scale spatial adapter module for spatial adjustment to obtain multi-scale enhanced spatially adapted RGB features and multi-scale enhanced spatially adapted SWIR features; the multi-scale enhanced spatially adapted RGB features and multi-scale enhanced spatially adapted SWIR features are input into the cross-band alignment module for spectral adjustment to obtain RGB image features and SWIR image features; the RGB image features and SWIR image features are input into the dynamic spectral weighted fusion module for feature fusion to generate fused image features; the fused image features and the text features corresponding to the labels of each surface anomaly type are analyzed by cosine similarity to determine the surface anomaly type of the multi-spectral remote sensing image. Therefore, by introducing the multi-scale spatial adapter and the cross-band alignment module, the spatial and spectral adaptability of the multi-spectral remote sensing image is fine-tuned, which significantly improves the accuracy and generalization of surface anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a schematic diagram of the process of a method for detecting surface anomalies in remote sensing images based on collaborative processing of images and texts in one embodiment; Figure 2 A schematic diagram of the structure of a method for detecting surface anomalies in remote sensing images based on collaborative processing of images and texts in one embodiment; Figure 3 is a schematic diagram of the structure of a multi-scale space adapter in an embodiment; Figure 4 is a schematic structural diagram of a cross-band alignment unit in one embodiment; Figure 5 Schematic diagram of the structure of a dynamic spectral weighted fuser in an embodiment. Detailed implementation manners
[0022] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0023] It should be noted that the diagrams provided in the embodiments of the present application only illustrate the basic concept of the present application in a schematic manner. Therefore, only the components related to the present application are shown in the diagrams, rather than being drawn according to the number, shape and size of the components in actual implementation. The types, quantities and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0024] In the present application, it should also be noted that when terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. appear, the orientation or positional relationship indicated is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present application. In addition, when terms such as "first" and "second" appear, they are only used for descriptive and distinguishing purposes and cannot be understood as indicating or implying relative importance.
[0025] In addition, it should also be noted that the features of various embodiments of the present application can be partially or wholly combined or integrated, and as can be understood by those skilled in the art, they can interact and operate in different ways. Each embodiment can be implemented independently of each other or in an associated relationship.
[0026] In one embodiment, as Figure 1 and Figure 2 shown, a method for detecting surface anomalies in remote sensing images based on graphic and text collaborative processing is provided, which is characterized by including the following steps: Step S1, obtaining text features corresponding to each surface anomaly type label. The text features corresponding to the surface anomaly type label are pre-converted into corresponding natural sentences according to the surface anomaly type label and a custom text prompt template, and the natural sentences are feature-extracted by a text encoder to obtain text features represented by vectors.
[0027] Among them, the surface anomaly type label is determined according to the possible surface types, and the possible surface types include abnormal surface types such as forest fires and earthquakes and normal surface types.
[0028] Among them, the text encoder is used for text feature encoding. The encoder of the pre-trained Transformer model is used to process the surface anomaly type labels in the open set of surface anomaly type labels, combine each surface anomaly type label with the corresponding text prompt template to form a natural sentence, so as to extract a set of text features required for contrastive learning. The open set of surface anomaly type labels ensures the openness of the detection system to different anomaly types.
[0029] Among them, as Figure 2 shown, an open set of surface anomaly type labels in text form can be made based on common surface types (such as forest fires, earthquakes, normal, etc.); through the text prompt template, each surface anomaly type label is converted into a natural sentence with richer semantic information. For example, the surface anomaly type label with the surface anomaly type of "forest fire" is converted into "This is an abnormal remote sensing image of a forest fire." A text prompt template for multi-spectral remote sensing images of surface anomalies can also be designed to supplement relevant information, such as the irregular spatial characteristics and complex spectral characteristics of the surface anomaly coverage area, etc.; using the encoder of the pre-trained Transformer model as the text encoder, each natural sentence is input into the text encoder and converted into a series of text features represented by vectors.
[0030] Among them, the open set of surface anomaly type labels contains an indefinite number of surface types, which are converted into natural sentences through the text prompt template; the text encoder converts natural sentences, such as "This is an abnormal remote sensing image of a forest fire," into a series of text features represented by vectors , where \(R\) represents the real number space, \(n\) represents the number of sequences of text features \(t\) (i.e., the number of categories in the open set of surface anomaly type labels), and \(D\) represents the feature dimension of the text feature \(t\) (i.e., each data point in the feature sequence is represented by a vector composed of \(D\) numerical values).
[0031] Among them, this embodiment uses the text encoder of the Contrastive Language-Image Pre-training (CLIP) model to convert each natural sentence into a series of text features represented by vectors. Specifically, the text encoder based on the Transformer architecture in the CLIP model has the same structure as the encoder of the pre-trained Transformer model but different parameters; the text encoder of the CLIP model is pre-trained on the WebImage Text (WIT) dataset, which is constructed by OpenAI and contains 400 million image and text pairs, covering diverse natural images and natural language descriptions on the Internet and is specifically used for multi-modal contrastive learning.
[0032] Among them, during the use of the text encoder of the CLIP model, all parameters remain frozen, avoiding the overfitting risk caused by a small number of surface anomaly samples during training and the performance degradation caused by fine-tuning, and leveraging its powerful semantic prior to enhance the generalization ability of the model.
[0033] Step S2: Extract and fuse the bands of the multi-spectral remote sensing image to be detected to obtain the RGB image and the SWIR image of the multi-spectral remote sensing image.
[0034] In one example, Sentinel-2 Level 2A multi-spectral data products, i.e., the multi-spectral remote sensing image to be detected, are obtained. After preprocessing such as band extraction and fusion, the red (Red), green (Green), and blue (Blue) bands of the multi-spectral remote sensing image are visualized as an RGB image, i.e., a true-color image, and the short-wave infrared 2 (Short-Wavelength Infra-Red 2, SWIR2), short-wave infrared 1 (Short-Wavelength Infra-Red 1, SWIR1), and near-infrared (Near Infra-Red, NIR) bands of the multi-spectral remote sensing image are visualized as a SWIR image, i.e., a false-color image.
[0035] Among them, the central wavelength of the short-wave infrared 2 band of the multi-spectral remote sensing image is 2.190 μm.
[0036] Among them, the central wavelength of the short-wave infrared 1 band of the multi-spectral remote sensing image is 1.610 μm.
[0037] Step S3: Input the RGB image and the SWIR image into the image encoder module for feature extraction to obtain the initial RGB features and the initial SWIR features.
[0038] Among them, the image encoder module includes two image encoders with the same network structure. One image encoder is used to process the RGB image, and the other image encoder is used to process the SWIR image.
[0039] Among them, the RGB image and the SWIR image are sliced into U parts along the height and V parts along the width by the image encoder to obtain image patches, and the image patches are linearly embedded into vectors with a fixed number of channels C to form an image patch sequence. Among them, , the values of U and V are set in the image encoder; an additional learnable classification CLS (Classification, CLS) sequence is added for the classification task; the image patch sequence and the CLS sequence form a complete embedding with a length of N and a feature channel number of for forward processing while keeping the scale unchanged, .
[0040] In one example, an RGB image with a size of 512×512×3 is processed in an image encoder (the image encoder is set to be divided into 16 parts along the height and 16 parts along the width): First, it is divided into 16*16 = 256 image patches, each image patch having a size of 32×32×3. Then each image patch is encoded into a vector (feature sequence) with the number of channels C = 768 (the size is 1×768). An additional learnable classification CLS (Classification, CLS) sequence is added for the classification task. The CLS sequence and all the image patches are combined, and finally the encoder outputs the initial RGB features with the feature length N = 256 + 1 and the number of channels C = 768.
[0041] Among them, the image encoder is used for the preliminary extraction of image features of multi-spectral remote sensing images. The encoders of the pre-trained ViT model are used to process the RGB image and the SWIR image respectively to extract the initial image features, so as to enhance the feature generalization ability in the spatial and spectral fine-tuning stages.
[0042] Among them, the image encoder of this embodiment adopts the Contrastive Language-Image Pre-training (CLIP) model. Specifically, the ViT-based image encoder in the CLIP model has the same structure as the encoder of the pre-trained ViT model, only with different parameters. The image encoder of the CLIP model is pre-trained on the WebImage Text (WIT) dataset, which is constructed by OpenAI and contains 400 million image and text pairs, covering diverse natural images and natural language descriptions on the Internet and is specifically used for multi-modal contrastive learning.
[0043] Among them, the encoder of the pre-trained Vision Transformer (ViT) model is used as the image encoder.
[0044] Among them, all the parameters of the image encoder of the CLIP model remain frozen during use, avoiding the overfitting risk caused by training with a small number of surface anomaly samples and the performance degradation caused by fine-tuning, and using its powerful semantic prior to improve the generalization ability of the model.
[0045] Among them, the image encoder of the CLIP model divides the image (i.e., the RGB image or the SWIR image) into image patches, linearly embeds these image patches into vectors with a fixed number of channels C, and then adds an additional learnable classification (Classification, CLS) sequence for the classification task. The image patches and the CLS sequence form a length of A series of forward processes are performed on the complete embedding with the number of feature channels being C while keeping the scale unchanged to obtain the initialized image features (i.e., initialized RGB features or initialized SWIR features).
[0046] It should be understood that there are significant differences between multi-spectral remote sensing images and natural images, including spectral distribution differences and spatial scale differences, while the CLIP model is only pre-trained on datasets containing natural images. Therefore, a targeted cross-domain generalization mechanism needs to be designed for the initialized image features to bridge the distribution differences, which is divided into two stages, including the spatial fine-tuning stage and the spectral fine-tuning stage.
[0047] Among them, the two image encoders respectively receive and process the RGB image and the SWIR image, extract the initialized RGB features and the initialized SWIR features, and then enter the spatial fine-tuning stage and the spectral fine-tuning stage in sequence.
[0048] Step 4: Input the initialized RGB features and the initialized SWIR features into the multi-scale spatial adapter module for spatial adjustment to obtain the multi-scale enhanced spatially adapted RGB features and the multi-scale enhanced spatially adapted SWIR features.
[0049] Among them, the multi-scale spatial adapter module is used to perform spatial fine-tuning on the initialized RGB features and the initialized SWIR features.
[0050] Among them, the multi-scale spatial adapter module includes two multi-scale spatial adapters with the same network structure. The initialized RGB features are processed by a multi-scale spatial adapter (Multi-Scale Spatial Adapter, MSSA) to improve the ability to extract local details such as the morphology and texture of the abnormal coverage area on the surface of the RGB image, and obtain the multi-scale enhanced spatially adapted RGB features; at the same time, the initialized SWIR image features are processed by another multi-scale spatial adapter to improve the ability to extract local details such as the morphology and texture of the abnormal coverage area on the surface of the SWIR image, and obtain the multi-scale enhanced spatially adapted SWIR features.
[0051] Among them, the spatial fine-tuning stage of the cross-domain generalization mechanism is reflected in step 4 of this embodiment. A multi-scale spatial adapter is established after the image encoder in both the RGB path (i.e., the path composed of the image encoder for processing the RGB image, the multi-scale spatial adapter, the linear layer, and the cross-band alignment unit) and the SWIR path (i.e., the path composed of the image encoder for processing the SWIR image, the multi-scale spatial adapter, the linear layer, and the cross-band alignment unit) to improve the model's local detail modeling and global semantic understanding capabilities.
[0052] Among them, the initialized image features (including initialized RGB features and initialized SWIR features) are used as inputs, which are separated into a CLS sequence and an image patch sequence to process global and local information respectively. After the image patch sequence is reconstructed into a 2D feature map, multi-scale features are extracted and fused through a residual structure and dilated convolutions with different dilation rates. Subsequently, the fused features are used to guide the update of the CLS sequence through global average pooling, and finally, multi-scale enhanced spatially adapted features (including multi-scale enhanced spatially adapted RGB features and multi-scale enhanced spatially adapted SWIR features) are output.
[0053] In one embodiment, the multi-scale spatial adapter module includes two multi-scale spatial adapters with the same network structure. The initialized RGB features and initialized SWIR features are input into the multi-scale spatial adapter module for spatial adjustment to obtain multi-scale enhanced spatially adapted RGB features and multi-scale enhanced spatially adapted SWIR features, including: inputting the initialized RGB features into one of the multi-scale spatial adapters of the multi-scale spatial adapter module for spatial adjustment to obtain multi-scale enhanced spatially adapted RGB features; inputting the initialized SWIR features into the other multi-scale spatial adapter of the multi-scale spatial adapter module for spatial adjustment to obtain multi-scale enhanced spatially adapted SWIR features.
[0054] In one embodiment, the processing process of the multi-scale spatial adapter is as follows: The initialized features input into the multi-scale spatial adapter are separated into a CLS sequence and an image patch sequence. The image patch sequence is reconstructed into a 2D feature map that retains spatial information. The 2D feature map is input into a convolutional enhancement unit for multi-scale convolutional processing to obtain multi-scale fused features. The convolutional enhancement unit is a four-branch multiple convolutional structure. The multi-scale fused features are input into a global feature extraction unit for feature extraction to obtain global features. The global features and the CLS sequence are added together to obtain an enhanced CLS sequence. After the multi-scale fused features are restored to a sequence format, they are recombined with the enhanced CLS sequence to obtain multi-scale enhanced spatially adapted features.
[0055] It should be understood that if this multi-scale spatial adapter is used to process the initialized RGB features, correspondingly, the initialized features input into the multi-scale spatial adapter are the initialized RGB features, and the corresponding output multi-scale enhanced spatially adapted features are multi-scale enhanced spatially adapted RGB features; if this multi-scale spatial adapter is used to process the initialized SWIR features, correspondingly, the initialized features input into the multi-scale spatial adapter are the initialized SWIR features, and the corresponding output multi-scale enhanced spatially adapted features are multi-scale enhanced spatially adapted SWIR features.
[0056] Among them, the multi-scale spatial adapter (MSSA) is used for spatial fine-tuning of the cross-domain generalization mechanism of features. It enhances the local detail extraction ability through multi-branch convolution and integrates the differences in multi-scale receptive fields to improve the expression ability and detection accuracy of the system.
[0057] In one embodiment, taking the multi-scale spatial adapter (MSSA) in the RGB path as an example, as Figure 3 shown, the initialization of image features is the initialization of RGB features, and the multi-scale enhanced spatial adaptation features are the multi-scale enhanced spatial adaptation RGB features. The specific operations are as follows: Separate the initialized RGB features into a CLS sequence and an image patch sequence to enhance the modeling capabilities of the global background and local details respectively; To enhance the ability to capture spatial details at different scales, a convolution enhancement unit is used to achieve multi-scale enhancement: specifically, the image patch sequence is reconstructed into a 2D feature map that retains spatial information , and this process can be expressed as: , represents reshaping the array shape to facilitate subsequent convolution processing. Here, H and W represent the length and width of the 2D feature map respectively, and the value of H is equal to the value of U, and the value of W is equal to the value of V; Apply a convolution enhancement unit to the 2D feature map . This convolution enhancement unit is a four-branch multiple convolution structure: First, perform a 1×1 convolution operation in each branch to adjust the number of feature channels to 1 / 4; no further convolution is performed on branch 1, which serves as a residual structure to generate an equivalent feature map to retain the original local details; the remaining three branches respectively perform dilated convolution operations. The convolution kernel sizes of the three branches are all 3×3, and the dilation rates of the three branches are 1, 2, and 3 respectively, generating multi-scale features with different local receptive fields, which are represented as: , , , ; Among them, , , , represent the convolution results of the first branch, the second branch, the third branch, and the fourth branch respectively, , , , ; , , , have receptive field sizes of 1, 3, 5, and 8 respectively; represents a standard convolution, i.e., a 1×1 convolution, used to adjust the channel dimension; , , respectively represent the dilated convolutions with dilation rates of 1 in the second branch, dilation rates of 2 in the third branch, and dilation rates of 3 in the fourth branch. Directly fuse the multi-scale features of the four branches to obtain: ; Among them, represents the multi-scale fusion feature, represents concatenation along the channel dimension; To enhance the model's ability to model global context information, by processing the multi-scale fusion feature , to update the CLS sequence representing global semantics : Introduce a global feature extraction unit. Extract global features through global average pooling GAP of the global feature extraction unit, and add the global features to the CLS sequence to obtain the enhanced CLS sequence, denoted as: ; Among them, GAP represents global average pooling, represents a fully connected layer, used to integrate global context information, represents the activation function, represents the enhanced CLS sequence, encoding both local anomaly details and global scene semantics at the same time; Finally, restore the multi-scale fusion feature to the sequence format, and recombine it with the enhanced CLS sequence to finally form the multi-scale enhanced spatially adapted RGB feature , denoted as: .
[0058] Similarly, the multi-scale spatial adapter MSSA in the SWIR path has the same structure and steps, and finally forms the multi-scale enhanced spatially adapted SWIR feature , which will not be elaborated here.
[0059] The multi-scale spatial adapter (MSSA) effectively integrates multi-scale spatial perception and global semantic modeling capabilities, not only enhancing the model's perception of fine-grained structures but also strengthening its cross-scale context understanding ability. In the task of surface anomaly detection, feature representations with scale invariance and semantic consistency are of great significance for complex scene recognition, and can significantly promote the generalization ability and adaptability of the model in heterogeneous multi-spectral image processing.
[0060] Step 5: Input the multi-scale enhanced spatially adapted RGB features and the multi-scale enhanced spatially adapted SWIR features into the cross-band alignment module for spectral adjustment to obtain RGB image features and SWIR image features.
[0061] Among them, the cross-band alignment module is used for spectral fine-tuning of the cross-domain generalization mechanism of features, achieving semantic alignment between the visible light, near-infrared, and short-wave infrared bands, suppressing redundant features with high spectral correlation, and improving the accuracy and interpretability of detection results.
[0062] Among them, the cross-band alignment module includes two linear layers with the same network structure and two cross-band alignment units with the same network structure; the multi-scale enhanced spatially adapted RGB features are processed through one linear layer and one cross-band alignment unit to supplement the response features of surface anomalies in the near-infrared, short-wave infrared 1, and short-wave infrared 2 bands, obtaining RGB image features that have been spatially and spectrally fine-tuned; meanwhile, the multi-scale enhanced spatially adapted SWIR features are processed through the other linear layer and the other cross-band alignment unit to supplement the response features of surface anomalies in the red, green, and blue bands, obtaining SWIR image features that have been spatially and spectrally fine-tuned.
[0063] In one embodiment, the cross-band alignment module includes two linear layers with the same network structure and two cross-band alignment units with the same network structure; Inputting the multi-scale enhanced spatially adapted RGB features and the multi-scale enhanced spatially adapted SWIR features into the cross-band alignment module for spectral adjustment to obtain RGB image features and SWIR image features includes: The multi-scale enhanced spatially adapted RGB features are input into one of the linear layers of the cross-band alignment module for processing to obtain RGB query features, RGB key features, and RGB numerical features; the multi-scale enhanced spatially adapted SWIR features are input into another linear layer of the cross-band alignment module for processing to obtain SWIR query features, SWIR key features, and SWIR numerical features; the SWIR query features, RGB key features, and RGB numerical features are input into one of the cross-band alignment units of the cross-band alignment module for spectral adjustment to obtain RGB image features; the RGB query features, SWIR key features, and SWIR numerical features are input into another cross-band alignment unit of the cross-band alignment module for spectral adjustment to obtain SWIR image features.
[0064] Among them, the spectral fine-tuning stage of the cross-domain generalization mechanism is reflected in step 5 of this embodiment. After the multi-scale spatial adaptation module, a cross-band alignment module (Cross-Band Alignment Module, CBAM) is established to achieve feature complementarity through differential modeling and residual fusion, improving the abnormal change perception ability and cross-band expression consistency. As Figure 4 shown, through linear projection of the multi-scale enhanced spatially adapted features (multi-scale enhanced spatially adapted RGB features or multi-scale enhanced spatially adapted SWIR features) by a linear layer, intermediate variables of the multi-spectral attention mechanism (i.e., query features, key features, and numerical features) are obtained, and then through the cross-band alignment unit, similarity measurement is achieved based on spectral information divergence (Spectral Information Divergence, SID), and a series of forward processes are implemented, finally obtaining the spatially and spectrally fine-tuned image features (i.e., RGB image features or SWIR image features). The linear layer in the RGB path projects the multi-scale enhanced spatially adapted RGB features through a learnable linear layer projection to generate the RGB query features , RGB key features , RGB numerical features , and the calculation formula is: ; where represents the linear layer; The linear layer in the SWIR path projects the multi-scale enhanced spatially adapted SWIR features through a learnable linear layer projection to generate the SWIR query features , SWIR key features , SWIR numerical features , and the calculation formula is: ; Subsequently, using the query feature, key feature, and numerical feature as the input of the cross-band alignment unit, the multi-spectral attention mechanism MSAttention is initiated.
[0065] In one embodiment, a series of forward processes are implemented: taking the RGB path as an example, the SWIR query feature, RGB key feature, and RGB numerical feature are input into one of the cross-band alignment units of the cross-band alignment module for spectral adjustment to obtain the RGB image feature, including: The SWIR query feature, RGB key feature, and RGB numerical feature undergo similarity measurement and weighted summation to obtain the multi-spectral attention weighted RGB feature with attention concentrated on the similar features of the RGB image and the SWIR image. The expression is: ; where, represents the multi-spectral attention weighted RGB feature, represents the Softmax activation function, is the attention weight coefficient of one of the cross-band alignment units, obtained by calculating the similarity between and ; represents the RGB numerical feature, R represents the real number space, N represents the length of the feature, and C represents the number of channels of the feature; where, the attention weight coefficient is generally obtained by matrix dot product: , the superscript T represents transpose; however, this similarity measurement method is difficult to effectively characterize the complex relationships between different bands and regions of the multi-spectral remote sensing images of surface anomalies. Therefore, in the multi-spectral attention mechanism MSAttention, the Spectral Information Divergence (SID) is used for the similarity measurement between the SWIR query feature and the RGB key feature to obtain the attention weight coefficient of one of the cross-band alignment units (i.e., the cross-band alignment unit on the RGB path). The expression is: ; where, refers to the similarity measurement method, is the natural exponential function, Denote the spectral information divergence matrix between the SWIR query feature and the RGB key feature. The spectral information divergence SID is a spectral classification method that measures the difference between two spectra based on information theory. Regarding the spectral vectors as random variables, it analyzes the similarity of two random vectors based on probability statistics theory, which is more in line with the actual physical process, is not sensitive to noise and illumination changes, and is more suitable for complex non-linear spectral analysis of multispectral remote sensing images. The spectral information divergence of the th column in the th row of the spectral information divergence matrix between the SWIR query feature and the RGB key feature is expressed as: ; where denotes the th sequence after probability normalization, , is the logarithmic function with base 10, denotes the th sequence after probability normalization, , and the superscript represents transpose; the smaller the spectral information divergence value of the th column in the th row of the spectral information divergence matrix between the SWIR query feature and the RGB key feature indicates that the two are more similar.
[0066] In order to more effectively characterize surface anomaly features, after subtracting the matrix of the RGB numerical feature from the multi-spectral attention weighted RGB feature and performing linear mapping, a residual connection is made with the SWIR query feature which is used as a residual structure to obtain the RGB difference feature ; In order to further enhance the cross-band feature complementarity, after the RGB difference feature is processed by layer normalization and a multi-layer perceptron, a residual connection is made with the RGB difference feature again to obtain the RGB image feature, which is expressed mathematically as: where denotes layer normalization, which is used to eliminate the offset of the feature distribution and accelerate convergence; denotes the multi-layer perceptron, which refines the features through local non-linear transformation and supplements the fine-grained information that may be ignored by the attention mechanism; denotes the RGB image feature, and this RGB image feature is the RGB feature adapted in space and spectrum.
[0067] In one embodiment, the RGB query feature, the SWIR key feature, and the SWIR numerical feature are input into another cross-band alignment unit of the cross-band alignment module for spectral adjustment to obtain the SWIR image feature, including: The RGB query feature, the SWIR key feature, and the SWIR numerical feature are subjected to similarity measurement and weighted summation to obtain a multi-spectral attention weighted SWIR feature with attention concentrated on the similar features of the SWIR image and the RGB image. The expression is: ; Wherein, represents the multi-spectral attention weighted SWIR feature, is the attention weight coefficient of another cross-band alignment unit, represents the SWIR numerical feature; Among them, the spectral information divergence is used for the similarity measurement of the RGB query feature and the SWIR key feature to obtain the attention weight coefficient of another cross-band alignment unit. The expression is: ; Wherein, represents the spectral information divergence matrix between the RGB query feature and the SWIR key feature. The spectral information divergence of the th row and the th column of the spectral information divergence matrix between the RGB query feature and the SWIR key feature is calculated as: ; Wherein, represents the th sequence after probability normalization, , represents the th sequence after probability normalization, ; The SWIR numerical feature is subjected to matrix subtraction and linear mapping with the multi-spectral attention weighted SWIR feature , and then subjected to residual connection with the RGB query feature serving as a residual structure to obtain the SWIR difference feature; the SWIR difference feature is subjected to layer normalization and multi-layer perceptron processing, and then subjected to residual connection with the SWIR difference feature to obtain the SWIR image feature.
[0068] Similarly, the cross-band alignment units in the SWIR path have the same structure and steps, from Query features , from Key features and numerical features are the inputs of the cross-band alignment unit in the SWIR path, which activates the multi-spectral attention mechanism MSAttention and performs a series of forward processes to finally form SWIR image features , that is, the SWIR features fine-tuned in space and spectrum, to achieve bidirectional information interaction between RGB images and SWIR images
[0069] The multi-spectral attention mechanism combines with the spectral information divergence SID metric, effectively enhancing the collaborative modeling ability between the multi-spectral bands of RGB images and SWIR images. Its advantages lie in being more in line with the physical distribution characteristics of multi-spectral data and having characteristics such as strong robustness and precise detail expression. Through bidirectional feature fusion and difference enhancement, the model has higher sensitivity and expressiveness to ground object changes and abnormal regions, providing a more adaptable feature modeling scheme for remote sensing scene semantic understanding
[0070] Step 6: Input the RGB image features and SWIR image features into a dynamic spectral weighted fuser for feature fusion to generate fused image features
[0071] Among them, the dynamic spectral weighted fuser DSWF is used for feature fusion after the cross-domain generalization mechanism, and realizes adaptive fusion of multi-spectral features based on the dynamic weights generated by the gating mechanism. The fused features are used as the image features required for contrast learning, thereby improving the overall performance of the system
[0072] Among them, the RGB image features and SWIR image features are simultaneously input into the DynamicSpectrally Weighted Fuser (DSWF), highlighting the key abnormal responses of multi-spectral remote sensing images and obtaining the adaptive fusion multi-spectral features fine-tuned in space and spectrum, that is, the fused image features
[0073] Among them, the feature fusion after the spatial and spectral fine-tuning stage of the cross-domain generalization mechanism is reflected in Step 6 of this embodiment. To further improve the fusion expression ability of RGB image features and SWIR image features in remote sensing anomaly detection, a dynamic spectral weighted fuser DSWF is established after the cross-band alignment module CBAM. As Figure 5As shown, the RGB image features fine-tuned in space and spectrum and the SWIR image features fine-tuned in space and spectrum are used as inputs. Through a gating mechanism, dynamic weights are generated to adjust the contributions of different modalities element-wise, and key anomaly response features are highlighted based on global and local channel interactions, thereby achieving more accurate cross-band feature fusion and finally outputting the adaptively fused multi-spectral features fine-tuned in space and spectrum (i.e., the fused image features).
[0074] In one embodiment, the RGB image features and the SWIR image features are input into a dynamic spectral weighting fusion device for feature fusion to generate the fused image features, including: The RGB image features are input into the dynamic spectral weighting fusion device and after layer normalization operation, feature separation is performed to obtain the CLS sequence of the RGB image features and the patch sequence of the RGB image features; the SWIR image features are input into the dynamic spectral weighting fusion device and after layer normalization operation, feature separation is performed into the CLS sequence of the SWIR image features and the patch sequence of the SWIR image features; the patch sequence of the RGB image features and the patch sequence of the SWIR image features are reconstructed into the 2D feature map of the RGB image features and the 2D feature map of the SWIR image features; after the 2D feature map of the RGB image features and the 2D feature map of the SWIR image features are concatenated along the channel dimension, they are input into the gating mechanism to generate an element-wise gating signal; according to the element-wise gating signal, multi-modal feature adaptive fusion is performed on the 2D feature map of the RGB image features and the 2D feature map of the SWIR image features to obtain the multi-modal fusion features, and the expression is: ; Where, represents reshaping the array shape, represents the element-wise gating signal, H and W respectively represent the length and width of the 2D feature map, represents a matrix with all elements being 1, represents element-wise multiplication; represents the multi-modal fusion features, represents the patch sequence of the RGB image features, represents the patch sequence of the SWIR image features; After performing global average pooling on the multi-modal fusion features, they are concatenated and flattened with the CLS sequence of the RGB image features and the CLS sequence of the SWIR image features to obtain the flattened image features; a channel re-weighting operation with a three-branch parallel structure is performed on the flattened image features to obtain the fused image features, and the expression of the channel re-weighting operation with a three-branch parallel structure is: ; Where, represents the Softmax activation function, represents a 1D convolution operation, represents the fused image features, represents the flattened image features, represents a fully connected layer.
[0075] Specifically: The RGB image features after spatial and spectral fine-tuning are separated into a CLS sequence and an image patch sequence after layer normalization operation; meanwhile, the SWIR image features after spatial and spectral fine-tuning are separated into a CLS sequence and an image patch sequence ; The image patch sequence is reconstructed into a 2D feature map and then concatenated along the channel dimension to generate element-wise dynamic weights (i.e., element-wise gating signals) based on a gating mechanism to achieve adaptive fusion of multi-modal features, expressed as: , ; where has an input channel number of 2C and an output channel number of C, is the Sigmoid activation function, represents the element-wise gating signal, represents a matrix with all elements being 1, and adaptive fusion of RGB and SWIR multi-modal features is achieved through element-wise multiplication ; represents the multi-modal fusion features; To integrate the global anomaly information of multi-spectral remote sensing images for surface anomaly detection, the multi-modal fusion features are subjected to global average pooling , and are concatenated and flattened with the CLS sequence that aggregates all image patch information, expressed as: ; where represents the flattening operation, represents the flattened image features, which fuse multi-modal information of spatial and spectral multi-scale interactions; To further enhance the accuracy of anomaly detection, for A three-branch parallel structure of channel reweighting operation is used to realize global and local cross-channel information interaction: one branch is not weighted and retains important information as a residual structure; the other two branches, one modeling global channel association through a fully connected layer and the other modeling local channel association through a 1D convolution operation, reweight the important channels that are sensitive to surface anomalies at different perspectives, expressed as: ; in, represents a 1D convolution operation, It represents the adaptively fused multispectral features (i.e., the fused image features) that have been fine-tuned spatially and spectrally, i.e., the image features that are finally passed to the contrastive learning stage, which can highlight the key abnormal responses.
[0076] Among them, the dynamic spectral weighted fusion DSWF effectively integrates the global and local cross-channel response characteristics through channel dimension splicing and parallel information interaction structure, and improves the recognition sensitivity of surface anomalies. The fusion features finally extracted have strong generalization anomaly expression capabilities, and can build accurate cross-modal alignment with text features, providing reliable support and semantic discrimination basis for anomaly detection tasks.
[0077] Step 7: Perform cosine similarity analysis on the fused image features and the text features corresponding to each surface anomaly type label to determine the surface anomaly type of the multispectral remote sensing image.
[0078] In one embodiment, a cosine similarity analysis is performed on the fused image features and the text features corresponding to each surface anomaly type label to determine the surface anomaly type of the multispectral remote sensing image, including: The fused image features and the text features corresponding to each surface anomaly type label are mapped to a shared semantic space with the same dimension, and the cosine similarity between the fused image features and the text features corresponding to each surface anomaly type is calculated; the cosine similarity between the fused image features and the text features corresponding to each surface anomaly type is processed through the Softmax activation function to obtain the probability that the fused image features belong to each surface anomaly type label, and the surface anomaly type label with the largest probability is used as the surface anomaly type of the multispectral remote sensing image.
[0079] The determination of the type of surface anomaly in the multispectral remote sensing image based on the acquired fused image features is embodied in step 7 of this embodiment. Figure 2 As shown in the figure, contrastive learning is used to establish cross-modal associations and generate surface anomaly type labels corresponding to images, which specifically includes the following steps: The fused image features With text features As input, both are mapped to a shared semantic space of the same dimension, represented as: , ; Among them, is the image feature after mapping, is the text feature after mapping, represents the dimensionality size of the shared semantic space, which is generally set to 512.
[0080] Calculate the cosine similarity matrix of the image feature after mapping and the text feature after mapping, and obtain the normalized probability through the Softmax activation function: ; Among them, is to calculate the cosine similarity, represents the probability that the fused image feature belongs to the labels of various surface anomaly types, which is a matrix with 1 row and
[0081] For example, the label with the highest probability is the "forest fire" label, and the "forest fire" label with the highest probability is used as the output of the surface anomaly type of this multispectral remote sensing image.
[0082] It should be understood that although each step in the flowchart of Figure 1 is shown in sequence according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1At least a part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed and completed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0083] In one embodiment, a remote sensing image surface anomaly detection device based on graphic-text collaborative processing is provided, which is characterized by including: A text feature acquisition module, configured to acquire text features corresponding to each surface anomaly type label. The text features corresponding to the surface anomaly type label are pre-converted into corresponding natural sentences according to the surface anomaly type label and a custom text prompt template, and the natural sentences are subjected to feature extraction by a text encoder to obtain text features represented by vectors; A band extraction module, configured to perform band extraction and fusion on a multi-spectral remote sensing image to be detected to obtain an RGB image and a SWIR image of the multi-spectral remote sensing image; An image encoder module, configured to perform feature extraction on the RGB image and the SWIR image to obtain an initial RGB feature and an initial SWIR feature; A multi-scale spatial adapter module, configured to perform spatial adjustment on the initial RGB feature and the initial SWIR feature to obtain a multi-scale enhanced spatially adapted RGB feature and a multi-scale enhanced spatially adapted SWIR feature; A cross-band alignment module, configured to perform spectral adjustment on the multi-scale enhanced spatially adapted RGB feature and the multi-scale enhanced spatially adapted SWIR feature to obtain an RGB image feature and a SWIR image feature; A dynamic spectral weighting fuser, configured to perform feature fusion on the RGB image feature and the SWIR image feature to generate a fused image feature; A similarity analysis module, configured to perform cosine similarity analysis on the fused image feature and the text features corresponding to each surface anomaly type label to determine the surface anomaly type of the multi-spectral remote sensing image.
[0084] For the specific limitations of the remote sensing image surface anomaly detection device based on graphic-text collaborative processing, reference can be made to the limitations on the remote sensing image surface anomaly detection method based on graphic-text collaborative processing in the above text, which will not be elaborated here. Each module in the above-mentioned remote sensing image surface anomaly detection device based on graphic-text collaborative processing can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0085] A computer device includes a memory and a processor. The memory stores a computer program. It is characterized in that when the processor executes the computer program, the steps of the above-mentioned remote sensing image surface anomaly detection method based on graphic and text collaborative processing are realized.
[0086] A computer-readable storage medium stores a computer program thereon. It is characterized in that when the computer program is executed by a processor, the steps of the above-mentioned remote sensing image surface anomaly detection method based on graphic and text collaborative processing are realized.
[0087] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0088] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0089] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for detecting surface anomalies in remote sensing images based on collaborative processing of images and texts, characterized in that It includes the following steps: Step 1: Obtain the text features corresponding to each surface anomaly type label. The text features corresponding to the surface anomaly type labels are pre-converted into corresponding natural sentences according to the surface anomaly type labels and a custom text prompt template, and the text encoder is used to extract features from the natural sentences to obtain text features represented by vectors; Step 2: Perform band extraction and fusion on the multi-spectral remote sensing image to be detected to obtain the RGB image and the SWIR image of the multi-spectral remote sensing image; Step 3: Input the RGB image and the SWIR image into the image encoder module for feature extraction to obtain the initial RGB features and the initial SWIR features; Step 4: Input the initial RGB features and the initial SWIR features into the multi-scale spatial adapter module for spatial adjustment to obtain the multi-scale enhanced spatially adapted RGB features and the multi-scale enhanced spatially adapted SWIR features; Step 5: Input the multi-scale enhanced spatially adapted RGB features and the multi-scale enhanced spatially adapted SWIR features into the cross-band alignment module for spectral adjustment to obtain the RGB image features and the SWIR image features; Step 6: Input the RGB image features and the SWIR image features into the dynamic spectral weighted fuser for feature fusion to generate the fused image features; Step 7: Perform cosine similarity analysis on the fused image features and the text features corresponding to each surface anomaly type label to determine the surface anomaly type of the multi-spectral remote sensing image.
2. The method for detecting surface anomalies in remote sensing images based on graphic and text collaborative processing according to claim 1, wherein, The multi-scale spatial adapter module includes two multi-scale spatial adapters with the same network structure; The step of inputting the initial RGB features and the initial SWIR features into the multi-scale spatial adapter module for spatial adjustment to obtain the multi-scale enhanced spatially adapted RGB features and the multi-scale enhanced spatially adapted SWIR features includes: Input the initial RGB features into one of the multi-scale spatial adapters of the multi-scale spatial adapter module for spatial adjustment to obtain the multi-scale enhanced spatially adapted RGB features; Input the initial SWIR features into the other multi-scale spatial adapter of the multi-scale spatial adapter module for spatial adjustment to obtain the multi-scale enhanced spatially adapted SWIR features.
3. The remote sensing image surface anomaly detection method based on graphic and text collaborative processing according to claim 2, characterized in that The processing process of the multi-scale spatial adapter is as follows; Separate the initial features input into the multi-scale spatial adapter into the CLS sequence and the image patch sequence; Reconstruct the image patch sequence into a 2D feature map that retains spatial information: Input the 2D feature map into the convolutional enhancement unit for multi-scale convolutional processing to obtain the multi-scale fusion features. The convolutional enhancement unit is a four-branch multiple convolutional structure; Input the multi-scale fusion features into the global feature extraction unit for feature extraction to obtain the global features; Add the global features and the CLS sequence to obtain the enhanced CLS sequence; After the multi-scale fusion features are restored to the sequence format, they are recombined with the enhanced CLS sequence to obtain the multi-scale enhanced spatially adapted features.
4. The method for detecting surface anomalies in remote sensing images based on graphic and text collaborative processing according to claim 1, characterized in that The cross-band alignment module includes two linear layers with the same network structure and two cross-band alignment units with the same network structure; The step of inputting the multi-scale enhanced spatially adapted RGB features and the multi-scale enhanced spatially adapted SWIR features into a cross-band alignment module for spectral adjustment to obtain RGB image features and SWIR image features comprises: Inputting the multi-scale enhanced spatially adapted RGB features into one of the linear layers of the cross-band alignment module for processing to obtain RGB query features, RGB key features, and RGB numerical features; Inputting the multi-scale enhanced spatially adapted SWIR features into another linear layer of the cross-band alignment module for processing to obtain SWIR query features, SWIR key features and SWIR numerical features; Inputting the SWIR query feature, the RGB key feature and the RGB numerical feature into one of the cross-band alignment units of the cross-band alignment module for spectral adjustment to obtain RGB image features; The RGB query feature, the SWIR key feature and the SWIR numerical feature are input into another cross-band alignment unit of the cross-band alignment module for spectral adjustment to obtain the SWIR image feature.
5. The method for detecting surface anomalies in remote sensing images based on graphic and text collaborative processing according to claim 4, characterized in that, The step of inputting the SWIR query feature, the RGB key feature, and the RGB numerical feature into one of the cross-band alignment units of the cross-band alignment module for spectral adjustment to obtain the RGB image feature includes: The SWIR query feature, the RGB key feature and the RGB numerical feature are subjected to similarity measurement and weighted summation to obtain a multispectral attention weighted RGB feature focusing on similar features of the RGB image and the SWIR image, and the expression is: ; Among them, represents the multi-spectral attention weighted RGB feature, represents the Softmax activation function, is the attention weight coefficient of one of the cross-band alignment units, represents the RGB numerical feature, where R represents the real number space, N represents the length of the feature, and C represents the number of channels of the feature; Among them, the spectral information divergence is used for the SWIR query feature and the RGB key feature for similarity measurement to obtain the attention weight coefficient of one of the cross-band alignment units , and the expression is: ; Among them, refers to the similarity measurement method, is the natural exponential function, represents the spectral information divergence matrix between the SWIR query feature and the RGB key feature. The spectral information divergence of the th column in the th row of the spectral information divergence matrix between the SWIR query feature and the RGB key feature is expressed as: ; Among them, represents the th sequence after probability normalization, , is the logarithmic function with base 10, represents the th sequence after probability normalization, , and the superscript represents transpose; The RGB numerical value feature and the multispectral attention weighted RGB feature are subjected to matrix subtraction and linear mapping, and then are subjected to residual connection with the SWIR query feature serving as a residual structure to obtain an RGB difference feature; The RGB difference feature is processed by layer normalization and multi-layer perceptron, and then connected with the RGB difference feature residual to obtain the RGB image feature; The step of inputting the RGB query feature, the SWIR key feature and the SWIR numerical feature into another cross-band alignment unit of the cross-band alignment module for spectral adjustment to obtain the SWIR image feature comprises: The RGB query feature, the SWIR key feature and the SWIR numerical feature are subjected to similarity measurement and weighted summation to obtain a multispectral attention weighted SWIR feature focusing on similar features of the SWIR image and the RGB image, and the expression is: ; Among them, represents the multi-spectral attention weighted SWIR feature, is the attention weight coefficient of another cross-band alignment unit, represents the SWIR numerical feature; Among them, the spectral information divergence is used for the RGB query feature and the SWIR key feature for similarity measurement to obtain the attention weight coefficient of another cross-band alignment unit , and the expression is as follows: ; Among them, represents the spectral information divergence matrix between the RGB query feature and the SWIR key feature. The spectral information divergence of the th column in the th row of the spectral information divergence matrix between the RGB query feature and the SWIR key feature is calculated as follows: ; Among them, represents the th sequence after probability normalization, , represents the th sequence after probability normalization, ; After subtracting the matrix and performing linear mapping on the SWIR numerical feature and the multi-spectral attention weighted SWIR feature and then performing residual connection with the RGB query feature as a residual structure a SWIR difference feature is obtained; The SWIR difference feature is processed by layer normalization and a multi-layer perceptron and then connected with the SWIR difference feature residual to obtain the SWIR image feature.
6. The method for detecting surface anomalies in remote sensing images based on graphic and text collaborative processing according to claim 5, wherein The step of inputting the RGB image features and the SWIR image features into a dynamic spectral weighted fusion device for feature fusion to generate fused image features includes: The RGB image features are input into a dynamic spectral weighted fusion device, and then subjected to layer normalization operation for feature separation, to obtain a CLS sequence of the RGB image features and an image block sequence of the RGB image features; The SWIR image features are input into a dynamic spectral weighted fusion device. After layer normalization operation, the features are separated into the CLS sequence of the SWIR image features and the image patch sequence of the SWIR image features; The image patch sequence of the RGB image features and the image patch sequence of the SWIR image features are reconstructed into the 2D feature map of the RGB image features and the 2D feature map of the SWIR image features; After the 2D feature map of the RGB image features and the 2D feature map of the SWIR image features are concatenated along the channel dimension, they are input into a gating mechanism to generate element-wise gating signals; According to the element-wise gating signals, multi-modal feature adaptive fusion is performed on the 2D feature map of the RGB image features and the 2D feature map of the SWIR image features to obtain multi-modal fusion features. The expression is: ; Among them, represents reshaping the array shape, represents the per-element gating signal, where H and W respectively represent the length and width of the 2D feature map, represents a matrix with all elements being 1, represents element-wise multiplication; represents the multi-modal fusion feature, represents the sequence of image patches of the RGB image feature, represents the sequence of image patches of the SWIR image feature; After global average pooling is performed on the multi-modal fusion features, they are concatenated and flattened with the CLS sequence of the RGB image features and the CLS sequence of the SWIR image features to obtain the flattened image features; Channel re-weighting operation with a three-branch parallel structure is adopted for the flattened image features to obtain the fused image features. The expression of the channel re-weighting operation with a three-branch parallel structure is: ; Among them, represents the Softmax activation function, represents the 1D convolution operation, represents the fused image features, represents the flattened image features, represents the fully connected layer.
7. The method for detecting surface anomalies in remote sensing images based on graphic and text collaborative processing according to claim 1, wherein, Performing cosine similarity analysis on the fused image features and the text features corresponding to each surface anomaly type label to determine the surface anomaly type of the multi-spectral remote sensing image, including: Mapping the fused image features and the text features corresponding to each surface anomaly type label into a shared semantic space with the same dimension, and calculating the cosine similarity between the fused image features and the text features corresponding to each surface anomaly type; Processing the cosine similarity between the fused image features and the text features corresponding to each surface anomaly type through the Softmax activation function to obtain the probability that the fused image features belong to each surface anomaly type label, and taking the surface anomaly type label with the highest probability as the surface anomaly type of the multi-spectral remote sensing image.
8. A remote sensing image surface anomaly detection device based on graphic and text collaborative processing, characterized in that, Including: A text feature acquisition module for acquiring the text features corresponding to each surface anomaly type label. The text features corresponding to the surface anomaly type label are pre-converted into corresponding natural sentences according to the surface anomaly type label and a custom text prompt template, and the natural sentences are feature-extracted by a text encoder to obtain text features represented by vectors; A band extraction module for performing band extraction and fusion on the multi-spectral remote sensing image to be detected to obtain the RGB image and the SWIR image of the multi-spectral remote sensing image; An image encoder module for performing feature extraction on the RGB image and the SWIR image to obtain initialized RGB features and initialized SWIR features; A multi-scale spatial adapter module for performing spatial adjustment on the initialized RGB features and the initialized SWIR features to obtain multi-scale enhanced spatially adapted RGB features and multi-scale enhanced spatially adapted SWIR features; A cross-band alignment module, which is used to perform spectral adjustment on the multi-scale enhanced spatially adapted RGB features and the multi-scale enhanced spatially adapted SWIR features to obtain RGB image features and SWIR image features; A dynamic spectral weighting fusion unit, which is used to perform feature fusion on the RGB image features and the SWIR image features to generate fused image features; A similarity analysis module, which is used to perform cosine similarity analysis on the fused image features and the text features corresponding to each surface anomaly type label to determine the surface anomaly type of the multi-spectral remote sensing image.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the remote sensing image surface anomaly detection method based on graphic-text collaborative processing according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the remote sensing image surface anomaly detection method based on graphic-text collaborative processing according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-modal power scene monitoring and early warning method based on image semantic fusion
CN117541863A
Weak supervision abnormal behavior detection system and equipment for children ADHD early screening and auxiliary diagnosis
CN118298430A
Improved unsupervised remote-sensing image abnormality detection method
WO2024055948A1
Cited By
Cultivated land change detection method based on cross-modal remote sensing data
CN121330499A
Geological disaster intelligent analysis method and system based on large-scale language model
CN121456331A