Remote sensing image cloud classification method, system and device based on multi-dimensional information dynamic coding and medium
By combining a remote sensing image cloud classification method based on multidimensional information dynamic encoding with multidimensional dynamic encoding cross-attention fusion and selective graph-enhanced neural networks, the problem of insufficient information fusion in remote sensing cloud type inversion is solved, achieving high-precision identification of complex cloud systems and improving model generalization ability.
Patent Information
- Application Number
- CN202510798485.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-11-07
AI Technical Summary
Existing remote sensing cloud type inversion methods are insufficient in processing multi-source satellite data fusion, making it difficult to effectively integrate non-image information such as time series and observation angles. This results in insufficient characterization of cloud field dynamics and unstable classification accuracy, especially in complex cloud structures where the effect is unsatisfactory.
A remote sensing image cloud classification method using multidimensional information dynamic encoding is proposed. By combining image data with multidimensional dynamic encoding cross-attention fusion module and selective graph enhancement neural network module, feature fusion and semantic perception are performed to construct cloud type inversion network, thereby improving the model's generalization ability and classification accuracy.
It achieves high-accuracy identification of clouds in complex weather systems and improves the model's generalization ability, breaking through the limitations of single image features, enhancing the ability to distinguish complex cloud types and the robustness of the model, and improving the accuracy and stability of remote sensing cloud type inversion.
Smart Images

Figure CN120913085A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of remote sensing image space-time fusion, and particularly relates to a remote sensing image cloud classification method, system, device and medium based on multi-dimensional information dynamic coding. BACKGROUND
[0002] In remote sensing applications, identifying and classifying clouds in the atmosphere is an important link to obtain information on the state of the Earth system. The type of cloud not only reflects its microphysical properties, but also is closely related to weather evolution and climate processes, so it has important application value in the fields of weather forecasting, natural disaster monitoring, energy assessment and environmental research. With the development of Earth observation satellites, multi-source remote sensing images provide rich data support for large-scale, long-term cloud monitoring, but how to extract effective cloud class semantic information from high-dimensional heterogeneous remote sensing data to complete the cloud type inversion task still faces many challenges.
[0003] Most traditional cloud type inversion methods rely on visible or infrared band images for texture or brightness temperature feature analysis, and their model expression ability and generalization ability are limited. In rule-driven methods, classification relies on threshold values set by human experience or expert knowledge bases, which is difficult to adapt to complex cloud system structures; and in end-to-end models based on deep learning, although frameworks such as convolutional neural networks (CNN) can learn certain spatial features, they often ignore the fusion of context dimensions such as observation perspective, temporal variation and geographical background, resulting in unsatisfactory classification results in areas with fuzzy boundaries and mixed cloud types. Especially in the aspect of multi-source satellite data fusion, existing methods often process a single image modality and lack sufficient processing of non-image information such as time series or orbit information and observation angles, limiting their ability to characterize and understand the dynamic characteristics of cloud fields. In addition, cloud images themselves have strong non-structural and high dynamic characteristics, further exacerbating the instability of classification accuracy.
[0004] Therefore, the above problems need to be solved. SUMMARY
[0005] The first object of the present application is to provide a remote sensing image cloud classification method based on multi-dimensional information dynamic coding, which greatly improves the accuracy of cloud body identification in complex weather systems and the generalization ability of the model.
[0006] The second object of the present application is to provide a remote sensing image cloud classification system based on multi-dimensional information dynamic coding.
[0007] The third object of the present application is to provide an electronic device.
[0008] The fourth object of the present application is to provide a computer storage medium.
[0009] Technical solution: To achieve the above purpose, the remote sensing image cloud classification method based on multi-dimensional information dynamic coding provided by the application comprises the following steps:
[0010] (1) The remote sensing cloud image and the cloud type label image are preprocessed and spatio-temporally matched to obtain a data set composed of cloud slice images, multi-dimensional coding information and corresponding cloud type labels, and the data set is divided into a training set, a validation set and a test set;
[0011] (2) A cloud type inversion network is constructed, the cloud type inversion network comprises a multi-dimensional dynamic coding cross attention fusion module and a selective graph enhancement neural network module, the multi-dimensional coding information and the cloud slice image are fused through the multi-dimensional dynamic coding cross attention fusion module to obtain a cloud type feature map, and the cloud type feature map is input into a plurality of serial selective graph enhancement neural network modules for learning and feature extraction, and finally a classification result is output;
[0012] (3) The training set and the validation set are used to train the cloud type inversion network, the classification result and the true value are used for loss function calculation to complete parameter optimization of the cloud type inversion network, the test set is used to evaluate the classification ability of the optimized cloud type inversion network, and the optimized network is used for cloud type inversion task of remote sensing cloud image of a region to be classified.
[0013] Optionally, the step (1) of preprocessing and spatio-temporal matching comprises the following steps:
[0014] (1.1) The acquired remote sensing cloud image and cloud type label image are subjected to radiation correction, atmospheric correction, image denoising and geometric correction processing, and the two types of images are subjected to unified resampling processing;
[0015] (1.2) Image information data of the remote sensing cloud image is extracted, the image information data comprising imaging time information, satellite orbit parameters and central latitude and longitude coordinates of the corresponding observation area; label information data is extracted from the cloud type label image, the label information data comprising imaging time information, profile scanning track and geographic projection information; the image information data and the label information data are combined to form multi-dimensional coding information;
[0016] (1.3) The cloud type label image is projected into the geographic coordinate system adopted by the original remote sensing cloud image in combination with the multi-dimensional coding information, and a spatial alignment operation is performed to obtain the cloud type label corresponding to the remote sensing cloud image;
[0017] (1.4) The registered remote sensing cloud image and the cloud type label image are cut and cropped to obtain cloud slice images, and the cloud slice images, the multi-dimensional coding information and the corresponding cloud type labels are combined to form a data set, and the data set is divided into a training set, a validation set and a test set according to a proportion.
[0018] Optionally, the multi-dimensional dynamic coding cross-attention fusion module in step (2) first encodes the multi-dimensional coding information to obtain a multi-dimensional coding information matrix, and then performs cross-attention fusion on the multi-dimensional coding information matrix and a randomly generated dynamic coding matrix to obtain a multi-dimensional dynamic coding matrix; and then performs cross-attention fusion on the multi-dimensional dynamic coding matrix and the cloud slice image to obtain a cloud type feature map.
[0019] Optionally, the multi-dimensional dynamic coding cross-attention fusion module in step (2) specifically includes the following steps:
[0020] encoding the multi-dimensional coding information to obtain a multi-dimensional coding information matrix E multi :
[0021]
[0022] wherein N is the number of dimensions of the multi-dimensional attribute, and d is the embedding dimension;
[0023] randomly generating a dynamic coding matrix E for semantic interaction with the multi-dimensional coding information matrix, M is the number of learnable dynamic vectors; inputting the multi-dimensional coding information matrix E multi and the dynamic coding matrix E dyn into a cross-attention mechanism, mapping the multi-dimensional coding information matrix E multi as a query Q, and mapping the dynamic coding matrix E dyn as a key K and a value V,
[0024] Q = E multi W Q , K = E dyn W K , and V = E dyn W V , multi
[0025] wherein W Q , W K , and W V are random matrices for mapping;
[0026] calculating the attention weight and fusing the value vector to obtain a calculation result A1 as follows:
[0027]
[0028] wherein Attention is an attention mechanism, softmax is an activation function, and T represents matrix transposition;
[0029] performing linear transformation Linear on the calculation result A1, and then performing residual connection and layer normalization LayerNorm to obtain a normalized result Norm1:
[0030] Norm1 = LayerNorm(X + Linear(A1))
[0031] Finally, through the multi-layer perception MLP and the residual connection, a multi-dimensional dynamic encoding matrix E is obtained fused The multi-dimensional dynamic encoding matrix fuses multi-dimensional encoding semantic information and dynamic context relationship;
[0032] E fused = X + Linear(A1) + MLP(Norm1),
[0033] The multi-dimensional dynamic encoding matrix E fused is cross-attention fused with the cloud slice image F , wherein H, W are the height and width of the image respectively, and D is the channel dimension; the cloud slice image F cloud is reconstructed into a two-dimensional sequence representation The multi-dimensional dynamic encoding matrix E fused is mapped into a query Q, and the two-dimensional sequence representation is mapped into a key K and a value V, and cross-attention operation is performed:
[0034] Q = E fused W Q , X = F cloud
[0035] Wherein W Q , W K , W V are random matrices for mapping;
[0036] The attention weight is calculated and the value vector is fused to obtain a calculation result A2:
[0037]
[0038] After the calculation result A2 is linearly transformed Linear, and then through the residual connection and the layer normalization LayerNorm, a normalized result Norm2 is obtained:
[0039] Norm2 = LayerNorm(X + Linear(A2))
[0040] Finally, through the multi-layer perception MLP and the residual connection, the tensor dimension is reconstructed back to the original image size, and a fused cloud type feature map T cloud is obtained:
[0041] T cloid = X + Linear(A2) + MLP(Norm2),
[0042] Optionally, the selective graph enhancement neural network module in step (2) first reconstructs the cloud type feature map dimension to form an input unit sequence, and inputs the saliency map guide module to construct a sparse semantic graph adjacency matrix. The sparse semantic graph adjacency matrix is processed by the graph enhancement attention mechanism to obtain a spatial calculation result. The spatial calculation result and the sparse semantic graph adjacency matrix are input into the local selective gate residual module to obtain a fusion result. The fusion result and the sparse semantic graph adjacency matrix are connected in residual to obtain a calculation result. Then, the calculation result is processed by layer normalization, multilayer perception, and residual connection to obtain a classification result.
[0043] Optionally, the selective graph enhancement neural network module in step (2) specifically includes the following steps:
[0044] First, the three-dimensional tensor cloud type feature map T cloud is reconstructed to a two-dimensional sequence vector form: N = H x W, the height H and the width W of the cloud type feature map are taken as the input unit sequence X;
[0045] A saliency map guide module is constructed to perform feedforward neural network on each input unit in the input unit sequence X to generate a saliency vector s i = MLP(x i ), x i represents the i-th input unit in the input unit sequence X, and a similarity matrix A ij = cos(s i , s j ) is calculated based on the cosine similarity of the saliency vectors between the input units. A sparse semantic graph adjacency matrix A graph is constructed by screening high correlation adjacency vectors using a Top-K strategy.
[0046] A graph enhancement attention mechanism is constructed to introduce graph structure information on the basis of a standard self-attention mechanism. The sparse semantic graph adjacency matrix A graph is taken as the query Q, the key K, and the value V to perform cross-attention processing to obtain a calculation result SA(A graph ), and the expression is:
[0047]
[0048] where γ is a trainable graph perception factor, softmax is an activation function, and D is an input vector dimension;
[0049] A local selective gate residual module is constructed, which generates a gate factor to introduce a gating mechanism to the sparse semantic graph adjacency matrix, and performs gate fusion with the graph enhancement attention to obtain a fusion result z, and the expression is:
[0050] g=σ(W g A graph +b)
[0051] z=g⊙SA(A graph )+(1-g)⊙A graph
[0052] Where ⊙ denotes element-wise multiplication, σ is the Sigmoid activation function, and W... g Both b and are trainable parameters;
[0053] The fusion result z is compared with the sparse semantic graph adjacency matrix A. grapj The calculation results are obtained by performing residual connections. Then, the calculation results are processed by LayerNorm, MLP, and residual connections to obtain the classification results.
[0054] FE = MLP(LayerNorm(A) graph +z))+A graph +z.
[0055] Based on the same inventive concept, the remote sensing image cloud classification system based on multidimensional information dynamic coding described in this invention includes:
[0056] The training data processing module is used to preprocess and spatiotemporally match remote sensing cloud images and cloud type label images to obtain a dataset consisting of cloud slice images, multidimensional coding information and corresponding cloud type labels, and divide the dataset into training set, validation set and test set.
[0057] The inversion network construction module is used to construct the cloud type inversion network. The cloud type inversion network includes a multidimensional dynamic coding cross-attention fusion module and a selective graph augmentation neural network module. The multidimensional dynamic coding cross-attention fusion module fuses multidimensional coding information and cloud slice images to obtain cloud type feature maps. The cloud type feature maps are then input into multiple cascaded selective graph augmentation neural network modules for learning and feature extraction, and finally output the classification results.
[0058] The cloud type inversion module is used to train the cloud type inversion network using training and validation sets, calculate the loss function between the classification results and the ground truth to optimize the parameters of the cloud type inversion network, evaluate the classification ability of the optimized cloud type inversion network using the test set, and use the optimized network to perform cloud type inversion tasks on remote sensing cloud images of the study area to be classified.
[0059] Optionally, the multi-dimensional dynamic coding cross-attention fusion module in the inversion network construction module first encodes the multi-dimensional coding information to obtain a multi-dimensional coding information matrix, and then performs cross-attention fusion on the multi-dimensional coding information matrix and a randomly generated dynamic coding matrix to obtain a multi-dimensional dynamic coding matrix; the multi-dimensional dynamic coding matrix and the cloud slice image are then subjected to cross-attention fusion to obtain a cloud type feature map; the selective graph enhancement neural network module first reconstructs the dimensions of the cloud type feature map to form an input unit sequence, and inputs the input unit sequence into the significant map guiding module to construct a sparse semantic graph adjacency matrix; the sparse semantic graph adjacency matrix is subjected to cross-attention processing by a graph enhancement attention mechanism to obtain a spatial calculation result; the spatial calculation result and the sparse semantic graph adjacency matrix are input into a local selective gate residual module to obtain a fusion result; the fusion result and the sparse semantic graph adjacency matrix are subjected to residual connection to obtain a calculation result; and the calculation result is subjected to layer normalization, a multilayer perceptron and residual connection to obtain a classification result.
[0060] Based on the same inventive concept, the electronic device provided by the application comprises a processor and a storage medium;
[0061] The storage medium is used for storing instructions;
[0062] The processor is used for operating according to the instructions to perform the steps of the method described above.
[0063] Based on the same inventive concept, the computer readable storage medium provided by the application has a computer program stored thereon, and the program is executed by a processor to implement the steps of the method described above.
[0064] Advantages: Compared with the prior art, the application has the following remarkable advantages:
[0065] (1) The application comprehensively utilizes image data and multi-dimensional dynamic coding information, performs feature fusion through a cross-attention mechanism, and constructs a semantic perception cloud type inversion network based on a selective graph enhancement neural network module, so as to break through the problem of single image feature dependence and insufficient context modeling in the existing cloud type inversion, realize multi-dimensional information fusion and spatial structure enhancement deep learning modeling, and greatly improve the accuracy of cloud body recognition in complex weather systems and the generalization ability of the model;
[0066] (2) The application breaks through the limitation of single image feature, performs deep fusion on multi-source coding information and image features through a multi-dimensional dynamic coding cross-attention fusion module, introduces a selective graph enhancement neural network module, establishes semantic adjacency relationships in non-Euclidean space, greatly improves the discrimination ability and model generalization ability of complex cloud types, and greatly improves the precision and robustness of remote sensing cloud type inversion;
[0067] (3) The multi-dimensional dynamic coding cross-attention fusion module in the application fuses the non-visual attributes (such as time, track parameters, geographic location information and other multi-dimensional coding information) of the remote sensing image with the image features through deep semantic fusion, introduces a learnable dynamic coding matrix to realize context semantic interaction, effectively enhances the perception ability of the model to the time and spatial changes in the remote sensing scene, improves the recognition and expression ability of the feature representation, optimizes the stability and efficiency of the training process through the residual connection and layer normalization mechanism, and thus significantly improves the accuracy of cloud type classification and the generalization ability of the model.
[0068] (4) The selective graph enhancement neural network module in the application constructs a sparse semantic graph adjacency matrix in a non-Euclidean space through a saliency map guiding mechanism, introduces a graph enhancement attention mechanism to strengthen the structural correlation between different pixels, so that the model has stronger discrimination ability when processing the boundary fuzzy and cloud type mixed areas. In addition, the graph structure information is dynamically fused through a gating mechanism, which effectively realizes selective information propagation, further suppresses redundant interference, and improves the adaptability and robustness of the network to complex cloud structures. BRIEF DESCRIPTION OF DRAWINGS
[0069] Figure 1 is a flowchart of the application;
[0070] Figure 2 is a flowchart of the multi-dimensional dynamic coding cross-attention fusion module in the application;
[0071] Figure 3 is a flowchart of the cross-attention mechanism in the multi-dimensional dynamic coding cross-attention fusion module in the application;
[0072] Figure 4 is a flowchart of the selective graph enhancement neural network module in the application. DETAILED DESCRIPTION
[0073] The technical solutions of the application will be further described below with reference to the accompanying drawings.
[0074] Example 1: As shown in the application, a remote sensing image cloud classification method based on multi-dimensional information dynamic coding includes the following steps: Figure 1
[0075] (1) The remote sensing cloud image and the cloud type label image are preprocessed and spatio-temporally matched to obtain a data set composed of cloud slice images, multi-dimensional coding information and corresponding cloud type labels, and the data set is divided into a training set, a validation set and a test set.
[0076] The preprocessing and spatio-temporal matching in step (1) specifically include the following steps:
[0077] (1.1) The remote sensing cloud image obtained by Fengyun-4 satellite and the cloud type label image provided by CloudSat satellite are processed by radiation correction, atmospheric correction, image denoising and geometric correction. The two types of images are uniformly resampled to ensure consistency in radiation characteristics, geometric structure and spatial resolution;
[0078] (1.2) Extract image information data of the remote sensing cloud image, which includes imaging time information, satellite orbit parameters and central latitude and longitude coordinates of the corresponding observation area. The satellite orbit parameters include orbit height, inclination and transit time. Label information data is extracted from the cloud type label image, which includes imaging time information, profile scanning track and geographic projection information. The image information data and the label information data are combined to form multi-dimensional encoding information;
[0079] (1.3) Combined with the multi-dimensional encoding information, the cloud type label image is projected into the geographic coordinate system used by the original remote sensing cloud image using a geographic registration algorithm. A high-order transformation method based on RPC model (Rational Polynomial Coefficients) is used for spatial alignment operation to ensure the consistency of the two images in space geometry, and the cloud type label corresponding to the remote sensing cloud image is obtained.
[0080] (1.4) The registered remote sensing cloud image and the cloud type label image are sliced and cropped to obtain cloud slice images, and the cloud slice images, multi-dimensional encoding information and corresponding cloud type labels are combined to form a data set. The data set is divided into training set, validation set and test set in the ratio of 7:2:1.
[0081] (2) Construct a cloud type inversion network, which includes a multi-dimensional dynamic coding cross attention fusion module and a selective graph augmented neural network module (Selective Graph-Augmented Transformer Block, SGATB). The multi-dimensional dynamic coding cross attention fusion module is used to fuse the multi-dimensional encoding information and the cloud slice image to obtain a cloud type feature map. The cloud type feature map is input into multiple serial selective graph augmented neural network modules for learning and feature extraction, and finally the classification result is output.
[0082] As Figure 2As shown, the multi-dimensional dynamic coding cross-attention fusion module first encodes the multi-dimensional coding information to obtain a multi-dimensional coding information matrix, and then cross-attention fuses the multi-dimensional coding information matrix with a randomly generated dynamic coding matrix to obtain a multi-dimensional dynamic coding matrix; then cross-attention fuses the multi-dimensional dynamic coding matrix and the cloud slice image to obtain a cloud type feature map; then inputs the cloud type feature map into a plurality of serial selective image enhancement neural network modules for learning and feature extraction, and finally outputs a classification result. The multi-dimensional dynamic coding cross-attention fusion module encodes non-image attributes such as time, geographical position, meteorological parameters, and sensor angle, and then performs semantic fusion with image features through a cross-attention mechanism.
[0083] As shown in Figure 3 , the multi-dimensional dynamic coding cross-attention fusion module specifically includes the following steps:
[0084] Encode the multi-dimensional coding information to obtain a multi-dimensional coding information matrix:
[0085]
[0086] Where N is the number of dimensions of the multi-dimensional attribute, d is the embedding dimension, E multi represents the multi-dimensional coding information matrix;
[0087] Randomly generate a dynamic coding matrix for semantic interaction with the multi-dimensional coding information matrix, M is the number of learnable dynamic vectors; input the multi-dimensional coding information matrix E multi and the dynamic coding matrix E dyn into a cross-attention mechanism, map E mu1ti as query Q, and map E dyn as key K and value V,
[0088] Q = E multi W Q , K = E dyn W K , and V = E dyn W V , X = E multi
[0089] Where W Q , W K , and W V are random matrices for mapping;
[0090] Calculate the attention weight and fuse the value vector to obtain the calculation result A1 as:
[0091]
[0092] Where Attention is an attention mechanism, softmax is an activation function, and T represents matrix transposition.
[0093] After the calculation result A1 passes through linear transformation Linear, residual connection and layer normalization LayerNorm, the normalized result Norm1 is obtained:
[0094] Norm1 = LayerNorm(X + Linear(A1))
[0095] Finally, through multi-layer perception MLP and residual connection, the fused multi-dimensional dynamic encoding matrix E is obtained fused The multi-dimensional dynamic encoding matrix fuses multi-dimensional encoding semantic information and dynamic context relationship.
[0096] E fused = X + Linear(A1) + MLP(Norm1),
[0097] The multi-dimensional dynamic encoding matrix E fused is cross-attention fused with the cloud slice image , where H and W are the height and width of the image respectively, and D is the channel dimension; F cloud is reconstructed into a two-dimensional sequence representation E fused is taken as the query Q, the key K and the value V, and cross-attention operation is performed:
[0098] Q = E fused W Q , X = F cloud
[0099] Where W Q , W K , and W V are random matrices for mapping.
[0100] The attention weight is calculated and the value vector is fused to obtain the calculation result A2:
[0101]
[0102] Subsequently, the calculation result A2 passes through linear transformation Linear, residual connection and layer normalization LayerNorm, and the normalized result Norm2 is obtained:
[0103] Norm2 = LayerNorm(X + Linear(A2))
[0104] Finally, through the multi-layer perception MLP and the residual connection, the tensor dimension is reconstructed back to the original picture size to obtain the fused cloud type feature map T cloud :
[0105] T cloud = X + Linear(A2) + MLP(Norm2),
[0106] Finally, the cloud type feature map fused with multi-dimensional semantic information is obtained, which provides a more context-aware high-quality representation for the subsequent classification module.
[0107] In the selective graph enhancement neural network module, the cloud type feature map is reconstructed in dimension to form an input unit sequence, which is input into the saliency graph guiding module to construct a sparse semantic graph adjacency matrix. The sparse semantic graph adjacency matrix is subjected to cross-attention processing through the graph enhancement attention mechanism to obtain a spatial calculation result. The spatial calculation result and the sparse semantic graph adjacency matrix are input into the local selective gate residual module to obtain a fusion result. The fusion result and the sparse semantic graph adjacency matrix are subjected to residual connection to obtain a calculation result. Then, the calculation result is subjected to layer normalization, multi-layer perception and residual connection to obtain a classification result. The selective graph enhancement neural network module SGATB introduces a saliency guiding mechanism and a graph structure enhancement attention mechanism, extracts cloud feature expression under non-Euclidean adjacency relationship, and cooperates with a gate residual mechanism to suppress redundant background interference.
[0108] As shown in Figure 4 , the selective graph enhancement neural network module specifically includes the following steps:
[0109] First, the three-dimensional tensor cloud type feature map T cloud is reconstructed in dimension to a two-dimensional sequence vector form: N = H x W, where H and W are the height and width of the cloud type feature map. This representation facilitates serving as an input unit token of the SGATB;
[0110] A saliency graph guiding module is constructed to perform a feedforward neural network on each input unit in the input unit sequence X to generate a saliency vector s i = MLP(x i ), where x i represents the i-th input unit in the input unit sequence X, and a similarity matrix A ij is calculated based on the cosine similarity of the saliency vectors between the input units, i.e., A i = cos(s j ). Then, a Top-K strategy is used to select high-correlation adjacency vectors to construct a sparse semantic graph adjacency matrix A graph .
[0111] A graph-aware attention (GAA) mechanism is constructed to introduce graph structure information on the basis of a standard self-attention mechanism, and a sparse semantic graph adjacency matrix A graph As a query Q, a key K and a value V, cross-attention processing is performed to obtain a calculation result SA(A graph ) The calculation formula of SA(A graph ) is as follows:
[0112]
[0113] where γ is a trainable graph-aware factor, softmax is an activation function, D is an input vector dimension, A graph is a sparse semantic graph adjacency matrix, which is used to enhance the modeling capability of the attention mechanism for semantic adjacency relationships in a non-Euclidean space; SA represents a calculation result processed by the graph-aware attention mechanism, which is used to retain semantic input unit regions with discriminative ability and suppress background or interference information.
[0114] A local selective gating residual module is constructed, in which a gating factor is generated to introduce a gating mechanism for the sparse semantic graph adjacency matrix, and the gating fusion result z is obtained by graph-aware attention weighting and gating fusion to enhance information selectivity, and the expression is as follows:
[0115] g=σ(W g A graph +b)
[0116] z=g⊙SA(A graph )+(1-g)⊙A graph
[0117] where ⊙ represents element-level multiplication, σ is a Sigmoid activation function, W g and b are trainable parameters,
[0118] The fusion result z and the sparse semantic graph adjacency matrix A graph are connected in a residual manner to obtain a calculation result, and then the calculation result is subjected to layer normalization LayerNorm, multi-layer perception MLP and residual connection to obtain a classification result FE, so as to complete further transformation and stable training of information.
[0119] FE=MLP(LayerNorm(A graph +z))+A graph +z
[0120] (3) using the training set and the validation set to train the cloud type inversion network, calculating the loss function of the classification result and the true value to complete the parameter optimization of the cloud type inversion network, using the test set to evaluate the classification ability of the optimized cloud type inversion network, and using the optimized network to perform the cloud type inversion task on the remote sensing cloud image of the to-be-classified research area; Specifically, the remote sensing cloud image of the to-be-classified research area is cut and cropped to obtain a 12x12 cloud slice image with each pixel point as a center point, and is input into the cloud type inversion network to obtain the cloud type inversion result of each center point. The present application processes the image of the to-be-classified area in a sliding window manner and outputs the cloud class of each pixel point, realizing fine-grained and high-precision cloud type inversion.
[0121] The present application effectively integrates the influence of time information, geographical position information, meteorological parameter information and observation angle information on the presentation of different cloud types in the cloud image, and uses the cloud type inversion network of the present application to perform cloud type inversion, which can complete the cloud type inversion task more quickly and accurately compared with the existing cloud type inversion method.
[0122] Embodiment 2: A remote sensing image cloud classification system based on multi-dimensional information dynamic coding in the present embodiment comprises:
[0123] The training data processing module is used for pre-processing and spatio-temporal matching of the remote sensing cloud image and the cloud type label image, obtaining a data set composed of cloud slice images, multi-dimensional coding information and corresponding cloud type labels, and dividing the data set into a training set, a validation set and a test set.
[0124] The pre-processing and spatio-temporal matching in the training data processing module specifically include the following steps:
[0125] The remote sensing cloud image obtained by the Fengyun-4 satellite and the cloud type label image provided by the CloudSat satellite are subjected to radiation correction, atmospheric correction, image denoising and geometric correction processing; and the two types of images are subjected to unified resampling processing to ensure consistency in radiation characteristics, geometric structure and spatial resolution;
[0126] The image information data of the remote sensing cloud image is extracted, the image information data including imaging time information, satellite orbit parameters and central latitude and longitude coordinates of the corresponding observation area, the satellite orbit parameters including orbit height, inclination and transit time; the label information data is extracted from the cloud type label image, the label information data including imaging time information, profile scanning track and geographical projection information; the image information data and the label information data are combined to form multi-dimensional coding information;
[0127] In combination with the multi-dimensional coding information, a geographic registration algorithm is used to project the cloud type label image into the geographic coordinate system used by the original remote sensing cloud image, a high-order transformation method based on the RPC model (Rational Polynomial Coefficients) is used for spatial alignment operation, and the consistency of the two images in space geometry is ensured, so as to obtain the cloud type label corresponding to the remote sensing cloud image;
[0128] The registered remote sensing cloud image and the cloud type label image are sliced and cropped to obtain a cloud slice image, and the cloud slice image, the multi-dimensional coding information and the corresponding cloud type label are combined to form a data set, and the data set is divided into a training set, a verification set and a test set in a ratio of 7:2:1.
[0129] The inversion network construction module is configured to construct a cloud type inversion network, which includes a multi-dimensional dynamic coding cross-attention fusion module and a selective graph enhancement neural network module SGATB. The multi-dimensional dynamic coding cross-attention fusion module is used to fuse the multi-dimensional coding information and the cloud slice image to obtain a cloud type feature map, and the cloud type feature map is input into a plurality of serially connected selective graph enhancement neural network modules for learning and feature extraction, and finally a classification result is output.
[0130] In the multi-dimensional dynamic coding cross-attention fusion module, the multi-dimensional coding information is first encoded to obtain a multi-dimensional coding information matrix, and then the multi-dimensional coding information matrix is cross-attention fused with a randomly generated dynamic coding matrix to obtain a multi-dimensional dynamic coding matrix. Then, the multi-dimensional dynamic coding matrix and the cloud slice image are cross-attention fused to obtain a cloud type feature map. Then, the cloud type feature map is input into a plurality of serially connected selective graph enhancement neural network modules for learning and feature extraction, and finally a classification result is output.
[0131] The multi-dimensional dynamic coding cross-attention fusion module specifically includes the following steps:
[0132] The multi-dimensional coding information is encoded to obtain a multi-dimensional coding information matrix:
[0133]
[0134] Wherein N is the number of dimensions of multi-dimensional attributes, d is the embedding dimension, E multi represents the multi-dimensional coding information matrix;
[0135] A dynamic coding matrix E is randomly generated for semantic interaction with the multi-dimensional coding information matrix E multi and the dynamic coding matrix E dyn is input into the cross-attention mechanism, so that E multi is mapped as the query Q, and Edyn The mapping is between key K and value V.
[0136] Q = E multi W Q K = E dyn W K V = E dyn W V X = E multi
[0137] Among them W Q W K W V All are random matrices used for mapping;
[0138] Calculate the attention weights and fuse the value vectors to obtain the result A1:
[0139]
[0140] Where Attention is the attention mechanism, softmax is the activation function, and T represents the matrix transpose;
[0141] The calculated result A1 is then subjected to a linear transformation (Linear), followed by residual connections and layer normalization (LayerNorm) to obtain the normalized result Norm1:
[0142] Norm1=LayerNorm(X+Linear(A1))
[0143] Finally, the fused multidimensional dynamic coding matrix E is obtained through a multilayer perceptron (MLP) and residual connections. fused The multidimensional dynamic coding matrix integrates multidimensional coding semantic information with dynamic contextual relationships;
[0144]
[0145] The multidimensional dynamic coding matrix E fused With cloud slice images Perform cross-attention fusion, where H and W are the height and width of the image, respectively, and D is the channel dimension; E cloud Reconstruct the tensor dimension to a two-dimensional sequence representation E fused As a query Q, Perform a cross-attention operation on key K and value V:
[0146] Q = E fused W Q , X = F cloud
[0147] Among them W Q WK , W V are random matrices for mapping;
[0148] The attention weight is calculated and the value vector is fused to obtain a calculation result A2:
[0149]
[0150] The calculation result A2 is then subjected to linear transformation Linear, residual connection and layer normalization LayerNorm to obtain a normalized result Norm2:
[0151] Norm2 = LayerNorm(X + Linear(A2))
[0152] Finally, the tensor dimension is reconstructed to the original image size through multi-layer perception MLP and residual connection to obtain the fused cloud type feature map T cloud :
[0153] T cloud = X + Linear(A2) + MLP(Norm2),
[0154] Finally, the cloud type feature map with fused multi-dimensional semantic information is obtained, which provides a more context-aware high-quality representation for the subsequent classification module.
[0155] In the selective graph enhancement neural network module, the cloud type feature map is reconstructed in dimension to form an input unit sequence, which is input into the saliency graph guiding module to construct a sparse semantic graph adjacency matrix. The sparse semantic graph adjacency matrix is subjected to cross-attention processing through the graph enhancement attention mechanism to obtain a spatial calculation result. The spatial calculation result and the sparse semantic graph adjacency matrix are input into the local selective gated residual module to obtain a fusion result. The fusion result is subjected to residual connection with the sparse semantic graph adjacency matrix to obtain a calculation result. Then, the calculation result is subjected to layer normalization, multi-layer perception and residual connection to obtain a classification result.
[0156] The selective graph enhancement neural network module specifically includes the following steps:
[0157] Firstly, the three-dimensional tensor cloud type feature map T cloud is reconstructed in dimension to a two-dimensional sequence vector form: N = H x W, the height H and the width W of the cloud type feature map, which is convenient as an input unit of SGATB;
[0158] A saliency graph guiding module is constructed for performing feedforward neural network on each input unit in the input unit sequence to generate a saliency vector s i = MLP(x i ), xi denotes the i-th input unit in the input unit sequence X, and a similarity matrix A is calculated based on the cosine similarity of the significance vectors between input units ij = cos(s i ,s j ), and a Top-K strategy is used to screen high-correlation adjacent vectors to construct a sparse semantic graph adjacency matrix A graph ;
[0159] A graph-aware attention (GAA) mechanism is constructed to introduce graph structure information on the basis of a standard self-attention mechanism, and the sparse semantic graph adjacency matrix A graph is taken as a query Q, a key K and a value V to perform cross-attention processing to obtain a calculation result SA(A graph ), and the calculation formula of SA(A graph ) is as follows:
[0160]
[0161] where γ is a trainable graph-aware factor, softmax is an activation function, D is the dimension of an input vector, A graph is the sparse semantic graph adjacency matrix, which is used to enhance the modeling capability of the attention mechanism for semantic adjacency relationships in non-Euclidean space; SA represents the calculation result after the graph-aware attention mechanism processing, which is used to retain the semantic input unit region with discriminative ability and suppress background or interference information;
[0162] A local selective gating residual module is constructed, in which a gating factor is generated to introduce a gating mechanism for the sparse semantic graph adjacency matrix, and a gating fusion result z is obtained by graph-aware attention weighting and gating fusion to enhance information selectivity, and the expression is as follows:
[0163] g = σ(W g A graph +b)
[0164] z = g o SA(A graph )+(1-g) o A graph
[0165] where o represents element-level multiplication, σ is a Sigmoid activation function, W g and b are trainable parameters,
[0166] The fusion result z and the sparse semantic graph adjacency matrix A graph are connected in a residual manner to obtain a calculation result, and then the calculation result is subjected to layer normalization LayerNorm, multi-layer perception MLP and residual connection to obtain a classification result FE, so as to complete further transformation and stable training process of information;
[0167] FE = MLP (LayerNorm (A graph +z)) + A graph +z
[0168] The cloud type inversion module is configured to train the cloud type inversion network using the training set and the verification set, to perform loss function calculation on the classification result and the true value to optimize the parameters of the cloud type inversion network, to evaluate the classification ability of the optimized cloud type inversion network using the test set, and to perform a cloud type inversion task on remote sensing cloud images of a to-be-classified research area using the optimized network. Specifically, the remote sensing cloud images of the to-be-classified research area are sliced and cropped with each pixel point as a center point to obtain 12x12 cloud slice images, and the cloud slice images are input into the cloud type inversion network to obtain a cloud type inversion result of each center point.
[0169] Embodiment 3: An electronic device in this embodiment includes a processor and a storage medium; the storage medium is configured to store instructions; the processor is configured to operate according to the instructions to perform the steps of the method as described above.
[0170] Embodiment 4: A computer-readable storage medium in this embodiment has a computer program stored thereon, and the program is executed by a processor to implement the steps of the method as described above.
Claims
1. A cloud classification method for remote sensing images based on dynamic coding of multi-dimensional information, characterized in that, The method comprises the following steps: (1) preprocessing and spatio-temporal matching of remote sensing cloud images and cloud type label images to obtain a dataset composed of cloud slice images, multi-dimensional coding information and corresponding cloud type labels, and dividing the dataset into a training set, a validation set and a test set; (2) constructing a cloud type inversion network, the cloud type inversion network comprising a multi-dimensional dynamic coding cross-attention fusion module and a selective graph enhancement neural network module, fusing the multi-dimensional coding information and the cloud slice images through the multi-dimensional dynamic coding cross-attention fusion module to obtain a cloud type feature map, and inputting the cloud type feature map into a plurality of serial selective graph enhancement neural network modules for learning and feature extraction, and finally outputting a classification result; (3) training the cloud type inversion network using the training set and the validation set, performing loss function calculation on the classification result and the true value to complete parameter optimization of the cloud type inversion network, evaluating the classification ability of the optimized cloud type inversion network using the test set, and using the optimized network to perform a cloud type inversion task on remote sensing cloud images of a to-be-classified research area. 2.The cloud classification method based on multi-dimensional information dynamic coding of remote sensing image according to claim 1, characterized in that: The preprocessing and spatio-temporal matching in step (1) specifically comprises the following steps: (1.1) performing radiation correction, atmospheric correction, image denoising and geometric correction processing on the obtained remote sensing cloud images and cloud type label images; and uniformly resampling the two types of images; (1.2) extracting image information data of the remote sensing cloud images, the image information data comprising imaging time information, satellite orbit parameters and central latitude and longitude coordinates of the corresponding observation area; extracting label information data from the cloud type label images, the label information data comprising imaging time information, profile scanning tracks and geographic projection information; and combining the image information data and the label information data to form multi-dimensional coding information; (1.3) projecting the cloud type label images into a geographic coordinate system adopted by the original remote sensing cloud images in combination with the multi-dimensional coding information, and performing spatial alignment to obtain cloud type labels corresponding to the remote sensing cloud images; (1.4) slicing and cutting the registered remote sensing cloud images and cloud type label images to obtain cloud slice images, and combining the cloud slice images, the multi-dimensional coding information and the corresponding cloud type labels to form a dataset, the dataset being divided into a training set, a validation set and a test set in proportion. 3.The cloud classification method based on multi-dimensional information dynamic coding of remote sensing image according to claim 1, characterized in that: In the multi-dimensional dynamic coding cross-attention fusion module in step (2), the multi-dimensional coding information is first encoded to obtain a multi-dimensional coding information matrix, then the multi-dimensional coding information matrix is cross-attention fused with a randomly generated dynamic coding matrix to obtain a multi-dimensional dynamic coding matrix, and finally the multi-dimensional dynamic coding matrix and the cloud slice images are cross-attention fused to obtain a cloud type feature map.
4. The cloud classification method of remote sensing image based on dynamic coding of multi-dimensional information according to claim 3, characterized in that: The multi-dimensional dynamic coding cross-attention fusion module in step (2) specifically comprises the following steps: The multi-dimensional coding information is encoded to obtain a multi-dimensional coding information matrix E multi : wherein N is the number of dimensions of multi-dimensional attributes, and d is the embedding dimension; Randomly generate a dynamic encoding matrix For semantic interaction with the multi-dimensional encoding information matrix, M is the number of learnable dynamic vectors; the multi-dimensional encoding information matrix E multi and the dynamic encoding matrix E dyn are input into the cross-attention mechanism, and the multi-dimensional encoding information matrix E multi is mapped as the query Q, and the dynamic encoding matrix E dyn is mapped as the key K and the value V, Q = E multi W Q K = E dyn W K V = E dyn W V X = E multi where W Q , W K , W V are random matrices for mapping; The attention weight is calculated and the value vector is fused to obtain a calculation result A1 as follows: wherein Attention is an attention mechanism, softmax is an activation function, and T represents matrix transposition. The calculation result A1 is subjected to linear transformation Linear, residual connection and layer normalization LayerNorm to obtain the normalized result Norm1: Norm1 = LayerNorm(X + Linear(A1)) Finally, through the multi-layer perception (MLP) and the residual connection, a multi-dimensional dynamic encoding matrix E is obtained fused The multi-dimensional dynamic encoding matrix integrates the multi-dimensional encoding semantic information and the dynamic context relationship. The multi-dimensional dynamic encoding matrix E fused cloud slice image cross-attention fusion, where H, W are the height and width of the image respectively, and D is the channel dimension; the cloud slice image F cloud reconstruct the tensor dimension to a two-dimensional sequence representation The multi-dimensional dynamic encoding matrix E fused is mapped to the query Q, and the two-dimensional sequence representation is mapped to the key K and the value V, and cross-attention operation is performed: where W Q , W K , W V are random matrices for mapping; The attention weight is calculated and the value vector is fused to obtain the calculation result A2: The calculation result A2 is subjected to linear transformation Linear, residual connection and layer normalization LayerNorm to obtain the normalized result Norm2: Norm2 = LayerNorm(X + Linear(A2)) Finally, through the multi-layer perception MLP and the residual connection, the tensor dimension is reconstructed back to the original picture size, to obtain the fused cloud type feature map T cloud :
5. The cloud classification method of remote sensing image based on dynamic coding of multi-dimensional information according to claim 1, characterized in that: In step (2), the selective graph enhancement neural network module first reconstructs the dimension of the cloud type feature map to form an input unit sequence, and inputs the saliency map guide module to construct a sparse semantic graph adjacency matrix. The sparse semantic graph adjacency matrix is subjected to cross attention processing by the graph enhancement attention mechanism to obtain a spatial calculation result. The spatial calculation result and the sparse semantic graph adjacency matrix are input into the local selective gate residual module to obtain a fusion result. The fusion result is subjected to residual connection with the sparse semantic graph adjacency matrix to obtain a calculation result. Then, the calculation result is subjected to layer normalization, multilayer perception and residual connection to obtain a classification result. 6.The cloud classification method of remote sensing image based on multi-dimensional information dynamic coding according to claim 5, characterized in that: The selective graph enhancement neural network module in step (2) specifically includes the following steps: First, the three-dimensional tensor cloud type feature map T cloud The reconstruction dimension is two-dimensional sequence vector form: N = H x W, the height H and width W of the cloud type feature map, as the input unit sequence X; A saliency map guided module is constructed for input unit sequence Each input unit performs a feedforward neural network to generate a saliency vector s i = MLP(x i ), x i represents the i-th input unit in the input unit sequence X, and a similarity matrix A is calculated based on the cosine similarity of the saliency vectors between the input units ij = cos(s i , s j ), and a Top-K strategy is used to screen high-correlation adjacent vectors to construct a sparse semantic graph adjacent matrix A graph ; The construction graph enhances the attention mechanism, which is used for introducing the graph structure information on the basis of the standard self-attention mechanism, and the sparse semantic graph adjacency matrix A grapg As the query Q, the key K and the value V, cross attention processing is performed to obtain a calculation result SA(A graph ), and the expression is: Wherein γ is a trainable graph perception factor, softmax is an activation function, and D is an input vector dimension; A local selective gate residual module is constructed, wherein a gate factor is generated to introduce a gating mechanism to the sparse semantic graph adjacency matrix, and the graph enhancement attention is weighted and gated fusion is performed to obtain a fusion result z, expressed as: g = σ(W g A graph +b) z = g O SA(A graph ) + (1 - g) O A graph where denotes element-wise multiplication, σ is a sigmoid activation function, W g and b are trainable parameters; fusing the result z with a sparse semantic graph adjacency matrix A graph The calculation result is obtained by performing residual connection, and then the classification result is obtained after the calculation result is subjected to layer normalization LayerNorm, multi-layer perception MLP and residual connection. FE = MLP(LayerNorm(A graph +z)) + A graph +z.
7. A remote sensing image cloud classification system based on multi-dimensional information dynamic coding, characterized in that, including: A training data processing module is configured to preprocess and spatiotemporally match the remote sensing cloud image and the cloud type label image to obtain a dataset composed of cloud slice images, multi-dimensional encoding information and corresponding cloud type labels, and divide the dataset into a training set, a validation set and a test set; An inversion network construction module is configured to construct a cloud type inversion network. The cloud type inversion network includes a multi-dimensional dynamic encoding cross attention fusion module and a selective graph enhancement neural network module. The multi-dimensional dynamic encoding cross attention fusion module is configured to fuse the multi-dimensional encoding information and the cloud slice image to obtain a cloud type feature map. The cloud type feature map is input into a plurality of selective graph enhancement neural network modules connected in series for learning and feature extraction, and finally outputs a classification result. A cloud type inversion module is configured to train the cloud type inversion network using the training set and the validation set, calculate the loss function of the classification result and the true value to complete the parameter optimization of the cloud type inversion network, evaluate the classification ability of the optimized cloud type inversion network using the test set, and use the optimized network to perform a cloud type inversion task on a remote sensing cloud image of a research area to be classified. 8.The remote sensing image cloud classification system based on multi-dimensional information dynamic coding according to claim 1, characterized in that: The multi-dimensional dynamic coding cross-attention fusion module in the inversion network construction module first encodes the multi-dimensional coding information to obtain a multi-dimensional coding information matrix, and then performs cross-attention fusion on the multi-dimensional coding information matrix and a randomly generated dynamic coding matrix to obtain a multi-dimensional dynamic coding matrix; and then the multi-dimensional dynamic coding matrix and the cloud slice image are cross-attention fused to obtain a cloud type feature map; The selective map enhancement neural network module first reconstructs the dimensions of the cloud type feature map to form an input unit sequence, and inputs the input unit sequence into the saliency map guiding module to construct a sparse semantic graph adjacency matrix; the sparse semantic graph adjacency matrix is cross-attention processed by a map enhancement attention mechanism to obtain a spatial calculation result; the spatial calculation result and the sparse semantic graph adjacency matrix are input into a local selective gate residual module to obtain a fusion result; the fusion result and the sparse semantic graph adjacency matrix are residual connected to obtain a calculation result; and then the calculation result is subjected to layer normalization, a multilayer perceptron and residual connection to obtain a classification result.
9. An electronic device, comprising: comprising a processor and a storage medium; the storage medium is configured to store instructions; the processor is configured to operate according to the instructions to perform the steps of the method of any one of claims 1-6.
10. A computer readable storage medium having stored thereon a computer program, characterized in that The program is executed by the processor to implement the steps of the method of any one of claims 1-6.
Citation Information
Patent Citations
Power channel corridor routing-inspection method based on tilt photography three-dimensional reconstruction technology
CN106441233A
Satellite remote sensing image cloud classification method based on space-time coding guidance
CN119495032A
Image classification method and apparatus, and method and apparatus for training image classification model
WO2025007941A1