ViT-based online incremental learning method for remote sensing image large model
Through the combination of the space-spectral dual-branch encoder and the LoRA module, the catastrophic forgetting problem of ViT large models in online incremental learning of remote sensing images is solved, efficient and scalable online incremental learning is achieved, adapting to the dynamic changes of remote sensing data flow, and improving the robustness and adaptability of the model.
Patent Information
- Application Number
- CN202510820334.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The existing ViT big models have catastrophic forgetting problems in online incremental learning of remote sensing images, and it is difficult to continuously learn new knowledge without accessing historical data, while keeping the knowledge learned unforgettable. The existing methods have limitations such as large storage overhead, high computational complexity, or the need to clarify task boundaries.
The space-spectral dual-branch encoder and hybrid connection encoder are adopted, combined with the LoRA module for online incremental learning, and cross-modal feature interaction and enhancement is achieved through adaptive model expansion strategies and regularization terms. The low-rank decomposition of the LoRA module is used to reduce calculation and storage requirements, and the model parameters are dynamically adjusted by real-time monitoring of loss changes.
Efficient, scalable and task-independent online incremental learning is realized, and new knowledge can be continuously learned in the remote sensing data flow while suppressing the forgotten knowledge of the learned, adapting to the changes in the data distribution, and improving the robustness and adaptability of the model.
Smart Images

Figure CN120339849A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to an online incremental learning method for a large remote sensing image model based on ViT. Background Art
[0002] With the rapid development of remote sensing technology, the continuously acquired massive and dynamic remote sensing image data has provided unprecedented opportunities for refined Earth observation. Large-scale pre-trained models such as ViT (Vision Transformer) have shown great potential in remote sensing image interpretation tasks due to their powerful feature extraction capabilities. However, in practical applications, remote sensing data often arrives in the form of continuous data streams, which may contain new ground object categories, changes in sensor characteristics, or data distribution drifts brought about by different time phases and different regions. This requires the model to have the ability of online incremental learning, that is, to continuously learn from new data without accessing historical data, while not forgetting the learned knowledge. Directly fine-tuning a large ViT model online usually leads to catastrophic forgetting, severely restricting its application in dynamic remote sensing scenarios.
[0003] Existing online incremental learning methods mainly include strategies such as experience replay, regularization-based, and parameter isolation-based. The experience replay method needs to store part of the historical data, facing storage overhead and data privacy issues, and has limited effectiveness when the data distribution changes drastically. The regularization-based method alleviates forgetting by penalizing the modification of important parameters for old tasks, but the cost of calculating the global parameter importance on a large ViT model is huge and it is difficult to meet the requirements of online real-time update. The parameter isolation-based method usually requires clear task boundary information, which is often not applicable in online remote sensing data streams with blurred or unknown task boundaries. Therefore, there is an urgent need to design an efficient, scalable, and task-agnostic online incremental learning method for remote sensing ViT large models.
[0004] ViT has an encoder-decoder architecture. The encoder is a Transformer encoder, including multiple stacked Transformer layers, and each Transformer layer includes multi-head self-attention (Multi-Head Self-Attention) and a feed-forward network (Feed-Forward Network). ViT can match multiple decoders with different functions, such as image classification, data reconstruction, etc. Its main function is to decode the output of the encoder. The Transformer decoder is a common decoder, which also includes multiple stacked Transformer layers.
[0005] InfoNCE loss (Information Noise Contrastive Estimation) is a contrastive learning loss function, which is widely used in unsupervised learning and self-supervised learning tasks. Its goal is to learn an effective representation of data by shortening the distance between positive sample pairs and increasing the distance between negative sample pairs. The core idea of the symmetric InfoNCE loss function is to symmetrize the InfoNCE loss function, that is, to consider the similarity in two directions at the same time.
[0006] LoRA (Low-Rank Adaptation) is an algorithm for efficient fine-tuning based on pre-trained models. Its core is to perform low-rank decomposition of the update amount ΔW of the model parameters into two low-rank matrices A and B. Only A and B need to be adjusted to achieve model fine-tuning, thereby significantly reducing the computational cost and storage requirements of fine-tuning while maintaining the original capabilities of the pre-trained model. Summary of the invention
[0007] The purpose of the present invention is to provide an online incremental learning method for a large remote sensing image model based on ViT, which solves the problem of catastrophic forgetting in the above-mentioned ViT large model online incremental learning scenario.
[0008] In order to achieve the above object, the technical solution adopted by the present invention is as follows: an online incremental learning method of a large model of remote sensing images based on ViT, comprising the following steps; S1, obtain remote sensing image dataset D, D={X1,X2,…,X i ,…,X N}, X i is the i-th image block labeled with category in D, , C is the number of spectral channels, P×P is the image block size, and N is the total number of image blocks; S2, constructing a dual-branch encoding and decoding network, including a space-spectrum dual-branch encoder, a hybrid connection encoder HCE and a hybrid connection decoder HCD, and the construction method includes S21~S23; S21, constructing a space-spectrum dual-branch encoder, including a space encoder and a spectrum encoder; The spatial encoder is used to i Perform random masking and spatial encoding to obtain the spatial feature e i,s , the spectral encoder is used to i Perform random masking and spectral encoding to obtain the spectral feature e i,c ; S22, constructing a hybrid connection encoder HCE; Obtain two Transformer encoders with the same structure, both of which are stacked by Z Transformer layers; Connect the input end of one Transformer encoder to the output of the spatial encoder, where the z-th Transformer layer is marked as E sz Connect the input end of the other Transformer encoder to the output of the spectral encoder, where the z-th Transformer layer is marked as E cz , 1 ≤ z ≤ Z; Modify the multi-head self-attention layer in E sz to the first hybrid connection attention module HCA1 to obtain the spatial encoding branch, and modify the multi-head self-attention layer in E cz to the second hybrid connection attention module HCA2 to obtain the spectral encoding branch. The two encoding branches form a hybrid connection encoder; The output features HCA(c, s) of HCA1 and HCA(s, c) of HCA2 are obtained according to the following formulas respectively; HCA(c, s) = Attention(Q c , K c , V c ) + Attention(Q c , K s , V s ), HCA(s, c) = Attention(Q s , K s , V s ) + Attention(Q s , K c , V c ), In the formula, Attention(∙) is the self-attention operation, Q c , K c , V c are the Q matrix, K matrix and V matrix in E cz respectively, and Q s , K s , V s are the Q matrix, K matrix and V matrix in E sz respectively; S23, construct a hybrid connection decoder HCD, including a spatial decoding branch and a spectral decoding branch. The spatial decoding branch is used to predict the masked area of the spatial encoder to obtain a spatial reconstruction map , and the spectral decoding branch is used to predict the masked area of the spectral encoder to obtain a spectral reconstruction map ; S3. Construct the loss function \(L\) of the dual-branch encoding and decoding network, and use the remote sensing image dataset \(D\) to train it to convergence by minimizing \(L\) to obtain the first basic model and freeze its parameters; S4. Delete the HCD in the first basic model, add the two outputs of HCE and send them into a classifier to obtain the second basic model, which is used to input image patches and output their corresponding classification results; S5. Online incremental learning of the second basic model; S51. Introduce a LoRA module into each of the spatial encoding branch and the spectral encoding branch of the second basic model; S52. Use the second basic model to online detect the data stream composed of image patches, update the parameters of the LoRA module based on the LoRA method with the data stream, and detect whether the data distribution of the data stream changes; S53. If it changes, incorporate the parameters of the LoRA module into the second basic model and repeat S51 - S52, otherwise do not process.
[0009] Preferably, in S23, the structure of the hybrid connection decoder HCD is as follows; Obtain two Transformer decoders with the same structure, both of which are stacked by \(Z\) Transformer layers; Connect the input end of one Transformer decoder to the output of the spatial encoding branch, where the \(z\)th Transformer layer is marked as \(D\) sz , and connect the input end of the other Transformer decoder to the output of the spectral encoding branch, where the \(z\)th Transformer layer is marked as \(D\) cz ; Modify the multi-head self-attention layer in \(D\) sz to the third hybrid connection attention module HCA3 to obtain the spatial decoding branch, modify the multi-head self-attention layer in \(D\) cz to the fourth hybrid connection attention module HCA4 to obtain the spectral decoding branch. The two decoding branches form a hybrid connection decoder, and the spatial decoder is used to predict the masked area of the spatial encoder to obtain the spatial reconstruction map , and the spectral decoder is used to predict the masked area of the spectral encoder to obtain the spectral reconstruction map ; The output features HCA(sD, sE) of HCA3 and HCA(cD, cE) of HCA4 are obtained according to the following formulas respectively; HCA(sD, sE)=Attention(Q sD , K sD , V sD )+Attention(Q sD , KsE ,V sE ), HCA(cD,cE)=Attention(Q cD ,K cD ,V cD )+Attention(Q cD ,K cE ,V cE ), wherein, Q sD , K sD , V sD are the Q matrix, K matrix, and V matrix in D sz respectively, K sE , V sE are the Q matrix and K matrix in E sz respectively; Q cD , K cD , V cD are the Q matrix, K matrix, and V matrix in D cz respectively, K cE , V cE are the Q matrix and K matrix in E cz respectively.
[0010] Preferably, in S3, the loss function L is calculated according to the following formula; , , are the spatial reconstruction loss and spectral reconstruction loss respectively, L SSC is the symmetric InfoNCE loss between the spatial feature and the spectral feature, and λ is the weight of L SSC .
[0011] Preferably, , , wherein, |M s | is the total number of pixels in the masked area M s of the spatial encoder, X s (p) and are the pixel values at position p in X i and respectively, |M c | is the total number of pixels in the masked area M c of the spectral encoder, X c (p) and are the spectral vectors at position p in X i and respectively; L SSCObtained according to the following formula; , where sim(∙,∙) is the cosine similarity calculation, exp(∙) is the exponential function, τ is the temperature hyperparameter, and e q,s , e q,c are the spatial feature and spectral feature corresponding to the q-th image patch in the remote sensing image dataset D, respectively.
[0012] Preferably, S51 is specifically to select any one layer of the Q matrix, V matrix, and K matrix in the spatial encoding branch, introduce the LoRA module, and select any one layer of the Q matrix, V matrix, and K matrix in the spectral encoding branch, and introduce the LoRA module.
[0013] Preferably, in S52, detecting whether the data distribution of the data stream changes is specifically; The data stream is divided into multiple batches in sequence, the classification loss value of each batch is calculated, and the batches are sampled with a sliding window. If the current t-th batch satisfies the following formula, it is determined that the data distribution of the t-th batch has changed; , where k is the sliding window size, and l j is the classification loss value of the j-th batch of the data stream.
[0014] Preferably, the spatial encoder includes a spatial mask layer, a first normalization layer, a two-dimensional convolutional layer, an attention layer, a first addition layer, a second normalization layer, a multi-layer perceptron, and a second addition layer connected in sequence, and the output of the spatial mask layer is skip-connected to the first addition layer, and the output of the first addition layer is skip-connected to the output of the second addition layer; The spatial mask layer is used to randomly apply a C×1×1 mask to the pixels in X i to obtain the first spatial feature , and only retain the annotations in the unmasked area; The first spatial feature is sequentially passed through the first normalization layer, the two-dimensional convolutional layer, and the attention layer to obtain the second spatial feature , The is added element-wise to to obtain the third spatial feature ; The is sequentially passed through the second normalization layer and the multi-layer perceptron to obtain the fourth spatial feature ; The is added element-wise to to obtain the spatial feature e i,s。
[0015] Preferably, the attention layer is a self-attention layer or a spatial attention layer; The spatial attention layer performs spatial attention operation according to the following formula; , In the formula, Attention1(Q, K, V) is the spatial attention operation, Softmax(∙) is the Softmax function, and Q, K, and V are the Q matrix, K matrix, and V matrix obtained by mapping the output of the two-dimensional convolutional layer; d k is the scaling factor.
[0016] Preferably, the spectral encoder includes a spectral masking layer, a first normalization layer, a linear projection layer, an attention layer, a first addition layer, a second normalization layer, a multi-layer perceptron, and a second addition layer, which are connected in sequence. The output of the spatial masking layer is skip-connected to the first addition layer, and the output of the first addition layer is skip-connected to the output of the second addition layer; The spectral masking layer is used to randomly apply a 1×P×P mask to the bands of X i to obtain the first spectral feature , and only retain the unmasked band information; The first spectral feature is sequentially passed through the first normalization layer, the linear projection layer, and the attention layer to obtain the second spectral feature , The is added element-wise to to obtain the third spectral feature ; The is sequentially passed through the second normalization layer and the multi-layer perceptron to obtain the fourth spectral feature ; The is added element-wise to to obtain the spectral feature e i,c。
[0017] Preferably, the optimization method of the LoRA module is as follows; Sa1, decompose the LoRA module into two low-rank matrices A and B; Sa2, maintain a hard buffer with a capacity of 4 to store the four image patches with the highest current classification loss values; Sa3, calculate the parameter importance Ω of the low-rank matrices A and B according to the following formula A and Ω B ; , , In the formula, x kis the k-th image block in the hard buffer, where 1 ≤ k ≤ 4, θ is the network parameter of the current second base model, and p(x k |θ) is the classification result of the current second base model for x k . ∇ is the gradient operation, ∘ is the element-wise product operation, is the element in the i-th row and j-th column of the weight matrix W of A A ; is the element in the i-th row and j-th column of Ω A ; is the element in the i-th row and j-th column of the weight matrix W of B B ; is the element in the i-th row and j-th column of Ω B ; Sa4, calculate the regularization value r and regularization term L of the current LoRA module according to the following formula LoRA ; , , where R is the set of regularization values of all current LoRA modules, and λ is the regularization coefficient; Sa5, construct the loss function L1 of the LoRA module and optimize the parameters of the LoRA module by minimizing L1; ; L cls (,∙,) is to calculate the classification loss, F(X;θ) is the predicted classification result of the current second base model for a batch of image blocks, Y is the true class label of this batch of samples, and F(X B ;θ) is the classification result of the current second base model for the samples in the hard buffer, and Y B is the true class label of the image blocks in the hard buffer.
[0018] The idea of the present invention is as follows: First, construct an air-spectrum double-branch encoder, a hybrid connection encoder HCE, and a hybrid connection decoder HCD to form a double-branch encoding and decoding network, and train this double-branch encoding and decoding network to obtain the first base model as shown in Figure 1 . Then, based on the first base model and the classifier, construct the second base model for the classification prediction of image blocks in the online data stream, and introduce two LoRA modules into the second base model. Continuously train the parameters of the LoRA modules with the online input data stream until it is monitored that the data distribution of the data stream changes. Incorporate the parameters of the LoRA modules into the network parameters of the second base model, then introduce two LoRA modules and continue learning until it is monitored again that the data distribution of the data stream changes.
[0019] Regarding the spatial-spectral dual-branch encoder: It is used to fully capture the inherent spatial context information and fine spectral discrimination features of image patches. The spatial encoder and the spectral encoder respectively extract features of the image patch in the spatial dimension and the spectral dimension, and enhance the perception ability of the spatial encoder and the spectral encoder for key features through the self-masking pre-training strategy. To achieve a more powerful feature expression ability, the attention mechanism adopted in the spatial encoder and the spectral encoder can use the multi-head attention mechanism, that is, the Q, K, and V matrices are divided into multiple independent attention heads along the feature channel dimension, so as to be able to learn different attention patterns in parallel and capture richer feature interaction information.
[0020] Regarding the Hybrid Connection Encoder HCE: It includes a spatial encoding branch and a spectral encoding branch. Both are based on the basic structure of the ViT encoder and belong to the Transformer encoder, which is stacked by Z Transformer layers. The Transformer layer itself includes a multi-head self-attention layer, and its input is the Q matrix, K matrix, and V matrix generated by the mapping of the previous stage output. Since the present invention has a dual-branch structure, the multi-head self-attention layer in the Transformer layer is modified into a hybrid connection attention module. In the present invention, the multi-head self-attention layer in the spatial encoding branch is replaced by the first hybrid connection attention module HCA1, and the multi-head self-attention layer in the spectral encoding branch is modified into the second hybrid connection attention module HCA2. According to the formula HCA(c,s)=Attention(Q c ,K c ,V c )+Attention(Q c ,K s ,V s ), it can be seen that the input of HCA1 is 5, namely Q s ,K s ,V s ,K c ,V c . The first three come from the spectral encoding branch, and the last two come from its own spatial encoding branch. The purpose is to guide the features of the spectral encoding branch to flow and aggregate towards the spatial encoding branch. Similarly, HCA2 can guide the features of the spatial encoding branch to flow and aggregate towards the spectral encoding branch. Since the two branches are symmetric, a symmetric information flow of the hybrid connection encoder HCE is constructed, realizing cross-modal information enhancement from spatial features to spectral features and improving feature representation.
[0021] Regarding the Hybrid Connection Decoder (HCD): Its overall architecture is the same as that of the Hybrid Connection Encoder (HCE), but the features processed by the hybrid connection attention module are different. The HCD is also a two-branch structure, including a spatial decoding branch and a spectral decoding branch, which are connected to the spatial encoding branch and the spectral encoding branch respectively. More specifically, the structure of the spatial decoding branch is the same as that of the spatial encoding branch, but HCA3 is used to replace HCA1 in the spatial encoding branch, and Q sD , K sD , V sD generated by fusing the output maps of the upper levels are generated according to the formula HCA(sD, sE), as well as K sE , V sE in the spatial encoding branch. Through this operation, the feature information of different levels in the spatial encoding branch of the HCE can be fully utilized to effectively guide the reconstruction process of the spatial decoding branch. In addition, the matrices introduced in the formula HCA(sD, sE) are the features of the corresponding layers of the spatial encoding branch and the spatial decoding branch. For example, the K sD , V sD of HCA3 in the fifth layer are from HCA1 in the fifth layer. This corresponding method enables the spatial decoding branch to directly utilize the feature details captured at the same abstract level of the spatial encoding branch during the reconstruction process, significantly enhancing its ability to obtain fine features. Similarly, for the spectral decoding branch, its structure is the same as that of the spectral encoding branch, but HCA4 is used to replace HCA2 in the spectral encoding branch, and feature fusion is performed according to the formula HCA(cD, cE).
[0022] Regarding incremental learning: In the learning scenario, the distribution of the data stream may change at any time, and there is no clear task boundary signal. To enable the model to automatically adapt to this change, the model adopts an adaptive model expansion strategy based on loss dynamics. The classification loss value of the model is monitored in real time when processing consecutive batches of remote sensing samples as input. Through the classification loss values within the sliding window, it is detected whether the data distribution of the data stream has changed. When a change occurs, the current LoRA module is frozen, and its weight parameters are incorporated into the second base model to avoid the continuous growth of the number of parameters caused by the infinite accumulation of LoRA modules. Then, 2 new LoRA modules are re-initialized to learn new knowledge that may appear in the subsequent data stream or adapt to the new data distribution.
[0023] Compared with the prior art, the advantages of the present invention are as follows: A new dual-branch encoding and decoding network is proposed, and a multi-level feature connection mechanism is introduced. The spatial and spectral features are extracted by using a parallelized spatial-spectral dual-branch encoder, and the hybrid connection attention (HCA) module is used to achieve in-depth interaction and enhancement of cross-modal features. In the pre-training stage, masked auto-encoding reconstruction and spatial-spectral contrastive learning are combined to enable the model to learn a robust and discriminative spatial-spectral joint feature representation. This design aims to provide a high-quality, decoupled, and fused feature basis for subsequent online incremental learning.
[0024] A new adaptive model expansion strategy is proposed: First, based on the first basic model trained by the dual-branch encoding and decoding network, the second basic model is constructed through a classifier. Second, the LoRA method is introduced in the second basic model for online incremental learning and network parameter update. This update method realizes parameter-efficient online model adjustment, that is, only a small number of LoRA parameters need to be trained while freezing the huge backbone network. Third, the present invention adopts an adaptive model expansion strategy based on the batch classification loss value. By real-time monitoring of the loss change of the online classification task, when a change is detected, the currently learned LoRA parameters are automatically frozen and their weights are merged into the backbone, and at the same time, new LoRA parameters are initialized to learn subsequent data, so as to dynamically adapt to the change of the data stream without task boundary information.
[0025] A regularization term is introduced in the LoRA module. To alleviate the catastrophic forgetting of the previously learned knowledge when learning the new LoRA module, a regularization term is introduced to punish the change of important parameters. Enable the model to continuously learn new knowledge in the continuously incoming remote sensing data stream, while effectively suppressing the forgetting of the learned knowledge, and achieve robust online incremental learning.
[0026] In summary, the present invention can achieve efficient, scalable, and task-agnostic online incremental learning. Brief Description of the Drawings
[0027] Figure 1 It is a diagram of the dual-branch encoding and decoding network architecture; Figure 2 It is a schematic diagram of constructing the second basic model and incremental learning based on the first basic model; Figure 3 It is a schematic diagram of the connection between HCA1 in the spatial encoding branch and HCA2 in the spectral-spatial encoding branch; Figure 4 It is a schematic diagram of the structure of the first hybrid connection attention module HCA1; Figure 5 It is a schematic diagram of the structure of the spatial encoder in the spatial-spectral dual-branch encoder; Figure 6Schematic diagram of the spectral encoder structure in the air-spectral dual-branch encoder; Figure 7 Comparison chart of the evaluation indexes between the present invention and the ViT image classification model; Figure 8 Comparison chart of the classification results between the present invention and the ViT image classification model. Specific implementation manner
[0028] The present invention will be further described below in conjunction with embodiments and drawings.
[0029] Embodiment 1: Refer to Figure 1 to to Figure 6 , an online incremental learning method for a large remote sensing image model based on ViT, comprising the following steps; S1, obtain a remote sensing image data set D, D = {X1, X2,..., X i ,..., X N}, X i is the i-th class-labeled image patch in D, , C is the number of spectral channels, P×P is the image patch size, and N is the total number of image patches; S2, construct a dual-branch encoding and decoding network, including an air-spectral dual-branch encoder, a hybrid connection encoder HCE, and a hybrid connection decoder HCD. The construction method includes S21~S23; S21, construct an air-spectral dual-branch encoder, including a spatial encoder and a spectral encoder; The spatial encoder is used to perform random masking and spatial encoding on X i to obtain a spatial feature e i,s , and the spectral encoder is used to perform random masking and spectral encoding on X i to obtain a spectral feature e i,c ; S22, construct a hybrid connection encoder HCE; Obtain two Transformer encoders with the same structure, both stacked by Z Transformer layers; Connect the input end of one Transformer encoder to the output of the spatial encoder, where the z-th Transformer layer is marked as E sz , and connect the input end of the other Transformer encoder to the output of the spectral encoder, where the z-th Transformer layer is marked as E cz , 1≤z≤Z; Modify the multi-head self-attention layer in E sz into a first hybrid connection attention module HCA1 to obtain a spatial encoding branch, and modify the multi-head self-attention layer in E czModify the multi-head self-attention layer into the second hybrid connection attention module HCA2 to obtain the spectral encoding branch, and the two encoding branches form a hybrid connection encoder; The output features HCA(c, s) of HCA1 and HCA(s, c) of HCA2 are obtained according to the following formulas respectively; HCA(c,s)=Attention(Q c ,K c ,V c )+Attention(Q c ,K s ,V s ), HCA(s,c)=Attention(Q s ,K s ,V s )+Attention(Q s ,K c ,V c ), In the formula, Attention(∙) is the self-attention operation, Q c ,K c ,V c are the Q matrix, K matrix and V matrix in E cz respectively, and Q s ,K s ,V s are the Q matrix, K matrix and V matrix in E sz respectively; S23. Construct a hybrid connection decoder HCD, including a spatial decoding branch and a spectral decoding branch. The spatial decoding branch is used to predict the masked area of the spatial encoder to obtain a spatial reconstruction map , and the spectral decoding branch is used to predict the masked area of the spectral encoder to obtain a spectral reconstruction map ; S3. Construct the loss function L of the dual-branch encoder-decoder network, and use the remote sensing image dataset D to train it to convergence by minimizing L to obtain the first basic model and freeze its parameters; S4. Delete HCD in the first basic model, add the two outputs of HCE and send them to a classifier to obtain the second basic model, which is used to input an image patch and output its corresponding classification result; S5. Online incremental learning of the second basic model; S51. Introduce a LoRA module into each of the spatial encoding branch and the spectral encoding branch of the second basic model; S52. Online detect the data stream composed of image patches using the second base model, update the parameters of the LoRA module based on the data stream using the LoRA method, and detect whether the data distribution of the data stream changes; S53. If it changes, incorporate the LoRA module parameters into the second base model and repeat S51 - S52; otherwise, do not process.
[0030] See Figure 3 , Figure 3 shows the hybrid connection method of the z - th layer Transformer layer E in the two spatial encoding branches sz and the z - th layer Transformer layer E in the spectral encoding branch cz . On the left is E sz . Therefore, when improving, replace the multi - head self - attention layer of the left - hand Transformer layer with HCA1, and on the right is E cz . Therefore, when improving, replace the multi - head self - attention layer of the right - hand Transformer layer with HCA2. The input increases from 3 channels to 5 channels, coming from two encoding branches respectively, so the interaction and enhancement of features are realized. Figure 3 The structures of the remaining layers are the same as those of the existing Transformer layer. See Figure 4 , taking HCA1 as an example, the input increases from 3 channels to 5 channels, and 2 self - attention layers are needed Figure 3 is the connection method of the 2 self - attention layers that form HCA1. There is one self - attention layer on the left and one on the right, and the outputs of the 2 self - attention layers are added through an addition layer.
[0031] Embodiment 2: See Figures 1 to 6 . Based on Embodiment 1, we give a method for constructing a hybrid connection decoder HCD: Obtain two Transformer decoders with the same structure, both stacked by Z Transformer layers; Connect the input end of one Transformer decoder to the output of the spatial encoding branch, where the z - th Transformer layer is marked as D sz , and connect the input end of the other Transformer decoder to the output of the spectral encoding branch, where the z - th Transformer layer is marked as D cz ; Modify the multi - head self - attention layer in D sz to the third hybrid connection attention module HCA3 to obtain the spatial decoding branch. Modify the multi - head self - attention layer in D cz to the fourth hybrid connection attention module HCA4 to obtain the spectral decoding branch. The two decoding branches form a hybrid connection decoder, and the spatial decoder is used to predict the masked area of the spatial encoder to obtain the spatial reconstruction map , the spectral decoder is used to predict the masked region of the spectral encoder to obtain a spectral reconstruction map ; The output features HCA(sD, sE) of HCA3 and HCA(cD, cE) of HCA4 are obtained according to the following formulas respectively; HCA(sD, sE) = Attention(Q sD , K sD , V sD ) + Attention(Q sD , K sE , V sE ), HCA(cD, cE) = Attention(Q cD , K cD , V cD ) + Attention(Q cD , K cE , V cE ), In the formula, Q sD , K sD , V sD are the Q matrix, K matrix, and V matrix in D sz respectively, K sE , V sE are the Q matrix and K matrix in E sz respectively; Q cD , K cD , V cD are the Q matrix, K matrix, and V matrix in D cz respectively, K cE , V cE are the Q matrix and K matrix in E cz respectively. The rest of Example 2 is the same as Example 1.
[0032] Example 3: Refer to Figures 1 to 6 , in S3, the loss function L is calculated according to the following formula; , , are the spatial reconstruction loss and spectral reconstruction loss respectively, L SSC is the symmetric InfoNCE loss between the spatial feature and the spectral feature, and λ is the weight of L SSC .
[0033] , , In the formula, |M s | is the masked region M of the spatial encoder sTotal number of pixels in the middle, X s (p) and are X respectively i and the pixel value at position p in the middle, |M c | is the masked area M of the spectral encoder c Total number of pixels in the middle, X c (p) and are X respectively i and the spectral vector at position p in the middle; L SSC is obtained according to the following formula; , where sim(∙,∙) is the cosine similarity calculation, exp(∙) is the exponential function, τ is the temperature hyperparameter, e q,s 、e q,c are the spatial feature and spectral feature corresponding to the q-th image patch in the remote sensing image dataset D respectively. The rest is the same as in Embodiment 1 and Embodiment 2.
[0034] Embodiment 4: Refer to Figures 1 to 6 , S51 is specifically, in the spatial encoding branch, select any layer of Q matrix, V matrix and K matrix, introduce the LoRA module, and in the spectral encoding branch, select any layer of Q matrix, V matrix and K matrix, introduce the LoRA module. In S52, detecting whether the data distribution of the data stream changes is specifically; Divide the data stream into multiple batches in order, calculate the classification loss value of each batch, and sample the batches with a sliding window. If the current t-th batch satisfies the following formula, it is determined that the data distribution of the t-th batch changes; , where k is the sliding window size, l j is the classification loss value of the j-th batch of the data stream. The rest is the same as in Embodiments 1 to 3.
[0035] Embodiment 5: Refer to Figures 1 to 6 , the spatial encoder includes a spatial masking layer, a first normalization layer, a two-dimensional convolutional layer, an attention layer, a first addition layer, a second normalization layer, a multi-layer perceptron, and a second addition layer connected in sequence, and the output of the spatial masking layer is skip-connected to the first addition layer, and the output of the first addition layer is skip-connected to the output of the second addition layer; The spatial masking layer is used to randomly apply a C×1×1 mask to the pixels in X i to obtain the first spatial feature , and only retain the annotations in the unmasked area; successively passes through the first normalization layer, the two-dimensional convolutional layer, and the attention layer to obtain the second spatial feature , The is obtained by element - by - element addition of the first addition layer and to obtain the third spatial feature ; The successively passes through the second normalization layer and the multi - layer perceptron to obtain the fourth spatial feature ; The is obtained by element - by - element addition of the second addition layer and to obtain the spatial feature e i,s。
[0036] The attention layer is a self - attention layer or a spatial attention layer. The spatial attention layer performs spatial attention operation according to the following formula; , where Attention1(Q, K, V) is the spatial attention operation, Softmax(∙) is the Softmax function, and Q, K, and V are the Q matrix, K matrix, and V matrix obtained by mapping the output of the two - dimensional convolutional layer; d k is the scaling factor.
[0037] The spectral encoder includes a spectral mask layer, a first normalization layer, a linear projection layer, an attention layer, a first addition layer, a second normalization layer, a multi - layer perceptron, and a second addition layer connected in sequence. The output of the spatial mask layer is skip - connected to the first addition layer, and the output of the first addition layer is skip - connected to the output of the second addition layer; The spectral mask layer is used to randomly apply a 1×P×P mask to the bands of X i to obtain the first spectral feature , and only retain the un - masked band information; Successively passes through the first normalization layer, the linear projection layer, and the attention layer to obtain the second spectral feature , The is obtained by element - by - element addition of the first addition layer and to obtain the third spectral feature ; The successively passes through the second normalization layer and the multi - layer perceptron to obtain the fourth spectral feature ; The is obtained by element - by - element addition of the second addition layer and to obtain the spectral feature e i,c . The rest is the same as in Embodiments 1 - 4.
[0038] Embodiment 6: Refer to Figures 1 to 6, the optimization method of the LoRA module is as follows; Sa1, decompose the LoRA module into two low-rank matrices A and B; Sa2, maintain a hard buffer with a capacity of 4 to store the four image patches with the highest current classification loss values; Sa3, calculate the parameter importance Ω of the low-rank matrices A and B according to the following formula A and Ω B ; , , In the formula, x k is the k-th image patch in the hard buffer, 1 ≤ k ≤ 4, θ is the network parameter of the current second base model, p(x k |θ) is the classification result of the current second base model for x k , ∇ is the gradient operation, ∘ is the element-wise product operation, is the element in the i-th row and j-th column of the weight matrix W A of A, is the element in the i-th row and j-th column of Ω A , is the element in the i-th row and j-th column of the weight matrix W B of B, is the element in the i-th row and j-th column of Ω B ; Sa4, calculate the regularization value r and regularization term L of the current LoRA module according to the following formula LoRA ; , , In the formula, R is the set of regularization values of all current LoRA modules, and λ is the regularization coefficient; Sa5, construct the loss function L1 of the LoRA module and optimize the parameters of the LoRA module by minimizing L1; ; L cls (, ∙, ) is to calculate the classification loss, F(X; θ) is the predicted classification result of the current second base model for a batch of image patches, Y is the true class label of this batch of samples, F(X B ; θ) is the classification result of the current second base model for the samples in the hard buffer, Y B is the true class label of the image patches in the hard buffer.
[0039] Example 7: See Figure 7 and Figure 8, in this embodiment, three general evaluation indicators in the field of image block classification are selected: Overall Accuracy (OA), Average Accuracy (AA), and Kappa Coefficient.
[0040] Overall Accuracy (OA), in English, represents the proportion of the total number of correctly classified samples to the total number of test samples, reflecting the overall classification performance of the model on the entire dataset. The formula is , where T TP represents the number of samples correctly classified as positive examples, T TN represents the number of samples correctly classified as negative examples, F FP represents the number of samples misclassified as positive examples, F FN represents the number of samples misclassified as negative examples.
[0041] Average Accuracy (AA), in English: represents the arithmetic mean of the classification accuracies of all classes, which can more fairly evaluate the performance of the model in the case of unbalanced class samples. Its calculation formula is: , where: n represents the total number of classes, and A n represents the accuracy of the
[0042] Kappa Coefficient: Measures the degree of consistency between the classification results of the model and the random classification results, reducing the influence of random consistency. Its calculation formula is: , where: p o is the sum of the number of correctly classified samples in each class divided by the total number of samples, that is, the overall classification accuracy, and p e is the random consistency. Assuming that the true number of samples in each class is a1, a2, ⋯, a c , and the number of samples predicted for each class is b1, b2, ⋯, b c , and the total number of samples is N, then there is .
[0043] Based on the above three evaluation indicators, the second basic model of the present invention and the ViT image classification model are respectively used for classification on the Berlin dataset, and the comparison chart of evaluation indicators is obtained as shown in Figure 7 . Among them, the second basic model of the present invention is the optimal model constructed based on Figures 1 to 6 . It can be seen from Figure 7 that the present invention is superior to the ViT image classification model in terms of overall classification accuracy OA, average accuracy AA, and Kappa coefficient.
[0044] To more intuitively express the classification results of the present invention and the ViT image classification model, a locally enlarged area of a sample is selected from the Berlin dataset, and classified using the second basic model and the ViT image classification model, resulting in a comparison diagram as shown in Figure 8 shown below. Figure 8 Among them, (a), (b), and (c) are the ground truth map, the classification result of the present invention, and the classification result of the ViT image classification model, respectively. The ground truth map represents the true situation of the land cover categories in the remote sensing image and is the benchmark for evaluating the performance of the classification algorithm. It can be seen from Figure 8 that the classification result of the present invention is closer to the ground truth map than the ViT classification result, showing better performance.
[0045] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An online incremental learning method for a large remote sensing image model based on ViT, characterized in that Including the following steps; S1. Obtain the remote sensing image dataset D, D = {X1, X2, …, X i , …, X N}, where Xi i is the i-th class-labeled image patch in D, , C is the number of spectral channels, P×P is the image patch size, and N is the total number of image patches; S2. Construct a dual-branch encoding and decoding network, including an air-spectrum dual-branch encoder, a hybrid connection encoder HCE, and a hybrid connection decoder HCD. The construction method includes S21~S23; S21. Construct an air-spectrum dual-branch encoder, including a spatial encoder and a spectral encoder; The spatial encoder is used to perform random masking and spatial encoding on X i to obtain the spatial feature e i,s , and the spectral encoder is used to perform random masking and spectral encoding on X i to obtain the spectral feature e i,c ; S22. Construct a hybrid connection encoder HCE; Obtain two Transformer encoders with the same structure, both of which are stacked by Z Transformer layers; Connect the input end of a Transformer encoder to the output of the spatial encoder, where the z-th Transformer layer is labeled as E sz Connect the input end of another Transformer encoder to the output of the spectral encoder, where the z-th Transformer layer is labeled as E cz where 1 ≤ z ≤ Z; Modify the multi-head self-attention layer in E sz to the first hybrid connection attention module HCA1 to obtain the spatial encoding branch, and modify the multi-head self-attention layer in E cz to the second hybrid connection attention module HCA2 to obtain the spectral encoding branch. The two encoding branches form a hybrid connection encoder; The output features HCA(c,s) of HCA1 and HCA(s,c) of HCA2 are obtained according to the following formulas respectively; HCA(c,s)=Attention(Q c ,K c ,V c )+Attention(Q c ,K s ,V s ), HCA(s,c)=Attention(Q s ,K s ,V s )+Attention(Q s ,K c ,V c ), Where Attention(∙) is the self-attention operation, and Q c , K c , V c are the Q matrix, K matrix, and V matrix in E cz respectively, and Q s , K s , V s are the Q matrix, K matrix, and V matrix in E sz respectively; S23. Construct a Hybrid Connection Decoder (HCD) including a spatial decoding branch and a spectral decoding branch. The spatial decoding branch is used to predict the masked region of the spatial encoder to obtain a spatial reconstruction map , and the spectral decoding branch is used to predict the masked region of the spectral encoder to obtain a spectral reconstruction map ; S3. Construct a loss function L for the dual-branch encoding and decoding network, and use the remote sensing image dataset D to train until convergence by minimizing L to obtain the first basic model and freeze its parameters; S4. Delete HCD in the first basic model, add the two outputs of HCE and send them to a classifier to obtain the second basic model, which is used to input image patches and output their corresponding classification results; S5. Online incremental learning of the second basic model; S51. Introduce a LoRA module into each of the spatial encoding branch and the spectral encoding branch of the second basic model; S52. Use the second basic model to online detect the data stream composed of image patches, update the parameters of the LoRA module based on the LoRA method using the data stream, and detect whether the data distribution of the data stream changes; S53. If it changes, incorporate the parameters of the LoRA module into the second basic model, and repeat S51~S52, otherwise do not process.
2. An online incremental learning method for a large remote sensing image model based on ViT according to claim 1, characterized in that, In S23, the structure of the hybrid connection decoder HCD is; Obtain two Transformer decoders with the same structure, both of which are stacked by Z Transformer layers; Connect the input end of a Transformer decoder to the output of the spatial encoding branch, where the z-th Transformer layer is denoted as D sz Connect the input end of another Transformer decoder to the output of the spectral encoding branch, where the z-th Transformer layer is denoted as D cz ; Modify the multi-head self-attention layer in D sz to the third hybrid connection attention module HCA3 to obtain the spatial decoding branch. Modify the multi-head self-attention layer in D cz to the fourth hybrid connection attention module HCA4 to obtain the spectral decoding branch. The two decoding branches form a hybrid connection decoder, and the spatial decoder is used to predict the masked region of the spatial encoder to obtain the spatial reconstruction map , and the spectral decoder is used to predict the masked region of the spectral encoder to obtain the spectral reconstruction map ; The output features HCA(sD,sE) of HCA3 and HCA(cD,cE) of HCA4 are obtained according to the following formulas respectively; HCA(sD,sE)=Attention(Q sD ,K sD ,V sD )+Attention(Q sD ,K sE ,V sE ), HCA(cD,cE)=Attention(Q cD ,K cD ,V cD )+Attention(Q cD ,K cE ,V cE ), Wherein, Q sD , K sD , V sD are respectively the Q matrix, K matrix, and V matrix in D sz . K sE , V sE are respectively the Q matrix and K matrix in E sz . Q cD , K cD , V cD are respectively the Q matrix, K matrix, and V matrix in D cz . K cE , V cE are respectively the Q matrix and K matrix in E cz .
3. An online incremental learning method for a large remote sensing image model based on ViT according to claim 1, characterized in that, In S3, the loss function L is calculated according to the following formula; , , are the spatial reconstruction loss and the spectral reconstruction loss respectively, and L SSC is the symmetric InfoNCE loss between the spatial feature and the spectral feature, and λ is the weight of L SSC .
4. An online incremental learning method for a large remote sensing image model based on ViT according to claim 3, characterized in that , , where, |M s | is the total number of pixels in the mask region M s of the spatial encoder, X s (p) and are the pixel values at position p in X i and respectively, and |M c | is the total number of pixels in the mask region M c of the spectral encoder, X c (p) and are the spectral vectors at position p in X i and respectively; L SSC Obtained according to the following formula; , where \(sim(\cdot,\cdot)\) is used to calculate the cosine similarity, \(exp(\cdot)\) is the exponential function, \(\tau\) is the temperature hyperparameter, \(e\) q,s , \(e\) q,c are respectively the spatial feature and spectral feature corresponding to the \(q\)-th image patch in the remote sensing image dataset \(D\).
5. An online incremental learning method for a large remote sensing image model based on ViT according to claim 1, characterized in that S51 is specifically to randomly select one layer of Q matrix, V matrix, and K matrix in the spatial encoding branch, introduce a LoRA module, and randomly select one layer of Q matrix, V matrix, and K matrix in the spectral encoding branch, and introduce a LoRA module.
6. An online incremental learning method for a large remote sensing image model based on ViT according to claim 1, characterized in that In S52, detecting whether the data distribution of the data stream changes is specifically; Divide the data stream into multiple batches in order, calculate the classification loss value of each batch, and sample the batches with a sliding window. If the current t-th batch satisfies the following formula, it is determined that the data distribution of the t-th batch has changed; , where k is the sliding window size, and l j is the classification loss value of the j-th batch of the data stream.
7. An online incremental learning method for a large remote sensing image model based on ViT according to claim 1, characterized in that, The spatial encoder includes a spatial mask layer, a first normalization layer, a two-dimensional convolutional layer, an attention layer, a first addition layer, a second normalization layer, a multi-layer perceptron, and a second addition layer connected in sequence. The output of the spatial mask layer is skip-connected to the first addition layer, and the output of the first addition layer is skip-connected to the output of the second addition layer; The spatial mask layer is used to randomly apply a mask of C×1×1 to the pixels in X i to obtain the first spatial feature , and only retain the annotations in the unmasked area; The second spatial feature is obtained through a first normalization layer, a two-dimensional convolutional layer, and an attention layer in sequence , The said is added element-wise with the output of the first addition layer and the third spatial feature is obtained through element-wise addition ; The fourth spatial feature is obtained through a second normalization layer and a multi-layer perceptron in sequence ; The is added to element-wise to obtain the spatial feature e i,s。 8. An online incremental learning method for a large remote sensing image model based on ViT according to claim 7, characterized in that, The attention layer is a self-attention layer or a spatial attention layer; The spatial attention layer performs spatial attention operation according to the following formula; , Wherein, Attention1(Q, K, V) is a spatial attention operation, Softmax(∙) is the Softmax function, and Q, K, and V are the Q matrix, K matrix, and V matrix obtained by mapping the outputs of the two-dimensional convolutional layer; d k is the scaling factor.
9. An online incremental learning method for a large remote sensing image model based on ViT according to claim 1, characterized in that, The spectral encoder includes a spectral mask layer, a first normalization layer, a linear projection layer, an attention layer, a first addition layer, a second normalization layer, a multi-layer perceptron, and a second addition layer that are connected in sequence. The output of the spatial mask layer is skip-connected to the first addition layer, and the output of the first addition layer is skip-connected to the output of the second addition layer; The spectral mask layer is used to randomly apply a 1×P×P mask to the i band of X to obtain the first spectral feature , and only retain the unmasked band information; The second spectral feature is obtained through a first normalization layer, a linear projection layer, and an attention layer in sequence , The said obtained through the first addition layer and element-wise addition to obtain the third spectral feature ; The said successively passes through a second normalization layer and a multi-layer perceptron to obtain a fourth spectral feature ; The said is added to the second addition layer and the element to obtain the spectral feature e i,c。 10. An online incremental learning method for a large remote sensing image model based on ViT according to claim 1, characterized in that, The optimization method of the LoRA module is; Sa1, decompose the LoRA module into two low-rank matrices A and B; Sa2, maintain a hard buffer with a capacity of 4 to store the four image patches with the highest current classification loss values; Sa3, calculate the parameter importance Ω of the low-rank matrices A and B according to the following formula A and Ω B ; , , where x k is the k-th image block in the hard buffer, 1 ≤ k ≤ 4, θ is the network parameter of the current second basic model, p(x k |θ) is the classification result of the current second basic model for x k , ∇ is the gradient operation, ∘ is the element-wise product operation, is the element at the i-th row and j-th column of the weight matrix W A of A, is the element at the i-th row and j-th column of Ω A , is the element at the i-th row and j-th column of the weight matrix W B of B, is the element at the i-th row and j-th column of Ω B ; Sa4, calculate the regularization value r and regularization term L of the current LoRA module according to the following formula LoRA ; , , In the formula, R is the set of regularization values of all current LoRA modules, and λ is the regularization coefficient; Sa5, construct the loss function L1 of the LoRA module, and optimize the parameters of the LoRA module by minimizing L1; ; L cls (,∙,) is used to calculate the classification loss. F(X;θ) is the predicted classification result of the current second base model for a batch of image patches, Y is the true class label of the samples in this batch, and F(X B ;θ) is the classification result of the current second base model for the samples in the hard buffer, and Y B is the true class label of the image patches in the hard buffer.
Citation Information
Patent Citations
Image classification pre-training model continuous learning method based on low-rank adaptive combination
CN117611913A
Image self-supervised learning method and system fusing contrast learning and feature mask modeling
CN119445379A
In-ground sensor systems with modular sensors and wireless connectivity components
US20200132658A1