A Multi-modal Remote Sensing Data Continuous Unknown Class Classification Method Based on Fast and Slow Learning Strategies for Inner and Outer Spaces

By adopting the fast and slow learning strategy of internal and external space in the classification of multimodal remote sensing data, the problem that the existing technology cannot continuously classify unknown classes is solved, and the continuous classification and robustness of multimodal remote sensing data is achieved.

CN119719891BActive Publication Date: 2025-06-10HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411761020.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-06-10
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

The existing multimodal remote sensing data classification methods cannot continuously classify unknown classes and cannot effectively process multimodal remote sensing data that arrives in incremental chronological order.

Method used

The multimodal remote sensing data continuous unknown class classification method based on internal and external space fast and slow learning strategies is adopted, and the continuous classification of unknown classes is achieved by building a cross-modal pixel-level spatial fusion module and a fast and slow strategy model of continuous learning.

Benefits of technology

The continuous unknown class classification of multimodal remote sensing data is realized, the ability to identify unknown classes is enhanced, and the robustness and efficiency of classification are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719891B_ABST
    Figure CN119719891B_ABST
Patent Text Reader

Abstract

A multi-modal remote sensing data continuous unknown class classification method based on the fast and slow learning strategy of internal and external spaces belongs to the technical field of remote sensing data classification. To solve the problem of realizing continuous classification of unknown classes in remote sensing data, the present invention includes obtaining multi-modal remote sensing data, performing block coding to obtain encoded hyperspectral data and encoded lidar data; constructing a cross-modal pixel-level spatial fusion module, and inputting the encoded hyperspectral data and the encoded lidar data into the cross-modal pixel-level spatial fusion module to obtain a cross-modal fusion output result; constructing a fast and slow strategy for continuous learning, including an internal space fast and slow learning strategy model and an external space fast and slow learning strategy model; inputting the cross-modal fusion output result and the incremental data of each round into the internal space fast and slow learning strategy model and the external space fast and slow learning strategy model constructed in step S4 for iterative learning, and then fusing the outputs of the models to obtain the multi-modal remote sensing data continuous unknown class classification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing data classification, and specifically relates to a multi-modal remote sensing data continuous unknown class classification method based on an internal and external space fast and slow learning strategy. Background Art

[0002] In the field of remote sensing, multi-modal data includes various data such as multi-spectral, hyperspectral, synthetic aperture radar, lidar, etc. This data fusion method can comprehensively analyze surface features from multiple dimensions such as spectral characteristics, temporal changes, and spatial distributions, thereby improving the analysis accuracy of remote sensing images. Existing methods for classifying multi-modal remote sensing data all assume that all modal data can be obtained at one time. However, in actual applications, this data arrives incrementally in chronological order, and there will be classes (unknown classes) in the multi-modal remote sensing data set that do not belong to the training set. Existing multi-modal remote sensing data classification methods can usually only classify known classes and cannot continuously classify unknown classes. Summary of the Invention

[0003] The problem to be solved by the present invention is to achieve continuous classification of unknown classes for remote sensing data, and a multi-modal remote sensing data continuous unknown class classification method based on an internal and external space fast and slow learning strategy is proposed.

[0004] To achieve the above object, the present invention is realized through the following technical solutions:

[0005] A multi-modal remote sensing data continuous unknown class classification method based on an internal and external space fast and slow learning strategy includes the following steps:

[0006] S1. Obtain multi-modal remote sensing data, including hyperspectral data and processed lidar data;

[0007] S2. Perform block coding on the hyperspectral data and the processed lidar data obtained in step S1 to obtain encoded hyperspectral data O HSI and encoded lidar data O LiDAR ;

[0008] S3. Construct a cross-modal pixel-level space fusion module, and input the encoded hyperspectral data and the encoded lidar data obtained in step S2 into the cross-modal pixel-level space fusion module to obtain a cross-modal fusion output result;

[0009] S4. Construct a fast and slow strategy for continuous learning, including an internal space fast and slow learning strategy model and an external space fast and slow learning strategy model;

[0010] S5. Input the cross-modal fusion output result obtained in step S3 and the incremental data of each round into the inner-space fast-slow learning strategy model and the outer-space fast-slow learning strategy model constructed in step S4 for iterative learning, and then fuse the outputs of the models to obtain the classification result of multi-modal remote sensing data for continuously unknown classes.

[0011] Furthermore, the hyperspectral data obtained in step S1 where H is the height, W is the width, and L is the length, is the set of real numbers, and the processed lidar data obtained The processing method of the lidar data is denoising processing and rasterization processing.

[0012] Furthermore, the specific implementation method of step S3 is to input O HSI and O LiDAR into two multi-head self-attention mechanisms respectively, and the specific implementation process is as follows:

[0013]

[0014] V HSI = [Noise V + O HSI , O LiDAR (3)

[0015] y HSI = MSA(O HSI , K HSI , V HSI ) + O HSI (4)

[0016]

[0017] V LiDAR = [Noise V + O LiDAR , O HSI (6)

[0018] y LiDAR = MSA(O LiDAR , K LiDAR , V LiDAR ) + O LiDAR (7)

[0019] y = (y HSI + y LiDAR ) / 2 (8)

[0020] where Noise K is the adaptive noise introduced by the key, Noise VThe adaptively introduced noise for the value, MLP is the multi-layer perceptron, and MSA is the multi-head self-attention mechanism. is the correlation score for the two modalities, O HSI , K HSI , V HSI respectively represent the query, key, and value of the multi-head self-attention mechanism corresponding to the hyperspectral data, O LiDAR , K LiDAR , V LiDAR respectively represent the query, key, and value of the multi-head self-attention mechanism corresponding to the lidar data, y HSI and y LiDAR are respectively the output results of the cross-modal pixel fusion module for the hyperspectral data and the lidar data, and y is the cross-modal fusion output result.

[0021] Furthermore, the specific implementation method of step S4 includes the following steps:

[0022] S4.1. Construct an inner space learning strategy model, and design a model M with a slower update speed 1 and a model M with a faster update speed 2 to achieve continuous classification of unknown classes:

[0023] First, use the discrete cosine transform on the input cross-modal fusion output result within the space to transform it into a representation in the orthogonal feature space The expression is:

[0024]

[0025] where T is the discrete cosine transform function and q is the total number of frequency components;

[0026] Secondly, introduce different regularization term weights for different frequency components to construct the regularization term L R as:

[0027]

[0028] where β i is the weight of the regularization term for the i-th frequency component, and t represents the t-th round of training;

[0029] S4.2. Construct an outer space learning strategy model:

[0030] Continuously generate M 1 and M 2 for learning, respectively realizing the preservation of old knowledge and the learning of new knowledge, and updating the parameters of the model by minimizing the loss function. The loss function L is expressed as:

[0031] L = L ML + λL R (12)

[0032] L ML = max(0, d + - d - + r) (13)

[0033] where L ML is the metric loss, λ represents the weight of the regularization term, d + represents the Euclidean distance between the sample and the positive instance, d - represents the Euclidean distance between the sample and the negative instance, r represents the margin between the positive and negative samples, and max is the maximum function;

[0034] Then the loss function L 1 of M slow is expressed as:

[0035]

[0036] Then the loss function L 2 of M fast is expressed as:

[0037]

[0038] where represents the regularization term of the M 1 model, represents the regularization term of the M 2 model;

[0039] S4.3. Define the slow learning strategy and the fast learning strategy:

[0040] In the inner space, different regularization term weights are introduced for different frequency components. The slow learning strategy applies higher weights to the low-frequency components to construct the regularization term The fast learning strategy applies lower weights to the low-frequency components to construct the regularization term

[0041] In the outer space, the learning rate set by the slow learning strategy is 1e- 6 ; The learning rate set by the fast learning strategy is 1e- 5 .

[0042] Furthermore, the specific implementation method of step S5 is: Set the data of the new category input for the t-th time as D (t) , and obtain the t-th outputs of the M 1 model and the M 2 model as and Fuse the t-th outputs of the two models to obtain the t-th multi-modal remote sensing data continuous unknown class classification result The expression of is:

[0043]

[0044] Among them, is the central feature representation of the c-th class, is the metric matrix, and the value of a represents the fusion degree of the features of the two models. a = 0 means that only the features output by M 1 are used; a = 1 means that only the features output by M 2 are used;

[0045] By successively inputting data of unknown classes, the recognition ability of the model for unknown classes is increased. The t-th input of data D( t ) of unknown classes will increase the recognizable classes from c( t-1 ) to c( t ).

[0046] Advantages of the present invention:

[0047] A multi-modal remote sensing data continuous unknown class classification method based on the internal and external space fast and slow learning strategy according to the present invention, while fusing cross-modal data, trains a model M with a slower update speed 1 and a model M with a faster update speed 2 , continuously inputs data of unknown classes to update the model parameters, and finally fuses the features output by M 1 and M 2 to achieve continuous classification of unknown classes.

[0048] A multi-modal remote sensing data continuous unknown class classification method based on the internal and external space fast and slow learning strategy according to the present invention realizes small sample incremental learning. It can not only combine data of different modalities to increase the robustness of classification, but also continuously learn to classify unknown classes. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a flowchart of a multi-modal remote sensing data continuous unknown class classification method based on the internal and external space fast and slow learning strategy according to the present invention;

[0050] Figure 2 is a block diagram of the method of the present invention;

[0051] Figure 3 is a schematic diagram of cross-modal fusion of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0052] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only a part of the embodiments of the present invention, rather than all of the specific embodiments. The components of the specific embodiments of the present invention usually described and shown in the drawings here can be arranged and designed in various different configurations, and the present invention can also have other embodiments.

[0053] Therefore, the detailed description of the specific embodiments of the present invention provided in the accompanying drawings below is not intended to limit the scope of the claimed invention, but merely represents selected specific embodiments of the present invention. Based on the specific embodiments of the present invention, all other specific embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present invention.

[0054] To further understand the content, features and effects of the present invention, the following specific embodiments are exemplified and are accompanied by the attached Figure 1 - Attached Figure 3 The details are as follows:

[0055] A multi-modal remote sensing data continuous unknown class classification method based on an internal and external space fast and slow learning strategy, comprising the following steps:

[0056] S1. Obtain multi-modal remote sensing data, including hyperspectral data and processed lidar data;

[0057] Furthermore, the hyperspectral data obtained in step S1 where H is the height, W is the width, L is the length, is the set of real numbers, and the processed lidar data obtained The processing method of the lidar data is denoising processing and rasterization processing;

[0058] S2. Perform block coding on the hyperspectral data and the processed lidar data obtained in step S1 to obtain the encoded hyperspectral data O HSI and the encoded lidar data O LiDAR ;

[0059] S3. Construct a cross-modal pixel-level spatial fusion module, and input the encoded hyperspectral data and the encoded lidar data obtained in step S2 into the cross-modal pixel-level spatial fusion module to obtain a cross-modal fusion output result;

[0060] Furthermore, the specific implementation method of step S3 is to use O HSI and O LiDARThey are respectively input into two multi-head self-attention mechanisms, and the specific implementation process is as follows:

[0061]

[0062] V HSI = [Noise V + O HSI , O LiDAR (3)

[0063] y HSI = MSA(O HSI , K HSI , V HSI ) + O HSI (4)

[0064]

[0065] V LiDAR = [Noise V + O LiDAR , O HSI (6)

[0066] y LiDAR = MSA(O LiDAR , K LiDAR , V LiDAR ) + O LiDAR (7)

[0067] y = (y HSI + y LiDAR ) / 2 (8)

[0068] Among them, Noise K is the adaptive noise introduced for the key, Noise V is the adaptive noise introduced for the value, MLP is the multi-layer perceptron, MSA is the multi-head self-attention mechanism, is the correlation score of the two modalities, O HSI , K HSI , V HSI respectively represent the query, key, and value of the multi-head self-attention mechanism corresponding to the hyperspectral data, O LiDAR , K LiDAR , V LiDAR respectively represent the query, key, and value of the multi-head self-attention mechanism corresponding to the lidar data, y HSI and y LiDAR are respectively the output results of the cross-modal pixel fusion module for the hyperspectral data and the lidar data, and y is the cross-modal fusion output result.

[0069] S4. Construct fast and slow strategies for continuous learning, including an inner-space fast and slow learning strategy model and an outer-space fast and slow learning strategy model;

[0070] Further, the specific implementation method of step S4 includes the following steps:

[0071] S4.1. Construct an inner space learning strategy model, and design a model M with a slower update speed 1 and a model M with a faster update speed 2 to achieve continuous classification of unknown classes:

[0072] First, use the discrete cosine transform on the cross-modal fusion output result input inside the space to transform it into a representation in the orthogonal feature space The expression is:

[0073]

[0074] where T is the discrete cosine transform function and q is the total number of frequency components;

[0075] Second, introduce different regularization term weights for different frequency components to construct the regularization term L R which is:

[0076]

[0077] where β i is the weight of the regularization term for the i-th frequency component, and t represents the t-th round of training;

[0078] S4.2. Construct an outer space learning strategy model:

[0079] Continuously generate M 1 and M 2 for learning, respectively realizing the preservation of old knowledge and the learning of new knowledge, and updating the parameters of the model by minimizing the loss function. The loss function L is expressed as:

[0080] L = L ML + λL R (12)

[0081] L ML = max(0, d + - d - + r) (13)

[0082] where L ML is the metric loss, λ represents the weight of the regularization term, d + represents the Euclidean distance between the sample and the positive instance, d - represents the Euclidean distance between the sample and the negative instance, r represents the margin between the positive and negative samples, and max is the maximum value function;

[0083] Then the loss function L of M 1 isslow Expressed as:

[0084]

[0085] Then M 2 's loss function L fast Expressed as:

[0086]

[0087] Wherein, Represents the regularization term of the M 1 model, Represents the regularization term of the M 2 model;

[0088] S4.3. Define slow learning strategy and fast learning strategy:

[0089] In the inner space, different regularization term weights are introduced for different frequency components. The slow learning strategy applies higher weights to low-frequency components to construct the regularization term The fast learning strategy applies lower weights to low-frequency components to construct the regularization term

[0090] In the outer space, the learning rate set by the slow learning strategy is 1e- 6 ; The learning rate set by the fast learning strategy is 1e- 5 ;

[0091] In the outer space, the slow learning strategy sets a smaller learning rate and updates the model parameters more slowly; the fast learning strategy sets a larger learning rate and updates the model parameters more slowly.

[0092] S5. Input the cross-modal fusion output result obtained in step S3 and the incremental data of each round into the inner space fast and slow learning strategy model and the outer space fast and slow learning strategy model constructed in step S4 for iterative learning, and then fuse the outputs of the models to obtain the multi-modal remote sensing data continuous unknown class classification result.

[0093] Furthermore, the specific implementation method of step S5 is: Set the data of the new category input at the t-th time as D (t) , and obtain the t-th outputs of the M 1 model and the M 2 model as and Fuse the t-th outputs of the two models to obtain the t-th multi-modal remote sensing data continuous unknown class classification result The expression of is:

[0094]

[0095] Wherein, It is the central feature representation of the c-th category, is the metric matrix. The value of a represents the fusion degree of the features of two models. a = 0 means only the features output by M 1 are used; a = 1 means only the features output by M 2 are used;

[0096] By successively inputting data of the unknown category, the recognition ability of the model for the unknown category is increased. When the data D of the unknown category is input for the t-th time (t) the recognizable categories are increased from c (t-1) to c (t) .

[0097] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0098] Although the present application has been described above with reference to specific embodiments, various improvements can be made to it and components thereof can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application can be combined with each other in any way, and the cases of these combinations are not exhaustively described in this specification only for the sake of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A method for classifying persistent unknown classes of multimodal remote sensing data based on fast and slow learning strategies in internal and external space, characterized in that: The steps include: S1. Acquire multimodal remote sensing data, including hyperspectral data and processed lidar data; S2. Block-encode the hyperspectral data obtained in step S1 and the processed lidar data to obtain the encoded hyperspectral data. HSI and the encoded lidar data O LiDAR ; S3. construct a cross-modal pixel-level spatial fusion module, input the encoded hyperspectral data and the encoded lidar data obtained in step S2 into the cross-modal pixel-level spatial fusion module, and obtain a cross-modal fusion output result; The specific implementation method of step S3 is to HSI and O LiDAR Input into two multi-head self-attention mechanisms respectively. The specific implementation process is as follows: V HSI =[Noise V +O HSI ,O LiDAR ] (3) y HSI =MSA(O HSI ,K HSI ,V HSI )+O HSI (4) V LiDAR =[Noise V +O LiDAR ,O HSI ] (6) y LiDAR =MSA(O LiDAR ,K LiDAR ,V LiDAR )+O LiDAR (7) and=(and HSI +y LiDAR ) / 2 (8) Among them, Noise K Adaptive noise introduced for the key, Noise V is the adaptive noise introduced by the value, MLP is a multi-layer perceptron, MSA is a multi-head self-attention mechanism, is the correlation score of the two modes, O HSI , K HSI 、V HSI Respectively represent the query, key, and value of the multi-head self-attention mechanism corresponding to the hyperspectral data. LiDAR , K LiDAR 、V LiDAR They represent the query, key, and value of the multi-head self-attention mechanism corresponding to the lidar data, y HSI and LiDAR are the output results of the cross-modal pixel fusion module of hyperspectral data and lidar data, respectively, and y is the cross-modal fusion output result; S4. Construct a fast and slow strategy for continuous learning, including an inner space fast and slow learning strategy model and an outer space fast and slow learning strategy model; Define slow learning strategy and fast learning strategy: In the inner space, different regularization term weights are introduced for different frequency components. The slow learning strategy imposes higher weights on low-frequency components to construct regularization terms. The fast learning strategy imposes lower weights on low-frequency components and constructs regularization terms. In the outer space, the learning rate set by the slow learning strategy is 1e -6 ; The learning rate set by the fast learning strategy is 1e -5 ; S5. Input the cross-modal fusion output results obtained in step S3 and the incremental data of each round into the inner space fast and slow learning strategy model and the outer space fast and slow learning strategy model constructed in step S4 for iterative learning, and then fuse the outputs of the models to obtain the continuous unknown class classification results of multimodal remote sensing data.

2. According to claim 1, a multimodal remote sensing data continuous unknown class classification method based on internal and external space fast and slow learning strategy is characterized in that: Hyperspectral data acquired in step S1 Where H is the height, W is the width, and L is the length. is a real number set, the processed lidar data obtained The processing methods of lidar data are denoising and rasterization.

3. According to claim 2, a multimodal remote sensing data continuous unknown class classification method based on internal and external space fast and slow learning strategy is characterized in that: The specific implementation method of step S4 includes the following steps: S4.

1. Construct an inner space learning strategy model, design a model M1 with a slower update speed and a model M2 with a faster update speed, and realize continuous classification of unknown classes: First, the input cross-modal fusion output is transformed into an orthogonal feature space using discrete cosine transform. The expression is: Where T is the discrete cosine transform function, q is the total number of frequency components; Secondly, different regularization term weights are introduced for different frequency components to construct the regularization term L R for: Among them, β i is the weight of the regularization term of the i-th frequency component, and t represents the t-th round of training; S4.

2. Constructing an outer space learning strategy model: Continuously generate M1 and M2 for learning, respectively realizing the preservation of old knowledge and the learning of new knowledge, and update the parameters of the model by minimizing the loss function. The loss function L is expressed as: L=L ML +λL R (12) L ML =max(0,d + -d - +r) (13) Among them, L ML is the measure of loss, λ represents the weight of the regularization term, and d + represents the Euclidean distance between the sample and the positive instance, d - represents the Euclidean distance between the sample and the negative instance, r represents the margin between the positive and negative samples, and max is the maximum value function; Then the loss function L of M1 is slow It is expressed as: Then the loss function L of M2 is fast It is expressed as: in, represents the regularization term of the M1 model, Represents the regularization term of the M2 model.

4. According to claim 3, a multimodal remote sensing data continuous unknown class classification method based on internal and external space fast and slow learning strategy is characterized in that: The specific implementation method of step S5 is: set the data of the new category input for the tth time to D (t) , and the t-th outputs of the M1 model and the M2 model are and The t-th output of the two models is fused to obtain the t-th multimodal remote sensing data continuous unknown class classification result The expression is: in, is the central feature representation of the cth class, is a metric matrix. The value of a indicates the degree of fusion of the features of the two models. a=0 means that only the features output by M1 are used; a=1 means that only the features output by M2 are used. By inputting unknown class data one by one, the model's ability to recognize unknown classes is increased. The data D of the unknown class is input for the tth time. (t) Change the identifiable categories from c (t-1) Increase to c (t) .

Citation Information

Patent Citations

  • Remote sensing image classification method and device based on multi-modal attention fusion technology

    CN116740422A

  • Hyperspectral and laser radar classification method for loop generation learning based on modal attention

    CN117893827A