A unified framework incremental learning remote sensing scene classification method and system
By constructing a unified framework for incremental learning remote sensing scene classification methods, the adaptability problem of remote sensing image classification models between different data sources is solved, the universality across data sources and the ability to quickly adapt to new categories are achieved, and the classification accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202411469952.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-21
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-10-21
AI Technical Summary
Most existing remote sensing image classification models are limited to training on specific datasets and cannot achieve universality across different data sources. They also lack a unified model framework for compatibility and comparison, resulting in the need for retraining when faced with new remote sensing datasets and an inability to adapt to new categories.
A unified framework incremental learning remote sensing scene classification method is constructed. Through the combination of a unified database, backbone network and classification head network, combined with the supervised training paradigm of linear programming incremental learning, it supports multiple model embeddings and various modal data inputs, achieves model accuracy, versatility and easy verification, and performs continuous fully supervised training on large-scale heterogeneous datasets.
It improves the versatility and generalization ability of the model, reduces computing resources and time consumption, avoids catastrophic forgetting problems, and enhances the model's rapid learning ability and classification accuracy when facing new data sets.
Smart Images

Figure CN119399525B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a unified framework incremental learning remote sensing scene classification method and system. Background Art
[0002] Current research on remote sensing image classification models is mostly limited to data training of specific datasets, which consumes a lot of time and resources and fails to achieve universality across different data sources. When faced with new remote sensing datasets, these deep models must be trained to adapt to new classes, so they cannot truly adapt to new classes; at the same time, these methods are developed and evaluated individually, and there is no universal model framework that can be used to uniformly carry and compare remote sensing classification algorithms.
[0003] Jilin University disclosed a remote sensing image classification method based on hyperspectral modality in its patent application document "A Hyperspectral Remote Sensing Image Classification Method and Related Equipment" (patent application number: 2023113853497). This method first obtains the hyperspectral remote sensing image to be classified, performs dimensionality reduction processing on the hyperspectral remote sensing image to be classified through principal component analysis to obtain the reduced-dimensional image data; secondly, the reduced-dimensional image data is input into the attention and multi-scale dense network to extract the spatial-spectral features of the image to obtain the spatial-spectral feature map of the image; then the obtained spatial-spectral feature map is input into the fully connected layer, the three-dimensional feature map output by the convolution is converted into a one-dimensional feature vector, and the feature space calculated by the previous layer is mapped to the sample space; finally, the extracted one-dimensional feature vector is optimized using the Adam optimizer, and the optimized one-dimensional feature vector is classified using the Sofmax classifier to obtain the classification result; this method is only applicable to model training for specific data sets and cannot be generalized to other data sets.
[0004] Beijing Institute of Technology has disclosed a remote sensing image classification method based on hyperspectral modality in its patent application "A hyperspectral image classification method based on full-spectral correlation learning network" (patent application number: 2023110735747). This method first constructs a three-dimensional data cube of each ground object target and its spectral grouping based on the attribute characteristics of rich spatial spectrum (space spectrum) information, strong spectral correlation and multimodal features in hyperspectral remote sensing images; secondly, the three-dimensional data spectral grouping is used as input data, and the group convolution and space spectrum convolution long short-term memory networks are combined to design a full-spectral correlation adaptive learning module, which dynamically Learn the spectral correlation of hyperspectral images; then use the full-spectral correlation adaptive learning module as the basic structural unit to build a backbone feature extraction network to extract deep semantic features that enhance spectral information and retain the intrinsic geometric structure; then use the three-dimensional data cube as input data to build an asymmetric fusion module based on the gated structure to align and integrate deep semantic features and traditional manual features, while suppressing noise and abnormal data interference; finally, integrate the above two modules to build a full-spectral correlation learning network to realize hyperspectral image classification; this method is trained and evaluated separately for a specific model, and the number of models that can be adapted for embedding is small.
[0005] It can be seen that the current research on remote sensing image classification models is mostly limited to data training of specific data sets, which consumes a lot of time and resources and fails to achieve universality across different data sources. Faced with new remote sensing data sets, these deep models must be trained to adapt to new classes and cannot truly adapt to new classes; secondly, the existing neural network algorithms are all developed and evaluated separately for specific scenarios, and there is no universal model framework that can be used to unify compatible embedding and compare remote sensing classification algorithms. Summary of the Invention
[0006] The purpose of the present invention is to provide a unified framework incremental learning remote sensing scene classification method and system to overcome the problems existing in the prior art. The present invention can construct a unified remote sensing image classification model framework paradigm. The framework supports the embedding of multiple models, making the model accurate, universal and easy to verify, and supports various types of modal data as model input, which significantly enhances the versatility of the framework; at the same time, a supervised learning training paradigm based on linear programming incremental learning is designed, which enables the model to perform continuous fully supervised training on large-scale heterogeneous data sets, thereby enhancing the generalization of the model.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] A unified framework incremental learning remote sensing scene classification method includes the following steps:
[0009] (1) Unified database construction: Unified database construction: Build a unified database containing several image datasets, where each image dataset consists of a training set, a validation set, and a test set, and then perform data enhancement on the unified database;
[0010] (2) Unified backbone network and classification head network model construction: The backbone network and classification head network are combined into a unified model;
[0011] (3) Unified hyperparameter construction: Construct unified optimizer, learning rate and loss function hyperparameters;
[0012] (4) Network training using the linear programming incremental learning supervised training paradigm: The training sets created in (1) are sequentially used as training data input into the unified backbone network and classification head network model constructed in (2). The hyperparameters constructed in (3) are applied to the linear programming incremental learning supervised training paradigm to train the unified backbone network and classification head network model. During the training process, the validation set created in (1) is used to calculate the classification performance of the unified backbone network and classification head network model in real time. Then, the weights of the unified backbone network and classification head network model are back-propagated using the loss function, optimizer, and learning rate hyperparameters to obtain the trained unified backbone network and classification head network model weights.
[0013] (5) Output of classification prediction results: Call the test_model_builder function to load the unified backbone network and classification head network model, then use the weights obtained in (4) as the parameters of the unified backbone network and classification head network model obtained in (2), input the test set constructed in (1) into the unified backbone network and classification head network model, and output the classification prediction results;
[0014] (6) Unified model verification: The classification prediction results output by (5) are used for performance evaluation to verify the performance of the unified backbone network and classification head network model in terms of classification accuracy and generalization ability;
[0015] Furthermore, (1) is specifically as follows: Indian Pines, Pavia University, PaviaCentre, Salinas, and Houston 2013 datasets are used for training, the datasets are randomly divided into training set and validation set in a ratio of 9:1, and then data augmentation is performed on the validation set;
[0016] Furthermore, the data enhancement methods include: horizontal flipping, normalization, size adjustment, color dithering, color blurring, image scale enhancement, image cropping enhancement or image translation enhancement;
[0017] Furthermore, (2) is specifically:
[0018] (2-1) Segmentation and embedding: Each image in the dataset is segmented into non-overlapping patches, and then the non-overlapping patches are linearly embedded into high-dimensional space vectors through the PatchEmbed linear layer;
[0019] (2-2) Position code addition: Add the position code to the high-dimensional space vector of each tile;
[0020] (2-3) Input Dropout layer: Input the high-dimensional space vector with position encoding into the Dropout layer and output the result vector A;
[0021] (2-4) Input the linear layer of the neural network classification model: Input the result vector A into the linear layer of the neural network classification model to obtain the result vector D;
[0022] (2-5) Merge: Downsample the result vector D and merge the downsampled data as the features extracted by the backbone network;
[0023] (2-6) Normalization and pooling: Normalize and pool the features extracted by the backbone network in sequence, output the result vector, and then input the result vector into the classification head network to complete the construction of a unified backbone network and classification head network model;
[0024] Furthermore, (2-4) is specifically:
[0025] (2-4-1) Construct a block-level multi-head self-attention mechanism: divide the result vector A into several non-overlapping image blocks, and calculate the attention feature information within each non-overlapping image block and between adjacent image blocks;
[0026] (2-4-2) Input the attention feature information into the multi-layer perceptron and output the result vector B;
[0027] (2-4-3) Normalize the result vector B and output the result vector C;
[0028] (2-4-4) Add result vector A to result vector C and output result vector D;
[0029] Furthermore, the constructing optimizer in (3) is to update the parameters according to the gradient of the loss function, and the optimizer is SGD, AdamW or Adam; the constructing learning rate includes polynomial decay, cosine annealing or step decay, polynomial decay adjusts the learning rate according to the total number of training iterations or cycles, cosine annealing smoothly adjusts the learning rate through the cosine function, and step decay decreases the learning rate at a preset number of iterations or training cycles;
[0030] Furthermore, the loss function constructed in (3) is specifically:
[0031] When the loss function type is the cross entropy loss function, it returns a standard cross entropy loss function;
[0032] When the loss function type is not the cross entropy loss function, the loss function is dynamically imported and the specified loss function class is instantiated. The loss function is constructed by setting parameters.
[0033] Furthermore, the above (4) is specifically:
[0034] (4-1) The training set constructed in (1) is used as training data in turn Input (2) to construct the unified backbone network and extract the feature vector list ,Will Input the classification head network constructed by (2) to obtain the weights of the old linear programming incremental learning classifier ;
[0035] (4-2) The training set constructed in (1) Input (2) to construct the unified backbone network and extract the feature vector list ,Will Input the classification head network constructed by (2) to obtain the weights of the new linear programming incremental learning classifier ;
[0036] (4-3) Use and (2) The unified backbone network and Perform adjustments, during which the classification performance of the unified backbone network and classification head network model is calculated in real time using the validation set created in (1), and then the weights of the model are backpropagated using the loss function, optimizer, and learning rate hyperparameters;
[0037] (4-4) Updated to .
[0038] A classification method for remote sensing scenes based on a unified framework incremental learning in a hyperspectral modal data environment, based on the above-mentioned unified framework incremental learning remote sensing scene classification method.
[0039] A unified framework incremental learning remote sensing scene classification system is based on the above-mentioned unified framework incremental learning remote sensing scene classification method, and is characterized by including:
[0040] A unified dataset construction module is used to build a unified database containing several image datasets and perform data enhancement on the unified database;
[0041] Unified network model construction module, used to adjust the model parameters in the experimental configuration file, realize the selection of different types of backbone networks, and combine the backbone network and classification head network into a unified model;
[0042] A unified hyperparameter building block for updating model parameters based on the gradient of the loss function to minimize the loss during training;
[0043] The network training module uses a supervised training paradigm based on linear programming incremental learning to train heterogeneous supervised datasets. It introduces a linear programming incremental learning classifier between each stage of incremental learning, adapting the deep learning classification model to new datasets while addressing the model's catastrophic forgetting problem.
[0044] The classification prediction result output module is used to input the test set data into the unified backbone network and classification head network model loaded with weights. The unified backbone network and classification head network model perform classification prediction on each image in the test set.
[0045] The unified model validation module is used to predict category labels and evaluate performance on the validation set to verify the performance of the unified model in terms of classification accuracy and generalization ability.
[0046] The above technical solution has the following advantages or beneficial effects:
[0047] The present invention provides a unified framework incremental learning remote sensing scene classification method. By constructing a unified image database (including training set, validation set and test set) and a unified backbone network and classification head network model, this method provides a standardized and reusable solution, greatly improving the efficiency of experiments and deployment, while also enabling the model to flexibly adapt to different remote sensing scene classification tasks; data enhancement is performed on the unified database, and by introducing more variants and complexity into the validation data, the model's generalization ability can be effectively improved when dealing with unseen or highly variable remote sensing scenes; through unified hyperparameter configuration (including optimizer, learning rate and loss function), the model training process can be systematically adjusted and optimized to maximize the model's classification accuracy and training efficiency, helping to find the optimal configuration that best suits the current task and reduce human intervention. and trial-and-error costs; the adopted linear programming incremental learning supervised training paradigm allows the model to gradually learn new data sets while retaining previously learned knowledge. It is suitable for processing large-scale or continuously updated remote sensing data, and there is no need to train the entire model from scratch, thus saving computing resources and time, and enabling existing deep learning classification models to adapt to new data sets, thereby achieving rapid learning of new classification capabilities, and also helping to avoid the catastrophic forgetting problem, that is, the model forgets old tasks when learning new tasks; the unified model validation step provides quantitative indicators for the classification accuracy and generalization ability of the model by predicting the category labels of the validation set and performing performance evaluation, which helps researchers or developers clearly understand the performance of the model and make further adjustments and optimizations as needed.
[0048] Furthermore, by using different data sets for training, the model can learn a wider range of richer scene features. The fusion of multi-source data helps the model to have stronger adaptability when processing actual remote sensing scenes. Different data sets represent different geographical conditions, lighting conditions and ground cover types. The model can learn these changing factors during the training process, so that it can more accurately classify new and unseen remote sensing images. The data set is divided into training set and validation set in a ratio of 9:1, which ensures that the model learns on a sufficient number of samples. At the same time, there are enough samples to verify the performance of the model, which helps to adjust the model parameters in time during the training process and improve the final classification accuracy. Data augmentation of the validation set can simulate various deformations and interferences that may be encountered in actual applications, so that the model can still maintain stable performance when facing these situations.
[0049] Furthermore, by introducing multiple data enhancement methods, the model can be exposed to more diverse data variants during training, thereby learning more robust feature representations, which helps the model make more accurate classification judgments when facing complex and changeable remote sensing scenes in practical applications, and improves the model's generalization ability; through data enhancement, the model can learn more comprehensive and rich feature information, which helps to improve the model's accuracy in classification tasks, especially for some subtle category differences or complex scene changes. Data enhancement can enhance the model's ability to capture these details, thereby improving classification accuracy.
[0050] Furthermore, through segmentation and embedding, remote sensing image data is effectively converted into high-dimensional space vectors, which not only reduces the complexity of the data but also improves the efficiency of data processing. Non-overlapping tile segmentation and PatchEmbed linear embedding ensure that each tile can be fully utilized while retaining the key information of the image; by adding position encoding, position encoding is added to the high-dimensional space vector, which solves the problem of position information loss that may be encountered by convolutional neural networks when processing sequence data, and helps the model better understand the relationship between different areas in the image, thereby improving the accuracy of classification; inputting the Dropout layer effectively prevents the overfitting of the model and enhances the generalization ability and robustness of the model; the result vector A is input into the linear layer of the neural network classification model and after merging, it is normalized and pooled, making the model structure more flexible and scalable.
[0051] Furthermore, by constructing a multi-head self-attention mechanism at the image block level, it is possible to capture the complex relationships within each non-overlapping image block and between adjacent image blocks, which helps the model to understand the image content more deeply and extract richer and more accurate feature information; inputting the attention feature information into the multi-layer perceptron can further refine and fuse these features, so that the output result vector B not only contains the information of the original image block, but also incorporates the correlation information between adjacent blocks, which helps to improve the accuracy of classification; normalizing the result vector B can stabilize the training process of the model and prevent the gradient disappearance or explosion problem caused by excessively large or small eigenvalues, which helps the model to better learn and adapt to new data; adding the result vector A and the result vector C not only retains the feature information of the original image block, but also incorporates the enhanced features processed by the self-attention mechanism and the multi-layer perceptron, so that the output result vector D has both the integrity of the original information and stronger expressive power.
[0052] Furthermore, by offering a choice of SGD, AdamW, or Adam optimizers, users can select the most appropriate optimization algorithm based on specific tasks and data characteristics. Learning rate configuration strategies such as polynomial decay, cosine annealing, and step decay enable dynamic adjustment of the learning rate as training progresses. This helps use a larger learning rate in the early stages of training to quickly reduce the loss, while gradually reducing the learning rate in the later stages to fine-tune model parameters, thereby avoiding overfitting and improving model generalization. Support for dynamic generation and customization of loss functions allows users to select or design the most appropriate loss function based on specific task requirements. Cross-entropy loss, a common loss function in classification tasks, directly measures the difference between the model's predicted probability distribution and the true label. By properly configuring the optimizer, learning rate, and loss function, the model's training stability and convergence speed can be significantly improved. By continuously optimizing model parameters to minimize the loss function, the model can gradually learn the inherent patterns and feature representations of the data, improving not only the model's performance on the training set but also, more importantly, enhancing the model's generalization ability on unseen test sets, enabling the model to better cope with complex scenarios in real applications.
[0053] Furthermore, when the cross-entropy loss function is used, the standard cross-entropy loss function is directly returned, which simplifies the operation; when other types of loss functions are needed, the loss function can be customized by dynamically importing and instantiating the specified loss function class and setting the corresponding parameters, so that the method of the present invention has higher flexibility and scalability when processing different types of remote sensing scene classification tasks.
[0054] Furthermore, by sequentially feeding the training set into a unified backbone network as training data, the new category data can be efficiently processed, reducing the amount of data required for each training session and enabling the model to gradually learn and adapt to the features of the new category, thereby improving training efficiency and classification accuracy. When old category data is accessible, the feature extractor and classifier are jointly fine-tuned using the current and new category data. This helps the model better adapt to new category data while maintaining its ability to recognize old categories, thereby improving the model's stability and generalization ability. When old category data is inaccessible, a linear programming method is used to update the classifier weight matrix, avoiding the need to retrain the entire model. This not only saves computing resources but also shortens the model update time, making it more suitable for practical application scenarios. Through incremental learning, the model can be gradually updated and optimized, maintaining high performance in the classification task of constantly changing remote sensing scenes. Compared with other incremental learning methods, the linear programming incremental learning classifier reduces the forgetting of old category information by directly adding the feature vectors of the new category to the weight matrix. The weights of the old categories remain unchanged during the update process, and the model can continuously utilize this old category information to make classification decisions.
[0055] The present invention also provides a classification method for remote sensing scenes based on a unified framework incremental learning in a hyperspectral modal data environment. Hyperspectral data has extremely high spectral resolution and contains rich ground object information. Through a unified framework and incremental learning strategy, hyperspectral data can be efficiently processed and analyzed, and feature information of great value for remote sensing scene classification can be extracted. Due to the complexity and diversity of hyperspectral data, traditional classification methods often find it difficult to achieve ideal classification results. However, this method, by constructing a flexible unified framework and combining it with an incremental learning strategy, can continuously adapt to new data and scenarios, thereby improving the generalization ability of the model and enabling it to more accurately classify different remote sensing scenes.
[0056] The present invention also provides a unified framework incremental learning remote sensing scene classification system, which provides high configurability and flexibility through a unified data set construction module, a unified network model construction module and a unified hyperparameter configuration module. Users can easily adjust the data set, select different types of backbone networks, configure model parameters and hyperparameters according to specific task requirements to adapt to different remote sensing scene classification tasks; the linear programming incremental learning supervised training module can effectively train heterogeneous supervised data sets, and introduce the linear programming incremental learning classifier between each stage of incremental learning, which not only enables the deep learning classification model to quickly adapt to the new data set, but also effectively alleviates the catastrophic forgetting problem that the model may encounter in the incremental learning process, thereby maintaining The stability and continuous performance of the model are improved; the classification prediction result output module uses a unified backbone network and classification head network model loaded with weights to make accurate classification predictions for each image in the test set; the unified model verification module can comprehensively verify the performance of the unified model in classification accuracy and generalization ability by predicting category labels and evaluating performance on the verification set. The comprehensive verification and evaluation mechanism helps users to understand the performance characteristics of the model more accurately, and then carry out targeted optimization and improvement; the system adopts a unified framework and modular design, which has strong adaptability and scalability. With the continuous development of remote sensing technology and the emergence of new task requirements, the system can easily add new modules or expand the functions of existing modules to adapt to new application scenarios and task requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 This is a flowchart of a unified framework incremental learning remote sensing scene classification method and system of the present invention. DETAILED DESCRIPTION
[0058] The present invention will be further described in detail below with reference to specific embodiments, which are intended to explain the present invention rather than to limit it.
[0059] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0060] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0061] Example 1:
[0062] like Figure 1 As shown, the present invention provides a unified framework incremental learning remote sensing scene classification method, comprising the following steps:
[0063] (1) Unified database construction: A unified database containing several image datasets is constructed, where each image dataset consists of a training set, a validation set, and a test set. Data augmentation is then performed on the unified database. Specifically, the Indian Pines, Pavia University, Pavia Centre, Salinas, and Houston 2013 datasets are used for training. The datasets are randomly divided into training and validation sets in a ratio of 9:1, and the validation set is then augmented.
[0064] Preferably, the data enhancement methods include: horizontal flipping, normalization, size adjustment, color dithering, color blurring, image scale enhancement, image cropping enhancement or image translation enhancement;
[0065] (2) Unified backbone network and classification head network model construction: The backbone network and classification head network are combined into a unified model, specifically:
[0066] (2-1) Segmentation and embedding: Each image in the dataset is segmented into non-overlapping patches, and then the non-overlapping patches are linearly embedded into high-dimensional space vectors through the PatchEmbed linear layer;
[0067] (2-2) Position code addition: Add the position code to the high-dimensional space vector of each tile;
[0068] (2-3) Input Dropout layer: Input the high-dimensional space vector with position encoding into the Dropout layer and output the result vector A;
[0069] (2-4) Input the linear layer of the neural network classification model: Input the result vector A into the linear layer of the neural network classification model to obtain the result vector D;
[0070] (2-4-1) Construct a block-level multi-head self-attention mechanism: divide the result vector A into several non-overlapping image blocks, and calculate the attention feature information within each non-overlapping image block and between adjacent image blocks;
[0071] (2-4-2) Input the attention feature information into the multi-layer perceptron and output the result vector B;
[0072] (2-4-3) Normalize the result vector B and output the result vector C;
[0073] (2-4-4) Add result vector A to result vector C and output result vector D;
[0074] (2-5) Merge: Downsample the result vector D and merge the downsampled data as the features extracted by the backbone network;
[0075] (2-6) Normalization and pooling: Normalize and pool the features extracted by the backbone network in sequence, output the result vector, and then input the result vector into the classification head network to complete the construction of a unified backbone network and classification head network model;
[0076] (3) Unified hyperparameter configuration: building unified optimizer, learning rate, and loss function hyperparameters;
[0077] Preferably, the optimizer is constructed to update the model parameters according to the gradient of the loss function, and the optimizer is SGD, AdamW or Adam;
[0078] Preferably, constructing the learning rate includes polynomial decay, cosine annealing or step decay, polynomial decay adjusts the learning rate according to the total number of training iterations or cycles, cosine annealing smoothly adjusts the learning rate through a cosine function, and step decay decreases the learning rate at a preset number of iterations or training cycles;
[0079] Preferably, the loss function is constructed as follows:
[0080] When the loss function type is the cross entropy loss function, it returns a standard cross entropy loss function;
[0081] When the loss function type is not the cross entropy loss function, the loss function is dynamically imported and the specified loss function class is instantiated. The loss function is constructed by setting parameters.
[0082] (4) Network training using the linear programming incremental learning supervised training paradigm: The training sets created in (1) are sequentially used as training data input into the unified backbone network and classification head network model constructed in (2). The hyperparameters constructed in (3) are applied to the linear programming incremental learning supervised training paradigm to train the unified backbone network and classification head network model. During the training process, the validation set created in (1) is used to calculate the classification performance of the unified backbone network and classification head network model in real time. Then, the weights of the unified backbone network and classification head network model are back-propagated using the loss function, optimizer, and learning rate hyperparameters to obtain the trained unified backbone network and classification head network model weights, which are specifically:
[0083] (4-1) The training set constructed in (1) is used as training data in turn Input (2) to construct the unified backbone network and extract the feature vector list ,Will Input the classification head network constructed by (2) to obtain the weights of the old linear programming incremental learning classifier ;
[0084] (4-2) The training set constructed in (1) Input (2) to construct the unified backbone network and extract the feature vector list ,Will Input the classification head network constructed by (2) to obtain the weights of the new linear programming incremental learning classifier ;
[0085] (4-3) Use and (2) The unified backbone network and Perform adjustments, during which the classification performance of the unified backbone network and classification head network model is calculated in real time using the validation set created in (1), and then the weights of the model are backpropagated using the loss function, optimizer, and learning rate hyperparameters;
[0086] (4-4) Updated to ;
[0087] (5) Output of classification prediction results: Call the test_model_builder function to load the unified backbone network and classification head network model, then use the weights obtained in (4) as the parameters of the unified backbone network and classification head network model obtained in (2), input the test set constructed in (1) into the unified backbone network and classification head network model, and output the classification prediction results;
[0088] (6) Unified model verification: The classification prediction results output by (5) are used for performance evaluation to verify the performance of the unified backbone network and classification head network model in terms of classification accuracy and generalization ability.
[0089] The present invention also provides a unified framework incremental learning remote sensing scene classification system, based on the above-mentioned unified framework incremental learning remote sensing scene classification method, including:
[0090] A unified dataset construction module is used to build a unified database containing several image datasets and perform data enhancement on the unified database;
[0091] Unified network model construction module, used to adjust the model parameters in the experimental configuration file, realize the selection of different types of backbone networks, and combine the backbone network and classification head network into a unified model;
[0092] A unified hyperparameter building block for updating model parameters based on the gradient of the loss function to minimize the loss during training;
[0093] The network training module uses a supervised training paradigm based on linear programming incremental learning to train heterogeneous supervised datasets. It introduces a linear programming incremental learning classifier between each stage of incremental learning, adapting the deep learning classification model to new datasets while addressing the model's catastrophic forgetting problem.
[0094] The classification prediction result output module is used to input the test set data into the unified backbone network and classification head network model loaded with weights. The unified backbone network and classification head network model perform classification prediction on each image in the test set.
[0095] The unified model validation module is used to predict category labels and evaluate performance on the validation set to verify the performance of the unified model in terms of classification accuracy and generalization ability.
[0096] Example 2:
[0097] like Figure 1 As shown, the present invention provides a unified framework incremental learning remote sensing scene classification method, comprising the following steps:
[0098] (1) Unified database construction: A unified database containing several image datasets is constructed, where each image dataset consists of a training set, a validation set, and a test set. Data augmentation is then performed on the unified database. Specifically, the Indian Pines, Pavia University, Pavia Centre, Salinas, and Houston 2013 datasets are used for training. The datasets are randomly divided into training and validation sets in a ratio of 9:1, and the validation set is then augmented.
[0099] Preferably, the dataset can cover RGB modalities, SAR modalities, infrared, multispectral and hyperspectral modalities, as well as corresponding ground feature category labels;
[0100] Preferably, the data enhancement methods include: horizontal flipping, normalization, size adjustment, color dithering, color blurring, image scale enhancement, image cropping enhancement or image translation enhancement;
[0101] (2) Unified backbone network and classification head network model construction: The backbone network and classification head network are combined into a unified model, specifically:
[0102] (2-1) Segmentation and embedding: Each image in the dataset is segmented into non-overlapping patches, and then the non-overlapping patches are linearly embedded into high-dimensional space vectors through the PatchEmbed linear layer;
[0103] (2-2) Position code addition: Add the position code to the high-dimensional space vector of each tile;
[0104] (2-3) Input Dropout layer: Input the high-dimensional space vector with position encoding into the Dropout layer and output the result vector A;
[0105] (2-4) Input the linear layer of the neural network classification model: Input the result vector A into the linear layer of the neural network classification model to obtain the result vector D;
[0106] (2-4-1) Construct a block-level multi-head self-attention mechanism: divide the result vector A into several non-overlapping image blocks, and calculate the attention feature information within each non-overlapping image block and between adjacent image blocks;
[0107] (2-4-2) Input the attention feature information into the multi-layer perceptron and output the result vector B;
[0108] (2-4-3) Normalize the result vector B and output the result vector C;
[0109] (2-4-4) Add result vector A to result vector C and output result vector D;
[0110] (2-5) Merge: Downsample the result vector D and merge the downsampled data as the features extracted by the backbone network;
[0111] (2-6) Normalization and pooling: Normalize and pool the features extracted by the backbone network in sequence, output the result vector, and then input the result vector into the classification head network to complete the construction of a unified backbone network and classification head network model;
[0112] Preferably, by adjusting the model parameters in the experimental configuration file, different types of backbone networks can be selected to adapt to different data inputs and task requirements;
[0113] Preferably, the classification head network is responsible for converting the feature maps extracted by the backbone network into the final classification results. According to the parameters in the experimental configuration file, different types of classification head networks can be selected and configured as needed;
[0114] Preferably, this framework supports embedding almost all open source neural network classification models in the industry. The models that have been connected include: MiT, PVTv2, Vision Transformer (ViT), Swin Transformer, InternImage, VMamba, ConvNeXt, LSKNet, UniRepLKNet, and FocalNet.
[0115] (3) Unified hyperparameter configuration: building unified optimizer, learning rate, and loss function hyperparameters;
[0116] Preferably, an optimizer is constructed to update parameters according to the gradient of the loss function, and the optimizer is SGD, AdamW or Adam. The specific choice depends on the complexity of the model and the characteristics of the dataset.
[0117] Preferably, constructing the learning rate includes polynomial decay, cosine annealing or step decay. Polynomial decay adjusts the learning rate according to the total number of training iterations or cycles, cosine annealing smoothly adjusts the learning rate through a cosine function, and step decay decreases the learning rate at a preset number of iterations or training cycles. The learning rate determines the amplitude of parameter update in each iteration and is an important hyperparameter for regulating training speed and convergence effect. The learning rate can be set in a constant manner or dynamically adjusted according to the progress of training.
[0118] Preferably, the present invention supports three learning rate scheduling strategies: polynomial decay (poly), cosine annealing (cosine), and step decay (step); polynomial decay adjusts the learning rate according to the total number of training iterations or cycles, cosine annealing smoothly adjusts the learning rate through a cosine function, and step decay decreases the learning rate at a preset number of iterations or training cycles. Different strategies and parameter configurations can adapt to different training requirements and optimization goals;
[0119] Preferably, the loss function is constructed as follows:
[0120] When the loss function type is the cross entropy loss function, it returns a standard cross entropy loss function;
[0121] When the loss function type is not the cross entropy loss function, the loss function is dynamically imported and the specified loss function class is instantiated. The loss function is constructed by setting parameters to customize the loss function, allowing flexible selection of different loss functions to meet various training requirements.
[0122] Preferably, the loss function is an indicator that evaluates the difference between the predicted result and the actual label. It is the objective function in the training process of the unified backbone network and classification head network model. For different tasks and model designs, an appropriate loss function can be selected to optimize the training effect.
[0123] The optimal experimental hyperparameter design is as follows: training will last for 50 epochs, with a batch size of 32 (on 4 GPUs, the total batch size is 128), the optimizer is AdamW, the learning rate is set to 0.00006, the weight decay is 0.01, and there is a learning rate multiplication factor of 10.0; the learning rate scheduler uses polynomial decay (poly), the update type is iteration number (iter); the loss function is cross entropy loss (CELoss), and the class label with index 255 is ignored;
[0124] (4) Network training using the linear programming incremental learning supervised training paradigm: The training sets created in (1) are sequentially used as training data input into the unified backbone network and classification head network model constructed in (2). The hyperparameters constructed in (3) are applied to the linear programming incremental learning supervised training paradigm to train the unified backbone network and classification head network model. During the training process, the validation set created in (1) is used to calculate the classification performance of the unified backbone network and classification head network model in real time. Then, the weights of the unified backbone network and classification head network model are back-propagated using the loss function, optimizer, and learning rate hyperparameters to obtain the trained unified backbone network and classification head network model weights, which are specifically:
[0125] (4-1) The training set constructed in (1) is used as training data in turn Input (2) to construct the unified backbone network and extract the feature vector list ,Will Input the classification head network constructed by (2) to obtain the weights of the old linear programming incremental learning classifier ;
[0126] (4-2) The training set constructed in (1) Input (2) to construct the unified backbone network and extract the feature vector list ,Will Input the classification head network constructed by (2) to obtain the weights of the new linear programming incremental learning classifier ;
[0127] (4-3) Use and (2) The unified backbone network and Perform adjustments, during which the classification performance of the unified backbone network and classification head network model is calculated in real time using the validation set created in (1), and then the weights of the model are backpropagated using the loss function, optimizer, and learning rate hyperparameters;
[0128] (4-4) Updated to ;
[0129] (5) Output of classification prediction results: Call the test_model_builder function to load the unified backbone network and classification head network model, then use the weights obtained in (4) as the parameters of the unified backbone network and classification head network model obtained in (2), input the test set constructed in (1) into the unified backbone network and classification head network model, and output the classification prediction results;
[0130] (6) Unified model verification: The classification prediction results output by (5) are evaluated to verify the performance of the unified backbone network and classification head network model in terms of classification accuracy and generalization ability. Specifically:
[0131] The trained unified backbone network and classification head network models enter evaluation mode, initialize evaluation metrics, load validation data, and input images into the unified backbone network and classification head network models for inference. The predicted results are obtained using argmax to obtain classification labels. The intersection and union of the predicted results and the true labels are calculated using the intersectionAndUnion function to generate metric data. The metric data is synchronized across GPUs using torch.distributed to ensure consistency across all processes in a distributed environment. The evaluation metrics for the semantic segmentation, change detection, classification, and visual localization tasks are updated. The overall evaluation metric is calculated and the evaluation results and corresponding labels are returned.
[0132] Preferably, the evaluation indicators for the semantic segmentation task are IoU, precision, recall rate and F1 score; the evaluation indicators for the change detection task are IoU, precision, recall rate and F1 score; the evaluation indicators for the classification task are Overall Accuracy, Average Accuracy and Kappa Coefficient; the evaluation indicators for the visual localization task are PR and IoU.
[0133] The present invention also provides a unified framework incremental learning remote sensing scene classification system, based on the above-mentioned unified framework incremental learning remote sensing scene classification method, including:
[0134] A unified dataset construction module is used to build a unified database containing several image datasets and perform data enhancement on the unified database;
[0135] Unified network model construction module, used to adjust the model parameters in the experimental configuration file, realize the selection of different types of backbone networks, and combine the backbone network and classification head network into a unified model;
[0136] A unified hyperparameter building block for updating model parameters based on the gradient of the loss function to minimize the loss during training;
[0137] The network training module uses a supervised training paradigm based on linear programming incremental learning to train heterogeneous supervised datasets. It introduces a linear programming incremental learning classifier between each stage of incremental learning, adapting the deep learning classification model to new datasets while addressing the model's catastrophic forgetting problem.
[0138] The classification prediction result output module is used to input the test set data into the unified backbone network and classification head network model loaded with weights. The unified backbone network and classification head network model perform classification prediction on each image in the test set.
[0139] The unified model validation module is used to predict category labels and evaluate performance on the validation set to verify the performance of the unified model in terms of classification accuracy and generalization ability.
[0140] The system of the present invention adopts a unified framework and modular design, and has strong adaptability and scalability. With the continuous development of remote sensing technology and the emergence of new mission requirements, the system can easily add new modules or expand the functions of existing modules to adapt to new application scenarios and mission requirements.
[0141] Table 1 Unified dataset construction details
[0142]
[0143] Example 3:
[0144] The present invention provides a unified framework incremental learning remote sensing scene classification method, comprising the following steps:
[0145] (1) Unified database construction: A unified database containing several hyperspectral modal image datasets is constructed, where each hyperspectral modal image dataset consists of a training set, a validation set, and a test set. The unified database is then augmented with data. Specifically, the Indian Pines, Pavia University, Pavia Centre, Salinas, and Houston 2013 datasets are used for training. The datasets are randomly divided into training and validation sets in a ratio of 9:1. The validation set is then augmented with data. The patch size of the hyperspectral image is 7x7.
[0146] Preferably, the data enhancement methods include: horizontal flipping, normalization, size adjustment, color dithering, color blurring, image scale enhancement, image cropping enhancement or image translation enhancement;
[0147] (2) Unified backbone network and classification head network model construction: The backbone network and classification head network are combined into a unified model. Taking Swin Transformer as the model backbone network as an example, the specific steps are as follows:
[0148] (2-1) Segmentation and embedding: Swin Transformer divides each image in the hyperspectral modality image dataset into non-overlapping patches, and then linearly embeds the non-overlapping patches into high-dimensional space vectors through the PatchEmbed linear layer;
[0149] (2-2) Position code addition: Adding the position code to the high-dimensional space vector of each tile helps capture the location information of the tile;
[0150] (2-3) Input Dropout layer: Input the high-dimensional space vector with position encoding into the Dropout layer and output the result vector A;
[0151] (2-4) Input the linear layer of the neural network classification model: Input the result vector A into the linear layer of the neural network classification model to obtain the result vector D;
[0152] (2-4-1) Construct a block-level multi-head self-attention mechanism: divide the result vector A into several non-overlapping image blocks, and calculate the attention feature information within each non-overlapping image block and between adjacent image blocks;
[0153] (2-4-2) Input the attention feature information into the multi-layer perceptron and output the result vector B;
[0154] (2-4-3) Normalize the result vector B and output the result vector C;
[0155] (2-4-4) Add result vector A to result vector C and output result vector D;
[0156] (2-5) Merge: Downsample the result vector D and merge the downsampled data as the features extracted by SwinTransformer;
[0157] (2-6) Normalization and pooling: Normalize and pool the features extracted by the Swin Transformer in sequence, output the result vector, and then input the result vector into the classification head network to complete the construction of the unified backbone network and classification head network model;
[0158] (3) Unified hyperparameter configuration: building unified optimizer, learning rate, and loss function hyperparameters;
[0159] Preferably, an optimizer is constructed to update parameters according to the gradient of the loss function, and AdamW is selected as the optimizer. The specific selection depends on the complexity of the model and the characteristics of the dataset;
[0160] Preferably, constructing the learning rate includes polynomial decay, cosine annealing or step decay. Polynomial decay adjusts the learning rate according to the total number of training iterations or cycles, cosine annealing smoothly adjusts the learning rate through a cosine function, and step decay decreases the learning rate at a preset number of iterations or training cycles. The learning rate determines the amplitude of parameter update in each iteration and is an important hyperparameter for regulating training speed and convergence effect. The learning rate can be set in a constant manner or dynamically adjusted according to the progress of training.
[0161] Preferably, the present invention supports three learning rate scheduling strategies: polynomial decay (poly), cosine annealing (cosine), and step decay (step); polynomial decay adjusts the learning rate according to the total number of training iterations or cycles, cosine annealing smoothly adjusts the learning rate through a cosine function, and step decay decreases the learning rate at a preset number of iterations or training cycles. Different strategies and parameter configurations can adapt to different training requirements and optimization goals;
[0162] Preferably, the loss function is constructed as follows:
[0163] When the loss function type is the cross entropy loss function, it returns a standard cross entropy loss function;
[0164] When the loss function type is not the cross entropy loss function, the loss function is dynamically imported and the specified loss function class is instantiated. The loss function is constructed by setting parameters to customize the loss function, allowing flexible selection of different loss functions to meet various training requirements.
[0165] Preferably, the loss function is an indicator that evaluates the difference between the predicted result and the actual label. It is the objective function in the training process of the unified backbone network and classification head network model. For different tasks and model designs, an appropriate loss function can be selected to optimize the training effect.
[0166] The optimal experimental hyperparameter design is as follows: training will last for 50 epochs, with a batch size of 32 (on 4 GPUs, the total batch size is 128), the optimizer is AdamW, the learning rate is set to 0.00006, the weight decay is 0.01, and there is a learning rate multiplication factor of 10.0; the learning rate scheduler uses polynomial decay (poly), the update type is iteration number (iter); the loss function is cross entropy loss (CELoss), and the class label with index 255 is ignored;
[0167] (4) Network training using the linear programming incremental learning supervised training paradigm: The training sets created in (1) are sequentially used as training data input into the unified backbone network and classification head network model constructed in (2). The hyperparameters constructed in (3) are applied to the linear programming incremental learning supervised training paradigm to train the unified backbone network and classification head network model. During the training process, the validation set created in (1) is used to calculate the classification performance of the unified backbone network and classification head network model in real time. Then, the weights of the unified backbone network and classification head network model are back-propagated using the loss function, optimizer, and learning rate hyperparameters to obtain the trained unified backbone network and classification head network model weights, which are specifically:
[0168] (4-1) The training set constructed in (1) is used as training data in turn Input (2) to construct the unified backbone network and extract the feature vector list ,Will Input the classification head network constructed by (2) to obtain the weights of the old linear programming incremental learning classifier ;
[0169] (4-2) The training set constructed in (1) Input (2) to construct the unified backbone network and extract the feature vector list ,Will Input the classification head network constructed by (2) to obtain the weights of the new linear programming incremental learning classifier ;
[0170] (4-3) Use and (2) The unified backbone network and Perform adjustments, during which the classification performance of the unified backbone network and classification head network model is calculated in real time using the validation set created in (1), and then the weights of the model are backpropagated using the loss function, optimizer, and learning rate hyperparameters;
[0171] (4-4) Updated to ;
[0172] (5) Output of classification prediction results: Call the test_model_builder function to load the unified backbone network and classification head network model, then use the weights obtained in (4) as the parameters of the unified backbone network and classification head network model obtained in (2), input the test set constructed in (1) into the unified backbone network and classification head network model, and output the classification prediction results;
[0173] (6) Unified model verification: The classification prediction results output by (5) are evaluated to verify the performance of the unified backbone network and classification head network model in terms of classification accuracy and generalization ability. Specifically:
[0174] The trained unified backbone network and classification head network models enter evaluation mode, initialize evaluation metrics, load validation data, and input images into the unified backbone network and classification head network models for inference. The predicted results are obtained using argmax to obtain classification labels. The intersection and union of the predicted results and the true labels are calculated using the intersectionAndUnion function to generate metric data. The metric data is synchronized across GPUs using torch.distributed to ensure consistency across all processes in a distributed environment. The evaluation metrics for the semantic segmentation, change detection, classification, and visual localization tasks are updated. The overall evaluation metric is calculated and the evaluation results and corresponding labels are returned.
[0175] Preferably, the evaluation indicators for the semantic segmentation task are IoU, precision, recall rate and F1 score; the evaluation indicators for the change detection task are IoU, precision, recall rate and F1 score; the evaluation indicators for the classification task are Overall Accuracy, Average Accuracy and Kappa Coefficient; the evaluation indicators for the visual localization task are PR and IoU.
[0176] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A unified framework incremental learning remote sensing scene classification method, characterized by: The following steps are involved: S1, unified database construction: Build a unified database containing several image datasets, where each image dataset consists of a training set, a validation set, and a test set, and then perform data enhancement on the unified database; S2, unified backbone network and classification head network model construction: Adjust the model parameters in the experimental configuration file to select different types of backbone networks. The classification head network is responsible for converting the feature maps extracted by the backbone network into the final classification results. Based on the model parameters in the experimental configuration file, different types of classification head networks are selected and configured as needed. The backbone network and classification head network are combined into a unified model. S3, unified hyperparameter construction: build unified optimizer, learning rate and loss function hyperparameters; S4, linear programming incremental learning supervised training paradigm network training: the several training sets created in S1 are sequentially input as training data into the unified backbone network and classification head network model constructed in S2, and the hyperparameters constructed in S3 are applied to the linear programming incremental learning supervised training paradigm to train the unified backbone network and classification head network model. During the training process, the validation set created in S1 is used to calculate the classification performance of the unified backbone network and classification head network model in real time, and then the weights of the unified backbone network and classification head network model are back-propagated using the loss function, optimizer and learning rate hyperparameters to obtain the trained unified backbone network and classification head network model weights; S5, classification prediction result output: call the test_model_builder function to load the unified backbone network and classification head network model, then use the weights obtained in S4 as the unified backbone network and classification head network model parameters obtained in S2, input the test set built in S1 into the unified backbone network and classification head network model, and output the classification prediction result; S6, unified model verification: The classification prediction results output by S5 are used for performance evaluation to verify the performance of the unified backbone network and classification head network model in terms of classification accuracy and generalization ability.
2. The unified framework incremental learning remote sensing scene classification method according to claim 1, characterized in that: Specifically, S1 uses the Indian Pines, Pavia University, Pavia Centre, Salinas, and Houston 2013 datasets for training. The datasets are randomly divided into training and validation sets in a ratio of 9:1, and then data augmentation is performed on the validation set.
3. The unified framework incremental learning remote sensing scene classification method according to claim 2 is characterized in that: The data enhancement methods include: horizontal flipping, normalization, size adjustment, color dithering, color blurring, image scale enhancement, image cropping enhancement or image translation enhancement.
4. The unified framework incremental learning remote sensing scene classification method according to claim 1, characterized in that: The S2 is specifically: S2-1, segmentation and embedding: Each image in the dataset is segmented into non-overlapping patches, and then the non-overlapping patches are linearly embedded into a high-dimensional space vector through the PatchEmbed linear layer; S2-2, position code addition: add the position code to the high-dimensional space vector of each tile; S2-3, input Dropout layer: input the high-dimensional space vector with position encoding into the Dropout layer and output the result vector A; S2-4, inputting the linear layer of the neural network classification model: inputting the result vector A into the linear layer of the neural network classification model to obtain the result vector D; S2-5, merging: downsample the result vector D and merge the downsampled data as the features extracted by the backbone network; S2-6, normalization and pooling: Normalize and pool the features extracted by the backbone network in sequence, output the result vector, and then input the result vector into the classification head network to complete the construction of the unified backbone network and classification head network model.
5. The unified framework incremental learning remote sensing scene classification method according to claim 4 is characterized in that: The S2-4 is specifically: S2-4-1, construct a block-level multi-head self-attention mechanism: divide the result vector A into several non-overlapping image blocks, and calculate the attention feature information within each non-overlapping image block and between adjacent image blocks; S2-4-2, input the attention feature information into the multilayer perceptron and output the result vector B; S2-4-3, perform normalization operation on the result vector B and output the result vector C; S2-4-4, add result vector A and result vector C, and output result vector D.
6. The unified framework incremental learning remote sensing scene classification method according to claim 1, characterized in that: The constructing optimizer in S3 is to update the parameters according to the gradient of the loss function, and the optimizer is SGD, AdamW or Adam; the constructing learning rate includes polynomial decay, cosine annealing or step decay, polynomial decay adjusts the learning rate according to the total number of training iterations or cycles, cosine annealing smoothly adjusts the learning rate through the cosine function, and step decay decreases the learning rate at a preset number of iterations or training cycles.
7. The unified framework incremental learning remote sensing scene classification method according to claim 1, characterized in that: The loss function constructed in S3 is specifically: When the loss function type is the cross entropy loss function, it returns a standard cross entropy loss function; When the loss function type is not the cross entropy loss function, the loss function is dynamically imported and the specified loss function class is instantiated. The loss function is constructed by setting custom parameters.
8. The unified framework incremental learning remote sensing scene classification method according to claim 1, characterized in that: The S4 is specifically: S4-1, the training set constructed in S1 is used as training data in turn Input the unified backbone network constructed by S2 and extract the feature vector list ,Will Input the classification head network constructed by S2 to obtain the weights of the old linear programming incremental learning classifier ; S4-2, the training set constructed in S1 Input the unified backbone network constructed by S2 and extract the feature vector list ,Will Input the classification head network constructed by S2 to obtain the weights of the new linear programming incremental learning classifier ; S4-3, use and S2 builds a unified backbone network and Perform adjustments, using the validation set created in S1 to calculate the classification performance of the unified backbone network and classification head network model in real time, and then backpropagate the model weights using the loss function, optimizer, and learning rate hyperparameters; S4-4, will Updated to .
9. A classification method for remote sensing scenes based on a unified framework for incremental learning in a hyperspectral modality data environment, characterized by: A unified framework incremental learning remote sensing scene classification method based on any one of claims 1-8.
10. A unified framework incremental learning remote sensing scene classification system, based on a unified framework incremental learning remote sensing scene classification method according to any one of claims 1 to 9, characterized in that: include: A unified dataset construction module is used to build a unified database containing several image datasets and perform data enhancement on the unified database; The unified network model construction module is used to adjust the model parameters in the experimental configuration file to achieve the selection of different types of backbone networks. The classification head network is responsible for converting the feature maps extracted by the backbone network into the final classification results. According to the model parameters in the experimental configuration file, different types of classification head networks are selected and configured as needed to combine the backbone network and the classification head network into a unified model. A unified hyperparameter building block for updating model parameters based on the gradient of the loss function to minimize the loss during training; A network training module based on the supervised training paradigm of linear programming incremental learning is used to train heterogeneous supervised datasets. It introduces a linear programming incremental learning classifier between each stage of incremental learning, adapting the deep learning classification model to new datasets while addressing the model's catastrophic forgetting problem. The classification prediction result output module is used to input the test set data into the unified backbone network and classification head network model loaded with weights. The unified backbone network and classification head network model perform classification prediction on each image in the test set. The unified model validation module is used to predict category labels and evaluate performance on the validation set to verify the performance of the unified model in terms of classification accuracy and generalization ability.
Citation Information
Patent Citations
Small sample network intrusion detection incremental learning classification method based on branch strategy
CN117095243A
SAR (Synthetic Aperture Radar) target increment identification method based on feature fusion, increment classifier and representation learning
CN118115800A