Remote sensing image cross-scene classification method, device and equipment based on collaborative learning

Through a collaborative learning method, the source domain data with noise annotation and the target domain pseudo-label are used for bidirectional knowledge migration, which solves the problems of noise tags and cross-scene knowledge migration in the prior art, and improves the accuracy and robustness of cross-scene classification of remote sensing images.

CN120219846APending Publication Date: 2025-06-27SHENZHEN HAOJIE ZHILIAN TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510342745.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing deep learning methods perform poorly when dealing with noise labels and cross-scene knowledge transfer, resulting in reduced model performance and high computational complexity, making it difficult to meet practical application needs.

Method used

Using a collaborative learning method, the initial classification model is trained through the source domain data with noise annotation, the target domain pseudo-label is generated, and the initial model is trained through the two-way knowledge transfer mechanism to filter the high confidence pseudo-label until the model converges, and the target classification model is obtained.

Benefits of technology

It improves the accuracy and robustness of cross-scene classification of remote sensing images, reduces the computational complexity, and meets the practical application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219846A_ABST
    Figure CN120219846A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image cross-scene classification method, device and equipment based on collaborative learning, and the method comprises the steps: training and generating an initial classification model through employing source domain data with noise labeling, predicting unlabeled target domain data based on the initial classification model, generating a target domain pseudo-label, and carrying out the recognition of the target domain pseudo-label through a collaborative learning mechanism. Performing bidirectional knowledge migration on the target domain pseudo-label and the data with the noise source domain, and training the initial classification model to obtain a target domain prediction entropy value; screening a high-confidence pseudo label according to the target domain prediction entropy value, replacing the target domain pseudo label, and returning to the step of executing bidirectional knowledge migration on the target domain pseudo label and the data with the noise source domain through the collaborative learning mechanism until the model is converged to obtain a target classification model; and classifying the to-be-classified remote sensing image by using the target classification model to obtain a classification result. By adopting the method, the accuracy of cross-scene classification of the remote sensing image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image data processing, and in particular, to a remote sensing image cross-scene classification method, device and equipment based on collaborative learning. Background Art

[0002] In the field of remote sensing image cross-scene classification, traditional methods mainly rely on handcrafted feature extraction and shallow models. These methods usually extract features such as texture and color of images and then use shallow models for classification. However, these methods perform poorly in dealing with complex data distribution differences and noisy label problems. In recent years, deep learning methods have made significant progress in remote sensing image classification. These methods usually use deep convolutional neural networks (such as ResNet, VGG, etc.) to automatically extract image features and perform classification through supervised learning.

[0003] In the process of implementing the present invention, the inventors have realized that the prior art has at least the following technical problems: Most deep learning methods assume that the source domain data is clean and are not robust enough to noisy labels. The existing methods have the following disadvantages when dealing with noisy labels and cross-scene knowledge transfer problems: First, they are not robust enough to noisy labels, resulting in a decline in model performance; second, they lack an effective cross-scene knowledge transfer mechanism and cannot make full use of the complementary information between the source domain and the target domain; third, the computational complexity is relatively high and it is difficult to meet the actual application requirements. These disadvantages limit the application of existing methods in remote sensing image cross-scene classification tasks. Summary of the Invention

[0004] Embodiments of the present invention provide a remote sensing image cross-scene classification method, device, computer equipment and storage medium based on collaborative learning to improve the accuracy of remote sensing image cross-scene classification.

[0005] To solve the above technical problems, an embodiment of the present application provides a remote sensing image cross-scene classification method based on collaborative learning. The remote sensing image cross-scene classification method based on collaborative learning includes:

[0006] Training a generated initial classification model using source domain data with noisy annotations;

[0007] Predicting unlabeled target domain data based on the initial classification model to generate target domain pseudo-labels;

[0008] Performing bidirectional knowledge transfer on the target domain pseudo-labels and the source domain data with noise through a collaborative learning mechanism, and training the initial classification model to obtain a target domain prediction entropy value. The collaborative learning mechanism includes active learning and feedback learning;

[0009] Screen high-confidence pseudo-labels according to the predicted entropy value of the target domain, replace the pseudo-labels of the target domain, and return the steps of performing two-way knowledge transfer on the pseudo-labels of the target domain and the noisy source domain data through the collaborative learning mechanism, and continue to execute until the model converges to obtain a target classification model;

[0010] Use the target classification model to classify the remote sensing image to be classified to obtain a classification result.

[0011] Optionally, before training the initial classification model using the source domain data with noisy annotations, the cross-scene classification method for remote sensing images based on collaborative learning includes:

[0012] Extract the feature representations of the source domain and the target domain remote sensing images through a pre-trained deep convolutional neural network, where the pre-trained deep convolutional neural network includes at least one of a semantic feature extraction network and a spectral feature extraction network, and the semantic feature extraction network is at least one of ResNet-50 or Vision Transformer, and the spectral feature extraction network is at least one of a 3D convolutional network or a spectral attention network;

[0013] Use the feature representation of the source domain as the source domain data and the feature representation of the target domain remote sensing image as the target domain data.

[0014] Optionally, the initial classification model is a two-branch network architecture, including a shared feature extractor, a source domain classifier, and a target domain classifier, and the noise type of the noisy source domain data is label flip noise or uniform noise.

[0015] Optionally, the steps of performing two-way knowledge transfer through the pseudo-labels of the target domain and the noisy source domain data and training the initial classification model to obtain the predicted entropy value of the target domain include:

[0016] Update the parameters of the initial classification model through the joint training of the pseudo-labels of the target domain and the noisy source domain data;

[0017] Based on the pseudo-labels of the target domain, reverse-optimize the parameters of the source domain classifier, and use the optimized initial classification model to predict the distributions of the source domain and the target domain;

[0018] Align the predicted distributions of the source domain and the target domain through the symmetric KL divergence loss to obtain the predicted entropy value of the target domain.

[0019] Optionally, before training the initial classification model using the source domain data with noisy annotations, the cross-scene classification method for remote sensing images based on collaborative learning further includes:

[0020] Screen the initial sample data for reliable samples to obtain reliable sample data;

[0021] For the reliable sample data of each category, according to the number of samples in the category, the following formula is used to introduce a weight factor:

[0022]

[0023] where w c is the category weight of category c, and N c is the number of samples in category c;

[0024] The categories that appear in both the source domain and the target domain are used as common categories, and the following formula is used to determine the weights of the common categories:

[0025] w p = 1 (c ∈ K)

[0026] where w p is the category weight of the common category, and K is the set of categories that appear in both the source domain and the target domain;

[0027] According to the category weights corresponding to the reliable sample data, each reliable sample data is weighted and labeled to obtain the source domain data with noisy labels.

[0028] To solve the above technical problems, an embodiment of the present application further provides a remote sensing image cross-scene classification device based on collaborative learning, including:

[0029] A first training module, configured to train and generate an initial classification model by using the source domain data with noisy labels;

[0030] A label generation module, configured to predict the unlabeled target domain data based on the initial classification model to generate target domain pseudo-labels;

[0031] A second training module, configured to perform bidirectional knowledge transfer on the target domain pseudo-labels and the source domain data with noise through a collaborative learning mechanism, and train the initial classification model to obtain a target domain prediction entropy value. The collaborative learning mechanism includes active learning and feedback learning;

[0032] A loop iteration module, configured to screen high-confidence pseudo-labels according to the target domain prediction entropy value, replace the target domain pseudo-labels, and return to the step of performing bidirectional knowledge transfer on the target domain pseudo-labels and the source domain data with noise through the collaborative learning mechanism and continue to execute until the model converges to obtain a target classification model;

[0033] An image classification module, configured to classify the remote sensing image to be classified by using the target classification model to obtain a classification result.

[0034] Optionally, the second training module includes:

[0035] A first updating unit, configured to update the initial classification model parameters through joint training of the target domain pseudo-labels and the noisy source domain data;

[0036] A second updating unit, configured to reversely optimize the source domain classifier parameters based on the target domain pseudo-labels, and use the optimized initial classification model to predict the distributions of the source domain and the target domain;

[0037] A distribution prediction unit, configured to align the prediction distributions of the source domain and the target domain through a symmetric KL divergence loss to obtain the target domain prediction entropy value.

[0038] Optionally, the remote sensing image cross-scene classification device based on collaborative learning further includes:

[0039] A sample screening module, configured to perform reliable sample screening on the initial sample data to obtain reliable sample data;

[0040] A first weight module, configured to introduce a weight factor for the reliable sample data of each category according to the number of samples of the category by using the following formula:

[0041]

[0042] where w c is the category weight of category c, and N c is the number of samples of category c;

[0043] A second weight module, configured to use the categories that co-occur in the source domain and the target domain as common categories, and determine the weights of the common categories by using the following formula:

[0044] w p = 1(c ∈ K)

[0045] where w p is the category weight of the common category, and K is the set of categories that co-occur in the source domain and the target domain;

[0046] A data weighting module, configured to weight and label each reliable sample data according to the category weight corresponding to the reliable sample data to obtain the source domain data with noisy labels.

[0047] To solve the above technical problems, an embodiment of the present application further provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above remote sensing image cross-scene classification method based on collaborative learning are implemented.

[0048] To solve the above technical problems, an embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above cross-scene classification method for remote sensing images based on collaborative learning are implemented.

[0049] The cross-scene classification method, device, computer device, and storage medium for remote sensing images based on collaborative learning provided by the embodiments of the present invention train an initial classification model by using source domain data with noisy labels, predict unlabeled target domain data based on the initial classification model to generate target domain pseudo-labels, and perform bidirectional knowledge transfer on the target domain pseudo-labels and the source domain data with noise through a collaborative learning mechanism, and train the initial classification model to obtain the target domain prediction entropy value; screen high-confidence pseudo-labels according to the target domain prediction entropy value, replace the target domain pseudo-labels, and return to the step of performing bidirectional knowledge transfer on the target domain pseudo-labels and the source domain data with noise through a collaborative learning mechanism and continue to execute until the model converges to obtain a target classification model; use the target classification model to classify the remote sensing images to be classified to obtain a classification result. Through active learning and feedback learning, bidirectional knowledge transfer between the source domain and the target domain is realized, thereby improving the accuracy of cross-scene classification of remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0051] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;

[0052] Figure 2 is a flowchart of an embodiment of the cross-scene classification method for remote sensing images based on collaborative learning of the present application;

[0053] Figure 3 is an example diagram of the model training process of a cross-scene classification method for remote sensing images based on collaborative learning of the present application;

[0054] Figure 4 is a schematic structural diagram of an embodiment of the cross-scene classification device for remote sensing images based on collaborative learning according to the present application;

[0055] Figure 5 is a schematic structural diagram of an embodiment of the computer device according to the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.

[0057] Explanation of some proper nouns:

[0058] Collaborative learning: It refers to achieving better knowledge transfer and noise suppression through the mutual cooperation and information sharing between two models.

[0059] Robust adaptation: It refers to improving the robustness and adaptability of the model by screening reliable samples and introducing weight factors in the presence of noisy labels.

[0060] Cross-scene classification of remote sensing images: It refers to the task of image classification in the case where the data distributions of remote sensing images between different scenes are quite different.

[0061] Referring to "embodiments" herein means that the specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of this application. The phrase appearing at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0062] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0063] Please refer to Figure 1 , as Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0064] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc.

[0065] Terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, and so on.

[0066] Server 105 can be a server that provides various services, such as a background server that supports the pages displayed on terminal devices 101, 102, and 103.

[0067] It should be noted that the remote sensing image cross-scene classification method based on collaborative learning provided in the embodiments of the present application is executed by the server. Correspondingly, the remote sensing image cross-scene classification device based on collaborative learning is set in the server.

[0068] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0069] are merely illustrative. According to actual needs, there can be any number of terminal devices, networks, and servers. Terminal devices 101, 102, and 103 in the embodiments of the present application can specifically correspond to application systems in actual production. Figure 2 Figure 2 Please refer to Figure 1 which shows a remote sensing image cross-scene classification method based on collaborative learning provided by an embodiment of the present invention. Taking the application of this method in the

[0070] S201: Train and generate an initial classification model using source domain data with noisy annotations.

[0071] Among them, the source domain data refers to collecting source domain remote sensing images with noisy labels.

[0072] Preferably, the initial classification model is a dual-branch network architecture, including a shared feature extractor, a source domain classifier, and a target domain classifier. The noise type of the source domain data with noise is label flip noise or uniform noise.

[0073] ​In a specific optional implementation, before step S201, that is, before training an initial classification model using source domain data with noisy annotations, the cross-scene classification method for remote sensing images based on collaborative learning includes:

[0074] Extract the feature representations of the source domain and the target domain remote sensing images through a pre-trained deep convolutional neural network, where the pre-trained deep convolutional neural network includes at least one of a semantic feature extraction network and a spectral feature extraction network. The semantic feature extraction network is at least one of ResNet-50 or Vision Transformer, and the spectral feature extraction network is at least one of a 3D convolutional network or a spectral attention network;

[0075] Take the feature representation of the source domain as source domain data, and take the feature representation of the target domain remote sensing image as target domain data.

[0076] Specifically, in this embodiment, feature extraction is completed through a pre-trained deep convolutional neural network (ResNet-50), aiming to extract discriminative features from remote sensing images in the source domain and the target domain. The specific steps are as follows:

[0077] First, preprocess the remote sensing images in the source domain and the target domain to meet the input requirements of the pre-trained model. Then input these preprocessed images into ResNet-50, and use its convolutional layer and pooling layer to extract the high-level features of the images, capturing the spatial information and semantic information in the images.

[0078] During the feature extraction process, select the output of the middle layer of the network as the feature representation of the image. Select the feature maps before the fully connected layer, and these feature maps have a higher dimension and can fully express the feature information of the image.

[0079] Through the above steps, discriminative features can be extracted from the remote sensing images in the source domain and the target domain. These features can not only effectively represent the content of the images, but also improve the robustness of the model to noisy labels, thereby enhancing the performance of cross-scene classification.

[0080] Furthermore, in a specific implementation, perform label smoothing on the samples in the source domain with a prediction-annotation difference degree higher than the threshold γ. The label smoothing acts on the feature information after the above feature extraction, and the corrected label is used for subsequent cross-entropy loss calculation.

[0081] Preferably, the threshold γ ranges between 0.1 and 0.3.

[0082] In a specific optional implementation, before step S201, that is, before training an initial classification model using source domain data with noisy annotations, the cross-scene classification method for remote sensing images based on collaborative learning further includes:

[0083] Screen the initial sample data for reliable sample screening to obtain reliable sample data;

[0084] For the reliable sample data of each category, according to the number of samples in the category, use the following formula to introduce a weight factor:

[0085]

[0086] where w c is the category weight of category c, and N c is the number of samples in category c;

[0087] Take the categories that appear in both the source domain and the target domain as common categories, and use the following formula to determine the weights of the common categories:

[0088] w p = 1 (c ∈ K)

[0089] where w p is the category weight of the common category, and K is the set of categories that appear in both the source domain and the target domain;

[0090] According to the category weights corresponding to the reliable sample data, weight and label each reliable sample data to obtain the source domain data with noisy labels.

[0091] Specifically, this embodiment uses a robust adaptive method for sample screening, which specifically includes:

[0092] Screen reliable samples: Improve the robustness of the model to noisy labels by screening reliable samples and introducing weight factors. Specifically, first set a threshold to screen reliable samples. For each sample, if the highest probability predicted by the model is greater than the set threshold, then the sample is considered reliable and can be used for training; otherwise, the sample is considered unreliable and is not used for training temporarily. This method can effectively avoid the interference of noisy labels on model training.

[0093] Introduce weight factors: Balance the number of samples in different categories through class weights and common weight factors, and set class weights according to the reciprocal of the number of category samples to balance the number of samples in different categories. The specific formula is as follows:

[0094]

[0095] where w c is the category weight of category c, and N c is the number of samples in category c;

[0096] At the same time, consider the categories that co-occur in both the source domain and the target domain, and set higher weights for these categories to enhance the model's attention to common categories. The specific formula is as follows:

[0097] w p = 1(c ∈ K)

[0098] where w p is the category weight of the common category, and K is the set of categories that co-occur in the source domain and the target domain.

[0099] S202: Predict the unlabeled target domain data based on the initial classification model to generate target domain pseudo-labels.

[0100] Furthermore, extend the pseudo-label generation mechanism in step S201 to the object detection task for detecting target domain data. Screen the detection box pseudo-labels through the region proposal network. The generation of the object detection box depends on the above-extracted feature information, and reuse the feature extraction network to achieve cross-task migration.

[0101] S203: Through the collaborative learning mechanism, perform two-way knowledge transfer on the target domain pseudo-labels and the noisy source domain data, and train the initial classification model to obtain the target domain prediction entropy value. The collaborative learning mechanism includes active learning and feedback learning.

[0102] Figure 3 is an example diagram of the model training process of a cross-scene classification method for remote sensing images based on collaborative learning in this application. In a specific optional implementation, in step S203, perform two-way knowledge transfer on the target domain pseudo-labels and the noisy source domain data, and train the initial classification model to obtain the target domain prediction entropy value, including:

[0103] Update the parameters of the initial classification model through the joint training of the target domain pseudo-labels and the noisy source domain data;

[0104] Based on the target domain pseudo-labels, reverse-optimize the parameters of the source domain classifier, and use the optimized initial classification model to predict the distributions of the source domain and the target domain;

[0105] Align the prediction distributions of the source domain and the target domain through the symmetric KL divergence loss to obtain the target domain prediction entropy value.

[0106] Among them, the KL divergence (Kullback-Leibler Divergence, KL Divergence) is an index for measuring the difference between two probability distributions, and is also called relative entropy (Relative Entropy).

[0107] Furthermore, this embodiment also calculates the distribution difference between the source domain and the target domain in an adaptive loss manner to improve the cross-scene adaptation ability of the model.

[0108] Meanwhile, the contrastive divergence loss is introduced. By comparing the feature distributions of the source domain and the target domain, the cross-scenario adaptation ability of the model is optimized. The contrastive divergence formula is used to calculate the difference between the feature distributions of the source domain and the target domain:

[0109]

[0110] where z i is the source domain feature, z i is the target domain feature, and σ is the activation function.

[0111] By minimizing the contrastive divergence loss, the feature distributions of the source domain and the target domain are made closer, thereby improving the cross-scenario adaptation ability of the model. It should be noted that in this embodiment, the contrastive divergence loss function directly acts on the output of the classifier after feature extraction, constraining the alignment of the feature space, which helps to improve the accuracy of prediction.

[0112] In this embodiment, by introducing a collaborative learning mechanism and a robust adaptive strategy, the present invention can make full use of the complementary information between the source domain and the target domain in the presence of noisy labels to achieve efficient cross-scenario classification.

[0113] Furthermore, in this embodiment, the training process of the source domain classifier is as follows:

[0114] Calculate the cross-entropy loss of the source domain data to measure the difference between the model prediction and the noisy labels. By minimizing this loss, the model can learn the feature representation of the source domain data. The cross-entropy loss of the source domain is calculated as follows:

[0115]

[0116] where is the true label (with noise) of the source domain, is the prediction of the model for the source domain data, N s is the number of samples in the source domain. i is the i-th sample in the source domain samples, and s is the abbreviation of the source domain (source), indicating that the variable comes from the source domain.

[0117] Specifically, the training of the target domain classifier includes:

[0118] On the target domain, since there are no true labels, the generated pseudo-labels are used to guide the training of the model. Calculate the cross-entropy loss of the target domain data, which measures the difference between the model prediction and the pseudo-labels. By minimizing this loss, the model can learn the feature representation of the target domain data and gradually adapt to the distribution of the target domain. The cross-entropy loss of the target domain is calculated as follows:

[0119]

[0120] Among them, is the pseudo-label of the target domain, is the prediction of the model for the target domain data, and N t is the number of samples in the target domain. i is the i-th sample in the target domain samples, and t is the abbreviation of the target domain (target), indicating that the variable comes from the target domain.

[0121] S204: Screen high-confidence pseudo-labels according to the target domain prediction entropy value, replace the target domain pseudo-labels, and return to continue to execute the step of performing two-way knowledge transfer on the target domain pseudo-labels and the noisy source domain data through the collaborative learning mechanism until the model converges to obtain the target classification model.

[0122] Specifically, generate pseudo-labels based on the target domain classification values as the supervision signals for the target domain. Select the category with the highest probability in the target domain classification values as the pseudo-label, and at the same time consider the confidence threshold. To ensure the quality of the pseudo-labels, this embodiment sets a confidence threshold. Only when the highest probability of a certain image exceeds this threshold, it is determined that the pseudo-label is reliable and can be used for subsequent training. In this way, high-quality pseudo-labels are screened out to avoid the interference of noisy labels on model training.

[0123] Furthermore, output the final classification results, including the prediction probabilities and classification labels of each category.

[0124] S205: Classify the remote sensing image to be classified using the target classification model to obtain the classification result.

[0125] In this embodiment, an initial classification model is trained using the noisy-labeled source domain data, the unlabeled target domain data is predicted based on the initial classification model to generate target domain pseudo-labels, two-way knowledge transfer is performed on the target domain pseudo-labels and the noisy source domain data through the collaborative learning mechanism, and the initial classification model is trained to obtain the target domain prediction entropy value; high-confidence pseudo-labels are screened according to the target domain prediction entropy value, the target domain pseudo-labels are replaced, and return to continue to execute the step of performing two-way knowledge transfer on the target domain pseudo-labels and the noisy source domain data through the collaborative learning mechanism until the model converges to obtain the target classification model; the target classification model is used to classify the remote sensing image to be classified to obtain the classification result. Through active learning and feedback learning, two-way knowledge transfer between the source domain and the target domain is realized, thereby improving the accuracy of cross-scene classification of remote sensing images.

[0126] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0127] Figure 4The principle block diagram of a cross-scene classification device for remote sensing images based on collaborative learning corresponding to the above-mentioned cross-scene classification method for remote sensing images based on collaborative learning is shown. As Figure 4 shown, the cross-scene classification device for remote sensing images based on collaborative learning includes a data separation module 31, a missing simulation module 32, a local repair module 33, a secondary repair module 34, and a result summary module 35. The detailed description of each functional module is as follows:

[0128] The first training module 31 is used to train and generate an initial classification model using the source domain data with noisy labels;

[0129] The label generation module 32 is used to predict the unlabeled target domain data based on the initial classification model to generate target domain pseudo-labels;

[0130] The second training module 33 is used to perform bidirectional knowledge transfer on the target domain pseudo-labels and the source domain data with noise through a collaborative learning mechanism, and train the initial classification model to obtain the target domain prediction entropy value. The collaborative learning mechanism includes active learning and feedback learning;

[0131] The loop iteration module 34 is used to screen high-confidence pseudo-labels according to the target domain prediction entropy value, replace the target domain pseudo-labels, and return to the step of performing bidirectional knowledge transfer on the target domain pseudo-labels and the source domain data with noise through the collaborative learning mechanism and continue to execute until the model converges to obtain the target classification model;

[0132] The image classification module 35 is used to classify the remote sensing image to be classified using the target classification model to obtain the classification result.

[0133] Optionally, the second training module 33 includes:

[0134] The first update unit is used to update the parameters of the initial classification model through the joint training of the target domain pseudo-labels and the source domain data with noise;

[0135] The second update unit is used to reverse-optimize the parameters of the source domain classifier based on the target domain pseudo-labels, and use the optimized initial classification model to predict the distributions of the source domain and the target domain;

[0136] The distribution prediction unit is used to align the prediction distributions of the source domain and the target domain through the symmetric KL divergence loss to obtain the target domain prediction entropy value.

[0137] Optionally, the cross-scene classification device for remote sensing images based on collaborative learning further includes:

[0138] The sample screening module is used to perform reliability sample screening on the initial sample data to obtain reliable sample data;

[0139] The first weight module is used to introduce a weight factor for the reliable sample data of each category according to the number of samples of the category by using the following formula:

[0140]

[0141] where w c is the category weight of category c, and N c is the number of samples of category c;

[0142] The second weight module is used to take the categories that co-occur in the source domain and the target domain as common categories and determine the weights of the common categories by using the following formula:

[0143] w p = 1 (c ∈ K)

[0144] where w p is the category weight of the common category, and K is the set of categories that co-occur in the source domain and the target domain;

[0145] The data weighting module is used to weight and label each reliable sample data according to the category weight corresponding to the reliable sample data to obtain the source domain data with noisy labels.

[0146] For the specific limitations of the remote sensing image cross-scene classification device based on collaborative learning, reference can be made to the limitations of the remote sensing image cross-scene classification method based on collaborative learning in the above text, which will not be elaborated here. Each module in the above remote sensing image cross-scene classification device based on collaborative learning can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0147] To solve the above technical problems, an embodiment of the present application also provides a computer device. For details, please refer to Figure 5 , Figure 5 which is the basic structural block diagram of the computer device in this embodiment.

[0148] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are communicatively connected to each other via a system bus. It should be noted that only the computer device 4 with components connected to the memory 41, the processor 42, and the network interface 43 is shown in the figure. However, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0149] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through means such as a keyboard, a mouse, a remote control, a touchpad, or a voice control device.

[0150] The memory 41 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or D interface display memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 4. Of course, the memory 41 can also include both the internal storage unit and the external storage device of the computer device 4. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as the program code of the remote sensing image cross-scene classification method based on collaborative learning. In addition, the memory 41 can also be used to temporarily store various data that have been output or will be output.

[0151] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run the program code stored in the memory 41 or process data, such as running the program code of the cross-scene classification method of remote sensing images based on collaborative learning.

[0152] The network interface 43 may include a wireless network interface or a wired network interface, and this network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0153] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing an interface display program, and the interface display program can be executed by at least one processor to enable the at least one processor to execute the steps of the cross-scene classification method of remote sensing images based on collaborative learning as described above.

[0154] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions to enable a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0155] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The drawings of the present application show preferred embodiments, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements for some of the technical features. Any equivalent structure made by using the specification and drawings of the present application, directly or indirectly applied in other related technical fields, is equally within the scope of the patent protection of the present application.

Claims

1. A remote sensing image cross-scene classification method based on collaborative learning, characterized in that: include: Use source domain data with noise annotations to train and generate an initial classification model; Predicting the unlabeled target domain data based on the initial classification model to generate a target domain pseudo label; By means of a collaborative learning mechanism, bidirectional knowledge transfer is performed on the target domain pseudo-label and the noisy source domain data, and the initial classification model is trained to obtain a target domain prediction entropy value, wherein the collaborative learning mechanism includes active learning and feedback learning; Filter high-confidence pseudo labels according to the target domain prediction entropy value, replace the target domain pseudo labels, and return to the step of performing bidirectional knowledge transfer on the target domain pseudo labels and the noisy source domain data through the collaborative learning mechanism until the model converges to obtain a target classification model; The target classification model is used to classify the remote sensing image to be classified to obtain a classification result.

2. The remote sensing image cross-scene classification method based on collaborative learning according to claim 1, characterized in that: Before the initial classification model is generated by training the source domain data with noise annotations, the remote sensing image cross-scene classification method based on collaborative learning includes: Extracting feature representations of source domains and feature representations of target domain remote sensing images through a pre-trained deep convolutional neural network, wherein the pre-trained deep convolutional neural network includes at least one of a semantic feature extraction network and a spectral feature extraction network, the semantic feature extraction network is at least one of ResNet-50 or Vision Transformer, and the spectral feature extraction network is at least one of a 3D convolutional network or a spectral attention network; The feature representation of the source domain is used as source domain data, and the feature representation of the target domain remote sensing image is used as target domain data.

3. The remote sensing image cross-scene classification method based on collaborative learning as claimed in claim 1, characterized in that: The initial classification model is a dual-branch network architecture, including a shared feature extractor, a source domain classifier and a target domain classifier, and the noise type of the noisy source domain data is label flipping noise or uniform noise.

4. The remote sensing image cross-scene classification method based on collaborative learning as claimed in claim 3, characterized in that: The performing of bidirectional knowledge transfer through the target domain pseudo-label and the noisy source domain data, and training the initial classification model to obtain the target domain prediction entropy value, includes: Updating the initial classification model parameters by jointly training the target domain pseudo labels with the noisy source domain data; Reversely optimizing the source domain classifier parameters based on the target domain pseudo-label, and using the optimized initial classification model to predict the distribution of the source domain and the target domain; The predicted distributions of the source domain and the target domain are aligned through the symmetric KL divergence loss to obtain the predicted entropy value of the target domain.

5. The remote sensing image cross-scene classification method based on collaborative learning as claimed in claim 1, characterized in that: Before the initial classification model is generated by training the source domain data with noise annotations, the remote sensing image cross-scene classification method based on collaborative learning also includes: Conduct reliability sample screening on the initial sample data to obtain reliable sample data; For reliable sample data of each category, according to the number of samples in the category, the following formula is used to introduce weight factors: Among them, w c is the category weight of category c, N c is the number of samples of category c; The categories that appear in both the source domain and the target domain are regarded as common categories, and the weight of the common categories is determined by the following formula: In p =1(c∈K) Among them, w p is the category weight of the common category, K is the set of categories that appear together in the source domain and the target domain; According to the category weight corresponding to the reliable sample data, each reliable sample data is weighted and labeled to obtain the source domain data with noise annotation.

6. A remote sensing image cross-scene classification device based on collaborative learning, characterized in that: include: The first training module is used to train and generate an initial classification model using source domain data with noise annotations; A label generation module, used to predict the unlabeled target domain data based on the initial classification model to generate a target domain pseudo label; A second training module is used to perform bidirectional knowledge transfer on the target domain pseudo-label and the noisy source domain data through a collaborative learning mechanism, and train the initial classification model to obtain a target domain prediction entropy value, wherein the collaborative learning mechanism includes active learning and feedback learning; A loop iteration module is used to select high-confidence pseudo labels according to the target domain prediction entropy value, replace the target domain pseudo labels, and return to the step of performing bidirectional knowledge transfer on the target domain pseudo labels and the noisy source domain data through the collaborative learning mechanism until the model converges to obtain a target classification model; The image classification module is used to classify the remote sensing image to be classified using the target classification model to obtain a classification result.

7. The remote sensing image cross-scene classification device based on collaborative learning according to claim 6, characterized in that: The second training module comprises: A first updating unit, configured to update the initial classification model parameters by jointly training the target domain pseudo-labels with the noisy source domain data; A second updating unit, configured to reversely optimize the source domain classifier parameters based on the target domain pseudo-label, and use the optimized initial classification model to predict the distribution of the source domain and the target domain; The distribution prediction unit is used to align the predicted distributions of the source domain and the target domain through the symmetric KL divergence loss to obtain the predicted entropy value of the target domain.

8. The remote sensing image cross-scene classification device based on collaborative learning according to claim 6, characterized in that: The remote sensing image cross-scene classification device based on collaborative learning also includes: The sample screening module is used to perform reliability sample screening on the initial sample data to obtain reliable sample data; The first weight module is used to introduce weight factors for reliable sample data of each category according to the number of samples in each category using the following formula: Among them, w c is the category weight of category c, N c is the number of samples of category c; The second weight module is used to treat the categories that appear in both the source domain and the target domain as common categories, and use the following formula to determine the weight of the common category: In p =1(c∈K) Among them, w p is the category weight of the common category, K is the set of categories that appear together in the source domain and the target domain; The data weighting module is used to weight and label each reliable sample data according to the category weight corresponding to the reliable sample data to obtain the source domain data with noise labels.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the remote sensing image cross-scene classification method based on collaborative learning is implemented as described in any one of claims 1 to 5.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the cross-scene classification method of remote sensing images based on collaborative learning is implemented as described in any one of claims 1 to 5.

Citation Information

Cited By

  • Remote sensing data classification method based on transfer learning and label noise filtering

    CN120689757A

  • Cross-regional crop classification method and device based on domain adaptation dual mechanisms

    CN122368606A